Skip to main content

Scalable Enterprise Deployment Options

Hrida.ai's stateless, container-first architecture means the same application runs identically whether you deploy it as a Python process on a VM, a container in a managed service, or a pod in a Kubernetes cluster. The difference between deployment patterns is how you orchestrate, scale, and operate the application — not how the application itself behaves.

Model Inference Is Independent

How you serve LLM models is separate from how you deploy Hrida.ai. You can use managed APIs (OpenAI, Anthropic, Azure OpenAI, Google Gemini) or self-hosted inference (Ollama, vLLM) with any deployment pattern. See Integration for details on connecting models.


Application Processes​

The same container image runs as two distinct process types — you'll see both as separate services in docker-compose.yaml, and both should be accounted for in any multi-instance deployment:

ProcessCommandWhat it does
hrida-ai-studioImage default (start.sh → uvicorn/FastAPI)The main application process — handles every inbound HTTP/WebSocket request: chat, auth, RAG, admin, and the Agent Builder run endpoints. This is what your frontend and reverse proxy talk to directly.
hrida-ai-studio-workerpython -m hrida_ai_studio.worker.hrida_workflow_workerA background worker dedicated to Agent Builder workflow execution. It consumes the hrida:workflow_jobs Redis Stream (a consumer group, so multiple worker replicas share the queue safely) that webhook/cron/event triggers enqueue, and runs the workflow graph to completion. See Agent Builder → Run Infrastructure → Distributed execution.
Not Celery, and not an embedding/document-processing worker

A common assumption is that a second "worker" process must be Celery-based and handle background embedding generation or document processing. Neither is true here:

  • The queue is Redis Streams (consumer groups + XCLAIM for crash recovery) — chosen because Redis is already a required dependency and Streams give at-least-once delivery without adding a second broker. There is no Celery/RQ/SQS anywhere in this codebase.
  • Document processing and embedding generation run inline in the API process (hrida-ai-studio) via FastAPI BackgroundTasks — same event loop, right after the HTTP response. They are never queued to hrida-ai-studio-worker. If embeddings are missing after an upload, check the API process's logs, not the worker's.
  • hrida-ai-studio-worker's only job is executing Agent Builder workflow runs that a trigger enqueued. If you don't use Agent Builder triggers, this process has nothing to do.

Do you need to run it? Only if you use Agent Builder triggers (webhook/cron/event-driven workflows) or want triggered workflow runs to survive an API-process restart. Chat, RAG, and manually-clicking Run in the Agent Builder editor all execute directly in the hrida-ai-studio process and work with no worker running at all.

Why run it in production​

If you do use triggers, running them through the worker instead of in-process on the API tier buys you four things:

  1. Durability for triggered runs. A webhook/cron/event-triggered workflow is enqueued to Redis rather than executed in-process — if the API server crashes or gets redeployed mid-run, the job survives in the queue and resumes from its last checkpoint (WorkflowRuns.checkpoint(), written after every node) instead of being silently lost.
  2. Isolation from user-facing traffic. A workflow doing many parallel LLM calls, a slow MCP tool, or a deep sub_workflow chain runs on the worker, not the process serving chat/page requests — a workflow-heavy burst doesn't degrade responsiveness for everyone else.
  3. Independent scaling. API replica count tracks request volume; worker replica count tracks trigger volume. These are rarely the same number in practice — decoupling them means you're not over- or under-provisioning one tier to satisfy the other.
  4. Crash recovery + throughput via multiple replicas. Worker replicas share one Redis Streams consumer group, so a crashed replica's claimed-but-unfinished job is picked up by another (XCLAIM) — and you can add replicas purely to process triggers faster without touching the API tier at all.
None of this matters without triggers

Manual "Run" clicks in the Agent Builder editor always execute in-process on the API server, regardless of whether the worker exists. For a chat/RAG-only or manual-runs-only deployment, the worker is pure overhead — its own container, its own copy of the embedding model loaded into memory at startup, one more process to monitor — for zero benefit. It's worth running in production because you use triggers, not by default.

Running it:

# Docker Compose — already defined as a separate service
docker compose up hrida-ai-studio-worker

# Bare process / systemd
python -m hrida_ai_studio.worker.hrida_workflow_worker
# or, from backend/worker/:
./start_worker.sh

Not using triggers? Turning it off: the compose service has restart: unless-stopped, so a plain docker stop hrida-ai-studio-worker comes right back on the next docker compose up or host reboot. To actually keep it off:

# Scale it to zero — persists across `docker compose up` re-runs on this host
docker compose up -d --scale hrida-ai-studio-worker=0 hrida-ai-studio-worker

Stopping it for good only affects webhook/cron/event-triggered runs — chat, RAG, and manual "Run" clicks in the editor are entirely unaffected.

Env varDefaultPurpose
REDIS_URLredis://localhost:6379Must point at the same Redis instance as the API process
WORKER_CONCURRENCY5Max concurrent workflow runs this replica processes at once
WORKER_HEALTH_PORT8081Exposes GET /live, GET /ready, GET /health (alias of /ready), and GET /metrics — see below
WORKER_SHUTDOWN_TIMEOUT_SECONDS25Max seconds to wait for in-flight jobs to finish on SIGTERM before exiting anyway. Keep below your orchestrator's terminationGracePeriodSeconds.
MAX_WORKFLOW_DELIVERY_ATTEMPTS3Redelivery attempts before a job is moved to the dead-letter stream and its run marked failed
WORKFLOW_STALE_MESSAGE_MS300000 (5 min)How long an unacked message must be idle before another replica claims it (crashed-peer recovery)
ORPHAN_RECOVERY_MIN_AGE_SECONDS120Minimum age of a Postgres run stuck in running with no matching stream entry before it's re-enqueued
ORPHAN_RECOVERY_INTERVAL_SECONDS180How often the periodic orphan-recovery scan runs (in addition to once at startup)
WORKFLOW_QUEUE_PARTITION_BY_CATALOGfalseGive each LLM catalog its own stream (hrida:workflow_jobs:{catalog_id}) so one high-volume catalog can't starve others

Health endpoints (port WORKER_HEALTH_PORT):

EndpointChecksUse for
GET /liveEvent loop is responsive — no external callsLiveness probe. Never fails on a Redis blip, so a transient outage doesn't trigger an unnecessary pod restart.
GET /readyRedis is actually reachable (ping + queue depth)Readiness probe. Returns 503 when Redis is down.
GET /healthAlias of /readyBackward compatibility for probe configs from before the /live+/ready split.
GET /metrics—Prometheus-compatible text metrics: hrida_worker_jobs_processed_total, hrida_worker_jobs_failed_total, hrida_worker_active_tasks, hrida_worker_dead_letter_total

Scale this service independently from the API process — add replicas to increase workflow throughput, or scale it to zero if triggers aren't in use. Replicas join the same Redis Streams consumer group, so jobs are load-balanced across them with no extra coordination. A crashed replica's claimed-but-unfinished job is reclaimed by another replica (XCLAIM), and any Postgres run stuck in status='running' with no matching stream entry is periodically re-enqueued by the worker's own orphan-recovery loop.

Concurrency vs. database pool size

WORKER_CONCURRENCY concurrent runs each hold a database connection during execution. Make sure DATABASE_POOL_SIZE + DATABASE_POOL_MAX_OVERFLOW on the worker comfortably exceeds WORKER_CONCURRENCY, or runs will queue on connection checkout under load instead of executing in parallel.


Shared Infrastructure Requirements​

Regardless of which deployment pattern you choose, every scaled Hrida.ai deployment requires the same set of backing services. Configure these before scaling beyond a single instance.

ComponentWhy It's RequiredOptions
PostgreSQLMulti-instance deployments require a real database. SQLite does not support concurrent writes from multiple processes.Self-managed, Amazon RDS, Azure Database for PostgreSQL, Google Cloud SQL
RedisSession management, WebSocket coordination, and configuration sync across instances.Self-managed, Amazon ElastiCache, Azure Cache for Redis, Google Memorystore
Vector DatabaseThe default ChromaDB uses a local SQLite backend that is not safe for multi-process access.PGVector (shares PostgreSQL), Milvus, Qdrant, or ChromaDB in HTTP server mode
Shared StorageUploaded files must be accessible from every instance.Shared filesystem (NFS, EFS, CephFS) or object storage (S3, GCS, Azure Blob)
Content ExtractionThe default pypdf extractor leaks memory under sustained load.Apache Tika or Docling as a sidecar service
Embedding EngineThe default SentenceTransformers model loads ~500 MB into RAM per worker process.OpenAI Embeddings API, or Ollama running an embedding model

Critical Configuration​

These environment variables must be set consistently across every instance:

# Shared secret — MUST be identical on all instances
HRIDAAI_SECRET_KEY=your-secret-key-here

# Database
DATABASE_URL=postgresql://user:password@db-host:5432/hridaai

# Vector Database
VECTOR_DB=pgvector
PGVECTOR_DB_URL=postgresql://user:password@db-host:5432/hridaai

# Redis
REDIS_URL=redis://redis-host:6379/0
WEBSOCKET_MANAGER=redis
ENABLE_WEBSOCKET_SUPPORT=true

# Content Extraction
CONTENT_EXTRACTION_ENGINE=tika
TIKA_SERVER_URL=http://tika:9998

# Embeddings
RAG_EMBEDDING_ENGINE=openai

# Storage — choose ONE:
# Option A: shared filesystem (mount the same volume to all instances, no env var needed)
# Option B: object storage (see https://docs.hrida.ai/reference/env-configuration#cloud-storage for all required vars)
# STORAGE_PROVIDER=s3

# Workers — let the orchestrator handle scaling
UVICORN_WORKERS=1

# Migrations — only ONE instance should run migrations
ENABLE_DB_MIGRATIONS=false
Database Migrations

Set ENABLE_DB_MIGRATIONS=false on all instances except one. During updates, scale down to a single instance, allow migrations to complete, then scale back up. Concurrent migrations can corrupt your database.

For the complete step-by-step scaling walkthrough, see Deployment & Scaling. For the full environment variable reference, see Environment Variable Configuration.


Choose Your Deployment Pattern​

Hrida.ai supports three production deployment patterns. Each guide covers architecture, scaling strategy, and key considerations specific to that approach.

Python / Pip on Auto-Scaling VMs​

Deploy hrida-ai-studio serve as a systemd-managed process on virtual machines in a cloud auto-scaling group (AWS ASG, Azure VMSS, GCP MIG). Best for teams with established VM-based infrastructure and strong Linux administration skills, or when regulatory requirements mandate direct OS-level control.

Container Service​

Run the official Hrida.ai container image on a managed platform such as AWS ECS/Fargate, Azure Container Apps, or Google Cloud Run. Best for teams wanting container benefits — immutable images, versioned deployments, no OS management — without Kubernetes complexity.

Kubernetes with Helm​

Deploy using the official Hrida.ai Helm chart on any Kubernetes distribution (EKS, AKS, GKE, OpenShift, Rancher, self-managed). Best for large-scale, mission-critical deployments requiring declarative infrastructure-as-code, advanced auto-scaling, and GitOps workflows.


Deployment Comparison​

Python / Pip (VMs)Container ServiceKubernetes (Helm)
Operational complexityModerate — OS patching, Python managementLow — platform-managed containersHigher — requires K8s expertise
Auto-scalingCloud ASG/VMSS with health checksPlatform-native, minimal configurationHPA with fine-grained control
Container isolationNone — process runs directly on OSFull container isolationFull container + namespace isolation
Rolling updatesManual (scale down, update, scale up)Platform-managed rolling deploymentsDeclarative rolling updates with rollback
Infrastructure-as-codeTerraform/Pulumi for VMs + config mgmtTask/service definitions (CloudFormation, Bicep, Terraform)Helm charts + GitOps (Argo CD, Flux)
Best suited forTeams with VM-centric operations, regulatory constraintsTeams wanting container benefits without K8s complexityLarge-scale, mission-critical deployments
Minimum team expertiseLinux administration, PythonContainer fundamentals, cloud platformKubernetes, Helm, cloud-native patterns

Observability​

Production deployments should include monitoring and observability regardless of deployment pattern.

Health Checks​

  • /health — Basic liveness check. Returns HTTP 200 when the application is running. Use this for load balancer and auto-scaler health checks.
  • /api/models — Verifies the application can connect to configured model backends. Requires an API key.

OpenTelemetry​

Hrida.ai supports OpenTelemetry for distributed tracing and HTTP metrics. Enable it with:

ENABLE_OTEL=true
OTEL_EXPORTER_OTLP_ENDPOINT=http://your-collector:4318
OTEL_SERVICE_NAME=hrida-ai

This auto-instruments FastAPI, SQLAlchemy, Redis, and HTTP clients — giving visibility into request latency, database query performance, and cross-service traces.

Structured Logging​

Enable JSON-formatted logs for integration with log aggregation platforms (Datadog, Loki, CloudWatch, Splunk):

LOG_FORMAT=json
GLOBAL_LOG_LEVEL=INFO

For full monitoring setup details, see Monitoring and OpenTelemetry.


Next Steps​


Need help planning your enterprise deployment? Our team works with organizations worldwide to design and implement production Hrida.ai environments.

Contact Enterprise Sales → sales@hrida.ai

Hrida.ai is proprietary software of Zlabs Innovation. See the license for terms. © 2026 Zlabs Innovation.