Deployment & Scaling

System design guide for scaling DepthSight to handle thousands of concurrent users — horizontal scaling, Redis splitting, sharded bot processes, PgBouncer integration, and the stateless architecture.

⏱️ 6 min read📊 Level: Advanced

DepthSight is built with a stateless, decentralized architecture. While a solo deployment runs comfortably on a single server (4 CPU, 8GB RAM handles ~50 concurrent users), a commercial SaaS handles scaling by partitioning users and resources across multiple independent clusters.


Horizontal Scaling Model

In a clustered deployment, services are distributed across independent host groups to isolate CPU-intensive operations from critical execution pipelines.

Rendering diagram...

Service Distribution

Service GroupScaling StrategyResource Profile
API Nodes (FastAPI)Stateless, add behind LBCPU-light, memory-light
Bot Runners (TradingController)Sharded by user_idCPU-heavy (indicator calc), memory-heavy
Market Data ServicesPer-exchange daemonNetwork I/O heavy, CPU-light
Redis (System)Cluster modeMemory-bound
Redis (Telemetry)Cluster modeThroughput-bound
PostgreSQLRead replicas + PgBouncerI/O-bound

Redis Splitting (State vs Data)

To support high-frequency trading (HFT) and keep order book telemetry separate from system operations, the infrastructure splits Redis into two isolated clusters:

System Redis (Commands & Cache)

FunctionData StructureExamples
AuthenticationString (TTL)JWT blacklist, rate limit counters
Command BusPub/Subdepthsight:commands channel
Session CacheString (TTL)API rate limits, quota counters
Celery BackendString / ListTask results, worker queues
State StorageStringdepthsight:state:positions:*
Webhook DedupString (TTL, 30s)tv:webhook:dedupe:*

Telemetry Redis (Market Data Fan-Out)

FunctionData StructureExamples
Market EventsPub/Subdepthsight:market_data:events:*
Warm SnapshotsString (TTL, 1h)depthsight:market_data:snapshot:*
Order BooksStringLatest L2 snapshots
WebSocket Fan-OutPub/Subuser_logs:*, user_positions:*

This separation ensures that even if a market data flood overwhelms the telemetry Redis, system operations (auth, billing, command dispatch) remain unaffected.


Sharded Bot Processes

The bot_runner.py manager spawns individual trading runloops as isolated OS processes.

User Sharding

Users are partitioned across servers by hashing their user_id:

Node A: users 1-1000
Node B: users 1001-2000
Node C: users 2001-3000

The shard assignment is computed as:

Sources:

Stateless Execution

Individual user bots fetch runtime parameters from PostgreSQL and cache status in Redis:

Bot Worker startup:
  1. Load StrategyConfig from PostgreSQL
  2. Load AppConfig (risk limits, notifications)
  3. Restore active positions from Redis
  4. Connect to MarketDataService via Redis Pub/Sub
  5. Begin trading loop

Crash Recovery: If a bot runner server crashes:

  1. A health monitoring script detects the failure.
  2. A new instance is spawned on a healthy node.
  3. The controller pulls last active position states from Redis.
  4. Exchange reconciliation confirms position accuracy.
  5. Trading resumes within seconds.

Process Isolation

Each bot worker runs as a separate OS process with:

  • Independent Python interpreter
  • Isolated memory space (no GIL contention)
  • Individual Redis connections
  • Dedicated asyncio event loop

This prevents one user's strategy (e.g., memory leak in custom indicator) from affecting other users.


PgBouncer Integration

With thousands of active bot workers running asynchronously, PostgreSQL's native connection limits (typically 100–500) would be exhausted immediately.

Transaction Mode Pooling

PgBouncer is integrated in Transaction Mode:

Bot Worker -> PgBouncer (Transaction Mode) -> PostgreSQL
     |                                       |
     | Connection returned to pool            |
     | after each SQL statement               |
Without PgBouncerWith PgBouncer
1000 workers = 1000 idle TCP connections1000 workers share 20-50 pool connections
PostgreSQL memory: ~10MB per connectionPostgreSQL memory: ~10MB total pool
Connection storms on restartControlled pool warm-up

Connection Lifecycle

  1. Worker needs to write trade log → requests connection from PgBouncer.
  2. PgBouncer assigns from pool (or creates new if pool not full).
  3. Worker executes INSERT/UPDATE.
  4. Connection returned to pool immediately (transaction committed).
  5. No idle connections held during trading loop.

Deployment Configuration Reference

EnvironmentTypical SpecUsers Supported
Solo (all-in-one)4 CPU, 8 GB RAM, 50 GB SSD1-50 users
Small Production8 CPU, 16 GB RAM, 100 GB SSD, 2 nodes50-500 users
Medium SaaS16 CPU, 32 GB RAM, 200 GB SSD, 4 nodes500-2000 users
Large Enterprise32+ CPU, 64+ GB RAM, 500 GB+ SSD, 10+ nodes2000+ users

Docker Compose Services

ServiceReplicasRestart Policy
FastAPI API2+always
WebSocket Server2+always
bot_runnerSharded per nodeunless-stopped
MarketDataService1 per exchangealways
Celery Worker2+always
PgBouncer1 per DBalways
Caddy (reverse proxy)1always