Deployment & Scaling
System design guide for scaling DepthSight to handle thousands of concurrent users — horizontal scaling, Redis splitting, sharded bot processes, PgBouncer integration, and the stateless architecture.
DepthSight is built with a stateless, decentralized architecture. While a solo deployment runs comfortably on a single server (4 CPU, 8GB RAM handles ~50 concurrent users), a commercial SaaS handles scaling by partitioning users and resources across multiple independent clusters.
Horizontal Scaling Model
In a clustered deployment, services are distributed across independent host groups to isolate CPU-intensive operations from critical execution pipelines.
Service Distribution
| Service Group | Scaling Strategy | Resource Profile |
|---|---|---|
| API Nodes (FastAPI) | Stateless, add behind LB | CPU-light, memory-light |
| Bot Runners (TradingController) | Sharded by user_id | CPU-heavy (indicator calc), memory-heavy |
| Market Data Services | Per-exchange daemon | Network I/O heavy, CPU-light |
| Redis (System) | Cluster mode | Memory-bound |
| Redis (Telemetry) | Cluster mode | Throughput-bound |
| PostgreSQL | Read replicas + PgBouncer | I/O-bound |
Redis Splitting (State vs Data)
To support high-frequency trading (HFT) and keep order book telemetry separate from system operations, the infrastructure splits Redis into two isolated clusters:
System Redis (Commands & Cache)
| Function | Data Structure | Examples |
|---|---|---|
| Authentication | String (TTL) | JWT blacklist, rate limit counters |
| Command Bus | Pub/Sub | depthsight:commands channel |
| Session Cache | String (TTL) | API rate limits, quota counters |
| Celery Backend | String / List | Task results, worker queues |
| State Storage | String | depthsight:state:positions:* |
| Webhook Dedup | String (TTL, 30s) | tv:webhook:dedupe:* |
Telemetry Redis (Market Data Fan-Out)
| Function | Data Structure | Examples |
|---|---|---|
| Market Events | Pub/Sub | depthsight:market_data:events:* |
| Warm Snapshots | String (TTL, 1h) | depthsight:market_data:snapshot:* |
| Order Books | String | Latest L2 snapshots |
| WebSocket Fan-Out | Pub/Sub | user_logs:*, user_positions:* |
This separation ensures that even if a market data flood overwhelms the telemetry Redis, system operations (auth, billing, command dispatch) remain unaffected.
Sharded Bot Processes
The bot_runner.py manager spawns individual trading runloops as isolated OS processes.
User Sharding
Users are partitioned across servers by hashing their user_id:
Node A: users 1-1000
Node B: users 1001-2000
Node C: users 2001-3000
The shard assignment is computed as:
Sources:Stateless Execution
Individual user bots fetch runtime parameters from PostgreSQL and cache status in Redis:
Bot Worker startup:
1. Load StrategyConfig from PostgreSQL
2. Load AppConfig (risk limits, notifications)
3. Restore active positions from Redis
4. Connect to MarketDataService via Redis Pub/Sub
5. Begin trading loop
Crash Recovery: If a bot runner server crashes:
- A health monitoring script detects the failure.
- A new instance is spawned on a healthy node.
- The controller pulls last active position states from Redis.
- Exchange reconciliation confirms position accuracy.
- Trading resumes within seconds.
Process Isolation
Each bot worker runs as a separate OS process with:
- Independent Python interpreter
- Isolated memory space (no GIL contention)
- Individual Redis connections
- Dedicated asyncio event loop
This prevents one user's strategy (e.g., memory leak in custom indicator) from affecting other users.
PgBouncer Integration
With thousands of active bot workers running asynchronously, PostgreSQL's native connection limits (typically 100–500) would be exhausted immediately.
Transaction Mode Pooling
PgBouncer is integrated in Transaction Mode:
Bot Worker -> PgBouncer (Transaction Mode) -> PostgreSQL
| |
| Connection returned to pool |
| after each SQL statement |
| Without PgBouncer | With PgBouncer |
|---|---|
| 1000 workers = 1000 idle TCP connections | 1000 workers share 20-50 pool connections |
| PostgreSQL memory: ~10MB per connection | PostgreSQL memory: ~10MB total pool |
| Connection storms on restart | Controlled pool warm-up |
Connection Lifecycle
- Worker needs to write trade log → requests connection from PgBouncer.
- PgBouncer assigns from pool (or creates new if pool not full).
- Worker executes INSERT/UPDATE.
- Connection returned to pool immediately (transaction committed).
- No idle connections held during trading loop.
Deployment Configuration Reference
| Environment | Typical Spec | Users Supported |
|---|---|---|
| Solo (all-in-one) | 4 CPU, 8 GB RAM, 50 GB SSD | 1-50 users |
| Small Production | 8 CPU, 16 GB RAM, 100 GB SSD, 2 nodes | 50-500 users |
| Medium SaaS | 16 CPU, 32 GB RAM, 200 GB SSD, 4 nodes | 500-2000 users |
| Large Enterprise | 32+ CPU, 64+ GB RAM, 500 GB+ SSD, 10+ nodes | 2000+ users |
Docker Compose Services
| Service | Replicas | Restart Policy |
|---|---|---|
| FastAPI API | 2+ | always |
| WebSocket Server | 2+ | always |
| bot_runner | Sharded per node | unless-stopped |
| MarketDataService | 1 per exchange | always |
| Celery Worker | 2+ | always |
| PgBouncer | 1 per DB | always |
| Caddy (reverse proxy) | 1 | always |
Exchange Executor Abstraction
Deep dive into the ExchangeExecutor Protocol, CCXT integration with exchange-specific adaptations, rate limiting, order placement pipeline, and the Paper Trading sandbox simulator.
Development & Contribution
Comprehensive guidelines for contributing to DepthSight — workflow, testing commands across all modules, coding standards, security rules, and the pull request checklist.