V5.7 Distributed Bar Runtime¶
Outcome¶
The evaluator can now host hot Strategy V2 runtimes and drive them from Kafka bar events. Runtime initialization and state restoration happen once per evaluator ownership period. Later events wake the existing runtime instead of rebuilding the strategy for every candle.
The Kafka inbox is completed only after the scheduler acknowledges the externally triggered runtime cycle. Evaluation batches run concurrently on a bounded pool while each individual strategy remains serialized by its scheduler handle and Kafka shard order.
Ownership path¶
trading control worker
|
+--> register durable bar subscriptions
+--> retain one shared bar clock per unique venue/instrument/timeframe
+--> do not start an eligible local runtime after cutover
Kafka evaluation batch
|
+--> evaluator shard fence
+--> event inbox claim
+--> per-strategy runtime lease + fencing token
+--> restore or reuse hot runtime
+--> externally trigger one serialized cycle
+--> persist state and order intents
+--> complete inbox and commit Kafka offset
Grid, martingale, price-tick, daily-session, and other non-bar-event strategies remain on the local realtime runtime path. The first distributed cutover is intentionally limited to ordinary intraday closed-bar strategies.
Order fencing¶
Every order generated by an owned runtime carries its runtime_fencing_token into both strategy_order_intents and pending_orders.
The order gateway validates the token before queueing an order. The pending-order worker validates it again atomically when claiming the order for exchange submission. A stale worker can therefore neither enqueue new live work nor send previously queued work after another evaluator has taken ownership.
Manual stop-and-close operations use token zero because they are control-plane actions rather than runtime decisions.
Shared bar clock¶
Distributed strategies still need one source of deterministic closed-bar events. The control worker now keeps one in-process clock subscription per unique full market key, not one callback per strategy. Strategy registrations only increment a reference count.
The full key remains:
bar:{venue}:{market}:{market_type}:{instrument_id}:{timeframe}
This preserves exchange-specific prices for crypto and provider-specific prices and sessions for stocks.
The later shared-market-data service will take over this clock without changing the Kafka event protocol or evaluator code.
Cutover controls¶
The current deployment examples enable the distributed bar path by default:
STRATEGY_DISTRIBUTED_BAR_ENABLED=true
STRATEGY_EVALUATOR_MODE=active
Both values are required. Enabling only the evaluator leaves the control worker authoritative; enabling only distributed routing leaves batches unexecuted.
Set STRATEGY_DISTRIBUTED_BAR_ENABLED=false to roll ordinary bar strategies
back to the trading-worker path during incident recovery. Drain evaluator
workers before changing the mode, and verify there are no processing inbox rows
or uncommitted order intents. The rollback switch does not change the ownership
of grid, martingale, DCA, tick-driven, or session-driven strategies.
Remaining work toward 100,000¶
This completes most of the ordinary bar-strategy execution spine. The remaining major systems are:
- Extract shared market ingestion and the bar clock into independently replicated services.
- Continue extracting martingale, DCA, stop-loss, and price-level protection into explicitly owned realtime actors. Durable grid actors are already implemented in the trading-worker ownership domain.
- Replace the database pending-order poller with an account-partitioned execution consumer while preserving its reconciliation logic.
- Add PgBouncer, read replicas, event retention/partitioning, and dedicated PostgreSQL deployment controls.
- Build workload generators and pass 10k, 25k, 50k, and 100k capacity and failure tests.
