What the benchmark actually compares
Run heavy_20260708_140300 used Azure ACI, Gemini, seed 42, and 300 tasks per scenario. The model, prompts, tools, rubric, and workload are shared. ChorusGraph includes its productized cache, memory, Route Ledger, and deterministic routing; the competent LangGraph baseline does not include those ChorusGraph layers.
This is an integrated-stack comparison, not an engine-only microbenchmark. It does not establish universal superiority, throughput under concurrent load, or performance against CrewAI or Microsoft Agent Framework.
Results · Methodology · Fairness disclosure · Raw artifacts