Show me the receipts
These benchmarks compare how different engines handle the same byte-identical inputs with the same task text. They exist to show whether our platform earns its complexity — not to rank vendors.
Published cases
Atlas Air MIA — aviation ground-handling billing
Byte-identical inputs across every arm, two months apart. Our pipeline was exact on May and missed overtime on Aug ($707 shortfall, documented and root-caused). Raw frontier models were wrong on every trial.
| Engine | May 2026 | Aug 2026 |
|---|---|---|
| FinAdvantage pipeline | $6,422.28 Δ $0.00 | $9,047.51 Δ −$707.00 |
| Raw frontier model | $12,299.64 wrong | $11,668.77 wrong |
| Frontier model + SQL tool | $6,145.44 wrong | $8,686.90 wrong |
How cases are scored
Each case runs identical inputs through every arm with the same task text. Scoring is automated exact-match against a human-verified golden — never a model estimate. Results are published whether they show us winning or losing. A case goes live only when its inputs, outputs, and prompt can all be downloaded and reproduced by someone who was not part of the run.
