Skip to main content
Back to Blog
Technical

Teardown: Our Own $707 Miss on the Aviation Billing Case

TechnicalSeptember 7, 20266 min read

Teardown: Our Own $707 Miss on the Aviation Billing Case

FT

FinAdvantage Team

FinAdvantage

On our public benchmark, the aviation ground-handling billing case has an uncomfortable line: our own pipeline missed August by +$707.00, and our agent runner by +$706.99. We published it as a miss rather than re-grading against an assumption that would have made it a pass. This post is the teardown.

The setup

Byte-identical inputs across every arm, two months apart: the client contract, the rate card, and the timesheets — 39 rows for May, 52 for August. The registered goldens are May $6,422.28 and Aug $9,047.51. Month 1 was a clean win for both our arms, exact to the cent.

What went wrong in month 2

Both our arms overstated the registered golden by $707 because of an overtime threshold nobody has confirmed in writing. Month 2 is exact only against an 8-hour overtime assumption no contract document states — and that assumption is still in dispute with the client. Rather than silently re-grading against the assumption that would make it a pass, we published the miss with the dispute attached.

Why this failure class matters

Note what did *not* happen: the engine did not invent a confident wrong number and move on. The miss is documented, root-caused to a specific disputed input, and reproducible from the published bundle. Compare that with the failure class we track across the benchmark as "wrong, and confident" — an engine that is wrong without saying so. Our $707 is wrong with the receipt attached.

What 12 of 12 leading AI models did

On the same byte-identical inputs, 12 of 12 leading AI models got the total wrong — including a month-1 result nearly double the golden from a raw model. Giving the same models our tool set got them closer, but still wrong on every trial that finished. 0 of 12 frontier outputs were exact across both months.

The point of publishing this

A benchmark that only ever shows wins is marketing. Ours shows the $707 overrun, a −$496.01 silent reclassification on the inventory case, and two citation defects in our own Q&A answers — because the value of a published eval is that it constrains *us*, too. Every input, prompt, and trace for this case is downloadable. Reproduce the miss, then tell us where the overtime threshold should have come from.

← Back to all articles

Prove it on your books

What it costs vs. what it saves

Run your own numbers, not ours. Each workflow calculator is built from the inputs you enter — hours, rates, transaction volume — and shows the time and money saved on your books.

Bank Reconciliation

Compare the cost of the plan against the hours your team spends matching statements every month.

Try this calculator

Month-End Close

See the calendar-days and headcount-hours a faster close frees up for your team.

Try this calculator

Invoice Processing

Estimate the saving from handling AP invoices without manual data entry.

Try this calculator

Where does your finance team stand on AI readiness?

Take a free 3-5 minute assessment. Get a 0-100 readiness score, your top gaps, and a shareable PDF report — mapped to the workflows that close them.

Built for CPA firms and finance teams — pick your segment in the assessment's first question. Or see the benchmark first

Teardown: Our Own $707 Miss on the Aviation Billing Case | FinAdvantage