FinAdvantage Team
FinAdvantage
On our public benchmark, the aviation ground-handling billing case has an uncomfortable line: our own pipeline missed August by +$707.00, and our agent runner by +$706.99. We published it as a miss rather than re-grading against an assumption that would have made it a pass. This post is the teardown.
Byte-identical inputs across every arm, two months apart: the client contract, the rate card, and the timesheets — 39 rows for May, 52 for August. The registered goldens are May $6,422.28 and Aug $9,047.51. Month 1 was a clean win for both our arms, exact to the cent.
Both our arms overstated the registered golden by $707 because of an overtime threshold nobody has confirmed in writing. Month 2 is exact only against an 8-hour overtime assumption no contract document states — and that assumption is still in dispute with the client. Rather than silently re-grading against the assumption that would make it a pass, we published the miss with the dispute attached.
Note what did *not* happen: the engine did not invent a confident wrong number and move on. The miss is documented, root-caused to a specific disputed input, and reproducible from the published bundle. Compare that with the failure class we track across the benchmark as "wrong, and confident" — an engine that is wrong without saying so. Our $707 is wrong with the receipt attached.
On the same byte-identical inputs, 12 of 12 leading AI models got the total wrong — including a month-1 result nearly double the golden from a raw model. Giving the same models our tool set got them closer, but still wrong on every trial that finished. 0 of 12 frontier outputs were exact across both months.
A benchmark that only ever shows wins is marketing. Ours shows the $707 overrun, a −$496.01 silent reclassification on the inventory case, and two citation defects in our own Q&A answers — because the value of a published eval is that it constrains *us*, too. Every input, prompt, and trace for this case is downloadable. Reproduce the miss, then tell us where the overtime threshold should have come from.
Prove it on your books
Run your own numbers, not ours. Each workflow calculator is built from the inputs you enter — hours, rates, transaction volume — and shows the time and money saved on your books.
Compare the cost of the plan against the hours your team spends matching statements every month.
Try this calculatorSee the calendar-days and headcount-hours a faster close frees up for your team.
Try this calculatorEstimate the saving from handling AP invoices without manual data entry.
Try this calculatorTake a free 3-5 minute assessment. Get a 0-100 readiness score, your top gaps, and a shareable PDF report — mapped to the workflows that close them.
Built for CPA firms and finance teams — pick your segment in the assessment's first question. Or see the benchmark first