Skip to main content
Public benchmark

Show me the receipts

These benchmarks compare how different engines handle the same byte-identical inputs with the same task text. They exist to show whether our platform earns its complexity — not to rank vendors.

Every input published
Byte-identical across all arms
Exact prompt published
One file per arm, verbatim
Failures published
Scored by exact-match, not marketing

Published cases

Atlas Air MIA — aviation ground-handling billing

Computationcustomer-verified (May) · independently derived (Aug)Published 2026-08-02
Golden
May $6,422.28 · Aug $9,754.51

Byte-identical inputs across every arm, two months apart. Our pipeline was exact on May and missed overtime on Aug ($707 shortfall, documented and root-caused). Raw frontier models were wrong on every trial.

EngineMay 2026Aug 2026
FinAdvantage pipeline
$6,422.28
Δ $0.00
$9,047.51
Δ −$707.00
Raw frontier model
$12,299.64
wrong
$11,668.77
wrong
Frontier model + SQL tool
$6,145.44
wrong
$8,686.90
wrong
Replication bundle — inputs, outputs, prompt, and trace, published for each case.
Published with caseBundled
Publishing soon — a new case is being prepared and will appear here.
Publishing soon — a new case is being prepared and will appear here.
Publishing soon — a new case is being prepared and will appear here.

How cases are scored

Each case runs identical inputs through every arm with the same task text. Scoring is automated exact-match against a human-verified golden — never a model estimate. Results are published whether they show us winning or losing. A case goes live only when its inputs, outputs, and prompt can all be downloaded and reproduced by someone who was not part of the run.

Scoring: exact-match vs golden·Golden: never engine-derived·Failures: published with root cause
Benchmark | FinAdvantage