Measured, priced and dated, and reproducible from the repository
Your team pays a frontier API for every call. At what volume does a small model you own do the same job for less?
Plenty of AI work is the same narrow task, run thousands of times: read a document, pull out the same fifteen numbers, put them in a database. Paying the most capable model in the world per call for that is the default, and it is rarely questioned, because the alternative is rarely measured on the same documents. This project measured it. It fine-tuned three small open models to read company financial filings, graded every answer against the figure the company itself filed, and rented the hardware to serve them, so the comparison is quality against quality and dollars against dollars.
Every figure on this page is read from the project's own results file, written by the same code that writes the repository's tables.
The task, and how it is marked
Every US public company files its quarterly and annual reports with the SEC, and tags the key figures in them in a machine-readable format called XBRL. The model gets the financial statements as text and has to return fifteen fields: revenue, net income, total assets, earnings per share, the auditor, and so on.
The tags are the answer key. No model marks another model's work: each answer is compared with the figure the company filed, by code. Where a filing does not report a field, the right answer is "not reported", and a model that makes a number up is wrong.
The headline test set is - filings published after the small models' training data was collected, from companies none of the training filings came from, so a good score cannot be memory.
period_end date fiscal_period FY, Q1, Q2 or Q3 revenue US$ cost_of_revenue US$, or not reported operating_income US$, or not reported net_income US$ eps_basic US$ a share, or not reported eps_diluted US$ a share, or not reported shares_diluted a count, or not reported total_assets US$ total_liabilities US$, or not reported cash_and_equivalents US$ stockholders_equity US$ auditor_name a name, or not reported state_of_incorporation a code, or not reported
Quality against cost
Each mark is one model, its share of fields right against what 1,000 filings cost it, with its 95% interval as a whisker. The frontier APIs are what their vendors billed through the project's gateway. The fine-tunes are the card's hourly rent over the requests it served a second, measured with 32 requests in flight, with the card busy half the time. Cost runs along a log scale: each gridline is ten times the last.
Point at a mark for its figures.
Your break-even
An API is paid by the call. A card is rented by the hour, whether it serves one call or its fill. So below some volume the API is cheaper, and above it the card is. Put in what you pay and how hard you would run the card; the measured throughput does the rest.
What this leaves to you. The rent is the only cost counted unless you add one. The throughput is what the card was measured serving, on these filings, at 32 requests in flight; a card you rent for a different model or longer documents will serve a different number. And "how busy" is a choice: traffic that arrives in bursts, or a tight latency target, means a card that sits idle more of the month, and every call it does serve carries more of the rent.
Every card measured, across how busy it is
Calls a month at which one card first costs no more than the API, as the project's own code works it out. Where it says "never", the card at that load already costs more per call than the API does, so no volume helps.
Where it falls short
An average can hide a field that got worse. So every fine-tune went through a release gate field by field: it passes a field only if the whole 95% interval of the difference sits above a three-point loss, after correcting for testing fifteen fields at once.
What this does not show
- One task. Fifteen fields from US financial filings. The method transfers to any narrow, high-volume task with an answer key; the numbers do not.
- Prices move. Every rent and every API price here carries the day it was read. The calculator is there so you can put in today's.
- OpenAI's costs are recomputed. The gateway first used priced GPT-5.6's cache writes as plain input; they are billed at 1.25 times. The costs here are recomputed from the bytes each call returned, which raised OpenAI's runs by 16 to 22%.
- Labels have errors. A hand audit of 200 items found 3.0% wrong, all since fixed and none in the test sets; the repository says how they were found.
- Training is not counted. The fine-tune is a one-off cost, tens of dollars of rented time here rather than a monthly one, and so is labelling the data if you have no answer key. Spread it over the months you expect to run, as a fixed cost.