Peter Parker

Project 06: Frontier Quality at a Fraction of the Bill

Measured, priced and dated, and reproducible from the repository

Your team pays a frontier API for every call. At what volume does a small model you own do the same job for less?

Plenty of AI work is the same narrow task, run thousands of times: read a document, pull out the same fifteen numbers, put them in a database. Paying the most capable model in the world per call for that is the default, and it is rarely questioned, because the alternative is rarely measured on the same documents. This project measured it. It fine-tuned three small open models to read company financial filings, graded every answer against the figure the company itself filed, and rented the hardware to serve them, so the comparison is quality against quality and dollars against dollars.

- of fields right, from a 2B model on filings published after it was trained
- per 1,000 filings read on one rented card
- filings a month and owning it pays against the cheapest API that does the job
- a month against the API it matches the most accurate frontier result measured

Every figure on this page is read from the project's own results file, written by the same code that writes the repository's tables.

The task, and how it is marked

Every US public company files its quarterly and annual reports with the SEC, and tags the key figures in them in a machine-readable format called XBRL. The model gets the financial statements as text and has to return fifteen fields: revenue, net income, total assets, earnings per share, the auditor, and so on.

The tags are the answer key. No model marks another model's work: each answer is compared with the figure the company filed, by code. Where a filing does not report a field, the right answer is "not reported", and a model that makes a number up is wrong.

The headline test set is - filings published after the small models' training data was collected, from companies none of the training filings came from, so a good score cannot be memory.

period_end              date
fiscal_period           FY, Q1, Q2 or Q3
revenue                 US$
cost_of_revenue         US$, or not reported
operating_income        US$, or not reported
net_income              US$
eps_basic               US$ a share, or not reported
eps_diluted             US$ a share, or not reported
shares_diluted          a count, or not reported
total_assets            US$
total_liabilities       US$, or not reported
cash_and_equivalents    US$
stockholders_equity     US$
auditor_name            a name, or not reported
state_of_incorporation  a code, or not reported

Quality against cost

Each mark is one model, its share of fields right against what 1,000 filings cost it, with its 95% interval as a whisker. The frontier APIs are what their vendors billed through the project's gateway. The fine-tunes are the card's hourly rent over the requests it served a second, measured with 32 requests in flight, with the card busy half the time. Cost runs along a log scale: each gridline is ten times the last.

Point at a mark for its figures.

Your break-even

An API is paid by the call. A card is rented by the hour, whether it serves one call or its fill. So below some volume the API is cheaper, and above it the card is. Put in what you pay and how hard you would run the card; the measured throughput does the rest.

Or price the API at a run this project measured
-calls a month and the card is the cheaper
-per 1,000 on the card, at this load
-a month on the API at your volume
-a month on cards at your volume
The monthly bill on each against the volume. The cards' bill is a staircase: a whole card's month is paid as soon as the volume needs one more. The break-even is where the API's line first meets it.

What this leaves to you. The rent is the only cost counted unless you add one. The throughput is what the card was measured serving, on these filings, at 32 requests in flight; a card you rent for a different model or longer documents will serve a different number. And "how busy" is a choice: traffic that arrives in bursts, or a tight latency target, means a card that sits idle more of the month, and every call it does serve carries more of the rent.

Every card measured, across how busy it is

Calls a month at which one card first costs no more than the API, as the project's own code works it out. Where it says "never", the card at that load already costs more per call than the API does, so no volume helps.

Where it falls short

An average can hide a field that got worse. So every fine-tune went through a release gate field by field: it passes a field only if the whole 95% interval of the difference sits above a three-point loss, after correcting for testing fifteen fields at once.

What this does not show