MagLev.
Research & evaluation
Spotlight · Benchmark 03
← Back to Benchmark 3
Spotlight · The quadratic tax

The industry meters tokens.
Tokens are linear. The cost is not.

Two agentic coding systems ran the identical thirty-prompt benchmark on the identical frontier model, on the same day, under the same rules. Every API call on both sides was captured at per-call granularity and reconciled against the provider’s own billing export. The model was a constant. Everything measured here is the cost of how that intelligence was operated.

One system consumed 12.1× the attention compute of the other — and it is the 12.1× system whose product cannot be played. The other absorbed an overnight shutdown and two deliberate process kills, held its cost of work flat across the whole day, and delivered a working two-game arcade that opens on a double-click.

Identical model on both sidesPer-call ledgers, billing-reconciled30 sequential hidden promptsResults only — no mechanism disclosed
12.1×
The attention-compute tax

Its counterpart’s quadratic attention compute against MagLev’s, computed per call at the context each call actually carried.

7.7×
Context carried per call

490,429 tokens on the average call against 63,983 — same prompts, same model, same day.

5.6×
State read off memory

130 petabytes of key-value state streamed against 23 — the physical form the tax takes.

5.6×
Modeled attention energy

At frozen constants applied identically to both sides. The ratio is independent of the constants; the absolutes are not.

Every figure on this page is computed per request from each system’s own metered logs. Return to the Benchmark 3 results →

01 / The tax nobody is billed for

Attention is quadratic. The invoice is not.

The unit of account in commercial AI is the token, and tokens are counted linearly. But the dominant cost of a transformer holding a long context is not linear — the work of attending scales with the square of what the model must hold in view. That term appears on no invoice, no dashboard and no pricing page. It is paid in silicon time, memory bandwidth and electricity.

It is measurable. Charge every call for the context it actually carried, apply the identical formula to both systems, and the difference between two ways of operating the same model stops being a matter of opinion.

Cumulative quadratic attention · identical model, identical prompts

The compounding

12.1× at completion

At completion: 4,038 billion attention pairs against 334 billion. Both curves start at zero. Neither system is smoothed.

“The model was identical on both sides of every result. The intelligence was a constant. Everything measured is the cost of how that intelligence was operated.”Benchmark instrument: Signal Arcade, pre-registered rubric and validator manual

Input consumed: 289,843,671.0 prompt tokens against 131,484,378.0. Output produced: 803,740.0 against 1,076,583.0 — the system that ate more tokens produced less. MagLev’s output total includes user-facing conversational streaming, a category largely absent from a protocol-constrained coding agent, so every ratio quoted here is conservative in MagLev’s disfavour.

02 / What a working day does to the price of work

One system’s cost of work compounds as it works. The other’s does not.

Both systems began the day at a comparable price per unit of output. The chart below indexes each system against its own opening hour, so neither is measured against the other’s scale — only against where it started.

Compute per unit of output, indexed to self

peaks near 39× its own opening

A value of 1 means the system is spending what it spent at the start of the day to produce the same amount of work. Rising means the same output costs more as the day goes on.

Prompt tokens carried on every model call

Why it compounds

peak 974,327.0 vs 99,714.0

One line climbs toward the ceiling of the model’s window and then has to be rebuilt from nothing. The other holds a band all day. Everything else on this page follows from this single difference.

MagLev never loses the thread. Persistent memory means the project is remembered rather than re-read, so the thirtieth prompt is answered with the discipline of the first — and priced like it. A session that runs all day does not become a session that costs all day.

03 / The morning after

What it costs to remember what you were doing.

Both systems were shut down cold overnight and resumed the next morning — a symmetric, controlled test of resumption. MagLev was additionally subjected to two deliberate process kills inside a thirteen-minute window while under load; its counterpart was never killed.

MagLev — resumed

Picked the project back up and kept building. Every disruption it absorbed across the entire benchmark — the overnight shutdown, both kills, and every break in between — came to a small single-digit percentage of its total, with its output rate through the kill window indistinguishable from normal operation.

Its counterpart — re-read everything

Its largest single resumption call arrived carrying 833,108 tokens of state processed entirely from scratch, and returned a few hundred tokens of output. That one call, on its own, cost 1.36× MagLev’s entire thirty-prompt benchmark.

Calls of that shape — the whole working state rebuilt from nothing before any new work can begin — account for 79.1% of its total attention compute across the benchmark and 89% of every token it processed fresh. The single worst of them is 11.23% of its own run in one call. MagLev’s worst single call is under one percent of its own.

“Nothing was produced for it. It is the price of opening the file again.”The state rebuilt on resumption is the same state that was there the night before

04 / What the tax is, physically

Every unit of attention is state moved through memory.

Quadratic attention compute is not an abstraction: it is memory traffic. Each unit corresponds to a fixed quantum of key-value state moved through the serving hardware. To produce a single token of output, the machine must read the entire context that token is conditioned on — so the cost of one more word rises with everything already said.

State read to produce one token of output

15.0× at the end of the day

By the final calls of the benchmark, one system was reading 318 GB of its own accumulated state for every single token it wrote. The other was reading 21 GB — close to where it began.

Total state read across the benchmark

5.6×

129.8 petabytes against 23.4 petabytes, for the same thirty prompts on the same model — and only the smaller number shipped a product that runs.

Modeled energy, frozen constantsAttention / KVTotal modeled
Claude Code (Opus 4.8)2,595 kJ4,203 kJ
MagLev (Opus 4.8)467 kJ2,620 kJ

The compute ratio converts directly to a serving-energy ratio, because every architecture constant cancels between two systems running the identical model on the identical fleet. Absolute energy figures carry reference-architecture constants and are therefore estimates derived from measured token surfaces, not measured device or datacenter electricity. Against a serving-fleet reference architecture, the same measured delta for this one benchmark — one user, one day — falls in the band of a single household’s full day of electricity, against roughly a dishwasher cycle for the winning run.

“Both systems had the same brain. The one that burned the household’s day of electricity lost — it shipped the product that cannot be played.”The comparison that requires no conversion

05 / The same number, at the size of the industry

One measured delta, scaled one honest layer at a time.

Everything below this line is extrapolation, and is labelled as such. It takes the one measured per-user, per-day delta and multiplies it by stated user counts and duty cycles, with no other assumptions. It is offered because the measured number is small and the industry is not.

Layer 1 · measured

One user, one day

A single developer, a single agent, a single working session — the delta this benchmark actually measured.

Layer 2 · users

× millions

Applied across an installed base below current agentic-coding adoption. No other assumption changes.

Layer 3 · duty cycle

4h → 24h

Nobody’s roadmap is a four-hour agent. Full workday, then overnight, then always-on is the stated end state of every major lab.

Layer 4 · agents per user

1 → many

One agent per developer is already behind the curve; parallel agent fleets are the current direction of travel.

Compounded, the four layers put the waste removed by operating the same models this way on the order of tens of gigawatts continuous — the scale at which new generating capacity is planned, permitted and built over years. These are illustrative extrapolations of a single measured run pair, presented as arithmetic rather than forecast. They also assume zero growth in token consumption, which no industry forecast does: every figure here is the smallest version of itself that will ever be true.

“The fastest power plant ever built is the one nobody has to build.”Capacity is not the only way to get capacity

The intelligence was the same.
The cost of operating it was not.

Two systems, one model, thirty prompts, one day. One consumed 12.1× the attention compute of the other, ate more than twice the tokens to produce less, paid a full multiple of its rival’s entire benchmark simply to wake up in the morning — and delivered a product that cannot be played. The other absorbed an overnight shutdown and two deliberate process kills for a few percent of its budget, held its price of work flat all day, and shipped a working arcade.

All of it comes from one property: persistent memory. MagLev remembers the project instead of re-reading it. The thread is never lost, the prompt never swells, and hour nine is priced like hour one.

← Back to Benchmark 3 All benchmarks