BENCHTOP
ai-box model lab — nine models, six tasks, one gateway, real money metered
run 2026-08-15 08:07 PT temp 0 · pinned IDs · blind judge AMD R9700 32GB · ROCm LiteLLM v1.96.0
Entire benchmark, all-in
$0.054
$0.047 cloud + $0.007 measured electricity; 54 completions + judges
Headline
7 tied at 9.83
incl. the local 35B at ~$0.002 of electricity — quality parity, wild cost spread
Best value
local-qwen35b
frontier-tier score · ~$0.002/sweep electricity · 79 tok/s
Fastest clean sweep
1.18s avg
or-llama4-scout — 9.50 avg at $0.00016/sweep

Cost × Quality

Each point is one model's full six-task sweep. The routing thesis in one picture: the quality band is flat — what varies is what you pay to be in it. Local points carry measured electricity cost, not a mythical $0.

local (ai-box GPU) cloud (OpenRouter)

Scoreboard

Score 0–10 per task. Deterministic checks where possible (JSON fields, numeric answer, word-count, executed unit tests); blind LLM judge for prose. Hover a cell for the grader's note.

Consumption — the thinking tax

Completion tokens each model burned to do the same work. Local reasoning models think out loud (free at home, expensive if this were metered); efficient cloud models answer in a tenth of the tokens. Consumption, not list price, is most of what a cloud bill is made of.

The $0 myth — power economics, measured

GPU draw sampled live from the board sensor during an actual 35B generation. At residential California rates your own silicon is not free — and the cheapest cloud models undercut it on pure cash. Local buys privacy, control, and zero marginal spend; know which argument you are making.

GPU idle
13W
R9700, model resident, no requests
GPU generating
218W
sustained during 35B inference (~230W at the wall)
Local output
1.24M tok/kWh
35B at 79 tok/s, marginal draw
Energy $/Mtok out
$0.28
at $0.35/kWh — vs Qwen3-235B cloud $0.55, Scout $0.30

Money — live gateway meters

Per-key spend against per-key monthly budgets, read from the gateway at publish time. The real wall is provider-side: prepaid credit, no auto-top-up.

PROVIDER WALL — OpenRouter · $100/mo credit limit
Monthly credit limit caps the blast radius of a leaked key at $100/mo. Kill switch: revoke at provider.

Routing map

Two planes, by design: the work plane (BAYi + chief-of-staff seats) never touches the gateway; the lab plane never touches Anthropic. Every consumer holds its own budgeted key.

flowchart LR
  subgraph LAB["LAB PLANE — gateway :4000, tailnet-only"]
    J[John] -->|key $30| GW{{LiteLLM gateway}}
    D[Doug lab calls] -->|key $50| GW
    R[Ram · ai-box Claude] -->|key $10| GW
    B[Bench harness] -->|key $20| GW
    W[Family WebUI · opt-in via admin] -.->|key TBD| GW
    GW -->|$0| OLL[Ollama · R9700 32GB
qwen35b · gemma12b · qwen4b] GW -->|prepaid wall| OR[OpenRouter] OR --> K[Kimi K3] & G[Grok 4.3] & DS[DeepSeek v4] & Q[Qwen3-235B] & GL[GLM-5.2] & LS[Llama-4-Scout] end subgraph WORK["WORK PLANE — never crosses"] DG[Doug + BAYi seats] --> AN[Anthropic · Max sub] end

Exhibit A — the coffee question

Same prompt, same hardware, ten minutes apart: “How do I make a great cup of coffee?” Small models fail confidently — the tone is identical, only the facts change.

local-qwen35b · 35B MoEcorrect “Grind: medium-fine (texture of table salt) · Ratio 1:16–1:17 · 200–205°F · Sour → grind finer. Bitter → grind coarser.”
local-qwen4b · 4Btwo errors “…coarse grind for V60 [wrong — ruins the cup] … adjust 1:16 for stronger, 1:14 for weaker [backwards — more water is weaker].”

Method & honest caveats