Homeroom is a homeschool app we build with The Academy of Chaos. Photograph the schedule grid taped inside a binder and it becomes a plan. Snap a worksheet and it comes back graded. Type “we did science pages 40 to 42 and an hour at the park” and the day logs itself. Hand it an essay and it scores against a rubric. When the week falls apart, which is the resting state of a homeschool, a planner puts it back together.
Every family gets a five dollar monthly ceiling on AI. On the models we run, five dollars buys about seven hundred photo imports. On Claude Opus 4.8 it would buy forty-seven.
That is the whole decision, and none of it is ideology.
Two gaps moved in opposite directions.
In January 2024 the best open model trailed the best closed one by about eight points on the public arena, and that gap was the reason no serious product shipped on open weights. By that August it was half a point. By February 2025 it was gone, when DeepSeek's open reasoning model matched OpenAI's o1 within a rounding error, 97.3 against roughly 97 on a math benchmark, at about a twenty-seventh of the price. Reasoning models then pulled the frontier back out. As of March 2026 the arena gap averages 3.3 points, which is the number Mozilla's July assessment carries forward.
While that gap wandered, the other one fell out of the sky. The cheapest model at GPT-4-class quality dropped 112 times in three years, forty-five dollars per million tokens down to forty cents. The historical PC compute curve returns about three and a half times over a comparable stretch. Dotcom bandwidth returned about two and a half. The steepest single drop belongs to open weights: Llama 3.1 cut the floor by eleven times in one quarter, and DeepSeek kept cutting.
Put them on one neutral harness and the decision fits in three lines. GLM 5.2, open, scores 67.8 percent at forty-three cents a task. Claude Opus 4.7 scores 68.5 at a dollar ninety-eight. One point apart. Five times the price. The newer Opus 4.8 is genuinely better at 71.9, and those four points cost five and a half times.
What we actually pay.
Homeroom runs two models. GLM 5.2 reads the text and is open in the way the word is supposed to mean: MIT licensed, weights published, ours to run ourselves. A licensed model reads the photographs, and that half of the stack is not open, which is worth saying plainly, because it is the half we could not take with us.
Four metered calls, just under three cents, about a sixth of what Opus 4.8 would have charged for the same tokens. A real photo import costs seven tenths of a cent, a seventh of Gemini 3.1 Pro and a fifteenth of Opus 4.8. Vision is the expensive half of the roadmap, and the model choice there is worth more to the unit economics than everything else combined.
On text we sit at parity, not underneath it. Claude Haiku 4.5 would have come out one percent cheaper than us on the only workload we have measured. Price is not why we chose the text model. Latency is. It answers in a second and a half where a frontier alternative takes half a minute to start talking, and a parent typing a sentence about the park will not wait half a minute. We also stopped paying for reasoning nobody asked for, which cut output tokens by half, cost by a third, and latency by more than half.
Closed still wins, and where it wins is worth saying.
Reasoning. Long-context retrieval, where a search across a million tokens finds 89 percent of what it is looking for on the best closed model and 41 percent on the best open one. Agentic depth, a fourteen-point gap on terminal work, wider than on any static benchmark.
None of those are Homeroom's problems. And they are likely not your daily business problems either.
The meter sets the ceiling on your roadmap.
Microsoft cancelled most of its Claude Code licenses at the end of June, because token billing had eaten a division's annual AI budget in months. Uber's engineers billed five hundred to two thousand dollars a month until spend got capped, four months into the year. By June Microsoft was evaluating a self-hosted open model for its heaviest Copilot workload, routing around its own partner's meter.
The cloud already ran this experiment. Around a hundred thousand dollars to move a petabyte out of one bucket. Four in five enterprises now pulling something back. GEICO paying two and a half times what it expected before it repatriated. 37signals projecting ten million saved over five years by leaving. Uber's fares rose ninety-two percent once riders had rebuilt their lives around the subsidized price, which is the cleanest version of the lesson: the low price was never the product. The lock-in was. Introductory AI pricing runs out around 2027, when the providers are public and the discounts are spent.
Five dollars has to feel like enough.
That is the number we built toward, and it is worth exactly as much as our control over what an import costs. Nobody controls a price somebody else sets. The open model wins because at one point of capability and a fifth of the price, with weights we can run ourselves the day the number moves, five dollars still feels like enough in the year the discounts end.
Pick the model that lets you keep the promise you made. For the work we do, that model is open.