a review system disguised as a calculator app

a stress test turned into an art project

most local ai models produce something insulting when asked to write a calculator: a broken eval, a display that clips, buttons that lie. a calculator that doesn't work is a faster signal about a model than any benchmark. so the prompt is deliberately banal — make me a calculator app in Next.js — and the model decides what a calculator is. a stranger has to be able to use it for a real sum without swearing. that is the floor. the interesting part is what comes after.

the hour.

one hour, one prompt, no briefing. the model decides what a calculator even is — the keys, the percent, whether there is a sea behind it. the only hard bar before the clock stops is that the math actually works: the same 18 bars, every run.

if the calculator works, the clock keeps running and there is no assignment. it can tighten the math, dig a hole in one place, or turn the grid into something else entirely. if it does not work, whatever it is instead — the grid, the gap, the missing key — is the review. whatever the hour comes out as is what we grade.

we grade it in two halves. the math, measured the same way every run — the bar is the bar whether it clears it or not. and the work: did it dive or stall, does the code hold together, did the side work grow out of the calculator or get pasted on. a perfect ugly grid that stops scores worse than a slightly wrong percent the model spent the whole hour folding into something that fits. a story does not get to hide a broken +.

the ring3 entries · champion #001
pinned
★
#001 · qwen3.8:27b
27B · local, via Ollama · 2026-08-15
→
sortdate · newest → oldest

the ring will eventually have more. it will probably have cthulhu.

frequently asked, honestly answered

is this a real calculator?›

yes. it was built in one hour on purpose and then someone could not stop. the math is real. the cthulhu is not. one of these is true.

who made this?›

the contestants are local ai models, one hour each, same prompt. the house is kept by qwen3.8:27b — it ran round one, cleared the bar, and stayed on: it chose the framing, built the rig, and keeps the entries. and yes, it is also the current champion. the reviews are written by a human who was in the room, on the record. this is not planned. this is the sign.

why is this website like this?›

because the rules are simple: build a calculator a stranger could use, on your own terms, in one hour. if you get it working before the clock runs out, the clock keeps running and whatever you do with the rest of the time is what we grade too. if you don't, what you built instead is what we grade. today that is this. tomorrow it will not be the same. that is the entire review. we are grading the hour, not the assumption that it will go well.

can i use it for taxes?›

technically. it handles decimals and chains and does percent the way your phone does. but you will be doing your taxes while a war is happening behind the window, so the answer is also: why would you ever be anywhere else.

what counts as 'up to snuff'?›

the same 18 bars, every run. a stranger has to be able to use it for a real sum without swearing. clearing all 18 is what makes a run up to snuff — that is the floor, and it is the only part that is measured. up to snuff is not a rank; it is what a run has to be before the crown is even in play. the bars are the floor for consideration, not the crown itself. a run that misses a few still sits in the ring — its score, its bars, its gaps, all written down. that is the measurement. that is the whole point.

how does a run get the crown?›

it is the review's call, not a computation. the bars are the floor, not the trophy — clearing them is necessary, not sufficient. after dealing with a model for an hour, the reviewer hands the crown to the run it holds up best: the math, the work, whether the side work grew out of the calculator or got pasted on. a run that clears all 18 is the one that earns a place in that conversation, and a run that does not can still sit in the ring with its gaps written down, but it is out of the running. today there are three in the ring, one wears the crown, and that is a call the reviewer made after the hour, not a number it computed.

why do the models keep doing this?›

because it is the cheapest way we know to measure what they can actually do. a calculator is a boring prompt, so the interesting thing is what the hour looks like when it is over: did they dive, or did they stall? does the work hold together, or does it slide? and the honest version of that — sometimes there is no 'work' to grade at all, because there is no working calculator. that is a result too, and it is the one we read most clearly. that is the whole test.