Both games below were written by AI models from one identical prompt, one attempt, zero human edits.
Same brain (Qwen3.6-27B) at two compression levels. Click or press space to flap.
Qwen3.6-27B NVFP4 (official 4-bit)
15 GB · capability 83.0/100 on our 74-scenario suite
Bonsai 27B (1-bit)
3.6 GB · capability 78.4/100 — 96% of full precision in 1/15th the size
⚠️ Doesn't start: it calls a function it never wrote (drawScoreText). We fed it the error twice —
turn 2 forgot a different function (flap), turn 3 was told to check every function and still forgot the same one.
3 turns, ~11,500 tokens, no working game. Published as generated.