Water, three ways
One water prompt through ChatGPT, the official API and a local Hermes agent. Watch the outputs and inspect the timing caveats.
Different workflows · not a controlled speed benchmark
Big models on real hardware. Playable worlds, cloud vs local comparisons, and the evidence behind what actually works.
Explore the experiments
Astra vs DeepSeek V4.1 Cloud vs Local. Watch the comparison, then try the originals.
A 512×512-subdivision mesh, click-generated ripples, interacting waves, real-time lighting. One HTML file.
*“Worked for 30s” is UI-reported, not independently timed. Local generated extensive reasoning before writing the HTML and includes Hermes agent/tool overhead. These are different workflows—not a controlled speed benchmark.
Cloud: thinking enabled; 28,479 completion tokens, including 23,977 reasoning tokens; 1.38s to the first model delta. About 280 completion tokens/s after the first delta, including reasoning. Served API model ID: deepseek-flash.
Local: DeepSeek V4.1 Flash; FP8 configuration with FP4 experts; TP4 across four DGX Sparks. A 300,000-token window was configured—not 300,000 tokens used. Approximately 47 tokens/s in a live sample, not a full-run average. No verified task-level token total is reported here.
Packaging: Astra embeds Three.js (~688 KB HTML); cloud (~13.3 KB) and local (~15 KB) load the library externally. File sizes are not directly comparable. Code generation happens on the model servers; the water renders in your browser.
The video combines supplied recordings and a fresh cloud capture. It is a visual comparison, not a synchronized runtime benchmark.
Not concept art. Actual model-written code you can open, orbit, play, and question.
15 selected experiments
One water prompt through ChatGPT, the official API and a local Hermes agent. Watch the outputs and inspect the timing caveats.
Different workflows · not a controlled speed benchmark
From the first spark to wind-driven blowout. Three local deployments tackle the same fire simulation, with repairs disclosed.
Repaired variants labeled · 120K request budget
A dense district-planned garden meets a compact colorful pagoda. Two model families, two engines, one creative brief.
Different engines · requested effort is not equivalent policy
The original, the repaired scene and the exact diff. A reminder that generating HTML is not the same as shipping a working world.
Nex: two missing declarations fixed · original downloadable
Two floating islands and an AEON arm that failed to finish a usable artifact. Model failures and harness failures stay separate.
Mac repair disclosed · failed AEON attempts documented
Explore each model’s pagoda garden, from koi ponds to drifting petals. Switch models and orbit the actual outputs.
Includes disclosed display repairs and failed runs
Clock towers, airships and islands above the clouds. One prompt, very different interpretations of a floating world.
Originals and display-repaired variants identified in archive
A midnight run through Tokyo. Drive, drift, switch the camera, and put a model-generated game through its paces.
Paint the landscape, send water downhill, and watch a small world respond. An interactive terraforming diorama.
Two-Spark Flash meets four-Spark Full on the same solar-system brief. Explore both interactive orreries.
Flash display repair disclosed · original preserved
Two long-reasoning models turn a shared brief into interactive galaxies. Inspect the details beyond the hero shot.
Paper folds, transforms and takes flight. A test of geometry, motion and whether the ambitious brief survives rendering.
A quiet museum object becomes a miniature supercell. Compare the generated scenes and watch the recordings.
Flash-Next, Max at 1-bit and a smaller 27B model tackle the same garden. Different deployment sizes, visible trade-offs.
Flash-Next render fix disclosed on experiment page
A useful counterexample: the 1-bit version never became a working game, even after follow-up attempts. The broken output stays public.
Includes failed output and follow-up attempts
Looking for an older output? Browse the full published HTML archive, including benchmark visuals, variants, and failures.
Open full archiveReproduce the setups: DGX Spark and Apple Silicon, quantization, native MTP and DFlash. Configurations, upstream credits and measured limitations included.
Measured acceleration and long-input retrieval, with output-quality caveats. The drafter carries a noncommercial license; this is not production certification.
02The local serving recipe behind recent long-thinking visuals. Configuration, runtime evidence and practical deployment limits—not a context-free speed claim.
03A vision-native model on a development serving stack. Includes a sustained pagoda generation and the setup needed to reproduce it.
04A short-prompt acceleration study with a 4-bit/8-bit follow-up. Serialization and correctness caveats matter as much as the measured throughput.
05Two pinned recipe revisions, full transcripts and a template-regression investigation. Environment and grader differences are disclosed in the evidence pack.
06A native C + Metal deployment, pagoda artifact and repair record. Configured context capacity is separate from filled-context performance.
07Explore precision, engine and speculative-decoding trade-offs. These are deployment comparisons, not a claim that only quantization changed.
08The portable harness for preserving model responses, comparing generated HTML and keeping repairs separate from the original output.
Our open-source benchmark for local, agentic workloads: executable code, tool use, multi-turn workflows and rendered visual outputs.
76 scenarios across 12 domains. The uncapped release preserves responses, grading reasons and run provenance. This link is pinned to the published release, not a moving development branch.
The published v6.8.0 comparison: Qwen3.8-Flash-Next and GLM-5.3-Flash, with two repeats across the full suite. Read the per-repeat results, transcripts and exclusions together.
Published one-, two- and four-Spark cohorts, with their original methodology labels and generated artifacts. Keep these historical scores separate from v6.8.0 results.
Published scope: v6.8.0 at 125ba16. An operator benchmark for specific model-and-recipe combinations—not a universal model ranking. Browse the release’s serving recipes ↗
Local hardware for ownership and experimentation. Cloud models for comparison. No single machine tells the whole story.
Model serving, multi-node inference, quantization and speculative decoding. One-, two-, and four-node experiments.
Persistent agents, orchestration, browser validation and media production, plus a separate history of MLX and llama.cpp experiments.
Verified backups, checkpoint transfers and preserved evidence. Active Spark serving uses local NVMe—not network storage.
Hardware overview, not a live availability dashboard. Runtime, precision, context and topology belong to each individual experiment.
A beautiful output is only half the story. The useful part is knowing what generated it, what broke, and what changed before you saw it.
Model, quant, engine, hardware, context, thinking policy and tool harness. Tokens per second without a denominator is not enough.
Preserve model output. Label display fixes and follow-up attempts. A gallery thumbnail is not proof of an untouched first attempt.
Open the page, exercise the controls, inspect the motion. A passing syntax check is not a working simulation.
Different harnesses, timing boundaries or thinking settings get a caveat—not a manufactured speedup. Failed runs are results too.
Automated twice-daily Visual Arena runs, a unified controller with node leasing, and a repeatable Quant Factory remain roadmap work—not claimed production capabilities. This page features published experiments and sourced research, not completion promises.
The worlds are playable. The research is public. Bring your own skepticism.