I built WorldBuild Bench because, as we all know, llm bench scores often say something very different from what models actually feel like to use. It's really dependent on the type of tasks.I personally want to test spatial, temporal, and causal coherence in an interactive 3D world. Does the model understand where things are, world stays consistent over time and do the consequences make sense? There is a million people generating random games here and there on yt, but I want something that I can reproduce every time a model comes out and gets scored.For this first run, 9 models received the same three roughly 30-line game prompt. I just added Kimi k3 to the results.They all ran in high-thinking mode through the same open source harness, with the same sub-agent setup and access to Three.js, Rapier, and Playwright. There is currently one run per model per brief, producing 27 browser-playable games.Because the qualities I’m interested in are difficult to score automatically, the main evalu...
Want to discover more AI signals like this?
Explore Steek