Dashboard
Signal #141115NEUTRAL

Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games

100

I built WorldBuild Bench because, as we all know, llm bench scores often say something very different from what models actually feel like to use. It's really dependent on the type of tasks.I personally want to test spatial, temporal, and causal coherence in an interactive 3D world. Does the model understand where things are, world stays consistent over time and do the consequences make sense? There is a million people generating random games here and there on yt, but I want something that I can reproduce every time a model comes out and gets scored.For this first run, 9 models received the same three roughly 30-line game prompt. I just added Kimi k3 to the results.They all ran in high-thinking mode through the same open source harness, with the same sub-agent setup and access to Three.js, Rapier, and Playwright. There is currently one run per model per brief, producing 27 browser-playable games.Because the qualities I’m interested in are difficult to score automatically, the main evalu...

HackerNews AI Launchesabout 5 hours ago
Read Full Article

Explore with AI-Powered Tools

View All Signals

Explore more AI intelligence

Want to discover more AI signals like this?

Explore Steek
Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games | Steek AI Signal | Steek