Dashboard
Signal #165428NEUTRAL

Astra can do a concerning amount with no chain of thought

80

Work done in a personal capacityTLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleadingOne of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark[1] and tried it on there!Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning:No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability with...

AI Alignment Forum4 days ago
Read Full Article

Explore with AI-Powered Tools

View All Signals

Explore more AI intelligence

Want to discover more AI signals like this?

Explore Steek
Astra can do a concerning amount with no chain of thought | Steek AI Signal | Steek