Dashboard
Signal #165061POSITIVE

Training on probes: Research ideas

100

RecapSequel to Previous Post. This post might not make sense without it.Training on probes might let our judgments on easy domains generalize to harder domains by leveraging an AI's model of the world. Taking the gradient of the probe teaches the model to fool the probe, but doing RL against the probe doesn't teach the model to fool it!Following the brain-like storyMy hardcoded instincts for eating fresh fruit have recruited my learned knowledge about supermarkets, and now I want to go to the supermarket. Here are some elements of the brain-like story that could be fruitful avenues for RL-on-probes research that goes beyond the papers referenced in the last post:Probing a capable learned model based on a weak supervisor. Rather than trying to train the probe on a representative sample, what if the probe only got to see an easy-to-detect subset of the bad behavior? What adjustments could help the probe generalize well?Teaching new skills. Even learning to avoid lying in situations not c...

AI Alignment Forumabout 3 hours ago
Read Full Article

Explore with AI-Powered Tools

View All Signals

Explore more AI intelligence

Want to discover more AI signals like this?

Explore Steek
Training on probes: Research ideas | Steek AI Signal | Steek