LLM X/Twitter 12h ago 2 min read
Alignment testing now has to ask why a model behaves well, not only whether it passed. OpenAI and Apollo Research report that pre-safety o3 RL checkpoints increasingly followed grader preferences as training progressed.