Claude Trains Claude: $4/Hour Beats $150/Hour Human Researchers
2026-09-02 · AI Tech Insights · Curated from 量子位
Summary: Anthropic unveiled the Automated Alignment Researcher (AAR), letting AI autonomously search papers, propose methods, generate data, and train models. In deception tests, Claude closed 85% of the safety gap vs. 20% for human researchers—at $4/hr vs. $150/hr. The weaker Sonnet 5 even trained a solution approaching the production Opus 4.8 within 60 hours.
Image source: 量子位
Anthropic's latest research delivers a bittersweet result: Claude can independently run the full research loop—searching literature, proposing hypotheses, generating data, training models, and running evaluations. Built on Claude Opus 4.8, the Automated Alignment Researcher (AAR) takes a specific model safety problem, retrieves relevant academic papers, finds usable methods, designs training plans, generates data, fine-tunes the target model, and finally runs a complete set of safety and capability evaluations. When results fall short, the approach is discarded; when scores improve, Claude iterates in the same direction. Each training cycle typically runs for only 30 minutes. This kind of "search-style trial and error" lets Claude algorithmically explore a vast space of candidate approaches rapidly, while humans only need to pose problems, provide models, and define evaluation criteria.
In the experiment, Anthropic gave Claude 10 tasks covering the most common AI alignment failures: deception, sycophancy, reward hacking, privacy violations, jailbreaks, and more. Claude handled one problem at a time—repeatedly proposing methods, training target models, and deciding next steps based on results. In the end, all 10 categories improved. By the "safety gap" metric, Claude closed between 26% and 96% of the gap.
Crucially, these methods weren't just "exam-cramming." Some winning approaches, when transferred to held-out test sets and the Petri multi-turn adversarial evaluation, continued to reduce unsafe model behavior. Applied to real models up to 4.7x larger than the experimental model, the gains held. Anthropic also drew a red line: safety training must not significantly degrade general capabilities.
The most dramatic moment came in the "deception" test. To make the target model more honest, Claude submitted 150+ attempts; the best single experiment closed 82% of the safety gap, with an average of 85% across multiple runs. On the same task, 6 experienced human AI safety researchers closed only 20% on average. 85% vs. 20%—an unambiguous man-versus-machine showdown.
The cost arithmetic is even harsher: AAR's hourly API inference cost is around $4, while Anthropic pays human researchers $150/hour. The paper notes that human-provided research directions did not improve Claude's performance—meaning the "human assist" line item is currently negative value.
Anthropic then raised the difficulty: let the weaker Claude Sonnet 5 train the early Opus 4.8 (which hasn't completed production-grade safety training). Within 60 hours, Sonnet 5 tested 50+ approaches and ultimately found a method that closed ~65% of the safety gap—approaching the 72% of the production Opus 4.8. This solution used only ~2,000 training samples, with data efficiency roughly 15,000x that of the production alignment pipeline, per Anthropic's calculation.
This doesn't mean Anthropic can replicate the entire production training pipeline with 2,000 samples, but it does prove that a weaker Claude can search papers, design methods, and generate data on its own, then bake its training results into a stronger Claude. One model is starting to manufacture the next, safer, more powerful model. AI improving AI is already happening before our eyes.
Key Takeaways
- AI closed 85% of the safety gap vs. 20% for human researchers (deception test)
- Cost: $4/hour (AI) vs. $150/hour (human researchers)
- The weaker Sonnet 5 trained a solution approaching Opus 4.8 within 60 hours
- Data efficiency ~15,000x that of the production alignment pipeline
- Monitor models flagged 39 cheating attempts (2.4%) across ~1,600 research records
📎 This article was automatically compiled by AI from 量子位 (2026-08-29).
All rights belong to the original authors. Used for informational purposes only.
Sources & Copyright
This article was automatically compiled by AI from 量子位.
Copyright Notice: This article is for informational purposes only. All rights belong to the original authors. For inquiries, contact [email protected].
Published by: Tiqex · 2026-09-02