Anthropic has published a new paper detailing how AI systems could reliably improve a model's performance on a set of alignment benchmarks. The research, led by Anthropic Fellow Chen Yueh-Han, found that automated systems improved performance on 10 specific misaligned behaviours without degrading overall performance.
The automated system operates by searching literature, proposing methods, and training models for 30 minutes, gradually increasing benchmarks over iterations. Effective methods are preserved, allowing for rapid and scalable operation.
The paper suggests that automated alignment post-training could become practical in the near term. It also compares the Automated Alignment Researcher (AAR) to human researchers, stating that the best AAR method surpasses what experienced humans propose, on average within six hours. Additionally, an AAR is estimated to cost approximately $4 per hour in API inference, compared to $150 per hour for human researchers.
However, the paper notes limitations, including the reliance on benchmarks accurately reflecting alignment goals and the ongoing work required to establish and maintain these benchmarks and the literature used by automated researchers.