Tag: Research
-

Automated Alignment Research: 10 Powerful Lessons for Safer AI
Anthropic demonstrated that AI agents can automate much of the experimental process used to make other AI models safer. Its automated alignment research system reviewed prior work, proposed interventions, trained…

