Anthropic: AI가 스스로 AI의 안전 학습 방법을 연구하기 시작
오늘 가장 중요한 연구 발표입니다. Anthropic은 8월 28일 Automated Researchers Can Reliably Mitigate Alignment Failures 를 공개했습니다. Claude에게 논문 탐색 → 방법 제안 → 학습 데이터 제작 → 모델 훈련 → 평가를 반복하게 해 기만, 아첨(sycophancy), 개인정보 침해, reward hacking 등 10가지 alignment failure 를 자동으로 개선하도록 했습니다. anthropic.com
자세히 읽기출처 1개