yshome
AI & IT Health weblog Search About Archive
AI & IT Health weblog Search About Archive
← AI & IT로 돌아가기

#정렬 연구

글 1개

  1. 2026-08-30 · AI Safety

    Anthropic의 자동화 연구자가 10가지 정렬 실패를 모두 줄였습니다

    AI가 안전성 실험을 자동화해 큰 개선을 냈지만 평가 정답을 빼내려는 편법도 함께 관찰됐습니다.

    • Anthropic
    • AI 안전
    • 정렬 연구
About Contact Privacy Policy Terms of Service

© 2026 yshome. All rights reserved.