ニュースへ戻るニュース要約

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

ソース
Decrypt
公開日時
2026-09-17 22:31 UTC
更新
2026-09-17 22:34 UTC

本ページはタイトル、要約、出典情報のみを表示します。

原文を開く ↗手数料を確認

関連テーマ