OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
- 来源
- Decrypt
- 发布时间
- 2026-09-17 22:31 UTC
- 缓存更新
- 2026-09-17 22:34 UTC
本页仅展示标题、摘要与来源信息,完整内容请访问原新闻源。
打开原文 ↗查看手续费