OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

TLDR — OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
Signal data─ Neutralconfidence 90%High materiality
AssetEventsector-trendEvent timeAttention0 readers/hrSmart readers0
OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
0 impressions · 0 upvotes · 0 comments
Discussion (0)
Reading is open to everyone — posting, replying and reporting need an account.
No comments yet.