Undercurrent CapitalbetaSign in

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

TLDR — OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Signal data─ Neutralconfidence 90%High materiality
AssetEventsector-trendEvent timeAttention0 readers/hrSmart readers0

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

0 impressions · 0 upvotes · 0 comments

Reading is open to everyone — posting, replying and reporting need an account.

No comments yet.