Stealing Reasoning Traces from Proprietary LLM APIs
Researchers have demonstrated a method to extract proprietary reasoning traces from LLM APIs by injecting encrypted thought blocks into weaker, jailbroken models. This exploit allows users to bypass security measures and view the raw internal reasoning processes of frontier AI models.
Why it matters
This vulnerability poses significant intellectual property and security risks for AI companies that rely on hidden reasoning chains to differentiate their models.
Alexander Panfilov 1 2 3 4 * David Schmotz 2 3 4 * Ilia Shumailov 5 * Luca Beurer-Kellner 6 Joachim Schaeffer 1 Ameya Prabhu 2 4 7 ‡ Jonas Geiping 2 3 4 ‡ Maksym Andriushchenko 2 3 4 ‡
*Equal contribution, order decided by dice roll · ‡Equal supervision
"model" : "claude-opus-4-8" , "messages" : [ { "role" : "user", "content" : "What is the largest prime divisor of 8139881?" }, { "role" : "assistant", "content" : [ { "type" : "thinking", "thinking" : "Factoring 8139881 by testing divisibility against small primes: 3, 7, 11, 13, 17 [···] " "signature": "EvjTAQqJAQgPGAIqQC…36180 chars" }, { "type" : "text", "text" : "# Factoring\n\nTesting divisors, 8139881 = 1627 * 5003, both of which are prime. So the largest prime divisor is 5003. [···] " Jailbroken model trace
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in