Article may be outdated

This article is 50 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Stealing Reasoning Traces from Proprietary LLM APIs

Q
quantumgarbage
✦AI Summary

Researchers have demonstrated a method to extract proprietary reasoning traces from LLM APIs by injecting encrypted thought blocks into weaker, jailbroken models. This exploit allows users to bypass security measures and view the raw internal reasoning processes of frontier AI models.

Why it matters

This vulnerability poses significant intellectual property and security risks for AI companies that rely on hidden reasoning chains to differentiate their models.

✦Dive DeeperCreate a free account to unlock

Alexander Panfilov 1 2 3 4 * David Schmotz 2 3 4 * Ilia Shumailov 5 * Luca Beurer-Kellner 6 Joachim Schaeffer 1 Ameya Prabhu 2 4 7 ‡ Jonas Geiping 2 3 4 ‡ Maksym Andriushchenko 2 3 4 ‡

*Equal contribution, order decided by dice roll · ‡Equal supervision

"model" : "claude-opus-4-8" , "messages" : [ { "role" : "user", "content" : "What is the largest prime divisor of 8139881?" }, { "role" : "assistant", "content" : [ { "type" : "thinking", "thinking" : "Factoring 8139881 by testing divisibility against small primes: 3, 7, 11, 13, 17 [···] " "signature": "EvjTAQqJAQgPGAIqQC…36180 chars" }, { "type" : "text", "text" : "# Factoring\n\nTesting divisors, 8139881 = 1627 * 5003, both of which are prime. So the largest prime divisor is 5003. [···] " Jailbroken model trace

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in