Researchers show encrypted chain-of-thought blocks from Anthropic, OpenAI, and Google can be replayed to extract a stronger model’s hidden reasoning in plaintext.
Read the original at simonwillison.net→Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed...
Original headline: "Stealing Reasoning Traces from Proprietary LLM APIs"
Coverage timeline
- Aug 11, 11:00 UTC Wired AI A New Trick Reveals AI Models’ Inner Thoughts
- Aug 11, 22:40 UTC Simon Willison lead source Stealing Reasoning Traces from Proprietary LLM APIs
- Aug 11, 22:57 UTC Hacker News (AI) LLM Model-Swapping Trick Can Expose AI Reasoning Traces
- Aug 12, 10:59 UTC r/LocalLLaMA Hidden Reasoning from Claude and GPT are Decoded, and it is interesting