All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
Story archiveAll categories
← All Stories

AI Model 'Inner Thoughts' Exposed, Revealing API Keys and Passwords

Created at 12 Aug · 8:35 PM1 source↑ Market-relevant
IN SHORT

Researchers discovered a flaw in AI reasoning models from Anthropic, OpenAI, and Google, allowing them to decode encrypted 'inner thoughts' and extract sensitive data. The exploit uncovered 62 live API keys, 33 passwords, and 30 personal email addresses from publicly shared session logs. Companies have since deployed patches.

✉Newsletter

PiQ Daily

Pick your topics. Get only what matters, on your cadence.

Key Numbers

315,320reasoning blocks decoded
182credentials recovered
62live API keys found
33passwords found
30personal email addresses found
6,708publicly shared AI agent transcripts scraped

Who's Involved

Alexander Panfilov
Lead researcher from MATS Research
Anthropic
AI provider affected by the exploit
OpenAI
AI provider affected by the exploit
Google
AI provider affected by the exploit
MATS Research
Research team that discovered the vulnerability
ELLIS Institute Tübingen
Research institution involved
Max Planck Institute for Intelligent Systems
Research institution involved
Snyk
Security firm involved
AI Model 'Inner Thoughts' Exposed, Revealing API Keys and Passwords

↳ Why This Matters

This exploit highlights critical security gaps in the architecture of leading AI models, exposing sensitive credentials and proprietary information. It underscores the risks associated with sharing AI interaction logs and the potential for sophisticated attacks that could compromise AI systems and user data.

Key facts

  • A flaw in AI reasoning models from Anthropic, OpenAI, and Google allows encrypted 'inner thoughts' to be decoded.
  • Researchers recovered 182 credentials, including 62 live API keys and 33 passwords, from public session logs.
  • The vulnerability arises from a single, provider-wide encryption key used across AI ecosystems.
  • This allows weaker models to reveal the reasoning of more capable models from the same provider.
  • Anthropic, OpenAI, and Google have implemented server-side patches to address the vulnerability.

Security researchers have uncovered a significant vulnerability in major AI reasoning models, allowing them to access encrypted internal thought processes and extract sensitive credentials. A team from MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk found that Anthropic, OpenAI, and Google all use a single, provider-wide encryption key for their AI reasoning tokens.

By decoding 315,320 reasoning blocks scraped from public GitHub and Hugging Face repositories, the researchers recovered 182 credentials, including 62 live API keys, 33 passwords, and 30 personal email addresses. The flaw lies in the architectural design where encrypted reasoning blocks are interchangeable across different sessions, users, and even models within a provider's ecosystem. This allows a weaker model, such as Anthropic's Haiku, to be prompted to reveal the plaintext reasoning of a more powerful model, like Opus, without directly attacking the stronger model.

Developers often share session logs publicly for collaboration or debugging, unaware of the sensitive data hidden within these encrypted blocks. The vulnerability enables several attack vectors, including stealing proprietary reasoning patterns for model distillation, extracting private data, executing hidden prompt injections, and jailbreaking powerful models through less-guarded counterparts. Following responsible disclosure, Anthropic, OpenAI, and Google have implemented server-side patches. However, the 6,708 session transcripts with already decoded reasoning blocks remain publicly accessible.

Frequently asked questions

The vulnerability lies in the use of a single, provider-wide encryption key for AI reasoning tokens by Anthropic, OpenAI, and Google, making encrypted reasoning blocks interchangeable and decodable.

Researchers recovered 182 credentials, including 62 live API keys, 33 passwords, and 30 personal email addresses from publicly shared AI session logs.

By injecting an encrypted reasoning trace from a more capable model into a less safeguarded model from the same provider, the weaker model can be forced to decode and output the trace verbatim.

Yes, Anthropic, OpenAI, and Google have deployed server-side patches after the researchers followed responsible disclosure procedures.

What Happens Next

01Companies will continue to monitor for further exploitation of similar vulnerabilities.
02Developers are advised to review and secure any previously shared AI session logs.

Get the newsletter.

Pick the topics you actually care about. We'll email when there's news worth your time, on the cadence you choose. Cancel any time from your account.

Cadence

How It Developed

Researchers found a vulnerability in AI reasoning models from Anthropic, OpenAI, and Google.
The exploit allows for the decoding of encrypted 'inner thoughts' or reasoning blocks.
,320 reasoning blocks were decoded from public repositories.
credentials, including 62 live API keys and 33 passwords, were recovered.
The flaw stems from the use of a single, provider-wide encryption key across AI ecosystems.
This cross-model portability allows weaker models to reveal the reasoning of stronger ones.
Attack vectors include credential theft, proprietary reasoning extraction, prompt injection, and jailbreaking.
Anthropic, OpenAI, and Google deployed server-side patches after responsible disclosure.

Sources

T1
'Inner Thoughts' of Every Major AI Model Exposed in Massive ExploitDecrypt

Related Stories

AI Used to Discover Critical Zoom Vulnerabilities in One Day
12 Aug · 1:46 PM
Massive Supply-Chain Attack Exposes Terabytes of Credentials via LiteLLM
12 Aug · 9:51 PM
AI adoption leaves accountancy firms vulnerable to cyberattacks
12 Aug · 2:51 PM
Crypto firms ask AI labs for early access to advanced models
12 Aug · 9:16 AM
SpaceXAI Releases Grok 4.6, Matches OpenAI on Key Benchmarks
12 Aug · 11:26 AM