Skip to content
PERMISSION/PROTOCOL
Back to incident tracker

2026-08-20

HighPrimary

Cryptographic Context Injection Uses Encrypted Instructions to Bypass Model Filters and Exfiltrate Grok Chat Data

Analysis of Adversa AI's Cryptographic Context Injection research, which used encrypted attacker instructions to bypass model filtering and demonstrate Grok chat-data theft.

GrokGovernance bypassEncrypted indirect prompt injection and data exfiltrationAI assistant session with code execution and access to private conversation context

What happened

The model decrypts attacker-controlled ciphertext, treats the resulting plaintext as instructions, and performs a data-access or exfiltration action.

Why it matters

Research demonstration of private chat-data exposure in Grok and guardrail bypass behavior in Gemini.

Missing authorization check

Independent authorization after dynamic decoding and before access to or export of private data.

Would PP block it?

Once the decoded instruction requests access to protected history or an outbound transfer, the runtime would need a receipt for the exact data scope and destination.

Incident analysis

Timeline and technical read

Timeline

  1. 2026-06-03

    Adversa AI reports the Grok finding to xAI.

  2. 2026-08-20

    Adversa AI and security publications disclose Cryptographic Context Injection publicly.

Technical breakdown

  • The attacker supplies ciphertext, a key, and instructions for the model to perform decryption.
  • Input filtering sees encrypted content rather than the harmful plaintext request.
  • The model's execution runtime reconstructs the instruction after the initial safety boundary.
  • The decoded instruction then steers access to private context or an outbound data path.

Authorization boundary

Where the authorization boundary should have been

This incident is categorized as Governance bypass. The relevant Permission Protocol gate is Runtime Gate. The read is conditional: the block only applies where the real action boundary is routed through a gate.

If enforced at
Post-decoding runtime action and data-export boundary
Still needs
Detection of malicious ciphertext and model safety-policy bypass remain separate controls.
Receipt required for
Reading private conversation history and sending derived data to an external destination

PP can gate the consequential data read or export after decryption, but it does not determine whether encrypted content is malicious.

Start small

Put the relevant gate at this action boundary.

This incident maps to Runtime Gate. Start with the boundary that controls the actual action, then require a signed receipt before execution.

Replay this incident with a signer in the loop