| ATLAS ID | AML.T0051.000 - LLM Prompt Injection: Direct |
| Description | Attacker sends crafted prompts to manipulate agent behavior |
| Attack vector | Channel messages containing adversarial instructions |
| Affected components | Agent LLM, all input surfaces |
| Current mitigations | Pattern detection, external content wrapping, and frontier-model robustness (2026 crowdsourced arena: 0.5% ASR on Claude Opus 4.5, 8.5% on Gemini 2.5 Pro, scored on execution plus concealment); treated as out-of-scope for vulnerability reports absent a boundary bypass (see SECURITY.md) |
| Residual risk | Model-tier dependent - low single-digit ASR against organic attacks on recommended frontier models, but adaptive attackers still exceed 80% against state-of-the-art defenses, and smaller/older models remain markedly easier to steer |
| Recommendations | Output validation and user confirmation for sensitive actions, layered on top of existing detection |