NVDA 229.28 ▼0.52%GOOGL 351.66 ▲0.97%MSFT 535.07 ▲2.38%AMD 608.10 ▼2.03%INTC 104.70 ▼2.22%TSMC 453.31 ▼1.02%AMZN 262.43 ▲3.29%META 718.67 ▼0.31%AAPL 336.64 ▼1.11%PLTR 209.05 ▲5.17%
Markets at last close

Microsoft · Security

Prompt injection reference maps jailbreak risks and defenses

·1 min read

A defensive knowledge base published as a GitHub gist catalogs prompt injection and jailbreak patterns, affected models and systems, and mitigations, with examples deliberately defanged. Prompt injection is framed as the broader problem of untrusted data and trusted instructions sharing one channel, while jailbreaking is a subset aimed at making a model ignore safety rules.

The taxonomy aligns the risk with OWASP LLM01:2025, MITRE ATLAS entries for prompt injection and LLM jailbreaks, and NIST AI 100-2e2025. Techniques span direct overrides, persona role-play, prefix manipulation, payload splitting, long-context many-shot attacks, multi-turn escalation, indirect injection through web pages, documents, email, repositories, tools, images, audio, Unicode obfuscation, and automated optimization methods.

Reported incidents include EchoLeak in Microsoft 365 Copilot, GitHub Copilot RCE, GeminiJack, Phishing for Gemini, and earlier ChatGPT plugin exfiltration, underscoring risk when agents can retrieve content, render links, or invoke tools. The knowledge base repeatedly cautions that attack success rates are version- and date-pinned, often overstated by lenient evaluation, and shift as vendors patch hosted systems.

Recommended mitigations emphasize defense in depth rather than relying on model alignment alone: instruction hierarchy, data delimiting, input and output classifiers, Unicode normalization, least-privilege tools, human confirmation for sensitive actions, capability sandboxing, and architectural separation between trusted instructions and untrusted content.

Originally reported by gist.github.comRead the source →
Related coverage
All Microsoft news →