Penetration testing for large language models, before attackers prompt them.
AI is now business-critical, and it introduces an attack surface traditional testing was never built for. SubRosa's AI red team probes your LLM applications for prompt injection, data leakage, and model and plugin abuse, the way a real adversary would.
Prompt injection · Data leakage · Model manipulation · Plugin abuse
What is LLM penetration testing?
LLM penetration testing is a hands-on security assessment of applications built on large language models, testing the model, its prompts, its data access, and its plugins and integrations for the ways an attacker could abuse them: prompt injection that hijacks behavior, jailbreaks that bypass guardrails, leakage of training data or system prompts, and plugin chains that reach systems the model should never touch. It goes beyond a model benchmark to prove real, exploitable impact in your deployment.
The AI attack surface.
We test every way an attacker could abuse an LLM application, from the prompt to the plugins.
Prompt injection & jailbreaks
Direct and indirect prompt injection and jailbreak techniques that hijack the model's behavior, bypass guardrails, or exfiltrate its instructions.
Data & prompt leakage
Testing for exposure of training data, system prompts, and other users' data through the model and its context window.
Model manipulation
Adversarial inputs that degrade, bias, or manipulate model outputs into producing harmful or unauthorized actions.
Plugin & integration security
Assessment of the tools, plugins, and integrations the model can call, where an injected prompt can pivot into real systems and data.
Testing an attack surface with no established playbook.
LLM testing is young, and much of what passes for it is a list of jailbreak prompts. This is what a substantive engagement covers.
- 01
Scoping the system, not just the model
The model is rarely the interesting part. We scope the system around it: the system prompt, retrieval sources, tools it can call, what it is permitted to do on a user's behalf, and where its output is trusted downstream.
- 02
Mapping trust boundaries
We establish where untrusted input can reach the model and what the model is trusted to do afterwards. Nearly every serious LLM finding comes down to a boundary that was assumed rather than enforced.
- 03
Direct and indirect prompt injection
Direct injection is the well-known case. Indirect injection is the dangerous one: instructions hidden in a document, a web page or a support ticket the model later ingests, executing when no attacker is present in the conversation.
- 04
Data and prompt leakage
We test whether the system prompt can be extracted, whether retrieval returns records the user is not entitled to, and whether prior users' data can surface. Retrieval-augmented systems frequently inherit permissions far broader than the user in front of them.
- 05
Tool and plugin abuse
Where the model can call tools, send email, query databases or execute code, we test whether injected instructions can drive those actions. This is where an LLM issue stops being embarrassing and starts being a breach.
- 06
Reporting and re-test
Findings are reported with reproduction prompts and concrete mitigations at the system level — input handling, permission scoping, output validation — rather than a recommendation to add more filtering. Retest is included.
Offensive security, applied to AI.
AI red team expertise
Our offensive team tests LLM applications with the same adversarial mindset we bring to networks and apps, adapted to how AI actually fails.
Mapped to OWASP LLM Top 10
Findings are mapped to the OWASP Top 10 for LLM Applications, so your risk is framed against the emerging industry standard.
Real deployment context
We test your real deployment, its prompts, data access, and integrations, not a generic model, so the results reflect the risk you actually carry.
From red team to remediation.
Your LLM pen test findings land in Sable, mapped to the OWASP LLM Top 10, prioritized, assigned, and tracked from open to retested, so AI risk becomes a managed program instead of a one-off report.
- CriticalOpenIndirect prompt injection via docLLM01
- HighIn progressSystem prompt disclosureLLM06
- HighRetestedPlugin call reaches internal APILLM07
- MediumOpenGuardrail bypass via role-playLLM01
Common questions
- What is LLM penetration testing?
- LLM penetration testing is a hands-on security assessment of applications built on large language models, testing the model, its prompts, its data access, and its plugins and integrations for the ways an attacker could abuse them: prompt injection, jailbreaks, leakage of training data or system prompts, and plugin chains that reach systems the model should never touch. It proves real, exploitable impact in your deployment rather than benchmarking a model in isolation.
- What kinds of LLM vulnerabilities do you test for?
- SubRosa tests for direct and indirect prompt injection, jailbreaks and guardrail bypass, training-data and system-prompt leakage, model manipulation, and insecure plugins and integrations where an injected prompt can pivot into real systems and data. Findings are mapped to the OWASP Top 10 for LLM Applications.
- What does LLM penetration testing actually cover?
- The system, not just the model. That means the system prompt, retrieval sources, the tools the model can call, what it is permitted to do on a user's behalf, and where its output is trusted downstream. Testing that consists of trying jailbreak prompts against a chat box is not an assessment — nearly every serious finding is a trust boundary that was assumed rather than enforced.
- What is indirect prompt injection and why does it matter more?
- Direct injection is a user typing malicious instructions. Indirect injection is instructions hidden in content the model later ingests — a document, a web page, a support ticket — which execute with no attacker present in the conversation. It is the more dangerous case because it needs no access to your interface and it scales.
- Can you test a model we did not build?
- Yes, and that is the common case. Most organisations build on a commercial model rather than training their own. The risk lives in your integration: what the model can reach, what it is trusted to do, and what happens to its output. That is testable regardless of who trained the underlying model.
- Is this a mature discipline?
- It is young, and we would rather say so. The techniques are evolving quickly and there is no settled standard equivalent to what exists for network testing. What we can say is that the failures we find are consistent and largely architectural — over-broad retrieval permissions, tools callable from untrusted input, output trusted downstream — and those are fixable at the system level rather than by adding another filter.
Secure your AI before attackers prompt it.
Book an LLM penetration test and find out exactly how an attacker could abuse your AI, and how to shut it down.