Penetration Testing

Penetration testing for large language models, before attackers prompt them.

AI is now business-critical, and it introduces an attack surface traditional testing was never built for. SubRosa's AI red team probes your LLM applications for prompt injection, data leakage, and model and plugin abuse, the way a real adversary would.

Prompt injection · Data leakage · Model manipulation · Plugin abuse

LLM penetration testing, defined

What is LLM penetration testing?

LLM penetration testing is a hands-on security assessment of applications built on large language models, testing the model, its prompts, its data access, and its plugins and integrations for the ways an attacker could abuse them: prompt injection that hijacks behavior, jailbreaks that bypass guardrails, leakage of training data or system prompts, and plugin chains that reach systems the model should never touch. It goes beyond a model benchmark to prove real, exploitable impact in your deployment.

What we test

The AI attack surface.

We test every way an attacker could abuse an LLM application, from the prompt to the plugins.

Prompt injection & jailbreaks

Direct and indirect prompt injection and jailbreak techniques that hijack the model's behavior, bypass guardrails, or exfiltrate its instructions.

Data & prompt leakage

Testing for exposure of training data, system prompts, and other users' data through the model and its context window.

Model manipulation

Adversarial inputs that degrade, bias, or manipulate model outputs into producing harmful or unauthorized actions.

Plugin & integration security

Assessment of the tools, plugins, and integrations the model can call, where an injected prompt can pivot into real systems and data.

How the engagement runs

Testing an attack surface with no established playbook.

LLM testing is young, and much of what passes for it is a list of jailbreak prompts. This is what a substantive engagement covers.

  1. 01

    Scoping the system, not just the model

    The model is rarely the interesting part. We scope the system around it: the system prompt, retrieval sources, tools it can call, what it is permitted to do on a user's behalf, and where its output is trusted downstream.

  2. 02

    Mapping trust boundaries

    We establish where untrusted input can reach the model and what the model is trusted to do afterwards. Nearly every serious LLM finding comes down to a boundary that was assumed rather than enforced.

  3. 03

    Direct and indirect prompt injection

    Direct injection is the well-known case. Indirect injection is the dangerous one: instructions hidden in a document, a web page or a support ticket the model later ingests, executing when no attacker is present in the conversation.

  4. 04

    Data and prompt leakage

    We test whether the system prompt can be extracted, whether retrieval returns records the user is not entitled to, and whether prior users' data can surface. Retrieval-augmented systems frequently inherit permissions far broader than the user in front of them.

  5. 05

    Tool and plugin abuse

    Where the model can call tools, send email, query databases or execute code, we test whether injected instructions can drive those actions. This is where an LLM issue stops being embarrassing and starts being a breach.

  6. 06

    Reporting and re-test

    Findings are reported with reproduction prompts and concrete mitigations at the system level — input handling, permission scoping, output validation — rather than a recommendation to add more filtering. Retest is included.

Why SubRosa

Offensive security, applied to AI.

AI red team expertise

Our offensive team tests LLM applications with the same adversarial mindset we bring to networks and apps, adapted to how AI actually fails.

Mapped to OWASP LLM Top 10

Findings are mapped to the OWASP Top 10 for LLM Applications, so your risk is framed against the emerging industry standard.

Real deployment context

We test your real deployment, its prompts, data access, and integrations, not a generic model, so the results reflect the risk you actually carry.

Every finding, tracked to closed.

From red team to remediation.

Your LLM pen test findings land in Sable, mapped to the OWASP LLM Top 10, prioritized, assigned, and tracked from open to retested, so AI risk becomes a managed program instead of a one-off report.

LLM findings in Sable
LLM findingsOWASP LLM Top 10
  • Critical
    Indirect prompt injection via doc
    LLM01
    Open
  • High
    System prompt disclosure
    LLM06
    In progress
  • High
    Plugin call reaches internal API
    LLM07
    Retested
  • Medium
    Guardrail bypass via role-play
    LLM01
    Open
Prompt · data · model · pluginsPrioritized · assigned

Common questions

What is LLM penetration testing?
LLM penetration testing is a hands-on security assessment of applications built on large language models, testing the model, its prompts, its data access, and its plugins and integrations for the ways an attacker could abuse them: prompt injection, jailbreaks, leakage of training data or system prompts, and plugin chains that reach systems the model should never touch. It proves real, exploitable impact in your deployment rather than benchmarking a model in isolation.
What kinds of LLM vulnerabilities do you test for?
SubRosa tests for direct and indirect prompt injection, jailbreaks and guardrail bypass, training-data and system-prompt leakage, model manipulation, and insecure plugins and integrations where an injected prompt can pivot into real systems and data. Findings are mapped to the OWASP Top 10 for LLM Applications.
What does LLM penetration testing actually cover?
The system, not just the model. That means the system prompt, retrieval sources, the tools the model can call, what it is permitted to do on a user's behalf, and where its output is trusted downstream. Testing that consists of trying jailbreak prompts against a chat box is not an assessment — nearly every serious finding is a trust boundary that was assumed rather than enforced.
What is indirect prompt injection and why does it matter more?
Direct injection is a user typing malicious instructions. Indirect injection is instructions hidden in content the model later ingests — a document, a web page, a support ticket — which execute with no attacker present in the conversation. It is the more dangerous case because it needs no access to your interface and it scales.
Can you test a model we did not build?
Yes, and that is the common case. Most organisations build on a commercial model rather than training their own. The risk lives in your integration: what the model can reach, what it is trusted to do, and what happens to its output. That is testable regardless of who trained the underlying model.
Is this a mature discipline?
It is young, and we would rather say so. The techniques are evolving quickly and there is no settled standard equivalent to what exists for network testing. What we can say is that the failures we find are consistent and largely architectural — over-broad retrieval permissions, tools callable from untrusted input, output trusted downstream — and those are fixable at the system level rather than by adding another filter.

Secure your AI before attackers prompt it.

Book an LLM penetration test and find out exactly how an attacker could abuse your AI, and how to shut it down.