LLM Security: Securing Large Language Models Before They Cost You

This white paper explains why large language models bring new security risks, with prompt injection as the main threat. In its case example, SubRosa found a law firm's AI assistant could be tricked into leaking privileged memos and skewing case-law summaries, then helped the firm add guardrails, access controls and monitoring that shut those attacks down.

JP
John Price
  • Reviewed by Kevin Schewe and Ratan Gupta
  • 5 min read
  • Download PDF
Share

Key Highlights

  • Prompt injection is the use of crafted or deceptive prompts to push an LLM into revealing sensitive data or acting unpredictably.
  • Warning signs include unusually detailed, repeated or confusing queries and requests for sensitive data dressed up as normal questions.
  • During testing, adversarial prompts pulled content from privileged internal memos out of a law firm's AI assistant.
  • Stricter input validation, role-based access, policy-based fine-tuning and real-time monitoring closed the gaps, and forensic review found no earlier breach.
  • Because AI systems keep changing, a single test is not enough: ongoing assessment and red-teaming are needed.

Artificial Intelligence (AI) has become a valuable tool for businesses, making tasks easier and faster. One of the most popular forms of AI today is Large Language Models (LLMs), like ChatGPT, which help generate text, answer questions, and automate conversations. However, along with their many benefits come significant security risks. This white paper will explore the hidden dangers of LLMs and show you practical ways to keep your AI tools safe and effective.

Understanding Large Language Models

Large Language Models are advanced computer programs that can understand and create human-like text.

Companies use LLMs for customer service, content creation, marketing, and even data analysis. These AI systems learn from reading lots of information from the internet and other sources. Once trained, they can respond to requests, produce reports, or even write software code.

Because they understand context and intent, LLMs can perform complex tasks, from drafting documents to analyzing customer feedback. Their ability to automate routine interactions frees employees to focus on more critical responsibilities, significantly improving business efficiency.

LLMs in ActionLLMs power virtual assistants, automated email replies, content creation tools, customer service chatbots, and data analysis platforms used by businesses globally.

Why People Love LLMs...

  • Streamlined customer interactions and support
  • Enhanced content creation and productivity
  • Improved efficiency through automation of repetitive tasks
  • Accelerated decision-making through fast data processing

Attack Focus: Prompt Injection

Prompt injection involves attackers manipulating LLM responses by crafting malicious or misleading prompts.

This technique can exploit the AI's responses, leading it to reveal sensitive data or behave unpredictably. For instance, an attacker might deceive a chatbot into disclosing confidential financial information or customer details.

These attacks pose significant threats, including misinformation, reputational harm, regulatory penalties, and compromised customer trust. As businesses increasingly rely on LLM-driven customer interactions, understanding and mitigating prompt injection risks become essential to secure operations.

Red Flag Alert:Beware of suspiciously detailed or unusual queries to your AI systems, as these could indicate attempts at prompt injection.

Common Prompt Injection Tactics

  • Overly detailed or repeated prompts
  • Prompts designed to confuse or mislead the AI
  • Requests for sensitive information disguised as legitimate queries

Case Example: A Realistic LLM Attack

SubRosa was engaged by this client to evaluate the security of a large language model (LLM) integrated into their internal knowledge management system. The firm relied on the AI assistant to draft case briefs, summarize filings, and surface past precedent. Our testing revealed critical vulnerabilities that could have led to data leakage and reputational harm.

  1. Prompt Injection Discovery During testing, SubRosa simulated adversarial prompts designed to override the assistant's instructions. The model revealed details from privileged internal memos — information that should have been inaccessible to end-users.
  2. Adversarial Input Handling Specialized test inputs were crafted to manipulate the assistant into delivering misleading summaries of case law. If left unchecked, these could have resulted in lawyers relying on inaccurate or adversarially biased information.
  3. Immediate Client Engagement Within hours of identifying the risk, SubRosa met with the firm's IT and compliance officers. Together, we reproduced the vulnerabilities and demonstrated how an external actor could exploit them through routine queries.
  4. Preventive Measures and Safeguards SubRosa recommended new guardrails: stricter input validation, role-based access controls, and fine-tuning the model with updated policy constraints. Forensic review confirmed no prior breaches, but the client implemented ongoing monitoring to detect and block injection attempts in real time.

The Result

SubRosa's timely intervention prevented the law firm's AI assistant from exposing confidential client information and misdirecting case research, averting both immediate reputational harm and long-term compliance consequences.

Risk Exposure Eliminated

The firm's LLM assistant no longer exposed privileged client data or internal memos. New guardrails and access controls stopped prompt injection attempts before they could yield sensitive results.

Operational Integrity Preserved

By mitigating adversarial input risks, attorneys could rely on the AI assistant for case research without fear of manipulation or misinformation. Productivity gains from AI adoption were preserved without introducing hidden liabilities.

Compliance Alignment

Remediation steps were mapped against the firm's regulatory obligations (client confidentiality, GDPR, ABA guidelines). This reduced potential liability in the event of regulatory review or legal discovery.

Client & Stakeholder Trust

By demonstrating proactive AI security testing, the firm reinforced its reputation for safeguarding sensitive information — a differentiator when winning and retaining high-value clients.

Foundation for Continuous Security

With monitoring and policy enforcement in place, the firm established a repeatable process for evaluating future AI integrations — ensuring that as the technology evolves, so does their security posture.

Key Takeaways

LLMs Expand the Attack Surface

Unlike traditional IT systems, large language models can be manipulated through natural language itself. That creates new, often overlooked risks for sensitive industries.

Real-World Impact is Tangible

Without testing, LLMs can expose confidential data, distort outputs, or be hijacked for malicious use. In legal, healthcare, and finance, these risks translate directly into liability and reputational damage.

Specialized Testing is Essential

Traditional pen testing stops at networks and endpoints. SubRosa's proprietary LLM penetration testing methodology closes this gap by simulating prompt injection, adversarial inputs, data leakage, and access control exploits.

Business Value is More than Security

Securing LLM deployments preserves productivity gains, ensures compliance, and strengthens client trust — delivering ROI beyond technical defense.

Continuous Monitoring is the Future

AI systems evolve constantly. One-off testing isn't enough. Building ongoing assessment and red-teaming into your security program is the only way to keep pace.

Put your AI assistant to the testSubRosa runs adversarial security testing on AI deployments to find prompt injection and data leakage risks.Explore AI security testing

Frequently asked questions

What is prompt injection?

It is an attack in which someone feeds a large language model deliberately malicious or misleading input so it gives up sensitive data or behaves in ways its owner never intended. A typical goal is getting a chatbot to hand over confidential financial or customer details.

What are the warning signs of a prompt injection attempt?

Watch for queries that are strangely detailed, repeated or designed to confuse the model, and for requests that seek sensitive information while posing as ordinary questions.

Why isn't traditional penetration testing enough for LLMs?

Conventional tests focus on networks and endpoints, but an LLM can be manipulated simply through the language it is given. Proper testing therefore imitates injected prompts, adversarial input, attempts to leak data and access control abuse.

How can a business protect its LLM from prompt injection?

In SubRosa's case example, the fixes were tighter input validation, role-based access controls, fine-tuning the model on updated policy constraints, and continuous monitoring to catch and block injection attempts as they happen.

How often should LLMs be security tested?

Continuously. AI systems change all the time, so one-time testing falls behind; ongoing assessment and red-teaming should be part of the wider security program.

Ready to strengthen your security posture?

Have questions about this article or need expert cybersecurity guidance? Connect with our team to discuss your security needs.