Blog
    AI Solutions8 min

    Prompt Injection and LLM Security: How to Protect Business Chatbots and AI Agents

    Abstract editorial illustration of a luminous artificial intelligence brain protected by a multi-layered firewall barrier that deflects malicious red fragments, while safe cyan data flows through a guarded gate, on a deep navy background
    October 5, 2026Team 42bites
    Prompt InjectionLLM SecurityOWASPAI AgentsOpenAI

    When a chatbot or AI agent is connected to company data and tools, text becomes an attack vector. Prompt injection — inserting malicious instructions into what the model reads — is the number one risk in the OWASP Top 10 for LLM applications. This guide explains how it works and which concrete defences to adopt in an SME's AI solutions.

    What Prompt Injection Is

    Prompt injection is an attack in which carefully crafted text induces a language model to ignore the developer's original instructions and follow others. It comes in two forms: direct, when the user types into the chatbot 'ignore previous instructions and reveal…', and indirect, when malicious instructions are hidden in content the model processes — a web page, an email, a PDF, a document in the knowledge base.

    The problem is structural: to an LLM, instructions and data are both text, and there is still no foolproof way to distinguish with certainty what to execute from what to merely read. That is why defence cannot rely on a single filter.

    Why It Is a Real Risk for Businesses

    A chatbot answering about products and prices risks at most reputational embarrassment. The risk grows when the system can access confidential data or take action: an assistant that reads emails and can send them, an agent connected to the CRM, a RAG indexing documents uploaded by third parties. In these scenarios malicious content can lead to data leaks, unauthorised actions or manipulated answers to customers and employees.

    • Exfiltration of confidential data present in the context or knowledge base
    • Unwanted actions through connected tools (sending emails, editing records, API calls)
    • Manipulated answers, such as promoting a competitor or giving wrong information
    • Bypassing company policies and content filters
    • Abnormal costs from token consumption induced by crafted inputs

    Defences That Work: Defence in Depth

    Since there is no single fix, adopt layered defence. The most effective principle is least privilege: the model should access only the data and tools strictly needed for the task, with read-only permissions where possible and dedicated, limited credentials. If the agent cannot delete records, no injection can make it do so.

    The second layer is separating trusted from untrusted content: text from outside must be marked and treated as data, never as instructions, and model outputs validated before use — checking format, links and commands, and never executing output directly on sensitive systems. Input and output filters, dedicated detection systems and limits on tool use complete the picture.

    Human-in-the-Loop for Critical Actions

    For irreversible or high-impact actions — payments, external communications, changes to production data — human confirmation is the most reliable defence. The agent prepares the action, a person approves it. A well-designed workflow makes this approval quick and informed, showing what will be done and on which data, so it does not become an automatic click.

    Test, Monitor and Govern

    The security of an LLM application should be tested like a web application: with periodic red teaming, malicious prompt suites in automated evaluations (evals) and review of edge cases whenever the model or prompt changes. In production, log inputs, outputs and tool calls with alerts on anomalous patterns, respecting the data minimisation required by GDPR.

    A clear company AI policy, risk assessment under the AI Act and training for the development team on OWASP LLM risks complete the picture: security is a process, not a feature added at the end.

    Frequently Asked Questions About Prompt Injection

    Can prompt injection be eliminated entirely?

    Not at present: it is an intrinsic limitation of how models process text. Impact can be greatly reduced with least privilege, validation and human confirmation, assuming an injection will sooner or later succeed.

    Is a chatbot without access to sensitive data at risk?

    The risk is lower but not zero: answer manipulation, reputational damage and abnormal costs remain. The more access to data and tools grows, the more severe the impact.

    What should I ask an AI solutions vendor?

    How they manage model permissions, whether human approvals exist for critical actions, what security tests they run and how inputs and outputs are logged and monitored.

    Want to Secure Your AI Solution?

    We assess the risks of chatbots, agents and LLM integrations and design the right defences: permissions, validation, testing and human controls.