Most companies started their AI journey with a simple website chatbot or a tool for drafting marketing copy. That initial phase is over. Today, AI applications process contracts and medical records, personalize customer experiences, approve or flag transactions, search internal wikis, and take direct action inside CRMs, ticketing platforms, and code repositories. When software gains the authority to read sensitive data and make production changes, security must be built in from day one. Retrofitting security after launch rarely works and almost always costs significantly more.
The financial pressure is substantial. IBM and the Ponemon Institute set the global average cost of a data breach at USD 4.99 million in their Cost of a Data Breach Report 2026, a record high representing a 12% jump over the previous year. That same research revealed that AI-driven attacks spiked by 56%. In a summary of the report, law firm Baker Donelson highlighted that incidents involving AI models, training data, or AI infrastructure accounted for 21% of all breaches, up from 13% the year prior.
Locking down databases and securing API endpoints remain essential, but with AI, those steps cover only a fraction of your exposure. Prompts, base models, training data, retrieved documents, third-party model providers, and generated outputs each introduce distinct risks. NIST highlights security and resilience among its core characteristics for trustworthy AI within its AI Risk Management Framework, while the OWASP GenAI Security Project maintains specialized guidance for LLM, generative, and agentic applications. Anyone reviewing an AI system should start with these two resources.
Why AI Applications Need a Different Security Approach
Conventional software is predictable: it processes structured input, runs fixed logic, and returns identical results for identical requests. Modern application security testing was built around that predictability.
AI systems operate differently. An AI model’s output depends on its prompt, base model, context window, runtime-retrieved documents, and available tool integrations. A single sentence hidden inside a retrieved file can influence the model just as strongly as developer instructions, and the model lacks a reliable method to distinguish between the two.
Consider an assistant connected to a corporate knowledge base containing HR records, board papers, and pricing strategies. If retrieval logic fails to strictly enforce user permissions, a cleverly phrased query can surface restricted files. The model will typically comply because answering queries using available context is precisely what it was designed to do.
Real-world incidents demonstrate this exposure:
- EchoLeak (CVE-2025-32711): Disclosed in June 2025 by researchers at Aim Security, this zero-click vulnerability in Microsoft 365 Copilot carried a CVSS score of 9.3. A case study on arXiv detailed how a single crafted email could prompt Copilot to exfiltrate internal data without requiring any victim interaction. The researchers bypassed Microsoft’s cross-prompt injection classifier, used reference-style Markdown links to avoid link redaction, relied on auto-loading images, and routed exfiltrated data through a Microsoft Teams domain permitted by the content security policy. Microsoft patched the issue on the server side before public disclosure. The key lesson for engineering teams is that Copilot was working as designed: it processed incoming content and acted with the full privileges of the active user.
- OWASP Top 10 for LLM Applications (2026 Edition): Published in early August 2026, this update ranks Prompt Injection first and Sensitive Information Disclosure second. According to Help Net Security, this edition marks the first time empirical incident data influenced the rankings, weighted 75% by practitioner feedback and 25% by 6,639 recorded incidents from public vulnerability and AI-harm databases. Excessive Agency rose to third place, Improper Output Handling moved to tenth, and System Prompt Leakage was expanded into Hidden Context Exposure. Supply chain risks and model/data poisoning remain prominent threats.
Securing an AI product requires maintaining your standard application security program while integrating dedicated AI controls. Teams that rely solely on conventional security typically discover critical gaps during their first penetration test.
Start With a Secure AI Architecture
Architectural choices made during design prevent far more vulnerabilities than post-launch patches. The primary objective is isolating the model from core business systems while restricting what it can view and execute. During design reviews, treat the LLM as an untrusted component—similar to a third-party library that requires isolation and should never receive database administrator rights.
A production-grade AI reference architecture relies on segmented operational layers:
| Architectural Layer | Core Responsibility |
| Authenticated Application Layer | Verifies user identity before requests reach model services. |
| API Gateway | Controls traffic flow, rate limiting, and token-based authorization. |
| AI Orchestration Layer | Assembles prompts, manages retrieval, and controls tool execution. |
| Isolated Model Services | Runs inference in isolated environments without direct access to core databases. |
| Controlled Knowledge Base Access | Filters search results so users only retrieve data they are authorized to view. |
| Input & Output Validation | Sanitizes user prompts, retrieved files, and generated model outputs. |
| Monitoring & Auditing | Logs user interactions, tool executions, and anomalous behavior for analysis. |
| Human-in-the-Loop Controls | Holds high-impact operations until an authorized human approves them. |
The foundational principle across every layer is least privilege: models must operate with only the precise data and permissions required for their specific task. A customer support bot needs access to product documentation and order statuses; it should never have network paths to payroll, financial ledgers, or administrative consoles. If those systems are unreachable at the network and identity layers, prompt manipulation cannot bridge the gap.
Effective segmentation also confines damage during security breaches. On September 21, 2026, the Dutch Institute for Vulnerability Disclosure (DIVD) reported an incident where an external AI agent exploited two zero-day vulnerabilities in Zammad helpdesk software. DIVD confirmed that robust network segmentation combined with rapid incident response contained the compromise. While AI was used offensively in that instance, the defensive takeaway remains identical: assume individual components will eventually be compromised and build boundaries to prevent lateral movement.
Protect Data Throughout the AI Lifecycle
For most applications, proprietary data represents far more value than the underlying model. Customer records, internal documentation, source code, financial figures, and private communications routinely pass through AI workflows, often via unmonitored channels.
The classic example remains Samsung’s 2023 exposure. Within three weeks of allowing staff to use ChatGPT, employees pasted proprietary source code and internal meeting notes into the tool on three separate occasions, as reported by The Register. According to Forbes (citing Bloomberg), Samsung initially capped prompt sizes at 1,024 bytes before banning generative AI tools on employee devices in May 2023. No malicious actor was involved; internal engineers simply transmitted confidential IP to an external cloud service.
Unintended exposure of this type has grown widespread. IBM’s 2026 research indicates that incidents involving shadow AI climbed to 43% (up from 20% in 2025), as summarized by Baker Donelson. Protecting organizational data requires securing the entire data pipeline: collection, storage, fine-tuning, retrieval, processing, and deletion.
Key data protection strategies include:
- Pervasive Encryption: Encrypt personal and proprietary data at rest and in transit across every hop between the application, orchestration layer, and model endpoints.
- Contextual Access Control: Implement role-based access control (RBAC) directly within the retrieval layer so queries return only content the specific user is permitted to see.
- Data Minimization & Redaction: Strip names, account numbers, and unnecessary identifiers using automated tokenization or redaction before assembling the final context window.
- Isolated Non-Production Environments: Never load real customer records into developer notebooks or test vector indexes. Development and testing stores require the same access controls and auditing applied to production systems.
- Data Poisoning Protections: Log and audit all edits to training sets and retrieval stores to prevent unauthorized modifications that could corrupt model behavior.
- Comprehensive Lifecycle Deletion: Ensure user deletion requests purge associated data from prompt logs, response caches, and vector embeddings.
- Early Data Classification: Tag data as public, internal, confidential, or restricted at index time, allowing the retrieval layer to programmatically enforce access policies during every search.
Address AI Security Risks Before Deployment
Threat modeling is an essential prerequisite for deployment. Engineering teams should map out how attackers might manipulate inputs, access restricted stores, abuse model agency, or pivot into connected infrastructure. The classic STRIDE framework offers a structured starting point, while MITRE ATT&CK / ATLAS supplies documented tactics targeting machine learning environments.
Four primary threats require specific technical mitigations:
1. Prompt Injection
In a prompt injection attack, malicious input overrides system instructions to force unintended model actions. Direct injections occur through interactive chat boxes. Indirect injections, the technique behind EchoLeak, hide malicious instructions inside emails, web pages, PDFs, or support tickets processed by the model. OWASP’s 2026 updates explicitly cover cross-modal attacks, where payloads are embedded inside image or audio files.
OWASP highlights a notable defense effect surrounding this threat: because enterprises invest heavily in blocking prompt injections, relatively few successful breaches appear in public databases. However, when an injection succeeds, damages are severe. Cycode’s analysis of IBM data places the average cost of a breach caused by prompt injection at USD 5.89 million.
Hardcoded system instructions cannot serve as access boundaries. Prompting a model to “never reveal salary data” merely guides its behavior; a persistent user can often bypass system instructions. True authorization must be enforced by the application layer and infrastructure. Maintain strict separation between system instructions and untrusted content, and constrain the actions a single response can execute.
2. Sensitive Information Disclosure
Confidential information can leak through model responses, application logs, search results, or improperly isolated context windows. OWASP’s decision to rename system prompt leakage to Hidden Context Exposure reflects an established operational reality: assume any data placed inside a context window can be extracted by an end user. Never place API keys, private credentials, or core authorization logic inside a prompt.
Enforce access controls before data reaches the model context window. Implementing permission-aware retrieval—where the search engine verifies user authorization before returning text chunks—is far more reliable than relying on model discretion. Breach metrics validate this approach: Alston & Bird’s analysis of IBM’s 2026 report (covering 602 breached organizations) revealed that 92% of organizations experiencing an AI-related breach lacked adequate AI access controls.
3. Supply-Chain Vulnerabilities
Modern AI applications rely on complex dependency chains: foundation model APIs, open-weight models from public repositories, third-party datasets, orchestration libraries, plugins, and Model Context Protocol (MCP) servers providing tool capabilities to autonomous agents.
Compromises can enter through any of these vector points:
- Unsafe File Formats: Models serialized in legacy formats like Python’s pickle can execute arbitrary code upon loading. Standardize on safer formats such as safetensors.
- Slopsquatting: Highlighted in Daxa’s security breakdown, AI coding assistants sometimes suggest plausible but non-existent package names. Attackers preemptively register these names on public package managers so developers inadvertently install malicious dependencies.
Maintain a real-time inventory of all external components, perform vendor security assessments, continuously monitor dependencies, and document incident response steps for compromised providers. Draft guidance in NIST’s Cyber AI Profile (NIST IR 8596) recommends tracking models, agents, APIs, access keys, datasets, and embedded integrations within an AI Bill of Materials (AIBOM) alongside standard software bill-of-materials tracking.
4. Improper Output Handling
Treat all model outputs as untrusted data. Passing generated text directly into SQL statements, shell scripts, HTML rendering engines, or downstream agents reintroduces classic vulnerabilities—such as SQL Injection, Cross-Site Scripting (XSS), and Remote Code Execution (RCE)—through a new interface.
Although this risk moved to tenth place in OWASP’s 2026 ranking, its scope was expanded to reflect how model outputs increasingly trigger automated tools, generate executable code, and authorize transactions. In these automated environments, a hallucinated or manipulated response can immediately trigger financial or operational damage.
Mitigate this risk by enforcing structured outputs:
- Mandate strict JSON Schema validation for model responses.
- Context-encode any dynamic text before rendering it in web browsers.
- Use parameterized queries for all database interactions.
- Apply strict least-privilege permissions to downstream operational services.
Apply Security Controls to AI Software
AI application security must address standard software vulnerabilities while protecting against AI-specific attack vectors. Core application security measures, strong authentication, access control, transport encryption, dependency scanning, and adherence to standards like the OWASP Application Security Verification Standard (ASVS), remain indispensable. An AI application running on an unpatched web framework inherits every underlying vulnerability of that web framework.
Traditional vulnerability scanners cannot evaluate runtime model behavior, necessitating dedicated adversarial red-teaming. Security testing should systematically evaluate:
- Direct and Indirect Prompt Injections: Testing payloads delivered through chat inputs, uploaded documents, emails, and web pages.
- Data Exfiltration: Attempting to extract system prompts, context window contents, API keys, and sensitive user data through responses and error logs.
- Retrieval Boundary Breaches: Verifying whether search components return documents outside the user’s explicit permissions.
- Agent Scope Escalation: Checking if autonomous agents can call unassigned tools or execute unauthorized actions.
- Supply Chain Attacks: Testing application resiliency against malicious model files, compromised dependencies, and poisoned retrieval stores.
The OWASP Artificial Intelligence Security Verification Standard (AISVS) 1.0, published in June 2026 at OWASP Global AppSec in Vienna, offers a structured verification framework. Modeled after the ASVS, it covers training pipelines, model development, deployment architecture, agent orchestration, monitoring, and decommissioning, with dedicated sections on supply chains, vector database security, and MCP integrations.
AISVS defines three verification levels. OWASP recommends Level 2 for production enterprise systems, customer-facing applications, solutions handling personal data, and systems making automated business decisions.
Follow Practical AI Security Best Practices
Securing AI applications does not require reinventing security fundamentals. Most organizations can expand their existing application security frameworks by embedding AI-specific controls at key decision points:
- Enforce Strict Least-Privilege Access: Limit model permissions to the exact operations required for the current task. Scope agent API tokens to individual, short-lived workflows.
- Validate Inputs and Outputs Continuously: Apply strict validation rules at every transition point, treating prompts, retrieved context, and generated text as untrusted content.
- Minimize Sensitive Context: Keep unnecessary sensitive data out of model prompts. Minimizing data exposure reduces the potential impact of a prompt injection or context leakage incident.
- Implement Purpose-Driven Logging: Log prompts, tool calls, and automated decisions to support incident forensics, but redact sensitive personal data to prevent log stores from becoming high-value targets.
- Conduct Regular Adversarial Testing: Run simulated attack scenarios tailored to AI failure modes whenever models, system prompts, or available tools change.
- Track the AI Supply Chain: Maintain an operational inventory covering model versions, third-party APIs, datasets, dependencies, and plugins, assigning clear internal ownership for each.
- Maintain Human-in-the-Loop Oversight: Require explicit human authorization for high-impact actions, such as executing financial transfers, modifying system records, or sending mass external communications.
Governance remains a major blind spot. Alston & Bird’s analysis indicates that 68% of breached organizations had no formal policies governing AI use or shadow AI management, up from 63% the previous year. However, the operational value of proactive defense is clear: organizations deploying security AI and automation extensively saved an average of USD 1.93 million per breach compared to those operating without those controls.
Make Security Part of AI Application Development
Security must be integrated into every phase of AI application Development, from early requirements gathering and architecture to deployment and continuous monitoring.
NIST’s AI Risk Management Framework (AI RMF) provides structured guidance for building trustworthy AI across four functions: Govern, Map, Measure, and Manage. To address generative AI risks directly, NIST published the Generative AI Profile (NIST AI 600-1) on July 26, 2024, helping organizations identify unique generative AI threats and align controls with operational priorities.
Regulatory frameworks are shifting rapidly:
- EU AI Act Timeline: The Digital Omnibus on AI was adopted by the European Parliament on June 16, 2026, and by the Council on June 29, 2026. As legal firm Freshfields notes, this update adjusts compliance deadlines for standalone high-risk AI systems to December 2, 2027, and to August 2, 2028 for AI integrated into products governed by existing EU safety laws. The core compliance requirements remain unchanged. Companies building recruitment tools, credit scoring algorithms, or educational platforms for European users should use this period to establish logging protocols, human oversight workflows, and technical documentation.
In daily engineering workflows, security reviews can no longer be treated as a single pre-launch checkpoint. Risk must be evaluated continuously as retrieval sources expand, models update, and agent toolsets evolve. Upgrading a base model, adding a vector data source, or assigning a new tool capability to an agent should immediately trigger a focused security review.
Successful execution requires cross-functional collaboration between developers, security engineers, data architects, and business owners from project inception. The goal is not merely preventing attacks, but delivering an auditable system where every operational capability has an assigned owner and high-risk actions can be safely halted.
When evaluating external implementation partners, verify their engineering depth by asking which AISVS level they test against, how they execute adversarial red-teaming, and where user data will be processed and stored.
Build for Security From the First Architecture Diagram
Securing an AI application becomes exponentially harder when treated as a final checklist item prior to launch. Security decisions should be mapped alongside system architecture, data flow design, model selection, and integration requirements from day one.
Start by documenting system boundaries: map what data the application can access, what records it processes, what actions it can execute autonomously, and where human intervention is required. Revisit this reference documentation continuously as system capabilities grow. This record serves as an essential alignment tool across engineering teams, internal auditors, and executive leadership.
For broader architectural guidance, consult the OWASP AI Exchange, which maintains a comprehensive directory of AI security threats and controls while contributing to international standardization efforts under ISO/IEC and the EU AI Act. Teams engineering autonomous agents should also review the OWASP Top 10 for Agentic Applications (published December 2025), which provides targeted defenses against agent goal hijacking, tool misuse, and privilege escalation.
Building secure AI is ultimately an engineering discipline. As the lead contributors to the OWASP 2026 project noted, organizations should stop trying to build a model that can never be tricked, and instead design resilient systems where nothing critical breaks when an error occurs. Strong access controls, disciplined data handling, robust output validation, structured threat modeling, and human oversight allow companies to deploy powerful AI applications safely at scale.



