AI Security Testing Gaps
SOC 2: Infrastructure Covered. Pentest: Web App Covered. Code Review: Application Covered. AI Model: Not Tested.
6 min read · 12 July 2026 · AI governance
An enterprise's TPRM team was assessing a vendor's AI-powered contract analysis product and requested evidence of security testing. The vendor's security team provided a comprehensive evidence package: their SOC 2 Type II report confirming controls around their cloud infrastructure, a penetration test report from a recognised security firm covering their web application layer, and a description of their secure development lifecycle including code reviews, static analysis, and DAST scanning during the CI/CD pipeline. The enterprise's TPRM analyst reviewed the evidence package and prepared to mark AI security testing as confirmed. A more experienced colleague reviewing the evidence package asked a clarifying question: of these three evidence items, which one tested the AI model itself , the model's resistance to adversarial inputs, its behaviour under prompt injection, its training data integrity, or its output accuracy and confabulation rate? The answer was none. The SOC 2 evaluated the operational controls around the infrastructure that hosted the AI model. The penetration test evaluated the web application that provided access to the AI model. The code reviews evaluated the application code that wrapped the AI model. All three were legitimate and valuable security assurance activities. None of them constituted AI-specific security testing , testing the model itself as a security surface.
What are AI Security Testing Gaps, Really?
AI security testing gaps are the omissions in a vendor's security testing programme that leave the AI model itself , as distinct from the infrastructure, application, and operational controls around it , untested against the specific attack categories that AI systems are vulnerable to. Traditional security testing frameworks , penetration testing, SAST/DAST, SOC 2 controls assessment , were designed for software applications and infrastructure, not for AI systems. Applying these traditional frameworks to AI products produces evidence of security testing that covers everything except the AI model itself.
The AI-specific attack surface is what traditional testing misses. AI models face a distinct set of security challenges that are not addressed by penetration testing web applications or auditing infrastructure controls: adversarial input attacks that cause the model to misclassify inputs, prompt injection attacks that cause the model to bypass its operational constraints, training data poisoning that compromises the model's learned behaviour, model extraction attacks that steal the model's decision logic, and output confabulation that introduces inaccurate information into the model's responses. None of these attack categories are assessed by a web application penetration test or a SOC 2 infrastructure audit.
The testing framework maturity gap is the industry context. AI-specific security testing frameworks are substantially newer and less standardised than traditional software security testing frameworks. The penetration testing methodology that applies to web applications has thirty years of development behind it. The adversarial ML testing methodology that applies to AI models is still developing. Vendor security teams that have mature traditional security testing programmes may have immature or absent AI-specific testing practices simply because the field is newer and the standards less established.
The confabulation testing gap is a specific AI security testing dimension that is absent from almost all vendor security testing programmes. Testing whether an AI model confabulates , produces confident, fluent, inaccurate outputs , in the specific domain the enterprise deploys it for is an AI quality assurance activity with direct security implications. An AI that confabulates legal analysis, medical information, or financial guidance creates downstream harm regardless of how well the application security is implemented.
Why this matters
AI security testing gaps matter for TPRM because the security assurance evidence that vendors provide for AI products often covers the traditional security dimensions comprehensively while leaving the AI-specific security dimensions completely unassessed. An evidence package that confirms SOC 2, penetration testing, and SDLC security provides genuine assurance about the infrastructure and application , and no assurance about the AI model.
Where most teams get this wrong
The most consistent failure is accepting traditional security testing evidence , SOC 2, penetration testing, SDLC , as equivalent to AI security testing evidence. The frameworks address different security surfaces. Both are necessary. Neither substitutes for the other.
- Traditional security evidence accepted as AI security evidence
- Adversarial input testing absent from vendor testing programme
- Prompt injection testing not conducted or not included in evidence
- Confabulation rate assessment not conducted
- Training data integrity testing not included in AI security programme
What good looks like
Mature AI security testing programmes conduct both traditional security testing (infrastructure, application, SDLC) and AI-specific security testing (adversarial inputs, prompt injection, confabulation rate, training data integrity) , and can provide evidence for both dimensions separately in response to TPRM requests.
- AI-specific testing programme alongside traditional security testing
- Adversarial input testing , model resistance to inputs designed to cause misclassification
- Prompt injection testing , model resistance to instruction override attempts
- Confabulation rate assessment , output accuracy validation in deployment domain
- Training data integrity testing , validation of training corpus against poisoning
Tooling
AI Security Testing , Garak, Adversarial Robustness Toolbox (IBM ART), Microsoft PyRIT
AI security testing frameworks specifically designed for LLMs and ML models test adversarial input resistance, prompt injection, jailbreak resistance, and output quality. Garak specifically covers LLM security testing across a range of attack categories. For TPRM practitioners, asking whether the vendor's AI security testing uses an AI-specific testing framework , distinct from traditional application security testing , provides a specific AI testing programme question.
Red Teaming , MITRE ATLAS adversarial ML framework, dedicated AI red team exercises
AI red team exercises that specifically target the model , not just the application wrapper , test adversarial categories that automated tools may miss. For TPRM practitioners, asking whether the vendor has conducted an AI red team exercise , where testers specifically attempt adversarial attacks against the model's decision logic, not just the application , provides a specific AI red team question.
Governance challenges
The governance challenge with AI security testing is the vendor knowledge gap. Many vendor security teams have strong traditional security testing expertise and limited AI-specific security testing expertise. Requesting AI-specific testing evidence may reveal that the vendor's security programme needs to expand its scope rather than that it has already done so. The governance resolution is including AI-specific testing requirements in vendor contracts as a forward-looking obligation.
- Require AI-specific security testing as a separate evidence category
- Ask for adversarial input test results specifically for the AI model
- Ask for prompt injection test results as a separate evidence item
- Ask for confabulation rate assessment in the enterprise's deployment domain
- Include AI security testing requirements in vendor contracts
If you are a small team
When reviewing AI vendor security evidence, ask one clarifying question before accepting the evidence package: of the security testing evidence you have provided, which specific test evaluated the AI model itself , specifically testing the model's resistance to adversarial inputs, prompt injection, or its output accuracy? If the answer is none, you have identified the AI security testing gap. Then ask what AI-specific testing the vendor plans to conduct and on what timeline. The question surfaces the gap. The follow-up establishes whether it is being addressed.
- Ask which evidence item specifically tested the AI model rather than the application or infrastructure
- Ask for adversarial input and prompt injection test results as separate evidence
- Ask about confabulation rate assessment in your deployment domain
- Include AI security testing requirements in vendor contracts
What to require
Ask directly:
"Of the security testing evidence you have provided , SOC 2, penetration testing, SDLC , which specific test evaluated the AI model itself against adversarial inputs, prompt injection, and output accuracy? Can you provide separate evidence of AI-specific security testing for the model?"
Expect as evidence
- AI-specific security testing evidence , adversarial inputs, prompt injection
- Confabulation rate assessment results
- Training data integrity testing
- AI red team exercise results
A vendor who provides SOC 2, penetration testing, and SDLC evidence should be asked for AI-specific testing evidence as a separate category. Traditional evidence covers the wrapper. AI testing covers the model. Both are needed.
How to evidence it
- AI-specific security testing records
- Adversarial input and prompt injection test results
- Confabulation rate assessment
- AI security testing requirements in vendor contracts
Key Takeaway
SOC 2: infrastructure covered. Penetration test: web application covered. Code review: application code covered. AI model: not tested. The evidence package was genuine and comprehensive for what it covered. What it covered was the wrapper around the AI. The AI itself , its adversarial input resistance, its prompt injection behaviour, its confabulation rate, its training data integrity , was not in any of the three evidence items. Traditional security testing frameworks were designed for applications and infrastructure. AI-specific testing frameworks address the model. Both are required. Neither substitutes for the other. The clarifying question , which of these tested the model itself , reveals the gap that reviewing the evidence package without asking it would miss.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association