AI Operational Risks
94% Accuracy. 6% Missed. Missed Records: Systematically the Most Complex. Human Review: Assumed Comprehensive.
5 min read · 25 July 2026 · AI governance
A clinical research organisation's data quality AI model screened incoming clinical trial records and flagged those requiring human data manager review before database lock. The model had been validated at 94% accuracy , meaning it correctly classified 94% of records as either requiring or not requiring review. The business process had been designed with the assumption that the model's 94% accuracy provided comprehensive coverage sufficient to replace manual screening of all records. After three months of production operation, a senior data manager noticed that the records reaching the database lock stage without flagging appeared to be systematically different from what was being flagged , the unflagged records that she was manually spot-checking seemed to have higher data complexity. An analysis confirmed the pattern: the 6% of records that the model failed to flag were systematically the most complex records in the dataset , edge cases with unusual data patterns, records from clinical sites that had not been well-represented in the model's training data, and records with subtle data integrity issues that required domain expertise and close reading to identify. The model was most accurate on routine records that a trained human reviewer would have identified quickly anyway. It was least accurate on the complex records where human review was most critical and most difficult to substitute. The 94% accuracy metric was real. The systematic pattern of where the 6% landed was the operational risk that the accuracy metric did not reveal.
What are AI Operational Risks, Really?
AI operational risks are the risks that arise from deploying AI systems in real-world operational contexts where the model's performance characteristics , accuracy, failure modes, confidence levels, and systematic biases , interact with business process assumptions, human oversight levels, and downstream consequence in ways that create operational vulnerability. An AI model that performs at 94% accuracy overall may have a 6% miss rate that is systematically concentrated in the highest-consequence cases, creating a risk profile that aggregate accuracy metrics do not capture.
The systematic error concentration problem is the core operational risk. AI models do not fail randomly , they fail systematically in patterns that reflect their training data limitations, their architectural assumptions, and the distribution of cases they are optimised for. A clinical data quality model trained predominantly on routine records will be less accurate on complex edge cases. A fraud detection model trained on historical fraud patterns will be less accurate on novel fraud methodologies. The systematic pattern of where a model fails is the operational risk information that aggregate accuracy metrics hide.
The business process assumption gap is the deployment risk. Business processes designed to incorporate AI screening or classification often assume a level of coverage that aggregate accuracy metrics appear to support but that the systematic error pattern undermines. A process designed assuming 94% overall accuracy provides adequate coverage may be designed inadequately if the 6% miss rate is concentrated in the records where human review is most critical. The process design and the model's actual performance characteristics need to be matched at the granular level, not the aggregate level.
The confidence calibration operational risk is related. AI models produce confidence scores for their classifications. Well-calibrated confidence scores can provide operational value by directing human review to the cases where the model is least confident , partially compensating for the systematic error pattern by flagging low-confidence cases for additional attention. Poorly calibrated confidence scores , where high-confidence scores do not reliably predict accuracy , eliminate this operational value and may even misdirect human attention away from high-confidence cases that are systematically more likely to be errors.
Why this matters
AI operational risks matter for TPRM because the business processes that depend on vendor AI products are typically designed around aggregate accuracy metrics rather than systematic error patterns. A vendor AI product that meets its stated accuracy specification may still create operational risk if the business process has been designed around assumptions that the systematic error pattern violates.
- Aggregate accuracy accepted as coverage characterisation
- Systematic error pattern not assessed , where in the data distribution failures concentrate
- Business process assumption alignment not validated against systematic error pattern
- Confidence calibration not assessed , do high-confidence scores reliably predict accuracy
- Human review scope not adjusted for AI systematic miss pattern
What good looks like
Mature AI operational risk programmes assess systematic error patterns alongside aggregate accuracy , specifically evaluating whether failures are concentrated in specific data segments, case types, or complexity ranges, and validating that business process assumptions are compatible with the systematic error pattern rather than only the aggregate accuracy.
- Systematic error pattern analysis , where in the data distribution do failures concentrate
- Business process assumption validation against systematic error pattern
- Confidence calibration assessment , high-confidence score reliability
- Segmented accuracy reporting , accuracy by case complexity, data source, or record type
- Human review scope adjusted for systematic miss pattern
Tooling
ML Model Analysis , SHAP for error analysis, Evidently AI for segmented performance reporting
Segmented model performance analysis tools break down accuracy metrics by data segment, case type, or complexity level , revealing systematic error patterns that aggregate metrics conceal. For TPRM practitioners, asking for segmented accuracy reporting , accuracy by case complexity level or data source , rather than only aggregate accuracy provides a specific systematic error assessment question.
Governance challenges
The governance challenge with AI operational risks is the metric selection problem. Aggregate accuracy is easy to measure and easy to communicate. Systematic error pattern analysis is more complex and may reveal risks that the business has already built processes around. The governance resolution is requiring segmented accuracy disclosure as a standard element of AI model performance reporting.
- Require segmented accuracy reporting , not just overall accuracy
- Assess business process assumptions against systematic error pattern
- Validate confidence calibration for operational use cases
- Adjust human review scope to cover systematic miss segments
- Monitor systematic error pattern stability as model drifts over time
If you are a small team
For any AI model used in a business process where the model's misses have consequences, ask for accuracy broken down by case complexity or data segment rather than only overall accuracy. Then ask which segments have the lowest accuracy and whether the business process has been designed accounting for elevated miss rates in those segments. If the process assumes overall accuracy is representative across all segments, you have identified a potential operational risk gap.
- Ask for accuracy broken down by case complexity or data segment
- Ask which segments have lowest accuracy
- Validate business process assumptions against segment-specific accuracy
- Assess human review scope coverage of high-miss segments
What to require
Ask directly:
"Can you provide accuracy metrics broken down by case complexity or data segment , and has your model's systematic error pattern been analysed to confirm that the cases most likely to be missed are the cases where your business process has human review coverage?"
Expect as evidence
- Segmented accuracy reporting by case type or complexity
- Systematic error pattern analysis
- Business process assumption validation against error pattern
- Human review scope coverage of high-miss segments
A vendor who confirms 94% accuracy should be asked for segmented accuracy. Overall accuracy is the average. Segmented accuracy reveals where the 6% concentrates. The concentration pattern determines the operational risk.
How to evidence it
- Segmented accuracy assessment records
- Systematic error pattern analysis
- Business process assumption validation
- Human review scope adjustment records
Key Takeaway
94% accuracy. 6% miss. 6% systematically concentrated in the most complex records. Human review assumed comprehensive based on 94% overall. The 94% was accurate for the aggregate. The complex records the 6% missed were exactly the records where human review was most critical and most difficult to substitute. Aggregate accuracy metrics average over the distribution. Systematic error pattern analysis reveals where in the distribution the failures concentrate. The business process assumption that needed validation was not whether the model was 94% accurate overall , it was whether the 6% it missed was distributed compatibly with the human review coverage the process provided. They were not.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association