Risk Scoring Subjectivity
Same Evidence. Same Vendor. Score of 34. Score of 67. Both From Your Team.
6 min read · 2 July 2026 · Compliance
A financial services company's TPRM team conducted an inter-rater reliability exercise , two assessors independently scored the same vendor using the organisation's standard risk assessment methodology with the same evidence package. The first assessor scored the vendor 34 out of 100 (low risk). The second scored the same vendor 67 out of 100 (moderate-to-high risk). Both assessors used the same scoring rubric, evaluated the same evidence, and applied the same weighting formula. The divergence came from interpretation of ambiguous evidence and compensating controls. The vendor lacked MFA on service accounts and offered network segmentation plus monitoring as compensating controls. Assessor 1 weighted the compensating controls significantly, reducing the access control finding's impact to minor. Assessor 2 considered network segmentation an inadequate compensating control for missing MFA and maintained the access control finding at significant impact. A difference of thirty-three points on the same vendor with the same evidence, both using the same methodology. Neither score was wrong by the methodology's rules. Both were valid interpretations. The organisation had no mechanism for identifying which was more accurate. They had been using the same methodology for three years without knowing it produced this range of inter-assessor variation.
What is Risk Scoring Subjectivity, Really?
Risk scoring subjectivity is the variation in assessment outcomes that results from differences in assessor interpretation of evidence, compensating controls, and scoring criteria within a defined risk assessment methodology. It is not a failure of the methodology's structure , the structure may be completely consistent and well-defined. It is a property of any assessment methodology that requires human interpretation of ambiguous inputs, which is any methodology that assesses complex security questions from evidence that is never perfectly complete or unambiguous.
The compensating control interpretation problem is the most significant source of scoring subjectivity in TPRM. Compensating controls are alternative security measures offered in place of a control that is specified by the assessment framework but not implemented by the vendor. Whether a specific compensating control adequately reduces the risk created by the missing control is a judgment that depends on the assessor's understanding of the threat model, the alternative control's actual effectiveness, and the context-specific factors that make a compensating control more or less adequate. Two assessors with different backgrounds , one from a network security background and one from an identity security background , may reach genuinely different conclusions about whether network segmentation adequately compensates for missing MFA on service accounts.
The evidence quality interpretation problem compounds the compensating control issue. Evidence that confirms a control exists , a policy document, a configuration screenshot, a third-party attestation , requires interpretation to assess its quality and its relevance to the specific risk question. An assessor who has seen many low-quality policy documents that are not implemented may discount policy documentation more heavily than an assessor who treats documentation as meaningful evidence of programme intent. The same document, evaluated by two assessors with different experience of the gap between documentation and implementation, may produce different findings.
- Compensating control interpretation variation , assessors with different security backgrounds reaching different conclusions about compensating control adequacy
- Evidence quality interpretation variation , different weightings of documentation versus technical evidence
- Threshold calibration differences , assessors calibrating what counts as satisfactory, partially satisfactory, and unsatisfactory differently
- No inter-rater reliability testing , organisations unaware of the scoring variation their methodology produces
- Score precision implying false accuracy , numerical scores suggesting precision that interpretive subjectivity cannot support
Why this matters
Risk scoring subjectivity matters for TPRM because risk scores drive consequential decisions , vendor approval or rejection, required risk mitigation measures, monitoring intensity, and contractual requirements. A risk score that varies by thirty-three points depending on which assessor conducted the assessment produces decisions that are as variable as the assessors. Vendors who happen to be assessed by a more lenient assessor receive better treatment than equivalent vendors assessed by a more strict one. Assessment decisions that appear to be based on objective methodology are substantially influenced by the specific assessor.
The portfolio comparability problem is the operational consequence of scoring subjectivity at scale. TPRM programmes that manage hundreds of vendor relationships need to compare risk scores across vendors to prioritise remediation, monitoring, and reassessment resources. If scores for different vendors were produced by different assessors with different interpretive tendencies, the scores may not be comparable , a score of 45 from one assessor may not represent the same risk level as a score of 45 from a different assessor. Portfolio-level risk analysis built on scores that are not comparable produces incorrect prioritisation.
Where most teams get this wrong
The most consistent failure is not testing inter-rater reliability before treating scores as comparable and actionable. Most TPRM programmes develop a scoring methodology and deploy it without measuring the variation that methodology produces across assessors. The variation exists and influences decisions before anyone knows it exists.
- No inter-rater reliability testing of scoring methodology
- Compensating control adequacy criteria not defined , left to assessor judgment
- Evidence quality criteria not specified , documentation versus technical evidence weighted by assessor experience
- Score variation undetected , different assessors producing different scores for the same vendor without comparison
- Portfolio decisions built on non-comparable scores
What good looks like
Mature risk scoring programmes define criteria that reduce interpretive freedom , specifying what counts as adequate compensating controls for specific control gaps, what evidence quality thresholds are required for different scoring levels, and conducting regular inter-rater reliability testing to measure and address scoring variation.
- Compensating control criteria defined , specific criteria for what compensating controls are adequate for specific control gaps
- Evidence quality thresholds specified , what evidence types are required for each scoring level
- Inter-rater reliability testing , regular testing of scoring variation across assessors
- Score calibration sessions , assessors reviewing the same cases together to align interpretive standards
- Automated scoring where feasible , reducing interpretive freedom by automating decisions that can be rule-based
Tooling
TPRM Platforms with Automated Scoring , BitSight, SecurityScorecard, RiskRecon
Security rating platforms provide automated, algorithmically derived risk scores based on observable external signals , reducing assessor interpretation by replacing human judgment with defined algorithms for specific signal types. For TPRM practitioners, incorporating automated scoring components alongside questionnaire-based assessment reduces overall scoring subjectivity for the signals that automated tools can measure.
GRC with Scoring Calibration , Archer, MetricStream
GRC platforms with scoring calibration features support assessor alignment through shared scoring rubrics, calibration case libraries, and inter-rater reliability tracking. For TPRM practitioners, using platforms that support assessor calibration and track scoring variation provides the organisational infrastructure for managing subjectivity.
Governance challenges
The governance challenge with risk scoring subjectivity is that eliminating it entirely is not possible , any assessment of complex security questions requires human judgment, and human judgment varies. The governance goal is reducing variation to an acceptable level through criteria specification, assessor training, calibration exercises, and automated scoring components , not eliminating it, but managing it so that decisions are sufficiently consistent to be defensible.
- Conduct inter-rater reliability testing annually , test scoring variation and track it over time
- Define compensating control adequacy criteria for the most commonly offered compensating controls
- Hold calibration sessions where assessors score the same vendor independently and compare
- Document interpretive decisions , how specific evidence was interpreted and why
- Use automated scoring for objective signals , reducing subjective interpretation where rule-based scoring is feasible
If you are a small team
Run one inter-rater reliability test with your assessment team. Take a completed assessment with its evidence package and have two assessors score it independently without seeing each other's work. Compare the scores. If the difference is more than ten points on a hundred-point scale, your methodology is producing more variation than is acceptable for consistent decision-making. Identify the specific questions where the scores diverged most and develop more specific criteria for those questions.
- Run inter-rater reliability test with two assessors on the same completed assessment
- Identify the questions with greatest score divergence , these need more specific criteria
- Define compensating control adequacy criteria for the most common compensating controls
- Hold quarterly calibration sessions for assessors
What to require
Ask directly:
"For any compensating controls you are offering in place of controls that our assessment framework requires , specifically [name control gap and offered compensating control] , can you provide technical evidence of the compensating control's implementation and explain specifically how it reduces the risk created by the missing control?"
Expect as evidence
- Technical evidence of compensating control implementation
- Specific explanation of how the compensating control reduces the missing control's risk
- Any independent validation of compensating control effectiveness
Assessors who disagree about compensating control adequacy need not both be wrong. They need a defined criterion for what adequacy means in this context. Define the criterion. Apply it consistently. The score that results is more defensible than the score that reflects one assessor's judgment.
How to evidence it
- Inter-rater reliability testing records
- Compensating control adequacy criteria documentation
- Assessor calibration session records
- Score variation tracking over time
Key Takeaway
The score of 34 and the score of 67 are both valid interpretations of the same evidence under the same methodology. The methodology is consistent. The interpretive space within the methodology is wide enough for two assessors to reach different conclusions about the adequacy of network segmentation as a compensating control for missing service account MFA. The number implies precision. The methodology does not deliver it. Testing inter-rater reliability before relying on scores for consequential decisions is the governance step that reveals how much precision the methodology actually provides. Define the compensating control criteria. Calibrate the assessors. Test the variation. The score that emerges from a calibrated team is more defensible than the score that reflects one assessor's interpretive tendency on a busy Tuesday.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association