INSTRUMENT 06 · ATTRIBUTE AGREEMENT STUDY

How do you validate this in a language quality already accepts?

Design a study with known cases, two appraisers, two trials and an explicit ceiling on system performance. Think of it like calibrating a scale against a known weight: you can't call a reading "wrong" if you never established how consistent the reference itself is.

Example: two senior inspectors each grade the same 50 welds, twice. They agree with each other (and themselves) 80% of the time — people are inconsistent too. That 80% is the ceiling: if the AI system scores 85% against that same reference, the extra 5% is noise, not proof of superiority, because you never showed humans could do better than 80% in the first place.

ACCEPTANCE CRITERIA · SET IN ADVANCE

Safety-relevant missUnacceptable at any rate
Unsupported citationZero tolerance
Silent coverage gapZero tolerance
Wrong classificationSet before testing
False alarmSet before testing
Cosmetic variationExplicitly not an error

AGREEMENT CEILING

80%

No system can score above 80% against a reference that is only 80% reproducible.

VERDICT

Restriction: welded structural packages, named suppliers, evaluated language, production approval support only

Cases: 50
Appraisers: 2 human + system
Trials: 2 each
Human agreement: 80%
System effectiveness: 85%
Achievable score ceiling: 80%

WHY THIS INSTRUMENT EXISTS

A tool that performs well against its own logic can still disagree with the reviewers who are supposed to accept its output — and a system nobody trusts in practice is not controlled, whatever its evaluation score says.

Measuring agreement between reviewers and between the system and reviewers, on the same cases, surfaces that gap before deployment instead of after a reviewer quietly stops reading the findings.