Failure-Mode Taxonomy and Coding Reliability

This section shifts from case counts to the coding scheme itself. The taxonomy is heterogeneous by design because the documented cases do not fail in only one way. It classifies governance-relevant loci of failure rather than forcing the corpus into a single ontology. That design is necessary because the retained rows include output failures, detection failures, longitudinal failures, framing-based bypasses, fixed-belief non-amplification failures, and recommender or moderation pathways.

Table: Taxonomy category prevalence and governance loci

Category Subcodes Full retained corpus (n/32) Human-harm subset (n/27) Human-harm higher-traceability subset (A1/A2, n/20) Governance locus
GEN GEN-E, GEN-C, GEN-M 19 15 10 Output safety and self-harm response boundaries
DET DET-FN 8 5 5 Detection-to-escalation coupling
LONG LONG-DEP, LONG-DEG, LONG-DEL, LONG-MEM 19 18 15 Multi-session design, memory behavior, dependency, and longitudinal guardrail drift
JB JB-RP, JB-MT, JB-AC 5 4 4 Framing-based bypass and adversarial evaluation
RTI RTI-CO 8 7 6 Non-amplification in fixed-belief exchanges
REC/MOD REC/MOD-AMP 4 4 4 Surfacing, ranking, and amplification pathways

In the outcome-based human-harm subset, LONG-DEP appears in 13 of 27 records, GEN-E in 11, GEN-C and RTI-CO in 7 each, and LONG-DEL in 6. In the higher-traceability human-harm subset (A1/A2), LONG-DEP appears in 12 of 20, GEN-E, GEN-C, RTI-CO, and LONG-DEL in 6 each, and DET-FN in 5. One retained death record remains intentionally uncoded at the subcode level because the public record is insufficient for conservative assignment.

The core ambiguity rules remain conservative. RTI-CO is assigned only when a discrete fixed-belief non-amplification failure is documented. LONG-DEL requires repeated or temporally extended reinforcement rather than a single quotable exchange. LONG-DEP requires more than high engagement; the record must support over-attachment, exclusivity, isolation, or erosion of offline protective factors. When evidence is insufficient, the rule is no code rather than forced coding.

Table: Reliability assets and current public status

Reliability asset Current public status
Retained incident summary audit Historical 8-incident agreement summary preserved from an earlier coding state; not in row-level parity with the current 2.0 released taxonomy coding.
Full-corpus incident audit matrix Archived as a forward scaffold with rater_1 populated and rater_2 / adjudicated_value reserved for future completion
Policy second-rater surface Archived as a forward scaffold; not yet completed

Appendix C (Reliability Appendix) reports the retained 8-incident summary and the current scaffold boundaries. The most interpretation-sensitive seams remain GEN-E versus no-code, LONG-DEL versus RTI-CO, and LONG-DEP versus high engagement without dependency features.

Reliability demands by evidence layer. This paper’s empirical claims rest on three evidentiary layers with different reliability burdens. First, the incident taxonomy has partial inter-rater support from the retained 8-incident audit (25% of the 32-row corpus; pooled kappa = 0.80), but that audit is concentrated in A1/A2 materials and excludes injury rows; full-corpus independent review remains incomplete. Second, the 16-document governance crosswalk codes document-level public-text properties. Because this layer remains a single-coder descriptive pass, findings such as first_class_equivalence = yes = 0 should be read as current corpus results rather than adjudicated zeroes. Third, the infrastructure-layer claim—that the inspected AI Incident Database and OECD schemas lack explicit fields for cross-episode dependency, degradation over time, and cross-session accumulation—is directly observable from public schema definitions and does not depend on inter-rater coding.