Failure-Mode Taxonomy and Codebook
Governance-oriented coding focused on evaluative loci rather than a single harm ontology.
Failure-Mode Taxonomy and Coding Reliability
This section shifts from case counts to the coding scheme itself. The taxonomy is heterogeneous by design because the documented cases do not fail in only one way. It classifies governance-relevant loci of failure rather than forcing the corpus into a single ontology. That design is necessary because the retained rows include output failures, detection failures, longitudinal failures, framing-based bypasses, fixed-belief non-amplification failures, and recommender or moderation pathways.
Table: Taxonomy category prevalence and governance loci
| Category | Subcodes | Full retained corpus (n/32) | Human-harm subset (n/27) | Human-harm higher-traceability subset (A1/A2, n/20) | Governance locus |
|---|---|---|---|---|---|
| GEN | GEN-E, GEN-C, GEN-M | 19 | 15 | 10 | Output safety and self-harm response boundaries |
| DET | DET-FN | 8 | 5 | 5 | Detection-to-escalation coupling |
| LONG | LONG-DEP, LONG-DEG, LONG-DEL, LONG-MEM | 19 | 18 | 15 | Multi-session design, memory behavior, dependency, and longitudinal guardrail drift |
| JB | JB-RP, JB-MT, JB-AC | 5 | 4 | 4 | Framing-based bypass and adversarial evaluation |
| RTI | RTI-CO | 8 | 7 | 6 | Non-amplification in fixed-belief exchanges |
| REC/MOD | REC/MOD-AMP | 4 | 4 | 4 | Surfacing, ranking, and amplification pathways |
In the outcome-based human-harm subset, LONG-DEP appears in 13 of 27 records, GEN-E in 11, GEN-C and RTI-CO in 7 each, and LONG-DEL in 6. In the higher-traceability human-harm subset (A1/A2), LONG-DEP appears in 12 of 20, GEN-E, GEN-C, RTI-CO, and LONG-DEL in 6 each, and DET-FN in 5. One retained death record remains intentionally uncoded at the subcode level because the public record is insufficient for conservative assignment.
The core ambiguity rules remain conservative. RTI-CO is assigned only when a discrete fixed-belief non-amplification failure is documented. LONG-DEL requires repeated or temporally extended reinforcement rather than a single quotable exchange. LONG-DEP requires more than high engagement; the record must support over-attachment, exclusivity, isolation, or erosion of offline protective factors. When evidence is insufficient, the rule is no code rather than forced coding.
Table: Reliability assets and current public status
| Reliability asset | Current public status |
|---|---|
| Retained incident summary audit | Historical 8-incident agreement summary preserved from an earlier coding state; not in row-level parity with the current 2.0 released taxonomy coding. |
| Full-corpus incident audit matrix | Archived as a forward scaffold with rater_1 populated and rater_2 / adjudicated_value reserved for future completion |
| Policy second-rater surface | Archived as a forward scaffold; not yet completed |
Appendix C (Reliability Appendix) reports the retained 8-incident summary and the current scaffold boundaries. The most interpretation-sensitive seams remain GEN-E versus no-code, LONG-DEL versus RTI-CO, and LONG-DEP versus high engagement without dependency features.
Reliability demands by evidence layer. This paper’s empirical claims rest on three evidentiary layers with different reliability burdens. First, the incident taxonomy has partial inter-rater support from the retained 8-incident audit (25% of the 32-row corpus; pooled kappa = 0.80), but that audit is concentrated in A1/A2 materials and excludes injury rows; full-corpus independent review remains incomplete. Second, the 16-document governance crosswalk codes document-level public-text properties. Because this layer remains a single-coder descriptive pass, findings such as first_class_equivalence = yes = 0 should be read as current corpus results rather than adjudicated zeroes. Third, the infrastructure-layer claim—that the inspected AI Incident Database and OECD schemas lack explicit fields for cross-episode dependency, degradation over time, and cross-session accumulation—is directly observable from public schema definitions and does not depend on inter-rater coding.