Appendices
Supporting materials for replication, verification, policy-registry auditability, and extended theoretical anchoring.
Appendices
The six appendices below form part of the manuscript. The downloads page lists the DOI-backed companion files preserved alongside the paper.
Search and Screening
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Release audit date: 2026-03-25
Public release date: 2026-03-25
Purpose
This supplement makes the registry build auditable at the screened-record level and documents the strongest search-history reconstruction still supportable from the archived materials. It records the source families searched, representative search strings, screening and deduplication rules, the retained screened ledger, the retained-corpus maintenance note required to reconcile the preserved 34-row screening ledger with the 32-row public registry corpus, the archive-reconstructed execution log in search_execution_log.csv, and the row-level screening_ledger.csv export.
The registry is reproducible from the screened-record layer forward. It does not preserve every preliminary search-engine impression, syndicated duplicate, or transient web result returned before incident-level screening. The new execution log is therefore an archive reconstruction of the search workflow, not a recovered raw-identification export from the original build materials.
Date Semantics
-
search cutoff: March 4, 2026. No incident or policy document first identified after this date enters the retained incident corpus or retained policy corpus. -
incident corpus freeze: March 4, 2026. Counts in the manuscript andincident_registry_coded.csvare keyed to this boundary. -
policy corpus freeze: March 4, 2026. Counts in the manuscript andpolicy_registry_coded.csvare keyed to this boundary. -
release audit date: March 25, 2026. Integrity-review date for the posted study materials. -
public release date: March 25, 2026. Publication date for the manuscript and standalone supplement files.
Search-Execution Log Status
search_execution_log.csv records one row per archive-reconstructed database or document-source workflow. The file includes:
-
source_family -
platform_or_database -
run_date -
search_string -
filters -
sort_order -
results_returned_n -
advanced_to_screening_n -
notesresults_returned_n = not_preservedindicates that the raw hit count was not archived in the original build materials and should not be reverse-engineered after the fact.
Screening Ledger and Source-Capture Files
Two additional CSV files now expose the screened layer more directly:
-
screening_ledger.csv— one row per candidate advanced to screening, withdedupe_cluster_id,final_disposition,exclusion_reason,source_capture_date,archive_reference, andlocal_snapshot_sha256where a local source copy exists. -
policy_source_index.csv— one row per retained policy document withdocument_genre,source_capture_date, and the public archive reference used in this release.source_capture_datein these files records the capture or verification date represented by the current public materials. It should not be back-interpreted as the original historical search date unless the file explicitly says so.
Source Families and Representative Searches
| Source family | Coverage | Representative queries or retrieval logic | Execution window | Notes |
|---|---|---|---|---|
| Legal and adjudicative sources | PACER, CourtListener, official state-court portals, docket follow-up | ("artificial intelligence" OR chatbot OR LLM OR Gemini OR ChatGPT OR Character.AI) AND (suicide OR self-harm OR wrongful death OR complaint OR order) |
Iterative collection through 2026-03-04 | Used for complaints, orders, petitions, and other formal filings. |
| Academic and benchmark sources | PubMed, Google Scholar, SSRN, arXiv, benchmark repositories | ("chatbot" OR "large language model" OR companion) AND (suicide OR self-harm OR mental health OR safety evaluation) |
Iterative collection through 2026-03-04 | Advanced only when the source contained a codable benchmark, evaluation, or incident anchor. |
| Journalistic and investigative sources | ProQuest, Factiva, Google News, outlet follow-up | (AI OR chatbot OR companion) AND (suicide OR self-harm OR death OR overdose OR delusion) |
Iterative collection through 2026-03-04 | Used for incident discovery and corroboration; syndicated duplicates were collapsed. |
| Legislative and regulatory sources | Hearing archives, agency releases, legislative records | (AI chatbot self-harm hearing); provider names plus hearing, FTC, attorney general, petition, transcript |
Iterative collection through 2026-03-04 | Used for hearing transcripts, agency complaints, and formal public record. |
| Provider-issued safety and governance materials | Preparedness frameworks, responsible-scaling policies, system cards, model cards, model specs, usage policies, governance posts | Provider-specific retrieval from official public documentation pages; one document row per codable document | Current through 2026-03-04 | Used for the 16-document policy corpus reported in the manuscript. |
Screening Workflow
Reviewer roles
-
Search, screening, and deduplication were performed by the author as a single-reviewer pass.
-
The public materials do not support a duplicated historical screen because no archived second screener log is available.
-
The preserved materials are sufficient to audit decisions from the screened-ledger layer forward.
Incident inclusion rule
Advance a candidate to the codable ledger only when all three conditions hold:
-
The record documents a specific AI-self-harm intersection rather than general commentary.
-
The record is attributable to a named platform, system, or evaluation object.
-
The public evidence is sufficient to code outcome and evidence fields under the registry rubric.
Exclusion rule
Exclude records that are:
-
unverifiable social-media claims or screenshots without attributable provenance;
-
duplicates or syndicated copies without new codable evidence;
-
general commentary or policy discussion without a specific incident or evaluation anchor;
-
too thin to support even Grade C coding.
Deduplication rule
Deduplicate first at the report level, then at the incident level. Preserve distinct platform or pathway exposures when the same harmed person encountered multiple systems in separately codable ways. This is why the Molly Russell inquest yields both 2017-IG-01 and 2017-PIN-01.
Registry Assembly Flow
| Stage | Count | Note |
|---|---|---|
| Candidate rows in preserved screening ledger | 34 | Incident-level evidence anchors advanced for codability review |
| Excluded at screening | 2 | Grade D unverifiable social-media claims |
| Included in preserved screening ledger | 32 | Retained rows before later retained-corpus maintenance |
| Reclassified out of codable retained corpus before public release | 1 | 2025-MLP-02 retained as contextual benchmark evidence only |
| New codable row added before the freeze | 1 | 2025-GEM-01 |
| Final public retained corpus | 32 | Matches incident_registry_coded.csv and the manuscript tables |
Preserved Screened Ledger
The table below preserves the screened ledger that underlies the retained-corpus maintenance note above. It contains the 34 records advanced to incident-level screening in the preserved ledger state. The machine-readable screening_ledger.csv carries the release-stage dispositions: 31 retained_in_registry_corpus rows, 1 retained_as_contextual_only row, and 2 excluded_at_screening Grade D rows.
| Log ID | Incident ID | Source type | Public identifier / access pathway | Decision | Exclusion reason |
|---|---|---|---|---|---|
| R001 | 2017-IG-01 |
Coroner/inquest | North London Coroner’s Court inquest record (2022) documenting Instagram amplification pathway | Included | — |
| R002 | 2017-PIN-01 |
Coroner/inquest | Same inquest record; Pinterest recommendation-email pathway (retained as separate pathway record) | Included | — |
| R003 | 2023-CHA-01 |
Journalism | La Libre Belgique (28 Mar 2023) and Vice/Motherboard (30 Mar 2023) reporting with described log excerpts | Included | — |
| R004 | 2023-CAI-01 |
Court filing | D. Colo. No. 1:25-cv-02907 (filed 15 Sep 2025) (complaint); CourtListener docket available | Included | — |
| R005 | 2024-CAI-02 |
Court order | M.D. Fla. No. 6:24-cv-01903 (Order, 21 May 2025); CourtListener docket available | Included | — |
| R006 | 2024-CAI-01 |
Court filings | E.D. Tex. No. 2:24-cv-01014 (filed 9 Dec 2024) (complaint; arbitration order Doc. 59, 23 Apr 2025); CourtListener docket available | Included | — |
| R007 | 2025-GPT-01 |
Court filings | JCCP No. 5431 / Case No. CGC-25-628528 (ChatGPT Product Liability Cases; filings dated 26 Dec 2025) with reproduced excerpts | Included | — |
| R008 | 2024-GPT-01 |
Court filing | Complaint excerpt without docket citation in the v1.1 source file (retained as Grade C excerpt) | Included | — |
| R009 | 2025-MAI-01 |
Evaluation report | Common Sense Media Meta AI Risk Assessment (Aug 2025), systematic testing with ages 13–17 | Included | — |
| R010 | 2025-NOM-01 |
Journalism | MIT Technology Review reporting (Jan–Apr 2025) with user screenshots regarding Nomi AI outputs | Included | — |
| R011 | 2025-REP-01 |
Complaint/filing | FTC complaint filing (Tech Justice Law Project, Young People’s Alliance, Encode) (2025) alleging dependency and crisis-detection gaps | Included | — |
| R012 | 2025-THR-01 |
Peer-reviewed study | Moore et al. (FAccT 2025) evaluation of therapy chatbots (systematic red-team) | Included | — |
| R013 | 2025-MLP-03 |
Red-team study | Schoene & Canca (2025), arXiv:2507.02990 (jailbreaking in self-harm contexts) | Included | — |
| R014 | 2025-MLP-02 |
Peer-reviewed study | McBain et al. (2025), Psychiatric Services (LLM alignment with expert clinicians) | Included | — |
| R015 | 2025-MHB-01 |
Peer-reviewed study | Pichowicz et al. (2025), Scientific Reports 15:31652 (mental-health chatbot benchmark) | Included | — |
| R016 | 2025-MLP-01 |
Government hearing | U.S. Senate Judiciary Subcommittee hearing transcript: Examining the Harm of AI Chatbots (16 Sep 2025) | Included | — |
| R017 | 2024-GPT-02 |
Journalism | The New York Times (18 Aug 2025) first-person account reproducing selected chat excerpts (paywalled access possible) | Included | — |
| R018 | 2025-ACC-01 |
Journalism | ABC News / triple j Hack interview (Aug 2025) (pseudonymized account; no transcript excerpts published) | Included | — |
| R019 | 2025-GPT-05 |
Journalism | ABC News / triple j Hack interview (Aug 2025) (pseudonymized account; no transcript excerpts published) | Included | — |
| R020 | 2024-CAI-03 |
Court filing | N.D.N.Y. No. 1:25-cv-01295 (filed 16 Sep 2025) (complaint) | Included | — |
| R021 | 2025-CAI-01 |
Court filing | D. Colo. No. 1:25-cv-02906 (filed 15 Sep 2025) (complaint); CourtListener docket available | Included | — |
| R022 | 2025-GPT-10 |
Court filing | Cal. Super. Ct. S.F. City & Cty. No. CGC-25-630808 (filed 6 Nov 2025) (complaint) | Included | — |
| R023 | 2025-GPT-02 |
Party statement | Social Media Victims Law Center press release (6 Nov 2025) (party statement; no reproduced excerpts) | Included | — |
| R024 | 2025-GPT-03 |
Party statement | Social Media Victims Law Center press release (6 Nov 2025) (party statement; no reproduced excerpts) | Included | — |
| R025 | 2025-GPT-06 |
Court filing | Cal. Super. Ct. S.F. City & Cty. (complaint dated 11 Dec 2025; case number not stated in the copy used) | Included | — |
| R026 | 2025-GPT-12 |
Court filing | Cal. Super. Ct. L.A. County No. 25STCV32382 (filed 6 Nov 2025) (complaint) | Included | — |
| R027 | 2025-GPT-08 |
Court filing | Cal. Super. Ct. S.F. City & Cty. No. CGC-25-630809 (filed 6 Nov 2025) (complaint) | Included | — |
| R028 | 2025-GPT-09 |
Court filing | Cal. Super. Ct. S.F. City & Cty. No. CGC-25-630811 (filed 6 Nov 2025) (complaint + exhibits; screenshot-verified messages) | Included | — |
| R029 | 2025-GPT-11 |
Court filing | Cal. Super. Ct. L.A. County No. 25STCV32383 (filed 6 Nov 2025) (complaint; screenshot-verified messages) | Included | — |
| R030 | 2025-GPT-07 |
Court filing | Cal. Super. Ct. L.A. County No. 25STCV32386 (filed 6 Nov 2025) (amended complaint) | Included | — |
| R031 | 2025-GPT-04 |
Court filing | Cal. Super. Ct. L.A. County No. 25STCV32379 (filed 6 Nov 2025) (complaint) | Included | — |
| R032 | 2025-GPT-13 |
Court filing | Los Angeles County Superior Court (complaint filed Jan 2026; case number not stated in the copy used) | Included | — |
| R033 | 2025-012 | Social media claim | Unverifiable social media claim (no attributable primary documentation recovered) | Excluded | Unverifiable; fails minimum reliability (Grade D) |
| R034 | 2025-013 | Social media claim | Unverifiable social media claim (no attributable primary documentation recovered) | Excluded | Unverifiable; fails minimum reliability (Grade D) |
Retained-Corpus Maintenance Update
The preserved screened ledger above is not itself the final public retained corpus. One older included contextual benchmark row was removed from the codable incident corpus, and one new A2 complaint row was added before the March 4, 2026 freeze.
| Change type | Record | Effect on the public retained corpus | Reason |
|---|---|---|---|
| Reclassified out of codable retained corpus | 2025-MLP-02 |
Removed from the 32-row incident retained corpus | Retained as contextual benchmark evidence rather than a codable incident row |
| New codable incident added | 2025-GEM-01 |
Added to the 32-row incident retained corpus | New A2 complaint record, filed March 4, 2026, before the search cutoff |
Policy-Corpus Assembly Rule
The parallel policy corpus was assembled separately from the incident ledger. A document entered the policy corpus only if it met all of the following conditions:
-
It was issued publicly by one of the five provider groups in the crosswalk.
-
It belonged to a codable governance genre: preparedness framework, responsible-scaling policy, system card, model card, model spec, usage policy, or comparable official safety/governance post.
-
It was current through the March 4, 2026 cutoff.
-
It was codable under
policy_registry_schema.csvand not merely a superseded draft or non-governance marketing page. -
The provider either appeared in the incident corpus or functioned as a major frontier-model governance comparator with a comparable public document set.
Companion-app vendors without a comparable public governance-document set by the freeze date remain out of scope for this crosswalk even when they appear in the incident registry.
Definitions and Coding Boundaries
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Public release date: 2026-03-25
Core Objects
| Term | Working definition |
|---|---|
AI-self-harm intersection |
The broad domain in which an AI system and a self-harm pathway co-occur in a documentable way. |
incident |
The underlying qualifying documented event. |
incident record |
The codable row tied to a specific harmed individual x platform/pathway exposure or a specific evaluation/demonstration record. |
pathway duplicate |
A retained second incident record for a distinct platform or pathway exposure involving the same harmed person. |
self-harm pathway |
A system-mediated sequence of exposure, reinforcement, instruction, escalation failure, or longitudinal dynamics that plausibly increases self-harm risk. |
trajectory-structured harm |
Harm that accrues across turns or sessions and depends on pathway dynamics, product affordances, or cross-session accumulation rather than a single output alone. |
human_harm_incident |
An incident record with a documented person-level death, injury, or harm-exposure outcome. |
evaluation_or_demonstration |
A benchmark, red-team, hearing-demonstration, or app-evaluation record retained because it documents unsafe behavior or governance-relevant failure signatures without a person-level injury/death event. |
Outcome Labels
Use the canonical incident labels:
| Outcome label | Definition |
|---|---|
DEATH |
A fatality is publicly reported. |
INJURY |
A non-fatal physical injury or self-harm injury is publicly reported. |
HARM_EXPOSURE |
Documented harm exposure, dependency, manipulation, or severe risk without a coded injury/death outcome. |
UNSAFE_OUTPUT |
Benchmark, red-team, hearing-demonstration, or app-evaluation evidence of unsafe output without a coded person-level injury/death event. |
Reliability Grades and Evidentiary Levels
Reliability grades describe the traceability of the incident’s public evidence anchor, not adjudicated merit.
| Grade / level | Definition |
|---|---|
A1 |
High-traceability primary-source record, such as a coroner finding, court order with reproduced excerpts, formal hearing transcript, or benchmark/evaluation report with directly inspectable methods and results. |
A2 |
Formal allegation or complaint record that remains unadjudicated. |
B |
Investigative or first-person journalism with meaningful evidentiary detail and inspectable excerpts or sourcing, but without the traceability of A1/A2. |
C |
Lower-traceability secondary reporting or party statements where the incident is still codable but evidentiary limits are material. |
Level 1 |
Documented incident record or benchmark/evaluation record anchored in inspectable materials. |
Level 2 |
Formal allegation/mechanism record that remains unadjudicated. |
Level 3 |
Paper-level governance inference; not assigned to seeded incident rows. |
Explicit B/C Boundary
-
Brequires inspectable evidentiary detail beyond bare assertion, such as reproduced excerpts, screenshots, or attributable first-person / investigative sourcing that allows a reviewer to trace why the row is codable. -
Cis used when the public record still supports row-level coding, but the available material is materially thinner: party statements, summary reporting, or limited excerpts without enough independent inspection to justifyB.
Taxonomy Key
The registry uses a six-category, 13-subcode, non-mutually-exclusive taxonomy. These codes identify loci of evaluative intervention rather than mutually exclusive harm types.
| Category | Subcode | Working definition |
|---|---|---|
GEN |
GEN-E |
Encouragement: validates or fails to discourage suicidal intent. |
GEN |
GEN-C |
Coaching: provides actionable self-harm instructions. |
GEN |
GEN-M |
Method provision: addresses lethality or means access. |
DET |
DET-FN |
Detection non-escalation: detects risk but does not transition to effective interruption, referral, or review. |
LONG |
LONG-DEP |
Dependency / erosion of offline protective factors over time. |
LONG |
LONG-DEG |
Worsening safety behavior across repeated turns or sessions. |
LONG |
LONG-DEL |
Belief reinforcement inconsistent with shared reality across turns or sessions. |
LONG |
LONG-MEM |
Memory resurfacing across sessions. |
JB |
JB-RP |
Role-play framing bypass. |
JB |
JB-MT |
Multi-turn bypass. |
JB |
JB-AC |
Academic or research framing bypass. |
RTI |
RTI-CO |
Fixed-belief non-amplification failure in a discrete exchange. |
REC/MOD |
REC/MOD-AMP |
Recommendation or platform-architecture amplification. |
Boundary Rules
RTI-CO Versus LONG-DEL
-
Code
RTI-COwhen the record documents a discrete fixed-belief non-amplification failure in one focal exchange. -
Code
LONG-DELwhen the record documents repeated or temporally extended reinforcement of that frame across turns or sessions. -
Code both when the public record shows a focal affirming exchange embedded in a longer reinforcement trajectory.
LONG-DEP Versus High Engagement
Do not code LONG-DEP for mere frequency of use. The record should show dependency, emotional exclusivity, erosion of offline supports, or a closely related attachment pattern.
DET-FN
DET-FN is not a general criticism of safety quality. It is reserved for cases where the system detects or is presented with acute risk cues but does not meaningfully interrupt, route, or escalate.
Derived Analysis View
incident_registry_analysis_view.csv is an analysis-only companion keyed by incident_record_id. It does not alter the canonical 20-field incident schema.
Legacy identifier note
-
incident_record_idis the row-level identifier used for the analysis companion. -
incident_idis retained in the same file as a backward-compatible legacy record identifier because downstream study materials already use that label. -
In the current public release,
incident_record_idandincident_idare identical string values.
record_type
-
human_harm_incidentfor person-level death, injury, or harm-exposure rows. -
evaluation_or_demonstrationfor rows whereuser_type = test.
evaluation_subtype
Use the following mapping for the six evaluation/demonstration rows:
| Incident ID | Evaluation subtype (evaluation_subtype) |
|---|---|
2025-MAI-01 |
App evaluation (app_evaluation) |
2025-MHB-01 |
Benchmark study (benchmark_study) |
2025-MLP-01 |
Hearing demonstration (hearing_demonstration) |
2025-MLP-03 |
Red-team exercise (red_team) |
2025-NOM-01 |
App evaluation (app_evaluation) |
2025-THR-01 |
Benchmark study (benchmark_study) |
trajectory_structured_strict
-
yeswhen one or more adjudicatedLONG-*subcodes are present. -
nootherwise.This field is fully mechanical: it is derived from the canonical taxonomy coding only, and it is the manuscript’s primary trajectory indicator.
trajectory_basis_strict
-
Use semicolon-separated
LONG-*subcodes whentrajectory_structured_strict = yes. -
Use
nonewhentrajectory_structured_strict = no.
trajectory_structured_broad
-
yeswhen the record satisfies the strict rule or the public record otherwise documents a cross-turn or cross-session pathway strongly enough to satisfy the prespecified appendix rule. -
nowhen only event-level evidence is supported. -
indeterminatewhen the public materials describe an extended pathway but remain too thin for a conservative broad-routeyes.This field is retained as a descriptive sensitivity check rather than the headline prevalence measure.
trajectory_basis_broad
-
Use semicolon-separated
LONG-*subcodes whentrajectory_structured_broad = yesbecause of adjudicated longitudinal codes. -
Use
documented_multi_turn_pathway_insufficient_for_subcodewhentrajectory_structured_broad = indeterminate. -
Use
nonewhentrajectory_structured_broad = no.
Legacy trajectory alias
-
trajectory_structured_flagis retained as a backward-compatible alias oftrajectory_structured_broad. -
trajectory_basisis retained as a backward-compatible alias oftrajectory_basis_broad.
One-page decision tree
-
Does the adjudicated registry row contain one or more
LONG-*subcodes? If yes, settrajectory_structured_strict = yes,trajectory_structured_broad = yes, and carry theLONG-*codes into both basis fields. -
If no
LONG-*code is present, does the public record still document repeated turns, repeated sessions, retained memory, or another clearly cumulative interaction pathway? If no, settrajectory_structured_strict = noandtrajectory_structured_broad = no. -
If the public record does document a repeated or cumulative pathway without adjudicated
LONG-*support, ask whether the evidence is specific enough to support a conservative broad-route longitudinal designation. If yes, settrajectory_structured_broad = yes. If no, settrajectory_structured_broad = indeterminate.
Worked examples
| Incident ID | Strict trajectory flag (trajectory_structured_strict) |
Broad trajectory flag (trajectory_structured_broad) |
Rationale |
|---|---|---|---|
2025-GEM-01 |
yes |
yes |
Adjudicated LONG-DEL, LONG-DEP, and LONG-MEM make the strict route mechanical. |
2025-GPT-04 |
yes |
yes |
LONG-DEL is present, so both flags resolve yes. |
2024-GPT-02 |
no |
indeterminate |
The public account suggests a repeated pathway, but the released evidence remains too thin for conservative LONG-* assignment. |
2025-MHB-01 |
no |
no |
The benchmark shows unsafe outputs without a retained multi-turn trajectory object at the released row level. |
2025-MAI-01 |
yes |
yes |
LONG-MEM is explicitly coded from the evaluation evidence. |
Cluster Fields
-
person_cluster_idgroups explicit pathway duplicates or clearly identical harmed-person rows for analytic deduplication. -
provider_cluster_idgroups rows by provider or platform family. -
case_cluster_idgroups rows by the same case family, inquest, hearing package, or benchmark study. -
pathway_duplicate_group_idgroups rows that document the same underlying harmed-person pathway but are intentionally retained as separate platform/pathway records.
Study Architecture
| Layer | Unit | What it contributes |
|---|---|---|
| Registry layer | 32 incident records | Incident metadata, outcome category, evidence grade, and evidence anchor. |
| Taxonomy layer | 13 multi-label subcodes | Governance-relevant failure loci across GEN, DET, LONG, JB, RTI, and REC/MOD. |
| Analysis layer | Derived analysis view | incident_record_id, strict and broad trajectory fields, and cluster-aware sensitivity fields. |
| Policy crosswalk layer | 16 public governance documents | Public placement of self-harm in preparedness, product-safety, and trust-and-safety documents. |
Policy Crosswalk Glossary
| Shorthand | Meaning |
|---|---|
Class IV |
Reactive incident-domain treatment, such as post-incident patches or case-specific remediation. |
native |
A policy document names self-harm as an evaluative object and pairs it with a usable evaluation procedure. |
proxy_represented |
A policy document addresses self-harm only through adjacent categories such as dangerous content or user well-being. |
operationally_underspecified |
A policy document names the domain but does not provide a usable prospective evaluation procedure. |
Absent / not represented |
Used as the absence label for both evaluative_object_status and risk_domain_status; self-harm is not present in codable form in the document. |
Ambiguous / insufficiently classifiable |
Self-harm is mentioned but the public text does not support a defensible Class I to Class IV placement. |
Recorded governance tier (risk_domain_status) |
The highest governance tier actually evidenced for self-harm in the document, coded as Class I, Class II, Class III, Class IV, Absent / not represented, or Ambiguous / insufficiently classifiable. |
Coder confidence (coder_confidence) |
Confidence based on document clarity, not on agreement with the organization’s framing. |
Document genre (document_genre) |
Broader analytic grouping used for within-genre policy comparison before provider-level rollups. |
All self-harm layers present (all_self_harm_layers_present) |
Semicolon-separated list of every governance layer in the document that explicitly carries self-harm treatment. |
One-Page Glossary for Frequent Shorthand
| Shorthand | Meaning |
|---|---|
full retained corpus |
All 32 incident rows in incident_registry_coded.csv. |
human-harm subset |
The 27 rows with person-level death, injury, or harm-exposure outcomes. |
A1/A2 primary robustness subset |
The 20 human-harm rows with higher-traceability A1 or A2 evidence. |
A1/A2/B secondary sensitivity subset |
The 22 human-harm rows with A1, A2, or B evidence. |
C-excluded mechanism sensitivity |
Human-harm rows after excluding C-grade records that still carry mechanism subcodes. |
Reliability Appendix
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Public release date: 2026-03-25
Source Basis
This appendix now distinguishes between two reliability components:
-
the retained 8-incident independent-rater summary audit already archived in earlier study materials; and
-
the new full-corpus audit scaffold archived in
incident_taxonomy_rater_matrix.csvandincident_taxonomy_adjudication_log.csv.Historical raw second-rater exports for the retained 8-incident audit are still not publicly available. The new full-corpus files therefore expose the primary-coder side of the audit matrix and the forward adjudication surface, but they should not be described as a completed independent second-rater re-audit.
Retained 8-Incident Audit Summary
Recoverable audited incident set
The retained audit summary covers the following incident records:
-
2017-IG-01 -
2023-CAI-01 -
2024-CAI-02 -
2025-GEM-01 -
2025-GPT-01 -
2025-MAI-01 -
2025-MLP-01 -
2025-REP-01
Coverage note
-
The recoverable audited set covers deaths, harm exposures, and unsafe-output rows.
-
The recoverable audited set does not include an injury row.
-
The recoverable audited set is concentrated in
A1andA2materials; it does not provide a raw archived audit surface for the most interpretation-sensitiveBandCrows.The retained 8-incident table is an archived historical agreement summary preserved from an earlier coding state. It should not be read as a row-level parity check or independent validation of the current 2.0 released taxonomy assignments, which have since been updated. Accordingly, the pooled and per-subcode agreement statistics below are historical audit artifacts, not current-release subcode validation metrics.
Retained pooled agreement summary
| Metric | Value |
|---|---|
| Incidents audited | 8 of 32 (25.0%) |
Total decisions (incident x subcode) |
104 |
| Both raters positive | 23 |
| Both raters negative | 73 |
| Rater 2 only | 8 |
| Rater 1 only | 0 |
| Raw agreement | 92.3% |
| Pooled Cohen’s kappa | 0.80 |
| Approximate 95% CI for pooled kappa | 0.67 to 0.93 |
The confidence interval is the asymptotic interval implied by the retained pooled 2 x 2 decision table (23 / 0 / 8 / 73), not a bootstrap interval from raw coder-level exports.
Retained per-subcode summary
| Subcode | Prevalence (either rater) | Raw agreement | Kappa | Positive agreement | PP | NN | R2 only | R1 only |
|---|---|---|---|---|---|---|---|---|
| GEN-E | 4/8 | 8/8 | 1.00 | 1.00 | 4 | 4 | 0 | 0 |
| GEN-C | 3/8 | 7/8 | 0.71 | 0.80 | 2 | 5 | 1 | 0 |
| GEN-M | 2/8 | 6/8 | 0.00 | 0.00 | 0 | 6 | 2 | 0 |
| DET-FN | 4/8 | 7/8 | 0.75 | 0.86 | 3 | 4 | 1 | 0 |
| LONG-DEP | 5/8 | 8/8 | 1.00 | 1.00 | 5 | 3 | 0 | 0 |
| LONG-DEG | 1/8 | 7/8 | 0.00 | 0.00 | 0 | 7 | 1 | 0 |
| LONG-DEL | 3/8 | 8/8 | 1.00 | 1.00 | 3 | 5 | 0 | 0 |
| LONG-MEM | 1/8 | 7/8 | 0.00 | 0.00 | 0 | 7 | 1 | 0 |
| JB-RP | 1/8 | 8/8 | 1.00 | 1.00 | 1 | 7 | 0 | 0 |
| JB-MT | 1/8 | 7/8 | 0.00 | 0.00 | 0 | 7 | 1 | 0 |
| JB-AC | 1/8 | 8/8 | 1.00 | 1.00 | 1 | 7 | 0 | 0 |
| RTI-CO | 5/8 | 7/8 | 0.75 | 0.89 | 4 | 3 | 1 | 0 |
| REC/MOD-AMP | 1/8 | 8/8 | 1.00 | 1.00 | 1 | 7 | 0 | 0 |
For sparse subcodes, the agreement counts and positive agreement are more informative than kappa alone.
Full-Corpus Audit Scaffold
Two new files now define the forward reliability surface:
-
downloads/incidents/incident_taxonomy_rater_matrix.csv -
downloads/incidents/incident_taxonomy_adjudication_log.csv
What the scaffold contains
-
One row per
incident_record_id x review_itemcombination across the full 32-record retained registry corpus. -
review_family = taxonomy_subcodefor the 13 multi-label taxonomy decisions. -
Additional review rows for
outcome_category,reliability_grade,eligibility_retention_status, andtrajectory_structured_broad. -
evidence_stratumpopulated asA1,A2,B, orCfor planned stratified reporting. -
blinding_requirement = blind_to_rater_1_and_expected_outcomesfor every row. -
rater_1populated from the current released coding. -
Blank
rater_2andadjudicated_valuecolumns reserved for a future independent second-rater completion pass. -
A status column marking every row as
awaiting_independent_second_rater.
What the scaffold does not contain
-
It does not reconstruct historical second-rater decisions that were not archived.
-
It does not justify stronger claims than the retained 8-incident summary audit supports.
-
It does not convert the current release into a completed full-corpus independent-rater package.
Intended reporting strata for a completed pass
If a true second-rater pass is later completed, report agreement separately for:
-
A1/A2 -
B -
CDo not collapse
BandCinto a single interpretive-risk stratum.
Planned review-surface rules
-
outcome_categoryandreliability_gradeshould be coded independently from the released CSV. -
eligibility_retention_statusshould distinguishretained_in_registry_corpusfrom any future contextual-only or excluded rows if the corpus changes. -
trajectory_structured_broadrequires independent review because it is not purely mechanical. -
trajectory_structured_strictdoes not require a second coder because it is mechanically derivable from adjudicatedLONG-*values.
Interpretation-Seams Still Worth Monitoring
The retained audit materials and the current codebook continue to identify three seams that deserve attention in any future completed re-audit:
-
GEN-Eversus no code when a response is ambiguous between encouragement and neutral acknowledgment. -
LONG-DELversusRTI-CO, where the practical distinction is temporal: focal exchange versus repeated reinforcement trajectory. -
LONG-DEPversus high engagement without a documented dependency or exclusivity signal.
What This Appendix Supports
This appendix supports a narrower claim than the earlier wording: the current taxonomy is operationalized and partially reliability-tested, but the public materials do not yet contain a completed full-corpus independent second-rater archive. The retained 8-incident summary audit remains informative; the new scaffold makes the next audit pass forward-completable and file-level auditable.
Policy Reliability Appendix
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Public release date: 2026-03-25
Purpose
This appendix defines the auditable reliability surface for the 16-document policy crosswalk and records what is and is not yet archived in the public release.
Archived Files
-
downloads/policy/policy_registry_second_rater.csv -
downloads/policy/policy_registry_adjudication_log.csv -
downloads/policy/policy_registry_coded.csv -
downloads/policy/policy_registry_schema.csv -
downloads/policy/policy_registry_coding_manual.md
Current Status
The policy crosswalk remains a single-coder descriptive pass in this release. The new files add the row-level second-rater scaffold and adjudication surface, but they do not contain completed independent second-rater decisions.
Independent-Review Surface
The intended non-derived analytic review surface is:
-
document_genre -
self_harm_named -
self_harm_domain_form -
risk_domain_status -
first_class_equivalence -
named_frontier_risk_domains -
self_harm_comparability_note -
threshold_mechanism -
deployment_consequence -
external_review_requirement -
specialized_red_teaming -
monitoring_obligation -
incident_response_layer -
all_self_harm_layers_present -
governance_fragmentation -
temporal_model -
multi_turn_evaluation -
memory_personalization_attention -
dependency_attention -
cross_session_accumulation -
mechanism_coverage -
evaluative_object_status -
epistemic_risk_attention -
architecture_level_attention -
detection_to_action_couplingDerived gap fields should be regenerated only after those base fields are independently reviewed and adjudicated.
Scaffold Design
policy_registry_second_rater.csv
-
One row per policy document.
-
record_idpopulated for all 16 documents. -
document_genreand all analytic coding fields left blank pending an actual independent second-rater pass. -
audit_status = awaiting_independent_second_raterfor every row.
policy_registry_adjudication_log.csv
-
Reserved for field-level disagreements.
-
Should be populated only after a completed second-rater pass exists.
Recommended Completion Rule
When a true second-rater pass is available:
-
Complete
policy_registry_second_rater.csvindependently from the canonical coded CSV. -
Compare coder 1 versus coder 2 on the non-derived analytic fields only.
-
Log every disagreement in
policy_registry_adjudication_log.csv. -
Recompute derived gap fields from the adjudicated base fields.
-
Report exact agreement and Cohen’s kappa for single-choice nominal fields.
-
Report exact-set agreement and label-level positive agreement for
mechanism_coverage. -
Report results within
document_genrebefore any provider-level rollup.
What This Appendix Supports
This appendix supports two bounded claims:
-
the public materials now provide a file-level, forward-completable reliability surface for the policy crosswalk; and
-
the policy crosswalk should still be described as a single-coder descriptive audit until the scaffold is actually completed by an independent second rater.
Policy Crosswalk Appendix
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Public release date: 2026-03-25
Purpose
This appendix makes the 16-document policy crosswalk directly auditable from the public study materials. It states the corpus-selection rule, summarizes the public-document field distributions, lists the fixed fields intended for independent review, and exposes the row-level coding ledger derived from policy_registry_coded.csv. It should be read together with policy_registry_second_rater.csv, which provides the forward second-rater scaffold, and policy_source_index.csv, which records public source-capture metadata for the retained policy corpus.
Corpus Selection Rule
A document enters the public policy corpus only if it satisfies all of the following conditions:
-
It is publicly issued by one of the five selected provider groups represented in the crosswalk.
-
It belongs to a codable governance genre: preparedness framework, responsible-scaling policy, system card, model card, model spec, usage policy, or comparable official safety/governance post.
-
It is current through the March 4, 2026 search cutoff and policy corpus freeze.
-
It is codable under
policy_registry_schema.csvusing the highest-evidenced-tier rule. -
It is not a superseded draft, marketing page, or non-codable announcement.
-
The provider either appears in the incident registry or functions as a major frontier-model governance comparator with a comparable public-document set by the freeze date.
The resulting corpus contains 16 documents across Anthropic, Google / Google DeepMind, Meta, OpenAI, and xAI. Companion-app vendors without a comparable public governance-document set by the freeze date remain out of scope even when they appear in the incident registry.
Provider-Group Freeze
Google and Google DeepMind are treated as one provider group for corpus selection and provider-level rollups, while their public documents retain separate issuing-organization labels at the row level.
| Provider group | Inclusion basis at the March 4, 2026 freeze |
|---|---|
| Anthropic | Included as a frontier-model provider with a public set spanning responsible-scaling and product-safety / governance-post genres by the freeze. |
| Google / Google DeepMind | Included because Gemini-related materials are both incident-relevant and part of the frontier-governance comparison set, with codable documents issued under both Google and Google DeepMind labels by the freeze. |
| Meta | Included because Meta appears in the incident record and also had codable public governance documents in the predefined genres by the freeze. |
| OpenAI | Included because OpenAI appears repeatedly in the incident record and had both preparedness and product-safety documents publicly available by the freeze. |
| xAI | Included as a major frontier-model governance comparator with codable frontier and product-safety documents publicly available by the freeze. |
Corpus Composition
By organization
| Organization | Documents |
|---|---|
| Anthropic | 5 |
| 1 | |
| Google DeepMind | 2 |
| Meta | 2 |
| OpenAI | 4 |
| xAI | 2 |
By document genre
| Document genre | Documents |
|---|---|
Frontier or scaling (frontier_or_scaling) |
5 |
Product-safety artifact (product_safety_artifact) |
5 |
Policy or governance post (policy_or_governance_post) |
6 |
By genre group
| Genre group | Documents | Tier profile | Native documents | Documents with any threshold mechanism | Documents with any multi-turn evaluation |
|---|---|---|---|---|---|
Frontier or scaling (frontier_or_scaling) |
5 | Absent / not represented in all 5 documents | 0 | 0 | 0 |
Product-safety artifact (product_safety_artifact) |
5 | Class II in 3 documents; Class III in 2 | 1 | 2 | 2 |
Policy or governance post (policy_or_governance_post) |
6 | Class II in 2 documents; Class III in 3; Absent / not represented in 1 | 1 | 0 | 3 |
High-level field summary
| Field | Distribution |
|---|---|
Self-harm named (self_harm_named) |
Explicit in 10 documents; absent in 6 |
Self-harm domain form (self_harm_domain_form) |
Subdomain in 6 documents; proxy domain in 2; native domain in 2; absent in 6 |
Recorded governance tier (risk_domain_status) |
Class II in 5 documents; Class III in 5; absent / not represented in 6 |
Evaluative-object status (evaluative_object_status) |
Native in 2 documents; proxy represented in 2; operationally underspecified in 6; ontologically absent in 6 |
First-class equivalence (first_class_equivalence) |
No in all 16 documents |
Threshold mechanism (threshold_mechanism) |
Partial in 2 documents; none in 14 |
Deployment consequence (deployment_consequence) |
Partial in 3 documents; none in 13 |
Multi-turn evaluation (multi_turn_evaluation) |
Explicit in 3 documents; partial in 2; none in 11 |
All self-harm layers present (all_self_harm_layers_present) |
Product safety in 8 documents; trust and safety in 2; absent in 6 |
Fixed Fields for Independent Review
The manuscript discusses the following fields as the minimum independent-review surface for the policy crosswalk:
-
document_genre -
self_harm_named -
self_harm_domain_form -
risk_domain_status -
evaluative_object_status -
first_class_equivalence -
threshold_mechanism -
deployment_consequence -
multi_turn_evaluation
Public-archive status of independent review
The policy crosswalk remains a single-coder descriptive pass. The public materials include policy_registry_second_rater.csv and policy_registry_adjudication_log.csv as forward-completable audit files, but they do not yet include a completed independent second-rater review for all 16 documents. This appendix therefore exposes the row-level evidence, excerpts, and field values needed for such a review, but it should not be described as a completed independent adjudication table.
Row-Level Public-Document Ledger
The table below is generated from the current coded CSV and is the authoritative row-level audit surface for the manuscript. Two retained Meta canonical URLs on ai.meta.com required login during the March 25, 2026 release audit; the study materials preserve their canonical URLs together with source-capture metadata and archived excerpt provenance.
| Record ID | Organization | Document | Date | URL | Key coded fields | Supporting excerpt | Confidence |
|---|---|---|---|---|---|---|---|
| ANTHROPIC-2025-01 | Anthropic | How people use Claude for support, advice, and companionship (blog_post; policy_or_governance_post) |
2025-06-27 | https://www.anthropic.com/news/how-people-use-claude-for-support-advice-and-companionship | named=explicit; form=proxy_domain; tier=Class III; genre=policy_or_governance_post; layers=product_safety; object=proxy_represented; first_class=no; threshold=none; deploy=none; multi_turn=partial |
when it does, it’s typically for safety reasons … refusing to provide dangerous weight loss advice or support self-harm. | moderate |
| ANTHROPIC-2025-02 | Anthropic | Building safeguards for Claude (blog_post; policy_or_governance_post) |
2025-08-12 | https://www.anthropic.com/news/building-safeguards-for-claude | named=explicit; form=subdomain; tier=Class II; genre=policy_or_governance_post; layers=product_safety; object=operationally_underspecified; first_class=no; threshold=none; deploy=partial; multi_turn=explicit |
We assess Claude’s adherence to our Usage Policy on topics like child exploitation or self-harm … including … extended multi-turn conversations. | high |
| ANTHROPIC-2025-03 | Anthropic | Protecting the well-being of our users (blog_post; policy_or_governance_post) |
2025-12-18 | https://www.anthropic.com/news/protecting-well-being-of-users | named=explicit; form=native_domain; tier=Class II; genre=policy_or_governance_post; layers=product_safety; object=native; first_class=no; threshold=none; deploy=none; multi_turn=explicit |
We focus on two areas: how Claude handles conversations about suicide and self-harm … we use a combination of model training and product interventions. | high |
| ANTHROPIC-2025-04 | Anthropic | Sharing our compliance framework for California’s Transparency in Frontier AI Act (blog_post; policy_or_governance_post) |
2025-12-19 | https://www.anthropic.com/news/compliance-framework-SB53 | named=absent; form=absent; tier=Absent / not represented; genre=policy_or_governance_post; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
Our FCF describes how we assess and mitigate cyber offense, chemical, biological, radiological, and nuclear threats … as well as the risks of AI sabotage and loss of control. | high |
| ANTHROPIC-2026-01 | Anthropic | Anthropic’s Responsible Scaling Policy: Version 3.0 (responsible_scaling_policy; frontier_or_scaling) |
2026-02-24 | https://www.anthropic.com/news/responsible-scaling-policy-v3 | named=absent; form=absent; tier=Absent / not represented; genre=frontier_or_scaling; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
Non-novel chemical/biological weapons production … High-stakes sabotage opportunities … Automated R&D in key domains. | high |
| GOOGLE-2024-01 | Generative AI Prohibited Use Policy (usage_policy; policy_or_governance_post) |
2024-12-17 | https://policies.google.com/terms/generative-ai/use-policy | named=explicit; form=subdomain; tier=Class III; genre=policy_or_governance_post; layers=trust_and_safety; object=operationally_underspecified; first_class=no; threshold=none; deploy=none; multi_turn=none |
Facilitates self-harm. | high | |
| GOOGLEDEEPMIND-2025-01 | Google DeepMind | Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2 (frontier_safety_framework; frontier_or_scaling) |
2025-11 | https://deepmind.google/models/fsf-reports/gemini-3-pro/ | named=absent; form=absent; tier=Absent / not represented; genre=frontier_or_scaling; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
It currently covers four risk domains … CBRN …, cybersecurity, machine learning R&D, and harmful manipulation, and also includes … misalignment risk. | high |
| GOOGLEDEEPMIND-2025-02 | Google DeepMind | Gemini 3 Pro - Model Card (model_card; product_safety_artifact) |
2025-12 | https://deepmind.google/models/model-cards/gemini-3-pro | named=explicit; form=proxy_domain; tier=Class III; genre=product_safety_artifact; layers=product_safety; object=proxy_represented; first_class=no; threshold=none; deploy=none; multi_turn=partial |
Dangerous content (e.g., promoting suicide, or instructing in activities that could cause real-world harm). | high |
| META-2023-01 | Meta | Llama 2 Acceptable Use Policy (usage_policy; policy_or_governance_post) |
2023-07-18 | https://ai.meta.com/llama/use-policy/ | named=explicit; form=subdomain; tier=Class III; genre=policy_or_governance_post; layers=trust_and_safety; object=operationally_underspecified; first_class=no; threshold=none; deploy=none; multi_turn=none |
Self-harm or harm to others, including suicide, cutting, and eating disorders. | moderate |
| META-2025-01 | Meta | Code World Model Preparedness Report (preparedness_framework; frontier_or_scaling) |
2025-09-24 | https://ai.meta.com/research/publications/code-world-model-preparedness-report/ | named=absent; form=absent; tier=Absent / not represented; genre=frontier_or_scaling; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
we conducted an automated assessment of CWM capabilities in two domains … namely Cybersecurity and Chemical & Biological risks. | high |
| OPENAI-2025-01 | OpenAI | Preparedness Framework (Version 2) (preparedness_framework; frontier_or_scaling) |
2025-04-15 | https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf | named=absent; form=absent; tier=Absent / not represented; genre=frontier_or_scaling; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
We currently focus this work on three areas of frontier capability … Biological and Chemical … Cybersecurity … AI Self-improvement capabilities. | high |
| OPENAI-2025-02 | OpenAI | Addendum to GPT-5 System Card: Sensitive Conversations (system_card; product_safety_artifact) |
2025-10-27 | https://cdn.openai.com/pdf/3da476af-b937-47fb-9931-88a851620101/addendum-to-gpt-5-system-card-sensitive-conversations.pdf | named=explicit; form=subdomain; tier=Class II; genre=product_safety_artifact; layers=product_safety; object=operationally_underspecified; first_class=no; threshold=partial; deploy=partial; multi_turn=none |
self-harm/intent … self-harm/instructions | high |
| OPENAI-2025-03 | OpenAI | Model Spec (2025/12/18) (model_spec; product_safety_artifact) |
2025-12-18 | https://model-spec.openai.com/2025-12-18.html | named=explicit; form=native_domain; tier=Class II; genre=product_safety_artifact; layers=product_safety; object=operationally_underspecified; first_class=no; threshold=none; deploy=none; multi_turn=none |
The assistant must not encourage or enable self-harm … always advising that immediate help should be sought if the user is in imminent danger. | high |
| OPENAI-2026-01 | OpenAI | GPT-5.3 Instant System Card (system_card; product_safety_artifact) |
2026-03-03 | https://deploymentsafety.openai.com/gpt-5-3-instant/gpt-5-3-instant.pdf | named=explicit; form=subdomain; tier=Class II; genre=product_safety_artifact; layers=product_safety; object=native; first_class=no; threshold=partial; deploy=partial; multi_turn=explicit |
we implemented dynamic multi-turn evaluations for mental health, emotional reliance, and self-harm that simulate extended conversations across these domains. | high |
| XAI-2025-01 | xAI | Grok 4.1 Model Card (model_card; product_safety_artifact) |
2025-11-17 | https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf | named=explicit; form=subdomain; tier=Class III; genre=product_safety_artifact; layers=product_safety; object=operationally_underspecified; first_class=no; threshold=none; deploy=none; multi_turn=none |
we employ input filters to reject specific classes of sensitive requests, such as those involving bioweapons, chemical weapons, self-harm, and child sexual abuse material. | high |
| XAI-2025-02 | xAI | xAI Frontier Artificial Intelligence Framework (frontier_safety_framework; frontier_or_scaling) |
2025-12-30 | https://data.x.ai/2025-12-31-xai-frontier-artificial-intelligence-framework.pdf | named=absent; form=absent; tier=Absent / not represented; genre=frontier_or_scaling; layers=absent; object=Absent / not represented; first_class=no; threshold=none; deploy=none; multi_turn=none |
This FAIF discusses two major categories of AI risk - malicious use and loss of control. | high |
Reading the Table
-
tieris the codedrisk_domain_status. -
genreis the codeddocument_genre, used for within-genre comparison before provider-level rollups. -
layersis the codedall_self_harm_layers_present. -
objectis the codedevaluative_object_status. -
first_classindicates whether self-harm is treated on frontier-comparable terms. -
thresholdanddeployreport whether self-harm findings trigger thresholded review or deployment consequences in the public document itself. -
multi_turnreports whether the document contains explicit, partial, or absent longitudinal evaluation procedures.
What This Appendix Supports
This appendix supports cautious claims about what the public documents themselves operationalize. It does not support claims about undisclosed internal practice, unpublished red-team protocols, or organization-wide governance measures not evidenced in the document row being coded.
Reference and Source Index
Version: 2.0
Search cutoff / incident corpus freeze / policy corpus freeze: 2026-03-04
Public release date: 2026-03-25
This appendix indexes the public evidence routes used in the registry and the public policy documents included in the crosswalk. The DOI-backed release files listed on the downloads page include the datasets, coding manuals, citation metadata, and audit-surface files preserved for public release.
Incident Citation and Evidence-Route Index
Archive status = present means a local evidence file is bundled in the release. Archive status = external_only means the public route is an external primary source and no local mirror is bundled.
| Incident ID | Platform | Outcome | Grade | Canonical primary citation / evidence anchor | Evidence route | Archive status |
|---|---|---|---|---|---|---|
2017-IG-01 |
Meta Instagram | DEATH | A1 | /incidents/2017-IG-01_Coroner_Report.pdf |
/evidence/incidents/2017-IG-01/2017-IG-01_Coroner_Report.pdf |
present |
2017-PIN-01 |
DEATH | A1 | /incidents/2017-PIN-01_Coroner_Report.pdf |
/evidence/incidents/2017-PIN-01/2017-PIN-01_Coroner_Report.pdf |
present | |
2023-CAI-01 |
Character.AI | DEATH | A2 | 1:25-cv-02907; /incidents/2023-CAI-01.pdf |
/evidence/incidents/2023-CAI-01/2023-CAI-01.pdf |
present |
2023-CHA-01 |
Chai Research (EleutherAI GPT-J fine-tuned) | DEATH | B | /incidents/2023-CHA-01_LaLibre.pdf; /incidents/2023-CHA-01_Vice.pdf |
/evidence/incidents/2023-CHA-01/2023-CHA-01_LaLibre.pdf |
present |
2024-CAI-01 |
Character.AI | INJURY | A2 | 2:24-cv-01014; /incidents/2024-CAI-01.pdf; /incidents/2024-CAI-01_Order.pdf |
/evidence/incidents/2024-CAI-01/2024-CAI-01.pdf |
present |
2024-CAI-02 |
Character.AI | DEATH | A1 | 6:24-cv-01903; /incidents/2024-CAI-02.pdf; /incidents/2024-CAI-02_Order.pdf |
/evidence/incidents/2024-CAI-02/2024-CAI-02.pdf |
present |
2024-CAI-03 |
Character.AI | INJURY | A2 | 1:25-cv-01295; /incidents/2024-CAI-03.pdf |
/evidence/incidents/2024-CAI-03/2024-CAI-03.pdf |
present |
2024-GPT-01 |
OpenAI ChatGPT | HARM_EXPOSURE | C | Christopher ‘Kirk’ Shamblin and Alicia Shamblin, individually and as successors-in-interest to Decedent, Zane Shamblin v. OpenAI, Inc., et al., No. 25STCV32382 (Cal. Super. Ct., Los Angeles County, filed Nov. 6, 2025); /incidents/2024-GPT-01.pdf |
/evidence/incidents/2024-GPT-01/2024-GPT-01.pdf |
present |
2024-GPT-02 |
OpenAI ChatGPT (GPT-4o, persona: “Harry”) | DEATH | B | https://www.nytimes.com/2025/08/18/opinion/chat-gpt-mental-health-suicide.html | https://www.nytimes.com/2025/08/18/opinion/chat-gpt-mental-health-suicide.html | external_only |
2025-ACC-01 |
AI companion chatbots (multiple; vendor unspecified) | HARM_EXPOSURE | C | /incidents/2025-ACC-01_ABC.pdf |
/evidence/incidents/2025-ACC-01/2025-ACC-01_ABC.pdf |
present |
2025-CAI-01 |
Character.AI | HARM_EXPOSURE | A2 | 1:25-cv-02906; /incidents/2025-CAI-01.pdf |
/evidence/incidents/2025-CAI-01/2025-CAI-01.pdf |
present |
2025-GEM-01 |
Google Gemini 2.5 Pro | DEATH | A2 | 5:26-cv-01849-VKD; /incidents/2025-GEM-01.pdf |
/evidence/incidents/2025-GEM-01/2025-GEM-01.pdf |
present |
2025-GPT-01 |
OpenAI ChatGPT (GPT-4o) | DEATH | A2 | CGC-25-628528; /incidents/2025-GPT-01.pdf; /incidents/2025-GPT-01_JCCP_Opposition.pdf |
/evidence/incidents/2025-GPT-01/2025-GPT-01.pdf |
present |
2025-GPT-02 |
OpenAI ChatGPT (GPT-4o) | HARM_EXPOSURE | C | https://techjusticelaw.org/2025/11/06/social-media-victims-law-center-and-tech-justice-law-project-lawsuits-accuse-chatgpt-of-emotional-manipulation-supercharging-ai-delusions-and-acting-as-a-suicide-coach/ | https://techjusticelaw.org/2025/11/06/social-media-victims-law-center-and-tech-justice-law-project-lawsuits-accuse-chatgpt-of-emotional-manipulation-supercharging-ai-delusions-and-acting-as-a-suicide-coach/ | external_only |
2025-GPT-03 |
OpenAI ChatGPT (GPT-4o) | HARM_EXPOSURE | C | https://techjusticelaw.org/2025/11/06/social-media-victims-law-center-and-tech-justice-law-project-lawsuits-accuse-chatgpt-of-emotional-manipulation-supercharging-ai-delusions-and-acting-as-a-suicide-coach/ | https://techjusticelaw.org/2025/11/06/social-media-victims-law-center-and-tech-justice-law-project-lawsuits-accuse-chatgpt-of-emotional-manipulation-supercharging-ai-delusions-and-acting-as-a-suicide-coach/ | external_only |
2025-GPT-04 |
OpenAI ChatGPT (GPT-4o) | DEATH | A2 | 25STCV32379; /incidents/2025-GPT-04.pdf |
/evidence/incidents/2025-GPT-04/2025-GPT-04.pdf |
present |
2025-GPT-05 |
OpenAI ChatGPT (model unspecified) | INJURY | C | /incidents/2025-GPT-05_ABC.pdf |
/evidence/incidents/2025-GPT-05/2025-GPT-05_ABC.pdf |
present |
2025-GPT-06 |
OpenAI ChatGPT (GPT-4o; ChatGPT Plus) | DEATH | A2 | CGC-25-631477; /incidents/2025-GPT-06.pdf |
/evidence/incidents/2025-GPT-06/2025-GPT-06.pdf |
present |
2025-GPT-07 |
OpenAI ChatGPT (GPT-4o Plus) | HARM_EXPOSURE | A2 | 25STCV32386; /incidents/2025-GPT-07.pdf |
/evidence/incidents/2025-GPT-07/2025-GPT-07.pdf |
present |
2025-GPT-08 |
OpenAI ChatGPT (GPT-4o) | HARM_EXPOSURE | A2 | Karen Enneking, individually and as successor-in-interest to decedent Joshua Enneking v. OpenAI, Inc., et al., No. CGC-25-630809 (Cal. Super. Ct., San Francisco County, filed Nov. 6, 2025); /incidents/2025-GPT-08.pdf |
/evidence/incidents/2025-GPT-08/2025-GPT-08.pdf |
present |
2025-GPT-09 |
OpenAI ChatGPT (GPT-4o) | INJURY | A2 | CGC-25-630811; /incidents/2025-GPT-09.pdf |
/evidence/incidents/2025-GPT-09/2025-GPT-09.pdf |
present |
2025-GPT-10 |
OpenAI ChatGPT (GPT-4o) | DEATH | A2 | CGC-25-630808; /incidents/2025-GPT-10.pdf |
/evidence/incidents/2025-GPT-10/2025-GPT-10.pdf |
present |
2025-GPT-11 |
OpenAI ChatGPT (GPT-4o) | INJURY | A2 | 25STCV32383; /incidents/2025-GPT-11.pdf |
/evidence/incidents/2025-GPT-11/2025-GPT-11.pdf |
present |
2025-GPT-12 |
OpenAI ChatGPT | DEATH | A2 | Christopher ‘Kirk’ Shamblin and Alicia Shamblin, individually and as successors-in-interest to Decedent, Zane Shamblin v. OpenAI, Inc., et al., No. 25STCV32382 (Cal. Super. Ct., Los Angeles County, filed Nov. 6, 2025); https://techjusticelaw.org/2025/11/06/social-media-victims-law-center-and-tech-justice-law-project-lawsuits-accuse-chatgpt-of-emotional-manipulation-supercharging-ai-delusions-and-acting-as-a-suicide-coach/; https://selfharm.ai/incidents/2025-GPT-12.pdf; /incidents/2025-GPT-12.pdf |
/evidence/incidents/2025-GPT-12/2025-GPT-12.pdf |
present |
2025-GPT-13 |
OpenAI ChatGPT (GPT-4o) | DEATH | A2 | /incidents/2025-GPT-13.pdf |
/evidence/incidents/2025-GPT-13/2025-GPT-13.pdf |
present |
2025-MAI-01 |
Meta AI (Instagram/WhatsApp/Facebook) | UNSAFE_OUTPUT | A1 | https://www.commonsensemedia.org/ai-ratings/meta-ai-risk-assessment; https://www.commonsensemedia.org/sites/default/files/featured-content/files/csm-ai-risk-assessment-metaai-08152025.pdf; /incidents/2025-MAI-01.pdf |
/evidence/incidents/2025-MAI-01/2025-MAI-01.pdf |
present |
2025-MHB-01 |
Mental health chatbots (29 agents) | UNSAFE_OUTPUT | A1 | Pichowicz, Kotas & Piotrowski (2025), Scientific Reports 15:31652, DOI: 10.1038/s41598-025-17242-4; https://doi.org/10.1038/s41598-025-17242-4; https://www.nature.com/articles/s41598-025-17242-4 | https://doi.org/10.1038/s41598-025-17242-4 | external_only |
2025-MLP-01 |
U.S. Senate Judiciary Committee, Subcommittee on Crime and Counterterrorism | HARM_EXPOSURE | A1 | https://www.judiciary.senate.gov/committee-activity/hearings/examining-the-harm-of-ai-chatbots; /incidents/2025-MLP-01_Raine_Testimony.pdf; /incidents/2025-MLP-01_Garcia_Testimony.pdf; /incidents/2025-MLP-01_Doe_Testimony.pdf; /incidents/2025-MLP-01_Torney_Testimony.pdf |
/evidence/incidents/2025-MLP-01/2025-MLP-01_Raine_Testimony.pdf |
present |
2025-MLP-03 |
Multi-LLM red-team (6 models) | UNSAFE_OUTPUT | A1 | https://arxiv.org/pdf/2507.02990.pdf; /incidents/2025-MLP-03.pdf |
/evidence/incidents/2025-MLP-03/2025-MLP-03.pdf |
present |
2025-NOM-01 |
Nomi AI (Glimpse AI) | UNSAFE_OUTPUT | B | https://incidentdatabase.ai/cite/1041/; https://futurism.com/ai-girlfriend-encouraged-suicide | https://incidentdatabase.ai/cite/1041/ | external_only |
2025-REP-01 |
Replika (Luka Inc.) | HARM_EXPOSURE | A2 | https://techjusticelaw.org/wp-content/uploads/2025/01/Complaint-and-Petition-for-Investigation-Re-Replika.pdf; /incidents/2025-REP-01.pdf |
/evidence/incidents/2025-REP-01/2025-REP-01.pdf |
present |
2025-THR-01 |
Therapy chatbots (multi-app evaluation) | UNSAFE_OUTPUT | A1 | Moore et al. (2025), FAccT ’25, DOI: 10.1145/3715275.3732039; https://doi.org/10.1145/3715275.3732039; https://facctconference.org/static/docs/facct2025-206archivalpdfs/facct2025-final197-acmpaginated.pdf | https://doi.org/10.1145/3715275.3732039 | external_only |
Policy Document Source Index
The 16 policy rows are directly indexed below. For quick lookup, the public-document identifiers are:
Two Meta canonical URLs are retained even though ai.meta.com required login during the March 25, 2026 release audit; corresponding source-capture metadata are preserved in policy_source_index.csv.
| Record ID | Document | Date | URL |
|---|---|---|---|
| ANTHROPIC-2025-01 | How people use Claude for support, advice, and companionship | 2025-06-27 | https://www.anthropic.com/news/how-people-use-claude-for-support-advice-and-companionship |
| ANTHROPIC-2025-02 | Building safeguards for Claude | 2025-08-12 | https://www.anthropic.com/news/building-safeguards-for-claude |
| ANTHROPIC-2025-03 | Protecting the well-being of our users | 2025-12-18 | https://www.anthropic.com/news/protecting-well-being-of-users |
| ANTHROPIC-2025-04 | Sharing our compliance framework for California’s Transparency in Frontier AI Act | 2025-12-19 | https://www.anthropic.com/news/compliance-framework-SB53 |
| ANTHROPIC-2026-01 | Anthropic’s Responsible Scaling Policy: Version 3.0 | 2026-02-24 | https://www.anthropic.com/news/responsible-scaling-policy-v3 |
| GOOGLE-2024-01 | Generative AI Prohibited Use Policy | 2024-12-17 | https://policies.google.com/terms/generative-ai/use-policy |
| GOOGLEDEEPMIND-2025-01 | Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2 | 2025-11 | https://deepmind.google/models/fsf-reports/gemini-3-pro/ |
| GOOGLEDEEPMIND-2025-02 | Gemini 3 Pro - Model Card | 2025-12 | https://deepmind.google/models/model-cards/gemini-3-pro |
| META-2023-01 | Llama 2 Acceptable Use Policy | 2023-07-18 | https://ai.meta.com/llama/use-policy/ |
| META-2025-01 | Code World Model Preparedness Report | 2025-09-24 | https://ai.meta.com/research/publications/code-world-model-preparedness-report/ |
| OPENAI-2025-01 | Preparedness Framework (Version 2) | 2025-04-15 | https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf |
| OPENAI-2025-02 | Addendum to GPT-5 System Card: Sensitive Conversations | 2025-10-27 | https://cdn.openai.com/pdf/3da476af-b937-47fb-9931-88a851620101/addendum-to-gpt-5-system-card-sensitive-conversations.pdf |
| OPENAI-2025-03 | Model Spec (2025/12/18) | 2025-12-18 | https://model-spec.openai.com/2025-12-18.html |
| OPENAI-2026-01 | GPT-5.3 Instant System Card | 2026-03-03 | https://deploymentsafety.openai.com/gpt-5-3-instant/gpt-5-3-instant.pdf |
| XAI-2025-01 | Grok 4.1 Model Card | 2025-11-17 | https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf |
| XAI-2025-02 | xAI Frontier Artificial Intelligence Framework | 2025-12-30 | https://data.x.ai/2025-12-31-xai-frontier-artificial-intelligence-framework.pdf |
Data Availability and Declarations
Data Availability
Version 2.0 study materials are available at https://selfharm.ai/downloads/ and are keyed to the DOI deposit at https://doi.org/10.17605/OSF.IO/49AGJ. The public release includes the coded incident registry, the analysis companion, the coded policy registry, schemas, coding manuals, search-execution log, screening and source-capture ledgers, incident and policy audit scaffolds, the manuscript PDF, and citation metadata.
Some records are intentionally routed to external primary sources rather than archived locally, and some legal materials remain paywalled or externally hosted. No private platform logs or non-public human-subject materials are included.
Ethics Review
This study used only publicly available records, court filings, hearing materials, public journalism, benchmark reports, and provider-issued governance documents. No direct participant recruitment, private records, or non-public human-subject data were used. On that basis, the work did not require IRB or human-subjects review under the applicable public-records standard.
Funding
This work received no external funding.
Competing Interests
The author declares no competing interests.
Sensitive-Material Handling
This article discusses suicide and self-harm but does not reproduce operational instructions. Means-specific details are redacted or omitted, quotations are minimized, and the focus remains on system behavior, evidentiary discipline, and prevention-oriented evaluation design. Because the corpus includes minors and deaths, the article relies only on already public materials and does not introduce new identifying detail beyond what is already part of the source record.
AI-Use Statement
Where generative AI assistance was used, it was not used to identify incidents, select sources, assign incident taxonomy codes, determine reliability grades, extract evidentiary claims, or resolve interpretive disputes. No AI-generated text was treated as evidence. AI assistance was used only for drafting, rephrasing, table formatting, and revision support, with all retained text checked against the cited source material and coded materials.
References
Artificial Intelligence Incident Database. (n.d.-a). CSETv1 charts. Retrieved March 2026, from https://incidentdatabase.ai/taxonomies/csetv1/
Artificial Intelligence Incident Database. (n.d.-b). Editor’s guide. Retrieved March 2026, from https://incidentdatabase.ai/editors-guide/
Artificial Intelligence Incident Database. (n.d.-c). GMF charts. Retrieved March 2026, from https://incidentdatabase.ai/taxonomies/gmf/
Anthropic. (2025, June 27). How people use Claude for support, advice, and companionship. https://www.anthropic.com/news/how-people-use-claude-for-support-advice-and-companionship
Anthropic. (2025, August 12). Building safeguards for Claude. https://www.anthropic.com/news/building-safeguards-for-claude
Anthropic. (2025, December 18). Protecting the well-being of our users. https://www.anthropic.com/news/protecting-well-being-of-users
Anthropic. (2025, December 19). Sharing our compliance framework for California’s Transparency in Frontier AI Act. https://www.anthropic.com/news/compliance-framework-SB53
Anthropic. (2026, February 24). Anthropic’s Responsible Scaling Policy: Version 3.0. https://www.anthropic.com/news/responsible-scaling-policy-v3
Centers for Disease Control and Prevention. (2024, November 29). Youth mental health: The numbers. https://www.cdc.gov/healthy-youth/mental-health/mental-health-numbers.html
Chatterji, A., Cunningham, T., Deming, D. J., Hitzig, Z., Ong, C. W., Shan, C. Y., & Wadman, K. (2025). How people use ChatGPT (Working Paper No. 34255). National Bureau of Economic Research. https://www.nber.org/papers/w34255
Faverio, M., & Sidoti, O. (2025, December 9). Teens, social media and AI chatbots 2025. Pew Research Center. https://www.pewresearch.org/internet/2025/12/09/teens-social-media-and-ai-chatbots-2025/
Garcia v. Character Technologies, Inc., No. 6:24-cv-01903 (M.D. Fla. May 21, 2025) (order on motions to dismiss).
Gavalas v. Google LLC and Alphabet Inc., No. 5:26-cv-01849 (N.D. Cal. filed March 4, 2026).
Google. (2024, December 17). Generative AI Prohibited Use Policy. https://policies.google.com/terms/generative-ai/use-policy
Google DeepMind. (2025, November). Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2. https://deepmind.google/models/fsf-reports/gemini-3-pro/
Google DeepMind. (2025, December). Gemini 3 Pro - Model Card. https://deepmind.google/models/model-cards/gemini-3-pro
Meta. (2023, July 18). Llama 2 Acceptable Use Policy. https://ai.meta.com/llama/use-policy/
Meta. (2025, September 24). Code World Model Preparedness Report. https://ai.meta.com/research/publications/code-world-model-preparedness-report/
Moore, J., Grabb, D., Agnew, W., Klyman, K., Chancellor, S., Ong, D. C., & Haber, N. (2025). Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (pp. 599–627). https://doi.org/10.1145/3715275.3732039
OECD. (2025). Towards a common reporting framework for AI incidents (OECD Artificial Intelligence Papers No. 34). OECD Publishing. https://doi.org/10.1787/f326d4ac-en
OECD.AI. (n.d.-a). Overview and methodology of the AI Incidents and Hazards Monitor. Retrieved March 2026, from https://oecd.ai/en/incidents-methodology
OECD.AI. (n.d.-b). Name it to tame it: Defining AI incidents and hazards. Retrieved March 2026, from https://oecd.ai/en/wonk/defining-ai-incidents-and-hazards
Ofcom. (2025, August 22). Protecting people from online suicide and self-harm material. https://www.ofcom.org.uk/online-safety/illegal-and-harmful-content/protecting-people-from-online-suicide-and-self-harm-material
OpenAI. (2025, April 15). Preparedness Framework (Version 2). https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
OpenAI. (2025, October 27). Addendum to GPT-5 System Card: Sensitive Conversations. https://cdn.openai.com/pdf/3da476af-b937-47fb-9931-88a851620101/addendum-to-gpt-5-system-card-sensitive-conversations.pdf
OpenAI. (2025, December 18). Model Spec (2025/12/18). https://model-spec.openai.com/2025-12-18.html
OpenAI. (2026, March 3). GPT-5.3 Instant System Card. https://deploymentsafety.openai.com/gpt-5-3-instant/gpt-5-3-instant.pdf
Pichowicz, W., Kotas, M., & Piotrowski, P. (2025). Performance of mental health chatbot agents in detecting and managing suicidal ideation. Scientific Reports, 15, 31652. https://doi.org/10.1038/s41598-025-17242-4
Robb, M. B., & Mann, S. (2025). Talk, trust, and trade-offs: How and why teens use AI companions. Common Sense Media. https://www.commonsensemedia.org/sites/default/files/research/report/talk-trust-and-trade-offs_2025_web.pdf
Spittal, M. J., et al. (2025). Can suicide and self-harm in children and adolescents be predicted by mental health and social care contact? A systematic review and meta-analysis. PLOS Medicine, 22(8), e1004581. https://doi.org/10.1371/journal.pmed.1004581
Tricco, A. C., Lillie, E., Zarin, W., et al. (2018). PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Annals of Internal Medicine, 169(7), 467-473. https://doi.org/10.7326/M18-0850
U.S. Senate Judiciary Committee, Subcommittee on Crime and Counterterrorism. (2025, September 16). Examining the harm of AI chatbots. https://www.judiciary.senate.gov/committee-activity/hearings/examining-the-harm-of-ai-chatbots
World Health Organization. (2019, June 24). Preventing suicide: A resource series. https://www.who.int/publications/i/item/preventing-suicide-a-resource-series
xAI. (2025, November 17). Grok 4.1 Model Card. https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf
xAI. (2025, December 30). xAI Frontier Artificial Intelligence Framework. https://data.x.ai/2025-12-31-xai-frontier-artificial-intelligence-framework.pdf