{"resourceId":"naic-ai-risk-evaluation-supplement","versions":[{"version":"legacy/2026-08-31/naic-ai-risk-evaluation-supplement","resource":{"id":"naic-ai-risk-evaluation-supplement","title":"Twelve-state pilot converts AI oversight into a risk-tiered regulatory evidence request","organization":"National Association of Insurance Commissioners","sector":"State insurance regulation","geography":"United States; 12 participating states","publishedAt":"August 31, 2026","sourceName":"AI Risk Evaluation Supplement version 5.0 and Pilot Project Summary","sourceLabel":"NAIC Big Data and Artificial Intelligence Working Group","sourceUrl":"https://content.naic.org/committees/h/big-data-artificial-intelligence-wg","evidenceClass":"standards-guidance","outcomeClass":"emerging","topics":["developers-agents","data-security","governance-procurement","operating-model"],"finding":"After field testing with California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin, NAIC exposed version 5.0 of its AI Risk Evaluation Supplement for public comment. The March–September pilot applies the tool in market-conduct reviews, financial analysis, and financial examinations while allowing jurisdiction-specific tailoring.","sledRelevance":"Although focused on regulated insurers, this is an operational model for state oversight of third-party AI: begin with a materiality-scoped inventory, escalate inquiry for direct consumer or financial impact, and request governance, model, data, limitation, explainability, and third-party evidence proportionate to risk.","evidence":"NAIC's change summary says pilot feedback led version 5.0 to make the model inventory explicit; distinguish AI systems from AI models; add materiality thresholds, inherent-risk language, direct-consumer and material-financial-impact fields, explainability and transparency questions, document-and-page evidence references, third-party model oversight, model-risk and limitation fields, data-to-model mapping, and an agentic-AI definition. The pilot is designed to assess usability and training needs; final effectiveness results are not yet available.","architectureImplications":"Maintain separate but linked inventories for systems, models, use cases, data sets, and third parties. Use the initial inventory to decide which systems warrant deeper technical and governance evidence, and support different thresholds for back-office, consumer-facing, and financially material workloads.","governanceImplications":"Adopt a proportional assurance process that can reuse existing audit and examination mechanisms, identify inherent risk before controls, require traceable evidence rather than unsupported questionnaire answers, coordinate requests across oversight bodies, and refine the instrument from operator and regulated-entity feedback.","securityPrivacyImplications":"Ask how sensitive and protected attributes enter models, how data sets map to model uses, how third-party models are governed, what limitations and failure modes are known, and how confidentiality is preserved when detailed model and control evidence is collected.","caveats":"Version 5.0 remains an exposure draft and the 12-state pilot is still underway. Participating jurisdictions may adapt the questions, results have not yet established inter-rater consistency or regulatory effectiveness, and insurance-specific materiality concepts will require translation for other SLED services."}},{"version":"enrichment/2026-09-05T02:42:45.193Z/naic-ai-risk-evaluation-supplement","resource":{"id":"naic-ai-risk-evaluation-supplement","title":"Twelve-state pilot converts AI oversight into a risk-tiered regulatory evidence request","organization":"National Association of Insurance Commissioners","sector":"State insurance regulation","geography":"United States; 12 participating states","publishedAt":"August 31, 2026","publicationDate":"2026-08-31","eventDate":null,"sourceName":"AI Risk Evaluation Supplement version 5.0 and Pilot Project Summary","sourceLabel":"NAIC Big Data and Artificial Intelligence Working Group","sourceUrl":"https://content.naic.org/committees/h/big-data-artificial-intelligence-wg","evidenceClass":"standards-guidance","outcomeClass":"emerging","topics":["developers-agents","data-security","governance-procurement","operating-model"],"finding":"After field testing with California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin, NAIC exposed version 5.0 of its AI Risk Evaluation Supplement for public comment. The March–September pilot applies the tool in market-conduct reviews, financial analysis, and financial examinations while allowing jurisdiction-specific tailoring.","sledRelevance":"Although focused on regulated insurers, this is an operational model for state oversight of third-party AI: begin with a materiality-scoped inventory, escalate inquiry for direct consumer or financial impact, and request governance, model, data, limitation, explainability, and third-party evidence proportionate to risk.","evidence":"NAIC's change summary says pilot feedback led version 5.0 to make the model inventory explicit; distinguish AI systems from AI models; add materiality thresholds, inherent-risk language, direct-consumer and material-financial-impact fields, explainability and transparency questions, document-and-page evidence references, third-party model oversight, model-risk and limitation fields, data-to-model mapping, and an agentic-AI definition. The pilot is designed to assess usability and training needs; final effectiveness results are not yet available.","architectureImplications":"Maintain separate but linked inventories for systems, models, use cases, data sets, and third parties. Use the initial inventory to decide which systems warrant deeper technical and governance evidence, and support different thresholds for back-office, consumer-facing, and financially material workloads.","governanceImplications":"Adopt a proportional assurance process that can reuse existing audit and examination mechanisms, identify inherent risk before controls, require traceable evidence rather than unsupported questionnaire answers, coordinate requests across oversight bodies, and refine the instrument from operator and regulated-entity feedback.","securityPrivacyImplications":"Ask how sensitive and protected attributes enter models, how data sets map to model uses, how third-party models are governed, what limitations and failure modes are known, and how confidentiality is preserved when detailed model and control evidence is collected.","caveats":"Version 5.0 remains an exposure draft and the 12-state pilot is still underway. Participating jurisdictions may adapt the questions, results have not yet established inter-rater consistency or regulatory effectiveness, and insurance-specific materiality concepts will require translation for other SLED services.","streamIds":["state-government"],"roles":{"sales":"Interpretation — Problem and stakeholders: State insurance regulators, examiners, model-risk specialists, and entity liaisons need proportionate evidence requests for diverse AI uses. Discovery: Can reviewers distinguish models from systems, identify direct consumer or financial impact, and locate support for questionnaire answers? Value hypothesis: Materiality-scoped inventory could direct examiner effort toward consequential uses while reducing repetitive requests. Potential engagement: A bounded examination-workflow pilot with training and evidence-traceability review. Evidence boundary: Version 5.0 was an exposure draft and the 12-state pilot remained underway in the archive. The record does not establish effectiveness, consistent scoring, or a mandatory requirement everywhere. Applicability beyond insurance remains limited until materiality concepts are translated and tested.","engineering":"Interpretation — Fit: Structure assurance evidence rather than treating the supplement as a model-validation product. Architecture: Link system, model, use-case, dataset, third-party, limitation, and document-page records, including data-to-model relationships. Prerequisites: Agreed materiality and inherent-risk definitions, entity evidence, and a common examination rubric. Constraints: Jurisdictional tailoring and draft revisions can change fields; preserve version provenance. Security: Restrict consumer data, protected attributes, and confidential model documentation while allowing authorized review. Proposed validation: Have trained reviewers independently assess the same examples, compare escalation decisions, and verify that cited pages support responses. Measure usability, training gaps, and reviewer agreement locally; the archived pilot had not yet demonstrated those results or regulatory effectiveness.","delivery":"Interpretation — Work and dependencies: Map questions into existing examinations, define evidence submission rules, and pilot bounded use cases. Ownership: Regulatory leads set scope; examiners assess evidence; entity owners explain models and data; security staff protect submissions. Skills and adoption: Provide common training on systems versus models, materiality, limitations, and third-party oversight, then capture reviewer and entity feedback. Governance checkpoints: Approve local adaptations and revisit procedures as the draft changes. Proposed acceptance: Traceable document references, consistent handling of sample risk tiers, recorded reviewer disagreements, and an actionable training-gap log. Risks: Inconsistent thresholds, excessive questionnaire burden, confidential evidence exposure, and premature transfer to non-insurance services can undermine a proportionate process even if every field is completed."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:42:45.193Z","enrichmentBasis":"archived evidence"}}]}