From the SLED-wide archive edition of August 31, 2026
Twelve-state pilot converts AI oversight into a risk-tiered regulatory evidence request
National Association of Insurance Commissioners · State insurance regulation · United States; 12 participating states
- Publisher
- AI Risk Evaluation Supplement version 5.0 and Pilot Project Summary
- Original publication
- August 31, 2026
- Source retrieved
- Not recorded in the historical archive
What happened
After field testing with California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin, NAIC exposed version 5.0 of its AI Risk Evaluation Supplement for public comment. The March–September pilot applies the tool in market-conduct reviews, financial analysis, and financial examinations while allowing jurisdiction-specific tailoring.
Why it matters
Although focused on regulated insurers, this is an operational model for state oversight of third-party AI: begin with a materiality-scoped inventory, escalate inquiry for direct consumer or financial impact, and request governance, model, data, limitation, explainability, and third-party evidence proportionate to risk.
Evidence and measured results
NAIC's change summary says pilot feedback led version 5.0 to make the model inventory explicit; distinguish AI systems from AI models; add materiality thresholds, inherent-risk language, direct-consumer and material-financial-impact fields, explainability and transparency questions, document-and-page evidence references, third-party model oversight, model-risk and limitation fields, data-to-model mapping, and an agentic-AI definition. The pilot is designed to assess usability and training needs; final effectiveness results are not yet available.
Limitations and uncertainty
Version 5.0 remains an exposure draft and the 12-state pilot is still underway. Participating jurisdictions may adapt the questions, results have not yet established inter-rater consistency or regulatory effectiveness, and insurance-specific materiality concepts will require translation for other SLED services.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Problem and stakeholders: State insurance regulators, examiners, model-risk specialists, and entity liaisons need proportionate evidence requests for diverse AI uses.
- Discovery
- Can reviewers distinguish models from systems, identify direct consumer or financial impact, and locate support for questionnaire answers?
- Value hypothesis
- Materiality-scoped inventory could direct examiner effort toward consequential uses while reducing repetitive requests.
- Potential engagement
- A bounded examination-workflow pilot with training and evidence-traceability review.
- Evidence boundary
- Version 5.0 was an exposure draft and the 12-state pilot remained underway in the archive. The record does not establish effectiveness, consistent scoring, or a mandatory requirement everywhere. Applicability beyond insurance remains limited until materiality concepts are translated and tested.
Pre-sales engineering
Role takeaway
- Fit
- Structure assurance evidence rather than treating the supplement as a model-validation product.
- Architecture
- Link system, model, use-case, dataset, third-party, limitation, and document-page records, including data-to-model relationships.
- Prerequisites
- Agreed materiality and inherent-risk definitions, entity evidence, and a common examination rubric.
- Constraints
- Jurisdictional tailoring and draft revisions can change fields; preserve version provenance.
- Security
- Restrict consumer data, protected attributes, and confidential model documentation while allowing authorized review.
- Proposed validation
- Have trained reviewers independently assess the same examples, compare escalation decisions, and verify that cited pages support responses. Measure usability, training gaps, and reviewer agreement locally; the archived pilot had not yet demonstrated those results or regulatory effectiveness.
Delivery
Role takeaway
Work and dependencies: Map questions into existing examinations, define evidence submission rules, and pilot bounded use cases.
- Ownership
- Regulatory leads set scope; examiners assess evidence; entity owners explain models and data; security staff protect submissions.
- Skills and adoption
- Provide common training on systems versus models, materiality, limitations, and third-party oversight, then capture reviewer and entity feedback.
- Governance checkpoints
- Approve local adaptations and revisit procedures as the draft changes.
- Proposed acceptance
- Traceable document references, consistent handling of sample risk tiers, recorded reviewer disagreements, and an actionable training-gap log.
- Risks
- Inconsistent thresholds, excessive questionnaire burden, confidential evidence exposure, and premature transfer to non-insurance services can undermine a proportionate process even if every field is completed.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Maintain separate but linked inventories for systems, models, use cases, data sets, and third parties. Use the initial inventory to decide which systems warrant deeper technical and governance evidence, and support different thresholds for back-office, consumer-facing, and financially material workloads.
Governance
Who approves, reviews and stays accountable for outcomes?
Adopt a proportional assurance process that can reuse existing audit and examination mechanisms, identify inherent risk before controls, require traceable evidence rather than unsupported questionnaire answers, coordinate requests across oversight bodies, and refine the instrument from operator and regulated-entity feedback.
Security and privacy
What data, permissions and controls need testing?
Ask how sensitive and protected attributes enter models, how data sets map to model uses, how third-party models are governed, what limitations and failure modes are known, and how confidentiality is preserved when detailed model and control evidence is collected.
The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.
Publication history
- 2026-08-31SLED-wide archive · Issue 044 resources
Stable resource ID: naic-ai-risk-evaluation-supplement