Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the SLED-wide archive edition of August 31, 2026

Standards or public-body guidanceEmergingNew this fortnight

Twelve-state pilot converts AI oversight into a risk-tiered regulatory evidence request

National Association of Insurance Commissioners · State insurance regulation · United States; 12 participating states

Publisher
AI Risk Evaluation Supplement version 5.0 and Pilot Project Summary
Original publication
August 31, 2026
Source retrieved
Not recorded in the historical archive
Read original source

What happened

After field testing with California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin, NAIC exposed version 5.0 of its AI Risk Evaluation Supplement for public comment. The March–September pilot applies the tool in market-conduct reviews, financial analysis, and financial examinations while allowing jurisdiction-specific tailoring.

Why it matters

Although focused on regulated insurers, this is an operational model for state oversight of third-party AI: begin with a materiality-scoped inventory, escalate inquiry for direct consumer or financial impact, and request governance, model, data, limitation, explainability, and third-party evidence proportionate to risk.

Evidence and measured results

NAIC's change summary says pilot feedback led version 5.0 to make the model inventory explicit; distinguish AI systems from AI models; add materiality thresholds, inherent-risk language, direct-consumer and material-financial-impact fields, explainability and transparency questions, document-and-page evidence references, third-party model oversight, model-risk and limitation fields, data-to-model mapping, and an agentic-AI definition. The pilot is designed to assess usability and training needs; final effectiveness results are not yet available.

Limitations and uncertainty

Version 5.0 remains an exposure draft and the 12-state pilot is still underway. Participating jurisdictions may adapt the questions, results have not yet established inter-rater consistency or regulatory effectiveness, and insurance-specific materiality concepts will require translation for other SLED services.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Problem and stakeholders: State insurance regulators, examiners, model-risk specialists, and entity liaisons need proportionate evidence requests for diverse AI uses.

Discovery
Can reviewers distinguish models from systems, identify direct consumer or financial impact, and locate support for questionnaire answers?
Value hypothesis
Materiality-scoped inventory could direct examiner effort toward consequential uses while reducing repetitive requests.
Potential engagement
A bounded examination-workflow pilot with training and evidence-traceability review.
Evidence boundary
Version 5.0 was an exposure draft and the 12-state pilot remained underway in the archive. The record does not establish effectiveness, consistent scoring, or a mandatory requirement everywhere. Applicability beyond insurance remains limited until materiality concepts are translated and tested.

Pre-sales engineering

Role takeaway
Fit
Structure assurance evidence rather than treating the supplement as a model-validation product.
Architecture
Link system, model, use-case, dataset, third-party, limitation, and document-page records, including data-to-model relationships.
Prerequisites
Agreed materiality and inherent-risk definitions, entity evidence, and a common examination rubric.
Constraints
Jurisdictional tailoring and draft revisions can change fields; preserve version provenance.
Security
Restrict consumer data, protected attributes, and confidential model documentation while allowing authorized review.
Proposed validation
Have trained reviewers independently assess the same examples, compare escalation decisions, and verify that cited pages support responses. Measure usability, training gaps, and reviewer agreement locally; the archived pilot had not yet demonstrated those results or regulatory effectiveness.

Delivery

Role takeaway

Work and dependencies: Map questions into existing examinations, define evidence submission rules, and pilot bounded use cases.

Ownership
Regulatory leads set scope; examiners assess evidence; entity owners explain models and data; security staff protect submissions.
Skills and adoption
Provide common training on systems versus models, materiality, limitations, and third-party oversight, then capture reviewer and entity feedback.
Governance checkpoints
Approve local adaptations and revisit procedures as the draft changes.
Proposed acceptance
Traceable document references, consistent handling of sample risk tiers, recorded reviewer disagreements, and an actionable training-gap log.
Risks
Inconsistent thresholds, excessive questionnaire burden, confidential evidence exposure, and premature transfer to non-insurance services can undermine a proportionate process even if every field is completed.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Maintain separate but linked inventories for systems, models, use cases, data sets, and third parties. Use the initial inventory to decide which systems warrant deeper technical and governance evidence, and support different thresholds for back-office, consumer-facing, and financially material workloads.

Governance

Who approves, reviews and stays accountable for outcomes?

Adopt a proportional assurance process that can reuse existing audit and examination mechanisms, identify inherent risk before controls, require traceable evidence rather than unsupported questionnaire answers, coordinate requests across oversight bodies, and refine the instrument from operator and regulated-entity feedback.

Security and privacy

What data, permissions and controls need testing?

Ask how sensitive and protected attributes enter models, how data sets map to model uses, how third-party models are governed, what limitations and failure modes are known, and how confidentiality is preserved when detailed model and control evidence is collected.

The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.

Publication history

  1. 2026-08-31SLED-wide archive · Issue 044 resources
Read preserved resource versions (JSON)

Stable resource ID: naic-ai-risk-evaluation-supplement