From the SLED-wide archive edition of September 3, 2026
National-lab system cuts grid-security data engineering from two months to hours while raising reported accuracy
U.S. Department of Energy CESER and Sandia National Laboratories · Critical infrastructure, public utilities, and cybersecurity · United States
- Publisher
- CESER and Sandia National Lab are Using AI to Safeguard the Electric Grid
- Original publication
- September 3, 2026
- Source retrieved
- Not recorded in the historical archive
What happened
DOE and Sandia report that their C2E2 research pipeline uses LLMs and generative AI to automate the collection, cleaning, and structuring of cyber and physical grid data before a conventional machine-learning model detects and locates threats. The team says the workflow reduced data engineering and model training from about two months to a few hours and increased reported threat-detection accuracy from 85% to 95%.
Why it matters
State and local governments oversee or operate electric, water, transit, emergency, and other cyber-physical systems with small specialist teams. The case shows a high-value role for generative AI upstream of operational decisions: preparing complex data for an established analytic model and producing actionable location information for a trained operator.
Evidence and measured results
The September 3 DOE update provides the 95% accuracy and hours-versus-two-months figures. Sandia's July 30 technical account explains that data engineering consumed about 90% of the prior pipeline's time, describes the LLM-to-clean-dataset-to-ML architecture, and notes that a related paper received an IEEE workshop best-paper award. The reported evaluation remains laboratory research; the team says utility-company testing and systematic hallucination analysis are next steps.
Limitations and uncertainty
The performance figures are project-team reported, and the public sources do not disclose sample size, class balance, confidence intervals, benchmark composition, or independent replication. A 95% aggregate accuracy rate may conceal operationally unacceptable misses. The system has not yet been reported as validated with a production utility, and hallucination behavior remains an acknowledged open question.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Problem and stakeholders: Utility operations, cybersecurity, data-engineering, emergency-coordination, and research teams may spend substantial effort preparing heterogeneous grid data before threat analysis.
- Discovery
- Where is time spent, which misses are consequential, and can experimentation remain separate from control?
- Value hypothesis
- Controlled generative preparation could reduce engineering effort while preserving detection quality.
- Potential engagement
- A laboratory or shadow-mode proof using approved representative data and utility specialists.
- Evidence boundary
- Reported months-to-hours savings and 85%-to-95% accuracy are project-team results, with utility testing and hallucination analysis pending in the archive. They are not production validation, independent replication, or transferable guarantees for electric, water, or other infrastructure, and require local confirmation before commitments.
Pre-sales engineering
Role takeaway
- Fit
- Keep generative AI upstream of specialized detection, producing validated structured data rather than grid commands.
- Architecture
- Separate authenticated ingestion, provenance-preserving preparation, detection/localization, and operator action; insulate operational control from model or cloud availability.
- Prerequisites
- Representative topology and attack data, conventional baseline, specialist reviewers, and false-positive/negative tolerances.
- Constraints
- Aggregate laboratory accuracy omits sample composition and may conceal dangerous misses.
- Security
- Test segmentation, privileges, poisoning, hallucinated fields, artifact integrity, and offline continuity.
- Proposed validation
- Compare preparation effort and severity-weighted detection across topologies and attacks, inspect generated data against inputs, and exercise generative-stage failure. Controlled success is evidence for a next stage, not authority for autonomous response or proof of production readiness.
Delivery
Role takeaway
Work and dependencies: Establish a non-production dataset, baseline preparation and detection, integrate validated outputs, and run shadow evaluation with operators.
- Ownership
- Operations retains control authority; cyber and data teams own threat review and preparation; specialists maintain evaluation; incident owners define escalation.
- Skills and adoption
- Train operators to inspect evidence, recognize unreliable outputs, and continue manually.
- Governance checkpoints
- Review benchmark coverage and residual risk before integration and after topology or model changes.
- Proposed acceptance
- Data provenance, local preparation-time comparisons, acceptable severity-specific misses and false alarms, working fallback, and operator-confirmed interpretation.
- Risks
- Hallucinated structure, unrepresentative data, and connectivity or supplier dependencies can undermine strong aggregate accuracy; public results do not supply universal release thresholds.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Keep the generative component upstream and separable: authenticated telemetry and topology data enter a controlled engineering pipeline, the LLM produces structured data with provenance, a specialized model performs detection and localization, and a human operator receives evidence for action. Hybrid or on-prem deployment may be necessary for operational-technology data, latency, and continuity; cloud connectivity should not become a control-plane dependency for grid operations.
Governance
Who approves, reviews and stays accountable for outcomes?
Define false-negative and false-positive tolerances, test across different grid topologies and attack types, require operator confirmation for consequential response, and rerun evaluation after model or topology changes. Procurement should require access to validation methods, benchmark data characteristics, version history, incident reporting, and safe degradation when the generative stage is unavailable.
Security and privacy
What data, permissions and controls need testing?
Grid topology, device telemetry, vulnerabilities, and response playbooks are highly sensitive. Use segmentation from operational control, least-privilege service identities, signed data and model artifacts, tamper-evident logs, adversarial testing, output validation, and controls against prompt or data poisoning. Do not allow an experimental model to issue autonomous control commands.
The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.
Publication history
- 2026-09-03SLED-wide archive · Issue 076 resources
Stable resource ID: sandia-c2e2-grid-security