{"resourceId":"fourcastnet-gray-swan-controlled-preprint-2024","versions":[{"version":"external-8da1a41cfe7a957a8928306158159666b436e98024a6aacb0744a10fe3588d79","resource":{"id":"fourcastnet-gray-swan-controlled-preprint-2024","title":"Controlled cyclone study exposes extreme-event failures hidden by global forecast scores","organization":"University of Chicago, University of California Santa Cruz and New York University researchers","sector":"Disaster forecasting evaluation","geography":"Global tropical cyclones; U.S. academic research","publishedAt":"October 19, 2024; inspected preprint version 1","publicationDate":"2024-10-19","eventDate":null,"sourceName":"arXiv:2410.14932v1","sourceLabel":"Original controlled research preprint","sourceUrl":"https://arxiv.org/html/2410.14932v1","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["infrastructure","governance-procurement","operating-model"],"finding":"A controlled experiment finds that ordinary forecast scores can conceal failure on stronger storms excluded from training.","sledRelevance":"Interpretation: Relevant to emergency managers evaluating the evidence behind hazard guidance, not proof that every current model fails on every extreme.","evidence":"The authors train 25 FourCastNet realizations across five dataset conditions, including full-data and matched-size random-removal controls. Training uses ERA5 1979–2015; testing covers 20 pressure-defined intense cyclones from 2018–2023. Removing strong tropical storms globally produces poor extreme forecasts despite similar global scores; basin-specific removal permits some transfer. Category labels are ERA5 pressure proxies, not literal observed wind categories.","architectureImplications":"Interpretation: Preserve training-domain and model-version documentation and evaluate event-conditioned errors alongside aggregate scores. Infrastructure planning should include the cost of independent validation.","governanceImplications":"Interpretation: Require an explicit extreme-event evidence review before expanding the decisions supported by a forecasting service.","securityPrivacyImplications":"Interpretation: Scientific weather inputs have limited personal-data relevance; protect data lineage, model artifacts and test-set integrity.","caveats":"One architecture and one hazard family in a controlled reanalysis setting. The inspected preprint is not the later PNAS text; journal and repository PDF access failed. Results do not directly test AIFS-TC, contemporary products or evacuation outcomes.","streamIds":["emergency-services"],"roles":{"sales":"Interpretation — Ask emergency managers and procurement evaluators how a forecast vendor supports claims about rare, consequential events. Include meteorological advisers, risk managers and continuity planners in discovery. A bounded engagement could review the evidence package and design an extreme-event challenge set. The value hypothesis is better identification of unsupported product scope before operational reliance. Avoid implying that this study disqualifies all AI forecasts or proves a particular vendor unsafe. Determine whether the agency needs independent technical support from a specialist forecasting institution. This is evaluation work, not a new warning product, and its usefulness depends on access to credible model and test documentation.","engineering":"Interpretation — Build a reproducible benchmark that separates ordinary conditions, intense storms and events poorly represented in training. Prerequisites include a pinned model version, appropriate reference data, consistent initialization rules and a specialist-approved definition of extremity. Compare with operational guidance and account for missing forecasts rather than silently excluding them. Test provenance and artifact integrity to prevent contamination of the evaluation. Where full retraining is impractical, document the narrower claim supported by available tests. A proof of value should report subgroup errors, false reassurance and uncertainty calibration. Neither a large ensemble nor a favorable global average should substitute for the required extreme-event assessment.","delivery":"Interpretation — A meteorological evaluation lead should own the review with emergency planners, data engineers and procurement staff. Inventory vendor claims, identify unsupported hazard regimes, obtain suitable reference cases and commission independent testing where necessary. Dependencies include access to model outputs and sufficient specialist time. Gate procurement and subsequent model upgrades on a documented evidence review. Proposed acceptance includes complete accounting of selected events, reproducible scoring, explicit unresolved extreme-event limits and an exercised fallback to established forecast channels. Train planners to ask what the model has actually been tested on. Results should narrow permitted uses when needed; the study does not supply a universal performance threshold or observed local safety benefit."},"retrievedAt":"2026-09-10T03:02:35Z","enrichedAt":"2026-09-10T03:05:03Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Train planners to distinguish aggregate accuracy from critical-event reliability; specialist review capacity remains necessary.","procurementImplications":"Interpretation: Require hazard-specific test evidence and explicit exclusions, with renewed evaluation after material model changes.","operatingModelImplications":"Interpretation: Forecast specialists evaluate technical claims and emergency leadership limits the authorized use of uncertain products.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2410.14932v1","referenceExcerpt":"Common metrics obscure poor performance for gray swans","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}