{"resourceId":"voice-prehospital-stroke-simulation-2025","versions":[{"version":"external-8c4fe0edad458d732fda1b52f15c0e7bb2b3e58dab0b9377ccef87cd89eeb6d8","resource":{"id":"voice-prehospital-stroke-simulation-2025","title":"VOICE stroke-assessment simulation exposes scoring and medication-recording failures","organization":"Harvard Medical School and collaborating researchers","sector":"Prehospital EMS assessment","geography":"U.S.-led research with international collaborators; simulated setting","publishedAt":"June 25, 2025 (arXiv submission history)","publicationDate":"2025-06-25","eventDate":null,"sourceName":"arXiv","sourceLabel":"Original feasibility preprint, version 1","sourceUrl":"https://arxiv.org/html/2507.22898v1","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"VOICE demonstrated a conversational assessment workflow but made consequential errors in simulated cases.","sledRelevance":"Interpretation: Useful for designing EMS copilot evaluations; it does not justify autonomous clinical triage.","evidence":"Three lay users assessed ten actor scenarios against neurologist-defined ground truth. Component scoring was correct in 42/50 observations; total scores matched in 5/10 cases. The system detected 6/7 strokes, flagged 2/3 mimics and hallucinated an anticoagulant. Mean completion was 6 minutes 15 seconds without a timed usual-care comparator. One reviewing physician rated decision confidence at least 3/5 in 4/10 cases.","architectureImplications":"Source: Azure-hosted voice agents shared assessment state and generated reports with video links. Interpretation: Pin components and validate every state update; historical API details are not current deployment instructions. Compare cloud latency with an explicit local fallback.","governanceImplications":"Interpretation: Medical directors should authorize intended use, adjudicate errors and gate model changes. A second model's review should not count as independent clinical sign-off.","securityPrivacyImplications":"Interpretation: Restrict audio, video and clinical-state access; define retention, consent or other lawful processing, vendor boundaries and deletion verification before using patient data.","caveats":"Preprint; simulated, small sample, no paramedic users or real patients, no power calculation. Local noise, language and time-pressure performance is unknown. Repository history gives June 25 despite the July-style identifier.","streamIds":["emergency-services"],"roles":{"sales":"Interpretation — EMS chiefs, medical directors and receiving stroke teams need dependable assessment handoffs under time pressure. Ask where current information is lost, how often staff correct records and whether reviewers can access original observations. A bounded engagement could map the handoff and test a voice copilot in an approved simulation lab. The value hypothesis is more complete, reviewable information with acceptable correction effort. This paper supplies failure cases for discovery, not a business case for replacing clinicians. Establish the agency's evaluation capacity and intended users before proposing a pilot. Do not promise improved survival, faster treatment or the same performance with contemporary models.","engineering":"Interpretation — Prototype a separate assessment sandbox with structured state, a reviewable transcript and explicit attribution of user observations versus generated conclusions. Prerequisites include clinician-approved scenarios, reference labels, secure recording storage and a pinned model configuration. Validate omitted questions, contradictory observations and medication ambiguity before testing broader case mixes. Keep generated scores advisory and prevent automatic destination changes. Test dropped connections, slow responses and interrupted sessions without delaying the established emergency path. A proof of value should compare record correctness and total reviewer effort against the current process, including all unsuccessful sessions. Cloud, hybrid and local designs need their own measured availability and privacy assessment.","delivery":"Interpretation — Assign an EMS medical director as clinical owner and pair clinical educators with application and security engineers. Map the existing stroke handoff, create consented simulation cases, train evaluators and run blinded record adjudication. Dependencies include reviewers, representative callers and an approved media-retention process. Gate any live study on documented clinical, privacy and workflow approval. Proposed acceptance requires every test case to have a traceable transcript, adjudicated component scores, recorded corrections and a demonstrated fallback; clinical error thresholds must be set prospectively by medical leadership. Measure staff confidence separately from correctness. Stop expansion if omissions, fabricated facts or review delays make the workflow unsafe."},"retrievedAt":"2026-09-14T03:01:23Z","enrichedAt":"2026-09-14T03:05:02Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Test voice interactions with diverse speech, hearing and motor needs. Measure correction burden and preserve an accessible human path.","procurementImplications":"Interpretation: Buy a bounded evaluation before operational reliance; require model-version notice, data-use terms and a demonstrable exit path.","operatingModelImplications":"Interpretation: Clinical quality owns decision validity, dispatch leadership owns workflow, and IT owns availability; none should infer approval from fluent output.","updateExplanation":"Newly catalogued historical evidence. Full-archive identifier and related-term searches found no canonical match. Adds a specific voice-agent failure analysis to existing EMS simulation coverage; no new release is claimed.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2507.22898v1","referenceExcerpt":"The total AI-guided FAST-ED score matched the ground truth score in 5 out of 10 cases (50%).","promptVersion":"sled-research-v3.2","model":null,"basis":"agent-reported inspection"}}}]}