{"resourceId":"planninglens-ai-planning-baseline-2026","versions":[{"version":"external-22d54ed05e636be4bb0be67f1d24f31ae196c27b747fa9efa6b1135d1a25bed3","resource":{"id":"planninglens-ai-planning-baseline-2026","title":"Independent planning scorecard establishes a baseline without claiming an AI effect","organization":"PlanningLens Ltd","sector":"Municipal planning evaluation","geography":"England; evaluation design has limited transfer to U.S. permitting","publishedAt":"August 3, 2026; data snapshot July 22, 2026","publicationDate":"2026-08-03","eventDate":"2026-07-22","sourceName":"PlanningLens","sourceLabel":"Commercial planning-data company's original independent analysis; not an official audit","sourceUrl":"https://planninglens.co.uk/ai-planning-scorecard.html","evidenceClass":"independent-research","outcomeClass":"emerging","topics":["knowledge-work","governance-procurement","operating-model"],"finding":"PlanningLens publishes a pre-trial decision-time baseline and comparator panels; it explicitly declines to attribute early changes to AI.","sledRelevance":"Interpretation: Helps localities define evidence needed before accepting throughput claims. Different planning laws and case mixes preclude importing the English baseline.","evidence":"The analysis reports 7,663 pilot-council decisions, a pooled median of 7.71 weeks, and 16 control councils. Baseline window: May 2024–April 2026. Timing runs from validation to decision, ignores deadline extensions and excludes cases over 364 days. May–July observations are explicitly noncausal.","architectureImplications":"Interpretation: Capture stage timestamps and model exposure separately; elapsed-time records alone cannot establish that a case used an assistant.","governanceImplications":"Interpretation: Predefine comparison rules and publish methodological changes before interpreting results.","securityPrivacyImplications":"Interpretation: Use minimal case identifiers for reproducibility and suppress personal details in public evaluation outputs.","caveats":"Commercial publisher, not part of the trial. Raw decision rows were not independently recomputed. Incomplete recent feeds, imperfect Camden matching and excluded long cases limit inference. Baseline is not an effectiveness verdict.","streamIds":["local-government"],"roles":{"sales":"Interpretation: Ask a permitting director, performance analyst and procurement officer what evidence would justify scaling an assistant. Clarify whether the problem is staff workload, applicant waiting time or missed deadlines, since these require different measures. Offer a bounded baseline-and-evaluation design using the customer's records. The value hypothesis is a better informed purchase or continuation decision, not immediate productivity. Do not use the 7.71-week English baseline as a customer benchmark. Confirm that case definitions and timestamps are reliable enough for comparison, and that someone independent of the supplier can challenge exclusions, missing records and an apparently favourable result.","engineering":"Interpretation: Assemble a versioned dataset linking case stage, application class, reviewer and actual AI exposure. Retain missingness flags and capture changes in staffing or policy. Prerequisites include stable identifiers, authorized records access and a pre-agreed comparison design. Test timestamp semantics and matching sensitivity before generating dashboards. Proposed validation should reproduce a sampled record's elapsed time from the source and show results with and without long-case exclusions. Treat a matched panel as observational evidence with residual confounding. The scorecard does not validate an agent, model or deployment architecture; any application integration requires its own security and reliability testing.","delivery":"Interpretation: A municipal performance lead should own the evaluation with planning staff and data engineering support. Freeze the initial protocol, record exclusions and establish a schedule for data refresh and discrepancy resolution. Dependencies include access to complete records and visibility into operational changes. Train stakeholders to distinguish component efficiency, resident waiting time and causal attribution. Proposed acceptance requires reproducible sampled calculations, visible data gaps and documented review of comparator suitability before any benefits claim. These are proposed controls, not measured improvements. Risks include selective follow-up, changing denominators and attributing concurrent process reforms to AI. Keep unresolved uncertainty in the published findings."},"retrievedAt":"2026-09-09T03:01:08Z","enrichedAt":"2026-09-09T03:04:04Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Add applicant experience and staff rework to timing data; neither is established by portal timestamps.","procurementImplications":"Interpretation: Require independently reproducible evaluation definitions, not just a supplier's chosen success metric.","operatingModelImplications":"Interpretation: Assign an analyst outside the deployment team to maintain comparators and record service changes.","updateExplanation":"New-to-archive August baseline; contributes an independent evaluation design to planning coverage. No September outcome or post-trial causal result claimed.","sourceVerification":{"openedUrl":"https://planninglens.co.uk/ai-planning-scorecard.html","referenceExcerpt":"This is a baseline, not a verdict.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}