From the NVIDIA edition of September 10, 2026
Cache-isolation preprint separates hardware timing evidence from simulated defenses
Tejasvi C. Addagada · AI inference platforms and shared-service assurance · Global technical applicability; no U.S. SLED field evaluation
- Publisher
- arXiv
- Original publication
- 2026-08-10
- Source retrieved
- 2026-09-11
What happened
The preprint measures a cache timing distinction and proposes principal-specific isolation; defense effectiveness is not established in production.
Why it matters
Relevant to government and education teams assessing shared copilots, coding assistants and document workflows. No measured SLED benefit or additional stream tag is asserted.
Evidence and measured results
Table 8 reports cold/cached latencies of 149.6/32.8 ms for a 2,119-token prefix on Qwen2.5-7B, vLLM 0.26.0 and A100, with 50 observations per arm across two blocks. Its stated 0.22 ratio is cached divided by cold, despite reversed wording. Defense experiments use 1,000 simulated trials; extreme success-rate columns are analytic controls.
Limitations and uncertainty
No NIM or Nemotron test. Boundary-salting efficiency is extrapolated, semantic-cache isolation unmeasured, and field adversarial testing remains future work. Table 4 and prose disagree on noise results; the load adversary differs from the theorem's payoff. Those numerical claims are excluded.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-11; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Discuss shared assistant confidentiality with security, data stewards and the service owner. Ask which groups share infrastructure and which documents enter prompts. Offer a bounded architecture and assurance review using synthetic content. The value hypothesis is identifying an unexamined information boundary before wider adoption. This paper can motivate questions but cannot quantify a customer's likelihood of compromise, demonstrate an institutional incident or certify a defense. Request production evidence from suppliers and avoid presenting simulation outputs as observed breach reduction.
Pre-sales engineering
Role takeaway
Map authenticated users to every cache and retrieval store in the proposed service. Require synthetic fixtures, isolated test identities and explicit permission boundaries. Compare allowed reuse within a principal with prohibited reuse across principals, recording server and client telemetry under realistic load. Inspect semantic retrieval separately from exact-match caching. Validate the actual engine and release instead of importing a paper's configuration. A useful proof of value demonstrates traceable identity enforcement and measures the resulting capacity cost; it does not establish universal protection against all side channels.
Delivery
Role takeaway
Security engineering should own the threat model with platform operators and application maintainers. Implement an inventory of shared state, change approval and incident escalation. Dependencies include trustworthy identity propagation, test capacity and staff able to interpret timing measurements. Train maintainers to revisit the boundary when routes or storage change. Proposed acceptance requires documented ownership for every shared store, reproducible tests of allowed and denied reuse, and a demonstrated containment procedure. Review before sensitive records enter the service. Risks include incomplete identity mapping and mistaking laboratory controls for operational assurance.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Validate the complete serving path and intended tenancy model before selecting placement or capacity.
Governance
Who approves, reviews and stays accountable for outcomes?
Assign separate approval owners for performance, answer quality and information boundaries.
Security and privacy
What data, permissions and controls need testing?
Protect prompts, retrieved records and telemetry; test authorization beyond the front-end login.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Include assistive-technology users in workflow validation and budget operator training; the source measures no accessibility outcome.
Procurement
What should contracts, pricing and exit terms secure?
Require workload-specific evidence, support obligations and recurring-cost assumptions.
Operating model
Which teams own the service once it runs?
Retain an accountable service owner, maintained test corpus and change-triggered revalidation.
What changed
New to the archive, not newly published today. Selected as scrutiny and implementation context for the September 10 cache-heavy serving benchmark; identifier and related-topic archive searches found no matching record.
Publication history
- 2026-09-10NVIDIA · Issue 054 resources
Stable resource ID: kvgov-cache-isolation-study-260809225-v1