Strategic Partners · Issue 05 ·
NVIDIA
Four sources cover September 10 NIM serving results and the Palantir supply-chain collaboration, alongside newly relevant cache-security guidance and independent research. One pattern links capacity acceptance to intended tenant isolation. Vendor claims, numerical inconsistencies and simulation limits remain explicit. No new measured U.S. SLED benefit or independently verified savings is established.
- Evidence records
- 4
- Cross-source patterns
- 1
- Evidence classes
- 2 vendor claim1 independent research1 standards or public-body guidance
- Outcomes
- 2 emerging2 cautionary
- Source freshness
- 2 new this fortnight1 recent1 older, newly relevant
- Research completed
- 2026-09-11
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
Measure capacity under the intended information-sharing policy
The NIM benchmark relies on substantial reuse, while NVIDIA guidance and independent research identify risks when reuse crosses information boundaries. Test performance with the controls intended for deployment. Neither security source demonstrates an exploit in the benchmarked NIM configuration.
Operating questionWhat reuse remains permissible after identity and data boundaries are enforced, and does that configuration still meet service targets?
Supporting evidenceNIM Ultra benchmark ties capacity gains to a cache-heavy workloadTejasvi C. AddagadaNVIDIA cache guidance puts prompt assembly and tenant boundaries in the design review
Full record · every source keeps its link and limitations
Evidence ledger
NIM Ultra benchmark ties capacity gains to a cache-heavy workload
NVIDIA reports improved Nemotron 3 Ultra serving throughput from a bundled NIM optimization stack.
Why it matters, evidence and limitations
- Why it matters
- Relevant to government and education teams assessing shared copilots, coding assistants and document workflows. No measured SLED benefit or additional stream tag is asserted.
- Evidence and measured results
- On four B200 GPUs, Table 1 reports 718 versus 1,997 output tokens/second for baseline versus NIM 2.0.12 at 50 tokens/second/user. Workload notation is 64K/400 with 76% KV reuse. These table values imply about 2.78x, whereas the headline says 2.5x. No independent replication, request sample size or variability estimate is supplied.
- Limitations and uncertainty
- The optimizations interact; individual contributions cannot be added. The source does not establish accuracy, institutional productivity or a transferable capacity multiplier. The numerical discrepancy remains unresolved.
Cache-isolation preprint separates hardware timing evidence from simulated defenses
The preprint measures a cache timing distinction and proposes principal-specific isolation; defense effectiveness is not established in production.
Why it matters, evidence and limitations
- Why it matters
- Relevant to government and education teams assessing shared copilots, coding assistants and document workflows. No measured SLED benefit or additional stream tag is asserted.
- Evidence and measured results
- Table 8 reports cold/cached latencies of 149.6/32.8 ms for a 2,119-token prefix on Qwen2.5-7B, vLLM 0.26.0 and A100, with 50 observations per arm across two blocks. Its stated 0.22 ratio is cached divided by cold, despite reversed wording. Defense experiments use 1,000 simulated trials; extreme success-rate columns are analytic controls.
- Limitations and uncertainty
- No NIM or Nemotron test. Boundary-salting efficiency is extrapolated, semantic-cache isolation unmeasured, and field adversarial testing remains future work. Table 4 and prose disagree on noise results; the load adversary differs from the theorem's payoff. Those numerical claims are excluded.
NVIDIA cache guidance puts prompt assembly and tenant boundaries in the design review
NVIDIA explains how shared prefix caching may disclose information through timing, including context added by applications.
Why it matters, evidence and limitations
- Why it matters
- Relevant to government and education teams assessing shared copilots, coding assistants and document workflows. No measured SLED benefit or additional stream tag is asserted.
- Evidence and measured results
- The guidance discusses prompt ordering, tenant isolation and suspicious-query monitoring. It acknowledges network, batching and tool-call noise. This is explanatory guidance without a measured deployment sample, control-effectiveness estimate or benefit baseline.
- Limitations and uncertainty
- Older guidance newly relevant to the current reuse-heavy benchmark. No specific vulnerability in NIM 2.0.12 is demonstrated, and mitigation suggestions are not a security certification.
NVIDIA and Palantir announce a supply-chain AI stack starting in NVIDIA operations
The companies announce integration of Nemotron with Foundry and AIP, starting in NVIDIA's own supply chain.
Why it matters, evidence and limitations
- Why it matters
- Strategic-partner ecosystem relevance is direct. Public-sector procurement and facilities teams may examine analogous workflows, but no SLED deployment or transferable outcome is established.
- Evidence and measured results
- The announcement describes cuOpt scenario planning, NeMo data and training libraries, human final decisions, and cloud or on-premises reference-architecture options. It supplies no measured before/after outcome, evaluation sample, cost baseline or independent confirmation.
- Limitations and uncertainty
- Deployment and future benefits are vendor/operator claims. Announced collaboration date is known; actual deployment start date is not. Sovereign branding does not establish a customer's compliance or control effectiveness.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Vendor claim
- A supplier-provided assertion that has not been upgraded to independent evidence.
- Independent research
- Research conducted outside the implementing organization.
- Standards or public-body guidance
- Normative or advisory guidance from a standards body or public institution.