Model Register

Last updated: 2026-09-20

19 models we built, 9 external tools we use. Each one states what it does, what it was measured on, and the limit past which we do not let it speak. 12 of the 19 answer a request in the product today, read from our serving register rather than claimed here.

Models we built

Trained or engineered by us, on our own data and our own compute.

Cargo underwriting model

Shipping safety
Answers requests today

Reads a consignment's chemistry, charge, packaging, route and paperwork, and proposes accept, refer or exclude with the facts that drove it.

Trained or measured on: 1,000,000 example shipments, 800,000 for training and 200,000 held out. Held-out macro AUC 0.973; agreement with the deterministic rules engine 0.991.

Where we stop it

Demonstration only, and the limit is the training data rather than the engineering: trained on simulated shipments rather than claims history, it reproduces our own rules engine 99.1% of the time. It is shown beside the score, never as the score — the rules engine still decides and the real-label severity model still handles consequence. It is not live underwriting and will not be on this training set. What would change that is claims history with observed labels.

Evidence: nvidia-lab/reports/cargo-underwriting-proxy-validation-latest.json · nvidia-lab/serving/registry.json

Incident severity (United States)

Shipping safety
Answers requests today

Given a reported hazardous-materials incident, estimates the chance it meets the US regulator's serious-incident test. How bad if it happens, never how likely.

Trained or measured on: 122,229 deduplicated US DOT Form 5800.1 hazardous-materials incident reports from 2020 to 2024, labelled by the regulator's own serious-incident determination rather than by anything this platform scored. Held out by event date, most recent quarter reserved; held-out ROC AUC 0.9491.

Where we stop it

The US head is the one we serve; the UK and Canadian heads are fitted but held back, and the three are never combined because each regulator defines severity differently.

Evidence: nvidia-lab/reports/incident-severity-validation-latest.json · nvidia-lab/serving/registry.json

Material property surrogate

Battery materials
Answers requests today

Predicts cycle life, energy density and capacity fade straight from a cell recipe, fast enough to sweep thousands of candidates before committing to a physics run.

Trained or measured on: 54,896 usable rows of a 56,700-point electrochemical (DFN) solver sweep. Held-out r-squared on the log scale: 0.965 for cycle life, 0.810 for energy density, 0.965 for capacity fade.77.2% of those rows carry a published literature constant for cycle life rather than a simulated ageing run, because six of the eight chemistries have no ageing parameters.

Where we stop it

Energy density and capacity fade are solid. Cycle life is not: it may not be used to rank recipes — a rank correlation of 0.11 between solver depths — and below 25 °C it is withheld.

Evidence: nvidia-lab/reports/material-surrogate-validation-latest.json · docs/CYCLE_LIFE_IMPROVEMENT_2026-08-07.md · nvidia-lab/serving/registry.json

Cell discharge-curve surrogate

Battery materials
Answers requests today

Predicts the whole shape of a cell's discharge curve in milliseconds, where the physics solve it stands in for takes minutes.

Trained or measured on: 43,916 training curves and 10,980 held-out curves recorded from the same physics sweep, held out by chemistry. Held-out curve error 48.6 mV root-mean-square, 22.5 mV mean absolute, 962 mV at the single worst point. The trained operator is 547,809 parameters.

Where we stop it

Serves the discharge curve in-process as ONNX from 2026-08-09, at 1.7 ms p95 on ordinary hardware. It is scoped: LCO, LFP, NCA, NMC532 and NMC622 only, at screening grade, with NMC532 and NMC622 carrying a nearest-set proxy caveat. NMC811, Sodium-ion and Solid-state are refused with a stated reason and no curve is returned. What used to stop it was a framework needing a newer Python — never the hardware, which had already been measured at 583 ms against a 200 ms target on CPU and 54 ms on a GPU. Exporting the graph removed the framework from the serving path entirely.

Evidence: nvidia-lab/reports/material-curve-surrogate-validation-latest.json · nvidia-lab/reports/serving-spine-benchmark-latest.json · nvidia-lab/serving/registry.json

Extreme-event exposure model

Site intelligence
Answers requests today

Scores a solar or battery site across seven perils — hail, wind, heat, outage, wildfire smoke, flood and ice — and attaches the operating action each implies.

Trained or measured on: Generated weather, established forensically on 2026-08-11 from the StandardScaler rather than from any record: ten of twelve features match textbook distributions with round parameters (month Uniform(1,12), hour Uniform(0,23), pressure Normal(1013,15), and mean == sd on aod_index, precip_mm and cape_jkg). Reproduce with scripts/audit/scaler-forensics.mjs.

Where we stop it

Illustrative, not a hazard estimate. It was not that the training set was lost — there was never an observed one, so the probabilities and impact scores show the shape of the output and say nothing about hazard at a location. The caveat ships on the API response as mustBeShownWith. Replacing it needs observed weather and observed events, which the platform does not yet hold.

Evidence: docs/MODEL_OPPORTUNITY_ANALYSIS_2026-08-07.md · nvidia-lab/serving/registry.json

Site ranker

Site intelligence
Answers requests today

Ranks candidate battery sites against land, demand and grid-connection evidence, so a developer knows early which are worth a full study.

Trained or measured on: Not a learned model — a deterministic spatial ranker. Validated on 17 July 2026 against a 10 km workload around Glasgow Airport of 441 land polygons, 390 demand signals and 299 substations: all 30 candidates from the accelerated path matched the reference implementation exactly. It is that reference implementation, in TypeScript, that answers a request here — held bit-identical to the Python ranker of record by a committed parity test that runs both on every CI pass.

Where we stop it

Screening evidence from public land and grid data — not a connection offer, a planning view or a land right. The accelerated cupy implementation is a separate entry in the serving register and is not reachable in any deployment here; the parity result says the two agree, not that the GPU one is running.

Evidence: nvidia-lab/reports/market-siting-h100-glasgow-20260717.json · apps/dashboard/app/bess-location-intelligence/market-siting-runtime-parity.test.ts · nvidia-lab/reports/inprocess-reader-latency-latest.json · nvidia-lab/serving/registry.json

Portfolio catastrophe model

Shipping safety
Evidence record, not a product surface

Simulates ten million years of battery-fleet losses to show how bad a bad year gets, and how often.

Trained or measured on: Not a learned model — a Monte Carlo engine over frequency and severity priors drawn from public incident data. Ten million simulated years across a 200-site book in three geographies, with probable maximum loss reported per cohort.

Where we stop it

The 200-site book is synthetic and is nobody's real exposure, so this is a methodology demonstration and not for pricing, capacity or regulatory use.

Evidence: nvidia-lab/reports/cat-portfolio-mc-latest.json

Cell ageing fits

Battery materials
Answers requests today

Fits fade curves to real laboratory cells, anchoring the platform's ageing assumptions to measured hardware rather than a rule of thumb, and reads back a state of health at a given cycle count.

Trained or measured on: 32 cells from NASA's public battery ageing datasets. A power law was the best-fitting shape for 27 of the 32. Data: NASA Ames Prognostics Center of Excellence battery ageing repository.

Where we stop it

12 of the 32 cells clear our quality floor, and the datasets never record their cathode chemistry — so no reading may be chemistry-matched and a request naming one is refused rather than answered. Beyond the cycle window a cell was actually discharged over (the longest ran 168 cycles) there is no figure, because extrapolating a fade curve past its evidence would turn a fit into a forecast.

Evidence: nvidia-lab/cell-ageing/artifacts/nasa-pcoe-ageing-fits-latest.json · nvidia-lab/cell-ageing/fit_ageing_models.py · apps/dashboard/app/api/models/cell-ageing/reader.ts · docs/CYCLE_LIFE_IMPROVEMENT_2026-08-07.md · nvidia-lab/serving/registry.json

Cell and pack simulator

Battery materials
Answers requests today

Runs the simplified battery physics behind the cell and pack experiment surfaces.

Trained or measured on: Nothing. It is hand-written physics, not a learned model.

Where we stop it

A fast proxy for the experiment surfaces, not a full electrochemical solve. Five chemistries hold parameters; anything else runs the NMC811-Graphite set, and every response says whether the requested chemistry was matched rather than letting its name stand over another chemistry's numbers.

Evidence: apps/dashboard/app/api/battmo-simulate/simulate-core.ts · apps/dashboard/app/api/battmo-simulate/route.ts · nvidia-lab/serving/registry.json

Terrain constraint sweep

Site intelligence
Answers requests today

Excludes candidate land parcels on slope using 30 m public elevation tiles, as an early screening filter. It answers for a parcel the sweep analysed and returns a stated absence for anywhere it did not — the register of parcels is not a continuous surface and nothing is interpolated between them.

Trained or measured on: Not a learned model. 3,152 of 3,153 declared British parcels analysed against 20 Copernicus GLO-30 elevation tiles; the one parcel that was skipped did not overlap the elevation data and is named in the record.

Where we stop it

Slope screening over public land-use data — a first filter, not a site assessment. What answers a request is the committed 2026-08-06 sweep, read back; the elevation tiles are not in this repository so no slope is recomputed for a location that was never swept, and a point more than 500 m from an analysed parcel gets the reason rather than a number.

Evidence: nvidia-lab/reports/gb-constraint-sweep-latest.json · nvidia-lab/terrain/run_gb_constraint_sweep.py · apps/dashboard/app/api/models/terrain-constraint/reader.ts · nvidia-lab/serving/registry.json

GB consenting outcome and duration

Siting & planning
Answers requests today

Given a candidate GB site, estimates whether a planning application is granted and how long the decision takes, conditioned on the planning authority, technology, capacity, co-location and whether it is a re-application.

Trained or measured on: 11,146 decided GB planning applications from the DESNZ planning register, across fourteen dated extracts spanning 2019 to 2026. The outcomes are the regulator's, not ours.

Where we stop it

It answers what happens to applications that reach a decision, so it does not yet speak for the 831 that are still open — every figure it serves is conditional on a decision being reached, and a request about an open application is refused rather than answered. Validated on a time-split, and the duration figure carries the report's own caveat that each split selects on the outcome in an opposite direction, so neither number is a clean accuracy.

Evidence: nvidia-lab/consenting/fit_consenting.py · nvidia-lab/consenting/export_consenting_serving.py · nvidia-lab/reports/consenting-validation-latest.json · apps/dashboard/app/api/models/consenting/reader.ts · docs/REPD_CENSORING_FINDING_2026-08-07.md · nvidia-lab/serving/registry.json

BESS dispatch optimiser

BESS dispatch
Answers requests today

Given a battery's rating, usable capacity and spot/ancillary prices, proposes a charge/discharge schedule across the horizon and prices it: arbitrage, ancillary and contract revenue against degradation and grid-charge cost. GB requests resolve against real Elexon day-ahead prices and real NESO ancillary clearing prices rather than a synthetic curve.

Trained or measured on: Not trained: a hand-specified greedy price-pairing rule (charge the cheapest hours, discharge the costliest, respecting SoC and cycle limits), not a fitted model. Deterministic and reproducible — the same request returns the same schedule every time — and measured in-process at p50 0.0749 ms / p95 0.121 ms / p99 0.184 ms, 4,000 calls after 500 warm-up.

Where we stop it

A stated greedy simplification, not a full LP/MILP solver — a real GPU LP path (NVIDIA cuOpt) exists but measured 9.8x SLOWER than this CPU heuristic on an identical 90-cell portfolio problem, so it stays unused rather than wired in for its own sake. Round-trip efficiency and the degradation cost per MWh remain flat, stated approximations pending a citable figure; a multi-asset portfolio's own price cannibalisation is not modelled.

Evidence: apps/dashboard/app/api/bess/dispatch-optimise/heuristic-dispatch.ts · apps/dashboard/app/api/bess/dispatch-optimise/portfolio-dispatch.ts · nvidia-lab/serving/registry.json

Planning-document constraint extraction

Due diligence
Answers requests today

Reads an uploaded environmental or permit document (flood, noise, landscape, ecology or heritage report) and returns a headline severity with the verbatim sentence it rests on, plus any numeric metrics found. Keyword rules always run; when a commercial AI provider key is configured, that model is asked too and its answer is used only if its quoted evidence is found in the document.

Trained or measured on: Not trained. The keyword rules are hand-written (packages/panels/src/due-diligence/constraint-extraction.ts). The gateway model is a vendor's general model, not tuned on our documents. The 2026-09-09 pilot measured a self-hosted Qwen2.5-7B on 20 documents, synthetic plus a small real set; that model is not what serves.

Where we stop it

Screening grade. Neither leg has been validated on this platform's own documents, and every surface line that shows a finding names which leg answered and says so. A high-severity headline escalates the due-diligence environmental requirement to critical, which a reader should treat as 'read this document', not as a verdict. The keyword rules are blind to negation ('no adverse impacts' can read as adverse). A gateway call sends the document's text to a commercial provider. The serving register recorded it as not serving until 2026-09-14; it had in fact served since 2026-09-09.

Evidence: nvidia-lab/reports/planning-constraints-validation-latest.json · nvidia-lab/serving/registry.json

BESS domain model

BESS market intelligence
Evidence record, not a product surface

Answers a question about GB Capacity Market, Frequency Response and wholesale arbitrage, the platform glossary and data-source licensing, grounded in this repository's own documents, using a LoRA adapter over Nemotron-3.5-Lightning-30B-A3B and a retrieval index of 649 chunks.

Trained or measured on: 1,811 training and 95 validation instruction–response rows generated from this repository's own glossary, its data-source licence notes and the BESS API module's doc comments; three epochs on eight H100s in 52 minutes. Validation loss fell from 0.97 to 0.78 and validation mean-token accuracy rose from 82% to 84%. Checked on five held-out questions: the adapter alone fabricated a licence answer on one of them; with retrieval it answered all five with no fabricated fact.

Where we stop it

Answers only from the indexed corpus of 649 chunks and only for the questions that corpus covers, and was evaluated on five hand-picked questions rather than a systematic set. Not connected to any product surface: nothing in the apps or packages reads the adapter, and it needs a GPU host carrying the 30-billion-parameter base model before anything could.

Evidence: nvidia-lab/model-cards/bess-domain-model-nemotron-3.5-lightning-lora-v1.md · nvidia-lab/bess-domain-model/adapter/adapter_model.safetensors · nvidia-lab/bess-domain-model/training-evidence/train.txt · nvidia-lab/bess-domain-model/data/train.jsonl · nvidia-lab/bess-domain-model/data/val.jsonl · nvidia-lab/bess-domain-model/rag-index/chunks.json

Thermal-runaway onset and severity (negative result)

Due diligence
Evidence record, not a product surface

Tested whether a learned model can predict thermal runaway in a mechanically indented lithium-ion cell from state of charge, capacity, chemistry and source family, with time to runaway and peak temperature for the tests that ran away. It was not promoted: held out by cell model, a regularised logistic model scored ROC AUC 0.924 with a lower bound of 0.880, which does not clear the 0.904 of a plain logistic fit on state of charge and chemistry, and a gradient-boosted model scored 0.824.

Trained or measured on: 226 independent indentation tests parsed from a public CC BY 4.0 record (Mendeley Data sn2kv34r4h), labelled from the measured temperature channels only: 34 runaway, 177 none and 15 unknown because the sensor saturated. The authors' computed 0 to 100 score is neither a target nor a feature. Evaluated by holding out one of 13 cell models at a time, with a verdict rule committed before any model was fitted. Time to runaway scored a concordance of 0.458 against 0.474 for a Kaplan-Meier baseline.

Where we stop it

226 tests of one abuse mode, mechanical indentation, on single cells at 23 to 27 °C; only 6 of 13 cell models hold a runaway, with 8 to 42 tests per cell model. The share of destructive tests that ran away is not a field failure rate and not a frequency. Nothing was exported and no product surface reads it.

Evidence: nvidia-lab/reports/mechanical-tr-runaway-validation-latest.json · nvidia-lab/model-cards/mechanical-tr-runaway-v1.md · nvidia-lab/runs/lab-month-a4-20260920T142125Z/SUMMARY.md · nvidia-lab/mechanical-tr/runaway_harness.py · nvidia-lab/serving/registry.json

GB battery realised revenue per MW (panels shipped, model a negative result)

BESS market intelligence
Evidence record, not a product surface

Tested whether a GB battery unit's observed revenue per MW per year can be predicted from what is public before it operates: the capacity market's duration class, registered MW, GSP group, connection level, commissioning year and the size of its lead party. It was not promoted: holding out one site at a time, a gradient-boosted model scored a mean absolute error of 17,378 GBP/MW/year with an upper bound of 24,395, which does not clear the 19,971 of a median by duration band, and ridge regression scored 18,846. What ships is the two measured panels the test was run on.

Trained or measured on: A panel of 232 GB balancing units identified as batteries, each with the evidence for it, and 122 units examined and rejected with a reason; and 5,061 unit-months of measured revenue for 229 of those units from 2023-11 to 2026-08, built from Elexon's balancing cashflow per unit and NESO's auction results by unit. Scored on 93 units across 70 sites for 2025-09 to 2026-08, with a verdict rule committed before any model was fitted. A unit's own previous year predicts its next with a rank correlation of 0.635 over 68 units, which the public attributes do not capture.

Where we stop it

Observed revenue is partial: balancing-mechanism net cashflow plus response and reserve auction availability revenue, with no wholesale trades, no capacity market revenue and no charges. The fleet reads 13,970 GBP/MW/year for 2025-09 to 2026-08 on measured streams, a fifth to a third of published total-revenue figures, and is not the same quantity as a published index. The identification panel is a floor, not a total. Nothing was exported and no product surface reads it.

Evidence: nvidia-lab/reports/gb-battery-revenue-validation-latest.json · nvidia-lab/model-cards/gb-battery-revenue-v1.md · nvidia-lab/runs/lab-month-a3-20260920T153456Z/SUMMARY.md · nvidia-lab/gb-battery-revenue/revenue_harness.py · data-platform/lake/features/gb-battery-fleet/identification-panel/gb-battery-unit-panel-manifest.json · data-platform/lake/features/gb-battery-revenue/unit-month-revenue-manifest.json · nvidia-lab/serving/registry.json

Large-format LFP calendar-fade regression (Offenburg)

Cells and packs
Evidence record, not a product surface

Predicts capacity retention for large-format (~200 Ah) stationary-storage LFP cells that are calendar-aged only (never cycled), from elapsed checkpoint, temperature and state of charge.

Trained or measured on: Ordinary least squares, 8 cells (2 temperatures x 2 SOC x 2 replicates) x 4 real checkups each = 32 rows, from Zenodo 22939281 (Offenburg University, CC-BY-4.0). R-squared 0.946; leave-one-cell-out MAE 0.886 percentage points vs. 3.088 for a constant-mean baseline (71% improvement), against a written-before-fitting bar of 20%.

Where we stop it

One manufacturer, one 180 Ah cell design, 8 cells total -- not a generalised large-format LFP calendar-fade rate. No per-checkup calendar date exists in the source record, so the fitted time coefficient is a fade-per-checkpoint rate, not a fade-per-day rate. Checkup_#1500 was dropped as a checked bit-identical duplicate of Checkup_#1000 in the source spreadsheet, leaving 4 real time points per cell rather than 5. Not wired into any customer-facing surface.

Evidence: nvidia-lab/bessler-calendar-fade/fit_bessler_calendar_fade.py · nvidia-lab/bessler-calendar-fade/results/result.json · nvidia-lab/bessler-calendar-fade/build_bessler_cell_table.py · nvidia-lab/bessler-calendar-fade/results/bessler-cells/cell-table.json

Large-format LFP cycle-fade rate and activation energy (Offenburg)

Cells and packs
Evidence record, not a product surface

Fits a temperature-specific capacity fade rate for large-format (~200 Ah) stationary-storage LFP cells under full cyclic aging, and derives an activation energy from the two temperatures tested as an independent cross-check against the source paper's own reported figure.

Trained or measured on: Linear least squares on each of the 4 fully-cycled cells' complete per-cycle capacity trajectory (up to 1,509 cycles), from Zenodo 22939281 (Offenburg University, CC-BY-4.0). Mean fade rate 0.0130 Ah/cycle at 35C and 0.0243 Ah/cycle at 50C; solving the Arrhenius equation on those two rates gives 34.3 kJ/mol against the paper's own reported 37.3 kJ/mol (8% disagreement, independent method).

Where we stop it

2 cells per temperature, one manufacturer, one 180 Ah cell design. The 50C cells' cycles-to-80% (1,289-1,326) is an observed crossing; the 35C cells' (2,508-2,749) is an extrapolation beyond the 1,500 recorded cycles. A straight-line fade rate understates acceleration near end-of-life. Cross-checked 2026-09-24 against an independent manufacturer's 280 Ah cell (Zenodo 20813753, TEMPEST): inconclusive, not a pass or fail -- a 1.74x observed/predicted rate disagreement is fully explained by TEMPEST's 4x higher C-rate (1C vs. Bessler's measured ~0.25C), a confound the temperature-only model was never fit to separate from. Not wired into any customer-facing surface.

Evidence: nvidia-lab/bessler-cycle-fade/fit_bessler_cycle_fade.py · nvidia-lab/bessler-cycle-fade/results/result.json · nvidia-lab/bessler-cycle-fade/cross_check_tempest.py · nvidia-lab/bessler-cycle-fade/results/tempest-cross-check.json

Capacity-normalised a2 features, tested on the Bessler large-format cell

Cells and packs
Evidence record, not a product surface

Retrains a2-LFP (leave-chemistry-out) with discharge capacity divided by each cell's own capacity at the early-cycle window's start before featurising, to test whether a2's raw-Ah features (two orders of magnitude off between small- and large-format cells) explain its catastrophic failure on large-format LFP cells.

Trained or measured on: Identical architecture, step count and 854-cell corpus/fold split as the original a2-LFP checkpoint, only the feature representation changed (a2_features.featurise() itself untouched). LFP leave-chemistry-out MAPE improved from 76.98% (below its 50.02% baseline, not promoted) to 39.45% (above a 65.38% baseline, promoted) on the model's own native small-format test cells.

Where we stop it

Fixes the model's own in-distribution LFP performance but NOT the cross-format transfer question it was built to test: scored against the same Bessler large-format cell (CA-181103-22, actual cycleLife 1,391.7), the normalised model predicts 184.5 cycles (86.7% error) -- improved from 66.2/6.6 (95.2%/99.5% error) before, but still a ~7.5x underprediction and not usable. Not adopted into the shared cell-ageing model; that decision belongs to whoever claims that row.

Evidence: nvidia-lab/bessler-a2-capacity-fix/build_a2_features_capacity_normalized.py · nvidia-lab/bessler-a2-capacity-fix/results/summary.json · nvidia-lab/bessler-a2-capacity-fix/results/lfp-fold-eval.json · nvidia-lab/bessler-a2-capacity-fix/results/bessler-score.txt

External tools we use

Models and toolkits built by others. Two of them we measured and retired; one gave us a result we make no product claim from.

Chronos-2

Amazon
Measured — no product claim made

A time-series foundation model, tested zero-shot against day-ahead electricity prices in Great Britain, Germany and California, and then routed through our own dispatch engine to see whether a better price forecast earns a battery more revenue.

Measured on: Great Britain, re-run 14 September 2026: 284 forecast origins between September 2025 and August 2026. Mean absolute error 13.72 £/MWh against the best baseline's 16.76 — an 18.1% improvement, checked with two leak detectors. Routed through the dispatch engine over the same 284 days it earned £888,599 against seasonal-mean's £999,326, so it was not promoted into the dispatch path: price error and dispatch revenue disagreed. Biasing the dispatch decision with its own quantile forecasts was tried in four designs, and every non-zero setting of every design earned less than the unbiased point forecast, so none was adopted. Germany (DE-LU day-ahead auction): 268 origins, mean absolute error 22.15 €/MWh against the best baseline's 26.90 — a 17.7% improvement — and the best dispatch revenue of the four models, the first market where both measures agreed. California (CAISO SP15): 268 origins, the best price error of the four models at 5.09 $/MWh, 11.2% better than persistence, but second on dispatch revenue at $734,208 against seasonal-mean's $759,104.

Where we stop it

+18.1% over the best calendar-aligned baseline on the 2026-09-14 re-run is what we cite for Great Britain; it supersedes the +20.2% of the earlier 286-origin scorecard, and an earlier +27.2% was a misaligned baseline and we retired it. A better price forecast did not mean more dispatch revenue in two of the three markets, so it serves no forecast in the product and is not in the dispatch path.

Evidence: nvidia-lab/reports/gb-dayahead-forecast-latest.json · nvidia-lab/reports/gb-dayahead-forecast-latest.md · nvidia-lab/reports/gb-dayahead-forecast-scorecard-latest.md · nvidia-lab/reports/gb-dayahead-dispatch-forecast-value-full-window-latest.json · nvidia-lab/reports/gb-dayahead-risk-aware-dispatch-bias-designs-latest.json · nvidia-lab/model-cards/chronos-2-gb-risk-aware-dispatch-eval.md · nvidia-lab/reports/de-dayahead-forecast-latest.json · nvidia-lab/reports/de-dayahead-dispatch-forecast-value-latest.json · nvidia-lab/model-cards/chronos-2-de-lu-dayahead-eval.md · nvidia-lab/reports/caiso-dayahead-forecast-latest.json · nvidia-lab/reports/caiso-dayahead-dispatch-forecast-value-latest.json · nvidia-lab/model-cards/chronos-2-caiso-sp15-dayahead-eval.md

Moirai

Salesforce
Trialled — not in the product

Trialled as a second time-series foundation model for British day-ahead electricity prices, chosen because it can take known-in-advance covariates — real Capacity Market and Frequency Response clearing prices — that the univariate Chronos-2 run could not use.

Measured on: Moirai-2.0-R-small zero-shot with the real covariates over 114 forecast origins between April and August 2026: mean absolute error 19.01 £/MWh against the seasonal-mean baseline's 16.95, 12.1% worse. Moirai-1.1-R-base fine-tuned on real GB price and covariates and scored on 41 held-out origins: 3.6% better than zero-shot Moirai but 15.4% worse than the best baseline, with both leak detectors clean. Routed through our dispatch engine on the same 41 days, both variants drove the simulated battery to a net loss — £2,662 fine-tuned and £19,537 zero-shot — against a £324,488 oracle.

Where we stop it

Neither variant beat a calendar baseline, so we did not adopt it and the line of investigation is closed as scoped. The measurement points at the problem shape rather than the model: monthly and yearly clearing prices carry little signal for half-hourly price moves a day ahead, and the forecast day curves came out nearly flat, which is a statement about this target and not about the model's quality on others. Listed because a register of only the tools we kept is not a register.

Evidence: nvidia-lab/reports/gb-dayahead-forecast-moirai-latest.json · nvidia-lab/model-cards/moirai-2.0-gb-dayahead-covariate-eval.md · nvidia-lab/reports/gb-dayahead-forecast-moirai-finetune-latest.json · nvidia-lab/model-cards/moirai-1.1-gb-dayahead-finetune-eval.md · nvidia-lab/runs/moirai-dispatch-revenue-eval-20260914T212500Z/SUMMARY.md

cuOpt

NVIDIA
Retired to ordinary hardware

GPU optimisation, trialled for scheduling when a grid battery charges and discharges against market prices.

Measured on: 90 independent dispatch problems, August 2026. cuOpt and the ordinary CPU solver produced identical answers — the worst cell £0.00 apart, with zero constraint violations on either side.

Where we stop it

Ten times slower than the CPU solver here — 3,131 seconds against 318 — because it is built for one large coupled optimisation and we brought it ninety small independent ones: a statement about shape, not about the product.

Evidence: nvidia-lab/reports/gb-portfolio-envelope-latest.json · nvidia-lab/runs/burst-c3-01/SUMMARY.md

RAPIDS cuGraph

NVIDIA
Retired to ordinary hardware

GPU graph analytics, trialled on which pieces of a project's evidence connect to each other.

Measured on: A nine-node, ten-edge project contract for a Glasgow site, 19 July 2026. Exact parity with the CPU path on every node and edge.

Where we stop it

2.68 seconds on the GPU against 0.03 on ordinary hardware — again problem shape, since a nine-node graph does not repay the start-up cost.

Evidence: nvidia-lab/reports/project-relationship-graph-glasgow-h100.json · nvidia-lab/reports/project-relationship-graph-glasgow-local.json · docs/NVIDIA_RELATIONSHIP_MODEL_PLAN_2026-07-19.md

PhysicsNeMo

NVIDIA
Toolkit we build with

The framework the discharge-curve surrogate is built with — a neural operator that learns a whole curve rather than a single number.

Measured on: Not scored on its own. What was scored is the model it produced, listed under the built models above.

Where we stop it

A training-time dependency; we cannot run it on our own machines, which is why the curve surrogate stays unreachable.

Evidence: nvidia-lab/serving/registry.json · nvidia-lab/reports/material-curve-surrogate-validation-latest.json

PyBaMM

Open source
Toolkit we build with

The battery physics solver that generates the truth data every materials surrogate here learns from.

Measured on: A 56,700-point sweep, of which 54,896 solved cleanly; 3.2% failed at physical extremes and each failure is recorded against its own point rather than dropped from the record.

Where we stop it

Two of our eight chemistries have a degradation parameter set and they share it, which is where the cycle-life limit above comes from.

Evidence: docs/CYCLE_LIFE_IMPROVEMENT_2026-08-07.md · nvidia-lab/reports/material-surrogate-validation-latest.json

Nemotron-Parse

NVIDIA
In progress

Reading land, grid and planning documents and proposing dated obligations for a person to confirm.

Measured on: A three-document corpus written for the test harness: 13 obligations proposed, 6 of them matching the intended answer, an F1 of 0.46.

Where we stop it

Extraction proposes and a person confirms — of the 13 obligations it proposed, none is auto-accepted.

Evidence: nvidia-lab/reports/obligation-extraction-latest.json

Nemotron-Nano-9B-v2

NVIDIA
Trialled — not in the product

Trialled for reviewing project evidence and returning a structured verdict a reviewer could check.

Measured on: Three evidence-review cases on our own eight-GPU node, 15 July 2026: all three returned valid structured output and invented no citations; mean score 0.733.

Where we stop it

It over-rejected the conditional cases, so we did not adopt it. Listed because a register of only the tools we kept is not a register.

Evidence: nvidia-lab/runs/cvp-nemotron-nano-9b-v2-20260715T113621Z-06be73da/result.json

Hosted language models

NVIDIA, OpenAI, Anthropic, Google
Runs only where an operator configures a key

One seam across four hosted providers for text extraction and drafting, so no vendor is hard-coded in.

Measured on: Not scored as a capability. Which provider answers depends on which key an operator has configured.

Where we stop it

Hosted services that need a key and run nothing without one; where configured, they draft and extract and never commit a number.

Evidence: nvidia-lab/serving/registry.json · docs/AI_MODELS_AND_COMPUTE_INVENTORY_2026-08-06.md

The external data our models learn from is listed separately, with its licences and attributions, on the Data Sources page. Questions about a model or a figure on this page? Contact us via the details on the main page.