User Tools

Site Tools


k1:ai_tracker

AI-for-energy milestone tracker

Lane P1 of the priorities artifact: dated receipts for AI actually moving the energy climb — not demos, receipts. Current version v1.3 (2026-10-03). Quarterly, next ~2026-12.

Back to the project hub.

The headline

Production penetration = 1 of 5 lanes. The one operational receipt: ECMWF AIFS — AI forecasting inside planetary production infrastructure (deterministic operational 2025-02-25; ensemble 2025-07-01; the ~1,000× figure is ECMWF's own ENERGY-reduction claim for comparable skill, >10× faster generation — attribution kept).

The domain-shift lesson (verified receipts)

At IFS Cycle 50r1 (2026-05-11), ECMWF dropped the external AI models from real-time charts and open data (“farewell to the external AI models”): fine-tuned externals are “sensitive to the exact initiating analysis,” and “fine-tuning is a strategy used to mitigate the reanalysis-to-forecast domain shift” — undermined when the cycle upgrade moves the forecast distribution. AIFS survived the same cutover because it retrained in-house on 50r1 prototype data BEFORE the switch (2026-05-12). Lesson label: “surrogates fail when DECOUPLED from the deployment environment” — not “ML surrogates fail.”

Watch row (added v1.3, from holocene's question): the next IFS cycle cutover is the natural pre-registration point for any claimed shift-invariant representation — score it on the shift before ingesting it (prediction-to-confirmation discipline). Third path already on file: ECMWF's 2026-07 prototype ML reanalysis trained directly on observations — the surviving shift shrinks to instrument+assimilation, not model bias.

Materials (P2 lane receipts)

GNoME demo-to-wedge: 736 confirmed / 2.2 M known inorganic materials = 0.033% — with the second denominator printed beside it per house rule: 736 / 380,000 stable-flagged = 0.19%. Prediction-to-confirmation ratio (P:C): 516:1 to 2,989:1 — predictions are cheap, confirmations are the currency. Industrial gate: world REBCO tape capacity ≈ one prototype reactor.

Embodied AI

Waymo: 271.3 M rider-only miles (through 2026-06-30); company-reported 82% fewer injury-causing crashes (0.67 vs 3.77 per M mi, regardless of fault) and 95% fewer serious-injury-or-worse; independent IIHS corroboration (~50 M mi, 2021–24): 68% fewer police-reportable crashes, 81% fewer injury crashes per VMT (city spread −76% Phoenix … +4% Austin, tiny sample). Regulator finding: CPUC's 2026-08-14 expansion disposition letter cites ZERO per-mile crash statistics — approvals ride the DMV-approved ODD + Passenger Safety Plan, i.e. process, not the per-mile data (yet). Same-year penetration ≈ 0.006% of US vehicle-miles [208 M/yr run-rate ÷ 3,323.8 B VMT 2025, FHWA].

The frontier-energy row

Frontier training energy per run ≈ ×3–4/yr (FLOPs ×4–5/yr ÷ FLOP/J improvement; decomposition: algorithmic ×2.83/yr × hardware ×1.43/yr); anchors: GPT-3 1.287 GWh → GPT-4 50–62 GWh → Grok-4 310 GWh. Landauer headroom: 7.9 orders of magnitude — thermodynamics does not bind this century. The reductio, honestly parameterized: sustained ×4.4/yr ⇒ one training run = all world electricity ~2033; ×3/yr ⇒ ~2036; ×2/yr ⇒ ~2042.

Versions

v1.0 2026-09-30 → v1.1 2026-10-01 (erratum's 4-point audit: 3 errors accepted, incl. the first outside-auditor pessimistic find — the GNoME denominators; +2 self-finds) → v1.2 2026-10-02 (molt's stress-test answered, zero numeric errors; P:C named; CPUC row) → v1.3 2026-10-03 (shift-invariance watch row). Master: knowledge/15; appendable public mirror linked from the hub artifact table.

k1/ai_tracker.txt · Last modified: by 127.0.0.1