1. Bottom line
This was one question pursued from four sides: is moving data, not arithmetic, what makes learning expensive — and can you build a scoring machine that rewards algorithms for moving fewer bytes? You fact-checked your own podcast claims the same day you made them (Tue 8 Sep 2026, week 37), turned the cost model into a public MNIST energy competition (Thu 10 Sep 2026, week 37), tested it against real chips (Esperanto, Cerebras) and against grid economics, and got the first measurement behind your no-backprop bet (reversible nets, Sun 13 Sep 2026, week 37). Your own week also shifted the thesis: the backprop tape wall is a capacity-and-dollars problem (~1.5–4.5% of the arithmetic energy at 7B width — derived, not measured), electricity is ~7% of an AI data centre's cost, and perf/W only decides purchases where power delivery is the binding limit — so "bytes and power delivery" is the framing that survives, not "joules". It serves the Lighthouse/Sutro goal strongly and the Anthropic / half-time-job decision and the runway hardly at all: no thread carries an ask, and hardware and paper research displaced the Anthropic block on Wed 9, Sat 12 and Sun 13 Sep. The decision it poses: box the energy work to Monday's Sutro room plus one evening so the Wed 16 / Thu 17 Sep (week 38) Anthropic blocks happen — or make it the job pitch by attaching one concrete ask to one of the people it produced.
2. The threads at a glance
| # | Thread | Dates | Main links | Status | Goal it serves |
|---|---|---|---|---|---|
| 1 | Podcast self-fact-check → memory-wall physics → Dally's heuristics | Sun 23 Aug (wk 34) root; peak Tue 8 Sep (wk 37) | Interlude doc · Fable fact-check · Memory Wall · Dally Heuristics | Shipped; Dally Heuristics has 3 confirmed errors; correction note to hosts not sent | Lighthouse (evidence for the cost model) |
| 2 | Why efficiency matters: grid power, electricity, the TCO paradox | Thu 27 Aug (wk 35) – Sun 13 Sep (wk 37) | Global electricity chat · PDF · GPU-hour note | Open; article promised to Ameen Patel not seen sent | Lighthouse "why now"; Resonance (SPC energy forum) |
| 3 | Dally grid model → MNIST energy competition | Fri 28 Aug (wk 35) – Mon 14 Sep (wk 38) | MNIST README · model 3 spec · redesign dossier | Public since Thu 10 Sep; PR #70 conflicting; MNIST-medium release due tonight | Lighthouse (first public learning benchmark) |
| 4 | Where bytes live: MoE, KV cache, FLOPs per byte, SRAM vs HBM | Sat 7 Feb (wk 6) root; Tue 8 – Mon 14 Sep | MoE chat · KV cache sizes · SRAM-vs-HBM cutoff | Answered by agents; several spoken numbers wrong | Lighthouse (which memory level an algorithm must live in) |
| 5 | Backprop vs the memory wall → reversible nets | Sat 31 Jan (wk 5) bet; Sun 30 Aug – Mon 14 Sep | Backprop brief · backprop-memorywall · batching report | Reversible result uncommitted (Intel only) | Lighthouse bet; the planned Anthropic artifact |
| 6 | Hardware landscape & co-design people | Wed 26 Aug (wk 35) – Mon 14 Sep | Chip landscape · reading page for Vijay Jain · SPC collaborators | Several loops open in both directions | Resonance |
| 7 | Esperanto / AI Foundry / Ainekko | Fri 11 – Mon 14 Sep | Esperanto doc · "espera" chat · AI Foundry research | Outreach drafted, not sent; a wrong story is circulating | Resonance; candidate real silicon |
| 8 | Cerebras wafer-scale dataflow (Natalia Vassilieva) | Sun 23 Aug (wk 34) – Sun 13 Sep | Natalia brief · notes · PR #7 | Met; no follow-up sent | Grid-model fidelity |
| 9 | Abstractions, evolvability, which hardware facts to freeze | Thu 30 Jul (wk 31) root; Fri 11 – Sun 13 Sep | Evolvability chat · CMOS survival · knol: engineering architecture | Parked; conclusion lives only in chats | Lighthouse (teachable core); competition rationale |
Adjacent, not counted as energy threads: influence functions (the Anthropic project — doc, Sat 12 Sep) and the Gradient Dissent layer-dropout review of Cerebras's "Don't Drop Dropout" (Wed 9 – Thu 10 Sep), which touches threads 5 and 8.
3. Thread by thread
Thread 1 — Podcast self-fact-check → memory-wall physics → Dally's heuristics
Trigger. The Interlude Show recording, Tue 8 Sep 2026 (week 37) 09:00–10:14 at SPC, hosts Vatsal Bajaj and Jake Kang. On air you said "Maybe data centers were using 5% of the electricity." The thesis is older: Bill Dally lecture notes from Tue 16 Dec 2025 (week 51), a Dally chat started Sun 29 Mar 2026 (week 13), and the Science Night version on Sun 23 Aug 2026 (week 34) ~19:38, where moving a word was ~100× adding it.
Timeline
- Thu 27 Aug 2026 (week 35) 13:52 — GPU-hour note (HBM pJ/bit, GPU latency ladder). Evening dinner report (after-math) corrects: register ~10×, DRAM ~6,400× an add (Horowitz 45 nm).
- Mon 31 Aug 2026 (week 36) 10:30 — Backprop's Days Are Numbered brief (Dally 64,000×, Gholami gap).
- Mon 7 Sep 2026 (week 37) 11:22–11:29 — Codex podcast brief warns that memory size alone does not give energy, and against the 420× claim.
- Tue 8 Sep 2026 (week 37)
- 09:00–10:14 recording (notes; the filename misdates it "08sep28").
- 10:58–11:12 Gemini Notebook decks The Algorithmic Reset and Beyond Backpropagation, plus a 5-minute video.
- 11:03 Fable fact-check chat starts (share); 11:17 Codex fact-check committed; 11:26 Fable page live; 11:35 posted in your #chat-yaroslav tagging both hosts (Fable page only).
- 11:38 "08sep26 - interlude show" doc created; the capacitance passage is in by 13:11.
- 13:04 Memory Wall and 13:44 Dally Heuristics published (unhashed URLs), both from Fable chat
5a230605on the second claude.ai account. - 14:32–15:25 coffee with Mark Saroufim; 16:02 self-factcheck doc.
- 16:37 you note that both Fable reports need fact-checking; 16:43 Intel Codex physical calibration (1.22 µm pitch, untracked); 17:08 "fable physical Dally calibration" doc — a partial fact-check.
- 17:41 you re-teach wire capacitance to Lucas Cassiano.
- By Sun 13 Sep 2026 (week 37) 16:43 — tailnet Memory Wall link and the chat link added to the Interlude doc.
What you learned.
- Dally's own numbers (CACM 65(9), Sept 2022, primary text read today): 32-bit add 20 fJ / 150 ps; two 32-bit words moved 1 mm 1.9 pJ / 400 ps; 40 mm corner to corner on a 400 mm² chip 77 pJ / 16 ns; going off chip 320 pJ per crossing; two DRAM words 1.3 nJ = 64,000× the add; a 256 MB memory access 58 pJ / 12.3 ns. A CPU add instruction is 80 pJ, so memory is only ~15× an instruction. "10 µm" follows: 20 fJ ÷ 1.9 pJ/mm ≈ 10.5 µm, and his AHA 2023 slide says "An add is worth 10um of movement."
- The mechanism: wire capacitance is stuck near 0.2 fF/µm, a bit costs ½CV², and supply voltage stopped falling around 2005. Your bolded takeaway in the doc: "capacitance = charge for 1 volt".
- The divergence (Gholami): accelerator FLOPs grew ~60,000× in 20 years (~3×/2 yr) vs DRAM bandwidth ~100× (~1.6×/2 yr) and interconnect ~30×.
- The Fable scorecard of your 57 on-air claims: 23 verified, 20 mostly right, 3 need nuance, 3 inaccurate, 4 unverifiable, 4 opinion — reliable on physics, loose on dates and specs.
- Your 2D-transistors / 1D-wires intuition got three grades the same day: Codex "useful motivation, inaccurate physical explanation", Memory Wall "not wrong, just incomplete", Fable "Verified".
- FLOPs per HBM byte (self-factcheck): H100 ≈ 990 TFLOPS ÷ 3.35 TB/s ≈ 300, so your "~1,000" at coffee was ~3× high — and your own 31 Aug brief already said ~295. Megakernels are "easy to beat weak baselines, hard to beat tuned ones".
Corrections to your own claims.
- The H800 export cut was NVLink (400 vs 900 GB/s), not HBM bandwidth — you said HBM on air and at Science Night.
- Horowitz's "Computing's Energy Problem" is ISSCC 2014; no 2012 talk exists. You said 2012 on air and again at the Vijay Jain lunch Thu 10 Sep.
- Name the source of any move/add ratio. 100× is ~1 mm of wire; 6,400× is Horowitz 45 nm (0.1 pJ add vs 640 pJ DRAM); 64,000× is Dally 2022.
- 30 mm under the 10 µm rule is 3,000 adds, not 30,000 (to Lucas Cassiano Tue 8 Sep, and at the Vijay Jain lunch).
- From the Codex fact-check: small peak memory does not by itself guarantee small energy; checkpointing stores activations, not Jacobians; MLPs reach 99.65% on MNIST without convolutions; a 1,000× toy speedup is not 1,000× energy.
- The TPU "cut" is an analyst forecast revision (~4M → ~3M units), not a confirmed production cut; your 27 Aug note, Fable ("mostly right") and Codex ("not confirmed") disagree on the cause.
- Your grid constant 1 fJ per byte per µm = 125 fJ/bit·mm, ~4× Dally's 29.7.
- The deck The Algorithmic Reset restates flagged claims as fact (arithmetic ~1% of energy; agents routinely find 1,000×). Don't share it uncorrected.
- Your published Dally Heuristics page is wrong in three places (details in "Filled in").
People. Vatsal Bajaj, Jake Kang, Mark Saroufim, Lucas Cassiano, Uliana Popov (you sent her the fact-check Tue 14:12). Sources: Bill Dally, Mark Horowitz, Amir Gholami.
Outputs. Fable fact-check · Codex fact-check · Memory Wall · Dally Heuristics · two decks and a video (Drive, 08sep26 day dir) · the recipe _request/request-claimresearch.md (written by the Codex agent in the 11:00 session) · the calibration doc.
Goal link (graded).
- Lighthouse / Sutro — medium-strong. It supplied calibration evidence (30 fJ/bit·mm, c/120) that later reached model 2. It did not supply the priced model itself: the competition ran on your own constants, then a Cerebras-style pitch.
- Learning as enjoyment — medium, with a cost. At 12:28 you said "I got really nerd-snagged by the numbers"; by 18:44, "I got too into reading all of these things", with headaches logged at 16:18, 18:02 and 20:44.
Filled in (survived verification).
- Which of your documents mis-state Dally (confidence high; primary CACM text via a course mirror, ACM itself returns 403).
- The backprop brief's Dally table is correct.
- Dally Heuristics (report = published page): (1) says 64,000× is 640 pJ ÷ 10 fJ — Dally writes 1.3 nJ ÷ 20 fJ; (2) says an add equals 5 µm — it is ~10.5 µm; (3) reads "only 15×" as an old-chip-era figure — it is today's instruction overhead (80 pJ vs 1.2 nJ). Its "Dally has used √n in talks for fifteen years" is unsupported; the √n cache cost is Ding & Smith's Data Movement Distance, and the physical access-energy ∝ √size heuristic is older (Amrutur & Horowitz). Its Python gives 37 fJ/bit·mm and 72.5 pJ / 13.17 ns with its own constants (disclosed), but its proposed fix gives 60.4 pJ, not the claimed 59.
- The 27 Aug after-math report credits √n correctly to Ding & Smith, but attributes "E ∝ d^β" to Dally (his model is linear) and quotes DRAM ~6,400× (likely Horowitz 45 nm; inferred).
- This also means the 10 µm figure Vijay Jain wrote into Slack is consistent with Dally. The number that spread wrong is Dally Heuristics' 5 µm.
- Memory Wall's checkable errors (confidence medium; the 675-line report was only partly checked): its claim that Moore's 1965 observation was about memory is wrong (it was components per circuit); a 1-mm wire going ~1 ns → ~500 ns is wrong by an RC estimate from its own dimensions (~13–20 ps → ~43–53 ns); Wulf & McKee assumed 80%/yr CPU speed-up, not 55%; internal contradictions on pin growth, the "two decades" flat row cycle, and 32-bit op energy. The Horowitz table (13 values) and Gholami's figures check out.
- Memory energy per bit, reconciled (confidence medium; vendor figures, no matched boundaries).
- SK hynix (IEDM 2023): HBM2E ~4.3 pJ/bit, HBM3E ~3.44 (20% lower), HBM4 projected ~1.7.
- O'Connor (MICRO 2017): HBM2 3.92 incl. ECC; GDDR5 14.0. Malladi (ISCA 2012): DDR3-1600 70 pJ/bit at full use, 160 typical.
- Verdicts: your 27 Aug note's HBM3E 3.44–4.05 is sourced (3.44 SK hynix; 4.05 Samsung via a secondary page). Memory Wall's DDR5 10–25 and LPDDR 5–15 have no primary source, and its "15–20×" DDR3→HBM3E gain holds only at peak DIMM use (~45× at typical use). The 14 Sep cutoff report's DDR5 ~80 has no source (nearest primary figure: DDR3's 70), its HBM link points at the wrong page, and it calls vendor claims "measured".
Still open. The rest of Memory Wall's history and roadmap claims; whether the H100 carries HBM2e or (as Memory Wall says) HBM3; whether the episode has aired.
Next step. Send Vatsal Bajaj and Jake Kang a five-line correction note before the episode airs: NVLink not HBM; Horowitz 2014; name the source of each ratio; TPU was a forecast; "5%" is the US, not the world.
Thread 2 — Why efficiency matters: grid power, electricity economics and the TCO paradox
Trigger. Your on-air 5% estimate and the argument that grid capacity becomes the ceiling. Roots: Thu 27 Aug 2026 (week 35) 12:10 you told Nikita Kopotun "half of the cost of GPU runtime is energy", then caught it yourself by 14:36 with the GPU-hour note.
Timeline
- Thu 27 Aug 2026 (week 35) 13:52 — GPU-hour note: electricity ~7% of TCO (Epoch 1 GW GB200 model). Mon 31 Aug (week 36): answered requests reconcile "half" vs 26% vs 7–12% (three denominators).
- Tue 8 Sep 2026 (week 37) 09:51 on air; 11:38–13:45 grid notes in the Interlude doc; 12:33 you open the Introl article on PJM's 6 GW shortfall and WhatsApp it to Eric Munsing at 12:37.
- Wed 9 Sep 14:30 — you create the Thu 17 Sep lunch invite on distributed AI boxes in homes. 17:00 SPC energy forum #2 (your attendance not established).
- Thu 10 Sep 11:02 — side chat before the SPC all-hands: electricity is ~10% of TCO; 11:17 the energy forum is announced. 14:46:03 you reopen the Introl article and one second later tell Christian Pehle "In 2027, there's going to be a 6-gigawatt shortfall." 15:50 you tell Ameen Patel "6 gigawatt more chips than we have power to power them" and promise him the article (walk summary).
- Fri 11 Sep 16:06–16:43 — "Global electricity usage and energy-efficient AI" chat, shared; 16:46 Introl reopened.
- Sat 12 Sep 13:12 — global_electricity_and_efficient_ai.pdf (7 pp) in the day dir. 14:09 Signal to Jason Yosinski: "electricity is only 10% of TCO in datacenters".
- Sun 13 Sep ~12:01 — you paste the same 10% line into claude.ai chat
5a230605; the partly captured reply says hyperscale is grid-power-limited, so perf/W buys capacity, not opex. ~21:51 at Science Night, Olia (surname unknown; identity from your 22:22 note) on methane from behind-the-meter gas-powered data centres.
What you learned.
- World vs US. World data centres used 485 TWh in 2025 ≈ 55 GW average ≈ 1.5% of electricity, rising to ~950 TWh (~3%) by 2030 (IEA). US data centres used 176 TWh = 4.4% in 2023 (LBNL), 6.7–12% by 2028. Your on-air 5% is right for the US, not the world.
- Electricity is ~7% of a GPU-hour's cost in the base case ($594M of ~$8.5B per GW-year, Epoch). 26% is a free-hardware counterfactual; "half" is Sequoia's non-GPU half. Energy does dominate cash opex at high utilisation.
- The chat's conclusion: the real constraint is power (GW) and connection time, not energy; >2,500 GW of projects wait in grid queues worldwide. The IEA high-efficiency vs base-case gap in 2035 (~230 TWh ≈ 26 GW) is, in the chat's words, "the largest 'power plant' on the table".
- Forecasting lesson: Andrae & Edler's 2015 data-centre forecast was ~3× too high on level because it assumed ~10%/yr efficiency gains vs ~20%/yr realised.
- PJM capacity price rose $28.92 → $329.17/MW-day (2024/25 → 2026/27); 63% of the 2025/26 rise is attributed to data centres. Your on-air point that your own electricity bill hasn't risen much is true in California and false in PJM.
- Jevons: energy per AI task falls ≥10×/yr while effective task volume grew ≥15× in a year.
Corrections to your own claims.
- The Interlude doc's grid lines (1.4 TW total, 2.2 TW in queue, data centres 25 GW and 4%) are US figures, not world. "4%" is an energy share, not a capacity share. "Only 13% of queue doesn't get withdrawn" is misworded: of 2000–2019 requests, 13% were built, 77% withdrawn, 10% still waiting.
- Energy as half of a GPU-hour's cost was ~5× too high (you caught it).
- The 6 GW shortfall has a source, and you told it wrong. It is PJM's 2027/28 capacity auction coming in ~6.6 GW short of its reserve target in 13 states plus DC — thinner reserves, not chips waiting for power, and not national. The "40%" in the Introl headline is a Gartner forecast.
- PJM's shortfall is ~0.2% of world average load: large locally, tiny globally.
People. Ameen Patel, Christian Pehle, Nikita Kopotun (organises the SPC energy forum; accepted Thu's lunch), Eric Munsing, Austin Diamond, Olia, Jason Yosinski. Sources: Anders Andrae, Jens Malmodin.
Outputs. The chat share, the PDF, and the Thu 17 Sep 2026 (week 38) 12:00 lunch you organised (two invitees still pending).
Goal link (graded).
- Lighthouse "why now" — medium. Tying the value of efficiency to joules per bit moved — the memory wall — is Claude's framing, not yet yours in writing.
- Career bridge — weak. No ask was attached.
- Resonance — medium. Nikita Kopotun's energy forum, the Thu lunch, and Olia are live contacts; the forum brings Andrew Cantino (Overview Energy, space-based solar) on Fri 18 Sep 2026 (week 38) 11:00 and offers a Panthalassa founder visit.
Filled in (survived verification).
-
US vs world, one table (confidence medium; IEA and LBNL pages returned 403, figures checked via snippets).
US World Installed capacity 1.34 TW, end-2025 (EIA) ~9.6 TW, end-2024 (derived from IRENA) Average load ~0.47 TW ~3.6 TW (derived) Grid queue 2.29 TW end-2024; 2.06 TW end-2025 (LBNL Queued Up) >2,500 GW incl. large loads (IEA, Apr 2026) Data centres now ~25 GW in 2024 (McKinsey; BNEF says ~35); 4.4% of electricity (LBNL) ~100 GW capacity; 415 TWh ≈ 1.5% (2024) 2030 >80 GW total (McKinsey) or +65–90 GW growth (Grid Strategies) ~950 TWh ≈ 108 GW average, ~3% One McKinsey (Sep 2024) line — 25 GW, 3–4% of US power, >80 GW by 2030 — probably explains your whole doc line. Slide rule: label every cell US or world, and capacity, peak or average.
-
The 6 GW source (confidence high: the article was open one second before you quoted it). Introl, "PJM Grid Crisis: 6GW Shortfall…" (Blake Crosley, Fri 6 Feb 2026, week 6), resting on PJM's own results (Wed 17 Dec 2025, week 51): 134,479 MW cleared at the $333.44/MW-day cap, 6,623 MW short; data centres are ~5,100 of 5,250 MW of forecast load growth. Morgan Stanley's "38 GW gap by 2028" is a different, national estimate (via Motley Fool, Thu 20 Aug 2026, week 34). No captured message sends Ameen the article (WhatsApp, Telegram, Messenger, Slack, Gmail, Signal checked; iMessage and in-person are not captured).
- What a saved watt is worth — the paradox resolved (confidence medium; assumptions are the agent's). Per watt of IT load per year, against Epoch's $8.55/W-yr total cost:
- Price is the limit: electricity $0.59 + facility sized in watts $1.43 ≈ $2.0/W-yr (~24%).
- Grid is the limit, on-site gas buys speed: ≈ $2.2/W-yr (~26%) (Rabobank: $1,700–2,000/kW engines, $107/MWh delivered; from excerpts).
- No workaround: extra compute revenue minus hardware ($6.2/W-yr). At CoreWeave's actual $6.9–8.2/W-yr (−2% GAAP margin) that leaves only $0.7–2.0; at Cleanview's projected $10–12M/MW, $3.8–5.8; for chips already bought and waiting for power (Nadella), $7–12/W-yr.
- So a 2× perf/W chip saves ~12–13% of cost in the first two cases — erased by any software or perf-per-dollar gap. Perf/W decides purchases only where power binds and revenue per MW sits well above cost. Both your statements are true, for different buyers: the grid ceiling is real for frontier buyers, who measure tokens per MW on LLMs; "electricity is 10% of TCO" describes buyers who are price-limited.
Still open. Whether "1.4 TW total" meant installed capacity or queued generation; whether you attended energy forum #2; the paywalled SemiAnalysis model.
Next step. Before Thursday's lunch, send Ameen Patel the Introl link with one line: PJM's reserve shortfall, not chips waiting for power.
Thread 3 — Dally grid model → the MNIST energy competition (Sutro's scoring machine)
Trigger. Sutro planning sprint #2, Wed 2 Sep 2026 (week 36), floated "maybe MNIST or pattern recognition as an intermediate" goal; it became the goal through the Fri 4 Sep sanity check and Sutro #30 (Mon 7 Sep). The handwritten origin is your Notability page Notability.2026/Mine/sutro/sutro ideas.pdf (exported Tue 8 Sep 11:27): MNIST tiers at 3×3 / 9×9 / 28×28 and an "AT²" time-energy-area scoring sketch.
Timeline
- Fri 28 Aug 2026 (week 35) 16:41–21:35 — NVML measurements on Modal A10G, B200, H100: the grid overstates the add/matmul gap ~7.5× (62.46× predicted vs 8.31× on B200).
- Tue 1 Sep 2026 (week 36) 17:06–23:42 — first A100 runs (Intel Codex): 8192³ matmul FP32 15.76 J, FP16 1.43 J, INT8 0.904 J; a 1-bit AND+POPCOUNT kernel 0.12269 J vs grid 0.118953 J.
- Wed 2 Sep — calibration against real silicon fails (24.4× die area, ~40× pitch); 12.82 J for MNIST-60k INT8 inference enters the "MNIST on grid model?" doc.
- Thu 3 Sep 13:52–17:11 — M5 Codex settles c/160, 1 µm, 1 fJ/byte (posted to the Sutro Telegram 17:18).
- Fri 4 Sep 16:40–17:06 — 9-agent sanity check: a 3×3 nearest-neighbour ceiling of 65.8%; recommends 8×8; not adopted.
- Sun 6 Sep — Boris Ginsburg hike: add synchronisation to the cost; 16:51 multiprocessor Grid VM session.
- Mon 7 Sep 2026 (week 37) 18:00 — Sutro #30: constants 1 µm / 1 fJ / 0.5 ps; A100 forward+backward is 2.85× forward in time, 2.74× in energy (the earlier "8×" was a BF16 artifact).
- Tue 8 Sep 16:00–17:58 — redesign dossier (58 forks); 17:03 Telegram: "the task of the submitter will be to create an appropriate programming language".
- Thu 10 Sep 16:48 Telegram: shift the IR choice to participants; 18:04–18:30 v4 tape ISA with a 50 fJ / 50 ps floor (PR #6 merged); 18:11 competition launched.
- Fri 11 Sep 08:10 Telegram: "I'm finding that 48 is too low." 14:29 / 14:35 emails to Bill Dally ("What's a good way to add multicore?"; out until Tue 22 Sep) and Ronny Krashinsky (who had forwarded you to Dally on Sun 30 Aug 19:04 — not cold). 15:52 you flag the big single-core-grid vs A100 mismatch; 18:03–18:15 pitch-128 spatial computer (PR #7 merged).
- Sat 12 Sep 04:06 answered requests: rank on grid energy; 07:34 PR #71 (Alex Varga) merged.
- Mon 14 Sep 2026 (week 38) 04:09 — outstanding PRs: PR #70 conflicts; MNIST-medium release promised to Lucas Cassiano for tonight.
What you learned.
- The model now has a published constants set (simplified-dally-model models 2/3): 1 µm per 32-bit word, 1 fJ per word per hop (31.25 fJ/bit·mm, 5% above Dally), c/120 = 0.4 ps per hop (model 2), reads and writes both charged, Manhattan distance, 50 fJ / 50 ps floor.
- No single pitch fits silicon: at 1.2 µm a 1,024-byte array is ~2× optimistic on energy and ~15× pessimistic on area.
- Multiprocessor finding — the tape is the wall: at 1 B/ns every streamed workload caps at 2×; 16 tiles do 64×64 int8 matmul 15× faster for 2.0× the energy; systolic arrays are 35–325× faster than a blocked single core.
- A100 noise: energy across 11 identical-work draws has CV 7.7% (23% spread) vs zero for grid scores, so decisions under ~15% measure the data centre. Best MNIST-medium entry is ConvNet v4 at 97.76% (PR #70).
- PR #69 (Andy Zhang): 3×3 MNIST as a 280-op Dally IR scores 4.625 pJ/inference vs 108 nJ on the A100 — a ~23,400× gap the PR itself calls launch-dominated at that size.
- Outside uptake: Alex Varga used GPT-6 Astra to optimise 16×16 matmul under your Dally cost model and reported 5.3× lower energy (X, seen Tue 8 Sep).
Corrections to your own claims.
- A micron is a thousandth of a millimetre, not a millionth (Sutro #30).
- "0.5 picosecond is 2 gigahertz" (Alex Varga, Sutro #30) is 2 THz.
- The pitch has drifted through 0.61 / 1.0 / 1.2 / 1.22 / 1.25 / 1.726 µm and the velocity through c/160, c/150, c/120, c/1200. Model 3's time is ~c/2,340 (1 ns per 128-node link) — calibrated to Cerebras's clock, not to Dally's wire.
- The Sutro #30 figure of ~1 J for an 8K matmul does not conflict with the 0.12269 J measurement: that was a 1-bit kernel; INT8 and FP16 were 0.9 and 1.4 J.
- The public Grid VM design page still says Chebyshev distance and prints "1.875 µm/ns" (a 1,000× typo).
People. Andy Zhang, Alex Varga, Lucas Cassiano, Mark Saroufim, Thomas Dybdahl Ahle, Boris Ginsburg, Cosmin Negruseri, Devrim Yasar, Bill Dally, Ronny Krashinsky, Ross Pantone.
Outputs. The competition (launched Thu 10 Sep 18:11), models 2 and 3, the multiprocessor page, the dossier, the Dally and Krashinsky emails.
Goal link (graded).
- Lighthouse — strong. It is the first public learning benchmark in the programme; matmul and sparse-parity leaderboards already had outside record-setters.
- Resonance — medium, not strong. Right now the waiting runs the other way: 8 of Andy Zhang's PRs have sat 9–11 days without your review.
- Coaching floor (5 focused h/week with an artifact) — plausibly met, estimated, not measured.
Filled in (survived verification).
- What is settled vs open in the constants (confidence high for the values, medium for the advice).
- Settled on main: writes charged, arithmetic free, units fJ/ps/ns, time scored, Manhattan distance.
- Open: velocity (model 2 wire-calibrated vs model 3 clock-calibrated); the 50 fJ floor, which matches no anchor (Dally's 8 KB sub-array is 320 fJ per word, 6.4× higher; Jouppi's 8 KiB SRAM ~3.75 pJ, ~75× higher); pitch (model 3 is 1.49–1.56× too dense; the area-exact 1.22–1.25 µm leaves energy per µm 14–17% low at 1 fJ/hop).
- Advice before the Tue 22 Sep 2026 (week 39) Dally resend: keep 1 µm / 1 fJ as normalised units; state that time is clock-calibrated; disclose the floor; retire the byte-grid ε, c/160, the typo and Chebyshev; mark model 1 as legacy.
- Does grid energy rank-order A100 energy? (confidence medium; n is tiny.)
- Pooled across all entries, rho = 0.98 for energy (n=9) and 0.90 for time (n=8) — but that pools three different cost models across 8 orders of magnitude, so it mainly says bigger jobs cost more.
- The spatial grid actually used for ranking has 2 scored entries, so it is untested.
- The one matched test (PR #71, three schedules of one MLP, rotating order, same container) used the retired single-core model: energy rho = 0.5, time rho = −1.0. The A100 orderings are reproducible (energy SD ~2%), so the single-core model inverts real A100 differences inside a tier. A100 idle-adjusted energy tracks A100 time exactly (rho = 1.0). Launch count and copies are ruled out as the cause.
Still open. Why PR #71's "no_slowdown" variant inverts (different CTA grids, inference batch 1 vs 30); re-scoring PRs #64–#66, #70, #71 under the spatial model.
Next step. Resolve PR #70's conflict, label its energy columns "single-core model", and release MNIST-medium at Sutro #31 tonight; on Tue 22 Sep send Bill Dally one constants table with the pitch-128 model.
Thread 4 — Where bytes live: MoE, KV cache, FLOPs per byte and the SRAM-vs-HBM cutoff
Trigger. The Boris Ginsburg hike, Sun 6 Sep 2026 (week 36): register 1 clock, SRAM ~10, HBM 500–1,000 (10:25), and, from Boris, that we are not limited by FLOPs (11:10, translated from Russian). On air Tue 8 Sep you said MoE lets you widen layers without paying the cost; at 11:58 you asked Claude about it. The "50 GB KV cache" line goes back to Mon 16 Feb 2026 (week 8).
Timeline
- Tue 8 Sep 2026 (week 37) 12:02–12:16 two MoE chats (public, public2), 12:03 filed in knol: suboptimization; 14:41 at coffee, "1,000 FLOPs per HBM byte" to Mark Saroufim; 14:43 MoE is great for training, not inference.
- Thu 10 Sep 12:49–12:50 at the Vijay Jain lunch: "10 gigabytes of SRAM is more than enough", "KV caches can be 50 gigabytes".
- Sat 12 Sep 10:06 at breakfast with Jason Yosinski: a chip "can only do, like, 500 megabytes" (the next sentence names Groq); 10:22 paging from DRAM "like 1000 cycles" (a latency figure); fetching SRAM >10 mm away loses to HBM. 14:41–15:08 Claude Code traces the 50 GB story; 14:48 ChatGPT KV Cache Sizes; a second share at 15:08 holds the Taalas answer.
- Sun 13 Sep 12:26 check-in asks for the cutoff; 18:44–20:10 Intel Codex sizes A100 memory levels; at 19:06 you invert the ratio to 300 bytes per FLOP; 19:16 you ask it to clarify sparse vs dense "15 vs 30".
- Mon 14 Sep 2026 (week 38) 04:15 — SRAM-vs-HBM cutoff report.
What you learned.
- MoE makes total parameters free, not active width. FLOPs saved in training become bytes owed at every inference replica: DeepSeek-V3's minimum decode unit is 320 GPUs, with >100 latency-bound all-to-alls per token. FLOPs are an honest currency for training and a dishonest one for decode; the fix is a shared cost model or a shadow price per resident byte charged to the training org.
- KV cache per token (BF16): Llama 3.1-8B 128 KiB (16 GiB at 128K); Llama 3.3-70B 320 KiB (40 GiB at 128K); DeepSeek V3 / Kimi (MLA) 68.6 KiB; Qwen3.5-35B-A3B hybrid 20 KiB. Your spoken "50 GB" checks out for 70B at 128K (~43 GB); your written "500k context can take 50GB" is ~3× low for 70B (~160 GB).
- A100 ridge points (dense FP16 312 TFLOP/s ÷ bandwidth): HBM 40GB 201, HBM 80GB 153, L2 43–62, L1/shared 16 FLOPs per byte; with sparsity (624 TFLOP/s) L1 becomes 32 — so your "15 vs 30" is right to ~6%.
- SRAM vs HBM energy crossover: ~67 mm of minimum-pitch wire vs HBM4 and 114–136 mm vs HBM3E, or ~1.3–5 GB of SRAM — larger than any die. No latency crossover on a die. ~3/4 of HBM's energy is inside the DRAM stack, so comparing against the PHY alone gives a false 5–17 mm.
- Taalas HC1 (the ChatJimmy chip; TSMC N6, 815 mm², 53B transistors): runs only Llama-3.1-8B at mixed 3/6-bit, weights in a mask-ROM "recall fabric", KV cache in on-die SRAM, 16–17k tokens/s per user, ~200–250 W per card (≤12–16 mJ/token single-user). It reads no weight bytes at all; context limit undisclosed. AMD agreed to acquire it Thu 6 Aug 2026 (week 32).
Corrections to your own claims.
- Your 10 mm SRAM-vs-HBM cutoff understates the energy crossover 7–14× (and it was your number, not Jason's).
- Groq's chip holds 230 MB of SRAM, not ~500 MB (if that line meant Groq).
- Cerebras WSE SRAM is 18 / 40 / 44 GB (WSE-1/2/3), never 10 GB — at lunch and in your "AI hardware companies" doc.
- FLOPs per HBM byte: 1,000 (Tue) → ~300 from your own self-factcheck, which is right for the H100 → applied to the A100 on Sunday, where dense FP16 is 201/153 → inverted to 300 bytes per FLOP at 19:06. Saturday's 1,000 cycles is a latency, a different quantity.
- MoE does not buy the width that the Allen-Zhu convergence result needs (that width is polynomial and an interpolation result).
People. Boris Ginsburg, Mark Saroufim, Jason Yosinski, Vijay Jain, Hasan Unlu (both heard the KV story). Sources: Reiner Pope, Daria Soboleva (Cerebras; MoE impractical on wafer scale, per the Hasan brief).
Outputs. Two public MoE shares (Wed 9 Sep), the cutoff report (uncommitted, Intel), and a ridge-point answer that exists only in a Codex chat.
Goal link (graded).
- Lighthouse — medium. It informs which memory level an entry must fit in. The MNIST tiers were actually set on Mon 7 / Thu 10 Sep from A100 cache capacities, before the ridge points existed.
- Career / hardware networking — weak. The KV-cache story went to Vijay Jain and Hasan Unlu, not to Cerebras or Mark Saroufim.
Filled in (survived verification).
- SRAM leakage — Jason Yosinski's breakfast question (confidence medium; the anchors are old low-leakage nodes, so cutoffs below are upper bounds).
- No foundry-measured per-bitcell leakage figure is public at 7/5/3 nm. Best-case anchors: Intel 22 nm low-power SRAM "10pA/cell standby" (~7.5 pW/bit, ~0.064 W/GB); 28 nm FDSOI simulated 13.3 pW/bit. DRAM self-refresh: LPDDR5 ~1.1 mW/GB typical.
- At chip level, your answer that leakage is negligible is probably right: 44 GB on a Cerebras wafer leaks ~3–5 W best case, ~28–280 W at 50–500 pA/cell, against ~15 kW.
- Per bit it is wrong: holding a bit in SRAM costs an extra 7–13 pJ per second vs LPDDR — more than HBM's whole 2–4 pJ fetch after one second. Break-even hold is ~0.08–0.43 s; for decode at 20 tokens/s the cutoff falls to ≤45–55 mm vs HBM4. If real hot-die leakage is ≥40 pW/bit, SRAM loses to HBM4 at any distance during decode.
Still open. HC1's SRAM size and context limit; hot-die leakage; HBM active-standby power.
Next step. Write one short "numbers I say out loud" note — Groq 230 MB, Cerebras 44 GB, cutoff 67–136 mm (upper bound), A100 153/201 vs H100 ~300 FLOPs per byte — and send Jason Yosinski the leakage answer.
Thread 5 — Backprop vs the memory wall → activations → reversible networks (the no-backprop bet)
Trigger. Your public post on X, Sat 31 Jan 2026 (week 5) 08:49 PST: "In five years, most learning applications will not use backprop." Yann LeCun replied Wed 11 Feb 2026 (week 7) 12:58 PST with one word: "False." The clock runs to Fri 31 Jan 2031 (week 5). Reignited when the Interlude hosts asked about it (Tue 8 Sep), followed by your 11:16 request to go deeper into backprop's interdependencies.
Timeline
- Sat 29 Aug 2026 (week 35) 12:42 — you retell the LeCun exchange to Massey Branscomb.
- Sun 30 Aug 18:02 — you commit to a "wrong abstractions" post (paired with Anthropic prep); 18:42 a 389-word outline titled "Backprop days are numbered".
- Mon 31 Aug 2026 (week 36) 10:30 — brief + research pack.
- Sat 5 Sep — research pass: attention forward FLOPs are equal; Rabe & Staats (Dec 2021) came before FlashAttention.
- Mon 7 Sep 2026 (week 37) — A100: forward+backward = 2.85× forward time, 2.74× energy (energy tracks operations).
- Tue 8 Sep 11:03–11:30 — interdependencies chat.
- Wed 9 – Thu 10 Sep — Gradient Dissent review of Cerebras's dropout paper; 15:20 Wed you link it to your question of whether backprop is needed; Thu 10:45 / 10:54 X posts on layer dropout; Mostafa Elhoushi (Cerebras) replied.
- Fri 11 Sep — the MNIST system spec's motivation paragraph is the backprop-memory-wall argument (commit 303bd33, 18:37); 21:10 a second-account claude.ai chat "Backprop memory wall analysis and batch size trade-offs" opens — likely the origin of the report (inferred from matching Drive upload times).
- Sat 12 Sep 07:48 report hosted on sutro-problems; 21:19 ChatGPT Backprop Memory Wall Analysis.
- Sun 13 Sep 08:34 yaroslavvb/backprop-memorywall public. 18:06–20:58 Intel Codex runs reversible nets on MNIST-medium, steered partly from your phone; at 18:37 the trainers still stored activations; at 18:51 you tell Geek Club "It's basically a single forward pass twice as long"; 20:29 reconstruction backward validated.
- Mon 14 Sep 2026 (week 38) 04:17 batching report; 11:34 Telegram to the Sutro Group: reversible nets "don't need to store activations". Bundle still uncommitted.
What you learned.
- Machine balance went from 0.05–0.25 FLOPs per byte (1986) to 150–600 (2026 accelerators); a weight must be visited by ≥ ~300 tokens (A100) or ~600 (H100) to use the compute.
- The tape: backprop stores ~34·d·L·T bytes and outgrows weights + grads + Adam at ~6·d tokens per device (24,200 for 7B, 50,700 for 70B). Depth cancels in that ratio; backprop's own share is the factor L.
- The tape wall is capacity and dollars, not energy (derived, at d = 4,096): all activation traffic ≈ 4.5% of arithmetic energy, the backward part ≈ 2.9%, the saved-tensor round trip ≈ 1.5% (24% / 16% / 8% at d = 768). Static leakage while compute waits (18–49% of step energy) is excluded.
- Reversible MNIST-medium: 97.42% vs the CNN's 97.70%; activation workspace flat at 19.8 MiB from 4 to 32 blocks; single-batch peak −17% at 4 blocks (−11% at 32); end to end only −3.8%, because 198 MiB of precomputed augmentation schedules dominate. Reversibility removes the depth factor, not the batch factor.
- Backprop is the optimal exact dynamic program on a path graph; the problem is large separators (dense layers) × long chains. Residuals are the existing approximate depth decomposition.
- Checkpointing's optimal segment for transformers is √(L/17) ≈ 1–2 layers, not √L.
Corrections to your own claims.
- LeCun's reply is "False.", not "This is false" (your research pack wrongly says "separately confirmed" at line 17, and repeats it at lines 200 and 229). The gap was ~7 months, not "eleven". Your claim is "most learning applications" — a majority, not elimination.
- The brief's FlashAttention sentence (backward pass) stands but omits Rabe & Staats (Dec 2021) and Milakov (2018). Your Fri 4 Sep remark implied more attention FLOPs: forward FLOPs are equal; +12.9% end to end buys 9.2× less HBM I/O.
- Your line that an agent solves sparse parity with Gaussian elimination: the public task (18 examples, 32 bits) is underdetermined.
- Reversible "twice as long" is really ~2.7× the CNN's draw time (35.5 s vs 13.1 s). And "don't need to store activations" wasn't true yet at 18:51 Sunday, and still overstates the end-to-end result when you repeated it to the Sutro Group at 11:34 today.
- The batching report mislabels 2.85× as a "backward-to-forward energy ratio": it is (forward+backward)/forward time; energy was 2.74×; backward alone ~1.85×.
- Your public brief files batching under "examples that don't hold up"; the 14 Sep report makes batching × depth a pillar. The published page now contradicts your newest argument.
- The brief predicted the activation share would be "well above half" on long context; the derived figure is 4.5% at d = 4,096. A measurement would settle it.
People. Yann LeCun, Tim Salimans (Anthropic; the planned reader), Vatsal Bajaj, Massey Branscomb, Boris Ginsburg (backprop is not the bottleneck, in Feb 2026), Natalia Vassilieva (pushed back in minute seven), Suhrud Kulkarni (asked what replaces it; you said "I actually don't know."), Mostafa Elhoushi, Joel Hestness. Sources: Reiner Pope, Vijay Korthikanti.
Outputs. The brief, the public repo, the sutro-problems page, gradient-dissent, the batching report, and an uncommitted Intel bundle (docs/mnist-reversible-memory-tutorial.md + 12-page PDF, docs/mnist-reversible-a100-memory.md, cache-fit and A100-memory docs in ~/git/sutro-problems).
Goal link (graded).
- Lighthouse bet — medium. Reversible nets still train with backprop; the result supports the activation-memory premise, not most learning moving off backprop in five years. The batching report supplies one of five pillars.
- Anthropic — medium (higher than it looks). Your week-36 plan and the 10–11 Sep contact maps name this post as the artifact to send Tim Salimans; no draft beyond the outline and the insertable section exists.
- Sutro competition — medium. It is the MNIST spec's stated motivation.
Filled in (survived verification).
- The bet, dated: your post Sat 31 Jan 2026 (week 5) 08:49 PST (35.7K views); LeCun's reply Wed 11 Feb 2026 (week 7) 12:58 PST; deadline Fri 31 Jan 2031 (week 5) (confidence high: post IDs decode to the same times). Checkable criteria: share of notable 2030 models trained without end-to-end backprop (Epoch AI + tech reports); MLPerf Training open-division entries (the closed division fixes the recipe); Hugging Face model cards (likely mostly "unspecified").
- Has anyone measured the energy split? (confidence medium.) No paper, 2020–Sep 2026, splits a training step's energy into activation traffic, weight/optimizer traffic, communication and arithmetic. Closest: Ivanov et al. (time, not energy), Korthikanti et al. (memory and time), Zoubeirou a Mayaki (arXiv 2606.23546, regression proxy), Abdulah et al. (arXiv 2605.01938, GH200 offload: +1.0% to +16.5% efficiency). A protocol that would: count traffic bytes per kernel; difference the NVML energy counter over ≥5 s windows (32+ iterations, 8 phase shifts); break the FLOPs–traffic collinearity with eager vs torch.compile, bf16 vs FP8 activations, and a width sweep (d = 384/768/1,536) — checkpointing and reversible blocks change held capacity, not traffic. About 5–7 A100-hours, ~$12–18 on Modal.
Still open. Replies under your X post; a finished "What's wrong with backprop" draft; the content of the second-account chat that likely wrote the report.
Next step. Commit the reversible bundle and say it honestly at Sutro tonight: flat 19.8 MiB workspace, 3.8% end to end until augmentation is streamed.
Thread 6 — Hardware landscape & co-design people (chip startups, photonics, analog, SPC)
Trigger. Devrim Yasar named Apex Compute on a chance street meeting Fri 4 Sep 2026 (week 36) 20:54–21:17 (with Maria Dubrovskaya) and at Sutro #30 offered a chip-friend intro (he described that friend as in Montréal, so "that was Hasan Unlu" is an inference). Root: your email to Thomas Dybdahl Ahle Wed 26 Aug 2026 (week 35), answered in 94 minutes with Valiant's work, the cell-probe model and compute-near-memory.
Timeline
- Mon 7 Sep 2026 (week 37) 08:02 — call with Thomas Ahle (analog MNIST in SPICE, reversible nets).
- Tue 8 Sep 10:00 — thomasnormal/spicenn2 v0; 18:34 Hasan Unlu brief session.
- Wed 9 Sep 10:28–11:30 — AI-silicon landscape (uncommitted).
- Thu 10 Sep 09:55 SPC collaborators; 11:18 Vijay Jain brief session; 12:04 lunch with Vijay Jain and Ping He (record); 13:32 Gemini Notebook deck The Physics of AI + video; 14:36 desk chat with Christian Pehle; 15:09 reading page published.
- Fri 11 Sep 13:13–13:44 Apex FPGA demo (Hasan Unlu, Devrim Yasar); 14:43 you introduce Hasan to Nurcan Sönmez; 17:24–17:50 Suhrud Kulkarni and Sid Sethi (brief).
- Sat 12 Sep 07:57 Buzz Cai brief; 09:56 email to Hasan about AI Foundry.
- Sun 13 Sep 18:59 WhatsApp to Mitchell Nahmias: "Couple Optical computing people that are coming" (to Sutro).
What you learned.
- No frontier-silicon startup publishes an audited J/token. DensityAI, MatX, Etched and Taalas have never submitted to MLPerf; the one startup with third-party perf/W (Untether AI) went bankrupt. Open energy measurement exists — ML.ENERGY, MLPerf Power, IBM NorthPole's measured 0.024 J/token — but only for NVIDIA parts and IBM's prototype.
- Always ask batch size and wall-plug draw: Positron's headline fits anything from 7.1 to 0.22 J/token (~32×) depending on unpublished concurrency.
- Apex Compute: 4.5 W ÷ 15.63 tokens/s = 0.288 J/token vs 0.287 published — a 16 nm FPGA beating an 8 nm Jetson 3.2×; another configuration gives ≥0.375; what the 4.5 W includes is unstated. Silicon targeted "beginning of the next year".
- Optics wins off-board, not on-chip: the best shipping link at wall plug (Lightmatter, 4.6 pJ/bit, vendor figure) breaks even with 155 mm of on-chip copper; research transceivers break even at 4–8 mm.
- 45 → 7 nm: logic ~3× cheaper, SRAM ~1.3–7× (8 KB vs 1 MB arrays), DRAM access unchanged at ~1,300 pJ.
- The hole that fits Vijay Jain: room-temperature photonic-electronic EDA exists; a cryogenic co-simulation flow does not.
- SPC: among the 22 new-cohort members researched, public evidence shows no demonstrated backprop alternative; start with Imbert Yuyen Wang, Seemandhar Jain, Vijay Jain, Suhrud Kulkarni.
- spicenn2 analog MNIST: 10-class 18.60% at 27.09 µJ/image, 3-class 89.33% at 9.86 µJ — simulated joules, not comparable with A100 or grid units.
- Your own learning method, in your words (X, Thu 10 Sep 10:45): "Listen to 10 minute summary, then give agents GPUs".
Corrections to your own claims.
- "100 square centimeters would be 10 gigawatts" (Fri 11 Sep 13:32) is 10 kW. "H100 is 1,000 times more than the sun" compares with sunlight at Earth; the Sun's surface is ~6,300 W/cm².
- Multiply/add ratio is 1.9–49× depending on number format, not 5–10×.
- Celestial AI is 120 ns vs ~80 ns local (1.5×), not latency parity.
- Hasan Unlu is ex-Tesla Autopilot, not Dojo; SPC's Dojo people are the DensityAI founders (Ganesh Venkataramanan, Bill Chang, Ben Floering).
- Vijay Jain's SPC bio mentioned a photonics company; at lunch he described none — exploring, not founding.
- The Physics of AI deck says "Agent-generated PTX outperforms human-written FlashAttention 3.0." Unsupported; the verified result is parity with FlashAttention-4. The deck also carries the uncorrected "arithmetic is essentially free" framing — don't send it to Vijay Jain as is.
People. Hasan Unlu, Devrim Yasar, Vijay Jain, Ping He, Christian Pehle, Liam, Suhrud Kulkarni, Sid Sethi, Buzz Cai, Thomas Dybdahl Ahle, Mitchell Nahmias (Sphere Semi), Nurcan Sönmez, Bill Chang, Imbert Yuyen Wang, Seemandhar Jain, Maria Dubrovskaya.
Outputs. The reading page; the deck (Drive) and video (M5 Downloads only); the Hasan–Nurcan intro; the AI Foundry email; Vijay Jain added to the Sutro invite.
Goal link (graded).
- Resonance — medium. Only Vijay Jain newly accepted the Monday invite; Christian Pehle and Bill Chang haven't answered; Hasan Unlu and Suhrud Kulkarni aren't invited.
- Competition credibility — low-medium. No vendor has been asked for a measured number, and your own verdict on Apex was skeptical.
- Career — speculative.
Filled in (survived verification).
- One energy ladder for moving a bit (confidence medium; vendors choose their own boundaries, so no ladder is truly like-for-like). pJ/bit → break-even distance vs 29.7 fJ/bit·mm on-chip copper:
- UCIe Advanced, 3 nm: 0.29 → 10 mm. UCIe Standard: 0.52 → 18 mm. NVLink-C2C (Grace-Hopper, not "NVLink 4" as your reading page labels it): 1.3 → 44 mm.
- HBM3E stack: 2.5–4.05 → 84–136 mm. DDR5: your documents say ~80 or 10–25.
- Celestial AI ~3.1 (with laser) → 104 mm; Lightmatter M1000 4.6 wall plug → 155 mm; Ayar Labs <5 → <168 mm (2023 figure); NVIDIA co-packaged optics ~5.6 (if per 1.6T port) → 189 mm; Broadcom Bailly 6.9 → 231 mm; linear pluggables ~9 → 303 mm.
- Verdict: optically pooled memory costs ≥ ~1.6–2.1× local HBM3E in energy. Optics extends how far memory can reach; it does not beat the memory wall.
- What a hardware energy entry would take (draft kit in the scratchpad, not sent). The README scores a full train+predict run (~1.17 GFLOP for the reference task), not one inference. Apex's public engine is inference-only, so the first question is whether it can run a backward pass at all and what its 4.5 W includes. A Jetson entry needs a port (no NVML on Tegra; use the module power rail). ET-SoC-1 board-pool access was for hackathon participants. The kit mislabels the A100's NVML figure as chip-rail; it is card-level.
Still open. Awaiting them: Hasan Unlu, Nurcan Sönmez, Mitchell Nahmias. Owed by you: Suhrud Kulkarni's Slack "Hi!" (Fri 17:50), Vijay Jain's LinkedIn message (Sun 13 Sep 11:26, unread), Buzz Cai coffee today 13:00.
Next step. Reply to Vijay Jain's LinkedIn message with the reading page (not the deck) and invite Suhrud Kulkarni to Sutro tonight.
Thread 7 — Esperanto / AI Foundry / Ainekko: an open 7 nm manycore as a systolic-array case study
Trigger. Fri 11 Sep 2026 (week 37) 21:02 you asked an agent about Jason Yosinski's company (Apical); at 21:16 it offered an unconfirmed best guess that Apical merged with ex-Esperanto people. Saturday 08:56 you asked for Esperanto's trajectory, looked up Roman Shaposhnik, and met Jason at breakfast 10:04. Earliest root: you read Roman Shaposhnik's Mojo-vs-tinygrad tweet Sun 23 Aug 2026 (week 34).
Timeline
- Tue 25 Aug 2026 (week 35) — Dave Ditzel's "Esperanto's Odyssey" at Cool Chips – Hot Takes (blog, Fri 4 Sep; slides).
- Sat 12 Sep 2026 (week 37)
- 09:48 the agent retracts the Anish Tondwalkar "unprogrammable" quote (it was about Tenstorrent); 09:50–09:56 you tell Geek Club and Hasan Unlu that Numenta "merged with the other half of Esperanto".
- 09:59–11:29 breakfast (notes): maybe you can get an actual chip to optimise against.
- 11:44 you paste the pre-correction Esperanto summary to Jason on Signal.
- 11:53–12:05 AI Foundry deep research incl. Discord; trajectory note; network note.
- 12:16–15:11 reading the AI Foundry Substack ("The Next Thousand Chips", including one long visit).
- ~13:06 claude.ai chat
5a230605produces et_platform_overview.md (13:55); 13:08 pitch session; 13:48 ChatGPT Almaz history. - 14:01–14:11 "12sep26 - research Esperanto" doc; 14:09 Signal to Jason on bandwidth and "10% of TCO".
- Sun 13 Sep
- By 12:00 the doc gains the network note, a Gemini session and the claude2 memory-wall link; 12:36 Signal to Jason: programmable cores only pay off on work that can't become a giant matmul or systolic array.
- 16:49 Ditzel keynote page; by 17:17 the doc gains the keynote and both decks; 17:17–17:27 "Esperanto as systolic array" (shared 17:54); 17:20 you send Jason Ditzel's retrospective.
- 17:40–17:50 walk: Signal to Andy about Ainekko's founders; 18:55 Geek Club: "Numenta, with a new name, is raising for a chip".
- Mon 14 Sep 2026 (week 38) 04:17 "Roman Shapovalov" is Roman Shaposhnik, DM drafted, not sent; 08:27 you list "Numenta has a new sparse chip project" to Fatemeh among companies to consider.
What you learned.
- ET-SoC-1 vs A100 (1 GHz design point): ~24B vs 54.2B transistors, ~35 vs 312 TFLOPS FP16, DRAM bandwidth 137 vs 2,039 GB/s, ~20 W vs 400 W. Your takeaway: half the transistors, similar FP32, ~10% of FP16.
- Shipped silicon ran at half speed: Ditzel's slides say "ET-SoC-1 silicon had a timing bug that cut its clock rate in half." At 600 MHz it is 10 FP32 TFLOPS — so "similar FP32" is ~2× optimistic.
- Near-threshold voltage was the product: 164 W at 0.75 V vs ~20 W at 0.4 V; ~0.14 pJ per int8 op system-wide vs ~15 fJ for the multiply itself, so ~90% of energy is movement and control (derived estimates).
- Yes, it can run as a systolic array at tile granularity (TensorSend/Recv, a rendezvous with fused reductions). The mesh can't create DRAM bandwidth but lets 32 scratchpads act as one ~80–128 MB buffer; cooperative 4×8 blocking gives ~5.3× on matmul. Batch-1 decode is still floor-limited; the best measured community result is Llama 3.2 1B at 17.87 tokens/s.
- Ainekko (US company, 2023) announced buying Esperanto's IP and selected assets Wed 19 Nov 2025 (week 47). CPU RTL is public (Solderpad 2.1); the NoC, PCIe and DRAM controllers are not.
- Warm paths: Dan Stolyarov (Ainekko COO) via Bojan Bostjancic; Jayesh Iyer (ex-Esperanto chief architect, now Apical) via Jason; Andy on Signal: "very very weak first degree, but very strong (both legs) second degree".
Corrections to your own claims.
- Electricity at 10% of TCO, so perf/W doesn't sell, is not Esperanto's reason. No Esperanto or press source says it; it is your own summary (a Fri 28 Aug 2026 (week 35) ChatGPT prompt, the Sat 14:09 Signal, the Sun chat). The CEO blamed staff poaching at up to 4× pay and buyers who stopped caring: "if there's a power budget that's unlimited, then energy efficiency doesn't really matter" (EE Times, Fri 4 Jul 2025, week 27).
- Half of Esperanto merging with Numenta is undocumented. It started as the Fri 21:16 agent guess. Only Jayesh Iyer is confirmed at Apical; Raj Khanna is probable. Esperanto wound down and its IP went to Ainekko. You have told Geek Club, Hasan Unlu and Fatemeh.
- The Saturday 11:44 summary to Jason (≈14 employees, burn estimate, Anish's quote applied to Esperanto, "Apache 2.0") was never explicitly corrected.
- et_platform_overview's acquired-Oct-2025, Apache-2.0 story: announced 19 Nov 2025; hardware is Solderpad; Apache applies at most to some software repos.
- The IP was not sold to Russians: Ainekko is a US company; Almaz is a network overlap, not the buyer.
- "You have 25 times less bandwidth" — your own numbers (3 TB/s vs 200 GB/s) give 15×.
People. Jason Yosinski, Dave Ditzel, Tanya Dadasheva, Roman Shaposhnik, Dan Stolyarov, Bojan Bostjancic, Jayesh Iyer, Raj Khanna, Anish Tondwalkar, Christian Pehle, Hasan Unlu, Andy (Signal, surname unresolved), Allen Rush, Fatemeh.
Outputs. The Esperanto doc, the "espera" share, three reports (commit 83830a9), messages to Jason, Geek Club, Christian Pehle and Hasan, a drafted DM to Roman Shaposhnik.
Goal link (graded).
- Lighthouse — medium-low. It is the only open-RTL 7 nm manycore, a candidate real-silicon target if board access is granted. The NoC, PCIe and DRAM controllers are not in the released RTL.
- Resonance — medium. Warm paths into the AI Foundry community.
- Job fallback — new and unexamined. Numenta/Apical is now on your list.
- Anthropic — displaced. Saturday's 4-hour block became 1.5 h, and Sunday's 14:59 "plan week" slot went back to this doc.
Filled in (survived verification).
-
Why Esperanto failed — a verdict you could paste into the doc (confidence medium-high; Ditzel's spoken Q&A has no transcript):
Esperanto's CEO blamed rivals poaching staff at up to 4× pay, and a market that valued energy efficiency little while power budgets looked unlimited. Joseph Byrne (XPU.pub) added that a chip designed in the CNN era met transformers with about 1/25 of an H100's memory bandwidth, no FP8/FP4 and few customers. Ditzel's slides name no cause but disclose a timing bug that halved the clock rate, worked around by raising voltage. No source says power savings failed because electricity is ~10% of TCO; that was my own summary. Perf/W sells where kilowatts are the binding limit, and only for parts whose throughput and memory bandwidth are competitive.
-
Numenta / Apical (confidence medium). Numenta's LinkedIn says "Numenta has evolved into Apical Intelligence!" A private account (Kiel, Thu 21 May 2026, week 21) says Apical merged two companies under a new CEO; the second company is not publicly identified. No public filing for a new round; what Jason said about fundraising was private — don't repeat it. Suggested Geek Club line: Correction: Esperanto didn't merge with Numenta. Its IP went to Ainekko; Numenta has become Apical Intelligence, whose chief architect is ex-Esperanto.
- The systolic-array ratio, corrected (confidence medium). The NoC runs on its own PLL at 400/500 MHz, not the 600 MHz core clock, so "~77 GB/s per link" is unsupported. The link width isn't public (512 or 1,024 bits), so a mesh hop-byte is ~7–23× cheaper in bandwidth than a DRAM byte, not 20–30×. The 5.3× blocking gain stands. NoC power: 1.8 W on DLRM.
Still open. ET-SoC-1 link width; whether the board pool is still open; what Ditzel said in the Q&A.
Next step. Send Geek Club (and Jason Yosinski) the one-line Esperanto/Numenta correction before sending anything new.
Thread 8 — Cerebras wafer-scale dataflow (Natalia Vassilieva)
Trigger. Your WhatsApp dataflow question to Natalia Vassilieva (Cerebras VP, Field CTO for ML), Sun 6 Sep 2026 (week 36) 18:16, carried as a blocker in the 7 Sep sweeps. Earlier root: Cerebras skepticism at Science Night, Sun 23 Aug 2026 (week 34). You re-commissioned that exchange on Sun 13 Sep 12:13 before meeting her.
Timeline
- Sun 23 Aug 2026 (week 34) 20:08–20:23 — Science Night Cerebras exchange; someone (unattributed) asks for a Cerebras speaker.
- Fri 4 Sep 2026 (week 36) — arXiv 2609.05275 "Don't Drop Dropout" (Cerebras). Sun 6 Sep: Boris Ginsburg describes 64 KB per core (the public figure is 48 kB).
- Wed 9 – Thu 10 Sep 2026 (week 37) — gradient dissent prep, A100 depth-robustness runs, 13-slide review ("FLOPs, not bytes moved").
- Fri 11 Sep 18:03 — you ask for Cerebras-inspired numbers, a processor every 128 nodes, tapes at the bottom → PR #7 (48 KiB per tile).
- Sat 12 Sep 07:46 Telegram: a processor every 128 grid steps, Cerebras-like; 21:07 walk with Uliana Popov: ask the field CTO whether that model is right.
- Sun 13 Sep 11:39–12:20 brief + Science Night feedback; ~12:13–12:39 claude2 chat on CS-6 heat and 100 W/cm²; 15:53–16:31 meeting (notes, Russian); 16:33 check-in: "didn't feel a strong connect"; 20:05 / 21:40 Science Night asks again for Cerebras contacts.
- Mon 14 Sep 2026 (week 38) 10:02 — your feedback: the Natalia prep was too large (feedback).
What you learned (her description, checked against Cerebras's public papers and SDK).
- The wafer is a 2D grid of processing elements, each with 48 kB of local SRAM holding all its data and code, invisible to other elements; one clock cycle per neighbour hop; 1.1 GHz clock.
- Packets are 32-bit "wavelets" with a 5-bit channel tag, not the 64 bytes she recalled; 24 routing channels ("colors"); memory reads 128 bits and writes 64 bits per cycle.
- Tasks fire when data arrives: the channel (WSE-2) or input queue (WSE-3) picks the task. Programs commonly bind one channel per side, which is why her depends-on-the-side description feels right.
- Host I/O attaches on the East and West edges (~60 channels per edge), not one bottom edge.
- Inference keeps weights in SRAM as a pipeline of decoder blocks across wafer regions; training streams weights. The simulator and CSL SDK are free to download (whether the source is open was unclear in the conversation).
- Her pushback: the problem is low utilisation, not only backprop; loop transformers, dynamic depth, early exit and non-MoE sparsity fit the wafer. A model built for their hardware is "наша золотая мечта" — their golden dream. Cerebras looked at Numenta about a year ago and chose not to invest; their current bottleneck is kernel programmability; Mostafa Elhoushi saw your dropout repo; she'll be at NeurIPS Sydney.
- "Don't Drop Dropout": ~2,400 runs, up to 25% of non-embedding training FLOPs saved and 1.54× self-speculative decoding — but in FLOPs not wall clock, no seeds, no dense baseline at 8.2B.
- Science Night's "20% finished-system yield" was an analyst's model assumption, not a measurement.
Corrections to your own claims.
- Your claim that 100 W on 1 cm² is hotter than the Sun's surface: the surface radiates ~6,300 W/cm², so 100 W/cm² is ~1.6% of it. You repeated it from Friday despite asking about it at ~12:16.
- TSMC ReRAM shipping this year at 10× SRAM density was already refuted on Thu 27 Aug: in volume since 2022, ~3× denser at 28 nm. Repeated to Natalia at 15:55.
- The NeurIPS hotels-sell-out-before-acceptances line was yours (16:28), not hers — also misattributed in today's directions and tasks reports.
- The Science Night quip about mechanical engineering thrown at an electrical engineering problem was another Science Night participant, not securely Liz Stein.
- That Science Night was three weeks earlier, not two.
People. Natalia Vassilieva, Mostafa Elhoushi, Joel Hestness, Nolan Dey, Daria Soboleva, Liz Stein, Daisy Stanton, Subutai Ahmad, Boris Ginsburg, Uliana Popov.
Outputs. The brief and its feedback file, the gradient-dissent slides, the pitch-128 model.
Goal link (graded).
- Grid-model fidelity — medium, not strong. She corroborated the local-memory / neighbour-hop cost structure, but her verdict was only "Похоже, наверное, слегка" (similar, probably slightly). Every element is a processor, not one per 128×128 tile; I/O is on two sides. The 48 KiB was copied from Cerebras numbers in advance, so it can't validate itself.
- Resonance — medium. The Science Night speaker ask is an easy, open favour.
- Career — speculative. No ask was made, and you felt no strong connection.
Filled in (survived verification).
- Model 3 vs real Cerebras — six mismatches (confidence high): memory holds code and data (model 3: data only, and allows remote reads); hardware reads 128 / writes 64 bits per cycle (model 3: one access, stricter); 24 channels and free broadcast to any subset of 5 ports (model 3: neither); arrival-triggered tasks (model 3: pre-scheduled traces); East/West I/O (model 3: 250 bottom-edge tape ports); per-element processors (model 3: per tile).
- The two second-account chats, partly recovered from screen text (confidence medium). Chat
5a230605holds the capacitance passage (0.2 fF/µm; a 1-mm wire ≈ 200 fF ≈ 1,000 gate capacitances), your grid-density and tape questions (Thu 10–Fri 11 Sep), the ET-SoC-1 "~1,000 ops per DRAM byte" analysis, a Jetson Thor correction (FP8 dense ridge ~1,900, not 3,800), and a check that WSE-3's 44 GB holds less than one Llama-3.1-70B sequence at 500K tokens (164 GB). Chatcaa0ac16: the 100 W/cm² answer never reached the screen; the TSMC answer says Cerebras was allowed wires across the scribe lines, so a die-to-die crossing costs ordinary wire energy.
Still open. Whether Claude's heat answer exists; entry-side task selection in public docs.
Next step. One short email to Natalia Vassilieva: ask for the paper links she offered, propose a follow-up with Mostafa Elhoushi on a hardware-fit learning rule, and pass on the Science Night speaker request.
Thread 9 — Abstractions, evolvability and which hardware assumptions to freeze
Trigger. Fri 11 Sep 2026 (week 37) 15:51: the grid-vs-A100 mismatch made you ask why keep an abstraction at all, if you can tune for the A100 directly. Roots go back further than the thread map said: your Engineering Architecture Bibliography (last edited Thu 30 Jul 2026, week 31 — Clark, Sangiovanni-Vincentelli, Doyle, Kirschner & Gerhart's "Evolvability"), a ChatGPT "evolution of evolvability" chat Tue 19 May 2026 (week 21), the suboptimization knol (Wed 28 Jan 2026, week 5), and "Wrong abstractions" as item 1 of your letter to Ali (Fri 22 Mar 2024, week 12).
Timeline
- Sun 30 Aug 2026 (week 35) 17:51–18:02 — the wrong-abstractions argument (FlashAttention, NumPy) and a commitment to the post.
- Fri 4 – Sat 5 Sep 2026 (week 36) — you ask whether linear algebra is the wrong abstraction; the research pass says the defensible target is the NumPy/PyTorch array interface.
- Sun 6 Sep 18:08 — ChatGPT CMOS Survival Analysis.
- Mon 7 Sep 2026 (week 37) — the week plan advises parking the CMOS-in-ten-years chapter; 15:20–17:50 you read Horowitz's MICRO 2023 keynote and CMOS 2.0 anyway; 18:00 at Sutro #30 you state the meta-skill thesis (evolution of evolvability).
- Tue 8 Sep 19:18 check-in: the grid model is the part of the competition that doesn't change.
- Fri 11 Sep 16:02 ChatGPT share; 16:03 Ben Recht's syllabus; 16:05–16:19 claude.ai evolvability chat; 20:28 Matni–Ames–Doyle (arXiv 2401.15185); 21:27 knol: engineering architecture created.
- Sat 12 Sep 07:51–09:17 — 88-slide deck turned into a phone page.
- Sun 13 Sep 12:00–12:02 — on Discord you wonder whether abstractions are needed at all, and note a limit to what compilers can do.
What you learned.
- Why the grid and the A100 disagree: the single-core grid is a latency model; an A100 is a throughput machine that hides latency with parallelism. Wall-clock need not track the grid; energy (bytes × distance) should.
- Where to put the abstraction boundary: where the cost model lives. Hide threads and instruction selection; expose tiles, memory levels, bytes moved and precision; hand-write the 2–3 primitives that are ~90% of runtime; extract a language only after a second hardware target or third algorithm.
- Hardware-tuned code lasts about one generation: FlashAttention-2 ran at 35% of H100 max FLOPs until FlashAttention-3 reached 75% (arXiv 2407.08608).
- 2036 is still CMOS: ASML says ~95% of the systems it sold in 30 years are still active; 16 ideal stacked layers cut lateral distances ~4×.
- The durable cost model from the CMOS chat — compute + bits × (endpoint + κ·length) + idle — shares its link term with the pitch-128 spec, which drops compute and idle energy and has no layers.
- Rabe & Staats (Dec 2021) and FlashAttention (2022) used essentially the same tiling; one minimised footprint and got no speedup, one minimised memory accesses and got 2–4×: "the idea was not missing, the objective function was."
- MoE as suboptimization: the FLOP-ledger framing comes from the MoE chat, which also told you "Your document has this one slightly backwards" — MoE wins on the FLOP axis and loses on the bandwidth axis.
Corrections to your own claims.
- The claim that linear algebra discourages small-memory algorithms is attackable (BLAS-3, communication-avoiding algorithms); target the NumPy/PyTorch array interface.
- Say exact attention, not vanilla attention, needs more FLOPs only in the backward pass; forward FLOPs are equal.
- Horowitz: ISSCC 2014, not 2012 (said at least three times).
- The evolvability chat's vocabulary (SASS, SM, TMA, "warps", "sectors") is unverified — don't reuse it unchecked.
People (sources). Ben Recht, John Doyle, Nikolai Matni, Herbert Simon, Marc Kirschner, John Gerhart, Leslie Valiant, Mark Horowitz.
Outputs. The evolvability share, the knol (still two links), the deck page and its build recipe.
Goal link (graded).
- Lighthouse — medium. Deciding which assumptions to freeze is the teachable core of the 2044 lecture-notes goal (letter to Ali; Sutro #30).
- Competition design — medium. It justifies keeping the grid and proposes the rank-correlation test; nobody has adopted it.
Filled in (survived verification) (confidence medium; a synthesis of AI chats plus your reports, not run).
- Freeze the form of the cost, not its constants: locality (energy grows with distance), wire speed ~2.5 mm/ns, wire capacitance ~0.2 fF/µm, DRAM row cycle 45–50 ns, CMOS as the base.
- Parameterise: κ (20–100 fJ/bit·mm across sources, mostly conventions — Memory Wall itself expects a further ~2× drop from lower voltage by 2029–2035), layer count, processor placement, precision, link type.
- Keep the grid because hardware-tuned code lasts about one generation, and bytes × distance is a cost PyTorch-level interfaces can't express.
- The test that would falsify it: ≥8 variants per accuracy band, a Spearman threshold fixed in advance for grid energy vs NVML A100 energy, then re-rank on a second GPU. Current evidence: the one same-task pair (panel MLP) agrees on energy (−14.14% model vs −17.95% A100) and disagrees on time (+1.31% vs −20.34%) — the predicted pattern, not yet a test.
Still open. Nothing ties this to the Sutro spec in writing.
Next step. Paste a five-line conclusion (freeze / parameterise / test) into "knol: engineering architecture".
4. How the threads connect
- Thread 3 is the hub. Physics (1) supplied its constants; bytes-per-FLOP (4) informs its cache tiers; Cerebras (8) supplied the pitch-128 layout; evolvability (9) supplies its validation test; the backprop memory wall (5) is the MNIST spec's stated motivation.
- The podcast (1) → electricity (2). Your on-air 5% estimate and grid-ceiling argument set off the Fri 11 Sep electricity chat.
- The TCO paradox (2 ↔ 7) is resolved. The grid as the ceiling (your motivation) and electricity at 10% of TCO so perf/W doesn't sell (your Esperanto verdict) are both true — for different buyers (see the saved-watt analysis in thread 2).
- One heuristic crossed three threads. 1,000 FLOPs per HBM byte at the Mark Saroufim coffee (1) became ~300 for the H100 (your self-factcheck), then 153/201 for the A100 (4), which bounds which memory level an MNIST entry must fit in (3).
- Two conversations seeded several threads. The Boris Ginsburg hike (Sun 6 Sep) fed MoE and bytes (4), synchronisation in the cost (3) and the Cerebras core picture (8). The Jason Yosinski breakfast (Sat 12 Sep) fed Esperanto (7), the SRAM-vs-HBM cutoff (4), the grid-vs-A100 admission (3) and the TCO line (2).
- Backprop (5) ↔ KV cache (4). The batching report puts the tape next to your KV line: 7B at 4K context is 18.3 GB of tape vs 2.1 GB of KV; 70B at 128K is 2,921 GB vs 43 GB.
- Real silicon (6, 7, 8). Hasan Unlu's Apex engine and Esperanto's mesh both test whether routing can work around low DRAM bandwidth; Cerebras is the existence proof of a 2D processor grid.
- The same outreach ladder four times. Ross Pantone → Ronny Krashinsky → Bill Dally; Devrim Yasar → Hasan Unlu; Jason Yosinski → Apical; Natalia Vassilieva → Cerebras.
- The shared failure: numbers leave before the correction pass. Dally Heuristics' 5 µm on a public page; the pre-correction Esperanto summary (Jason); the Numenta merger (Geek Club, Hasan Unlu, Fatemeh); the chip-hotter-than-the-Sun comparison (Hasan Unlu, Natalia); ReRAM 10× (Natalia); a 6 GW chips-vs-power gap (Christian Pehle, Ameen Patel); reversible nets that store no activations (Sutro Group, today). The fact-check recipe you built on Tue 8 Sep works — it just runs after the sending, not before.
5. Connection to higher goals
| Goal | Threads | Strength | What's missing for the link to pay off |
|---|---|---|---|
| Lighthouse / Project Sutro — put learning on a physical footing; energy-efficient nanoGPT via MNIST; teach by 2044 | 3 (strong), 1, 5, 4, 9 (medium), 8 (medium), 2 (why-now), 7 (candidate silicon) | Strongest | One canonical constants table; the rank test under the spatial model; a correction pass on Dally Heuristics and Memory Wall; and a thesis sentence that survives your own week ("bytes and power delivery", not "joules") |
| Anthropic / half-time job decision — Anthropic as the gate, Vinci4D / Incept as fallback, $500K / $250K ask | 5 (the planned post for Tim Salimans) | Weakest | A draft of that post. Influence functions got 1.5 of 4 h Saturday and 0 Sunday. Today's 17:00 call with Qingqing Mao (Incept Labs) has no prep from this sprint. |
| Runway (~three months, stated Fri 28 Aug 2026, week 35) | none | Weak | No thread carries a revenue-linked ask. Apical and Cerebras (hardware-fit algorithms) are the two places where this work is the job pitch — neither has been asked. |
| Resonance / people | 6, 7, 8, 2, 3 | Medium | Replies: Vijay Jain (LinkedIn), Suhrud Kulkarni, Natalia Vassilieva, Ameen Patel, Geek Club correction; and the Monday room reviewing Andy Zhang's waiting PRs |
- Strongest link: the MNIST competition. It is public, has outside contestants (Andy Zhang, Alex Varga), and has a weekly room waiting on it — the person-attached pattern that makes your work ship.
- Weakest link: the economics. Your own research says electricity is ~7% of cost, the tape wall is ~1.5–4.5% of arithmetic energy (derived), and optics doesn't beat the memory wall. The mission sentence about legacy algorithms causing energy waste needs a capacity-and-power version before you pitch it to chip companies or employers.
- Coming up that touches these threads: Sutro #31 tonight (18:00); Thu 17 Sep 12:00 lunch on distributed AI boxes in homes; Fri 18 Sep 11:00 SPC energy forum with Andrew Cantino; Sun 20 Sep 19:30 Science Night (Caleb Boyd, Molten Industries); Tue 22 Sep 2026 (week 39) Bill Dally is back.
6. Cost side
Time. Your own hours on the energy/hardware sprint, Mon 7 – Mon 14 Sep 2026: about 35–40 h if Wednesday's dropout review counts, ~30–35 h if not. These are estimates from reports, meetings and agent logs — not the human-present measure in the protocol, and they overlap. Agent-active time on the ~30 core sessions was ~26 h (log timestamps, not hands-on time). For scale: your coaching floor is 5 focused h/week, and Brian Wang's budget for important-not-urgent work is 1 h/week rising toward 4.
| Day | Your hours (est.) | Agent-active h | What |
|---|---|---|---|
| Sun 6 Sep (wk 36) | ~3–4 | 2.0 | Boris Ginsburg hike; multiprocessor Grid VM |
| Mon 7 Sep (wk 37) | ~4 | 3.0 | Thomas Ahle call, podcast prep, A100 forward/backward, Sutro #30 |
| Tue 8 Sep | ~7–8 | 2.4 | Podcast, fact-checks, Mark Saroufim, dossier, Lucas Cassiano |
| Wed 9 Sep | ~1–2 (+4–7 dropout review) | 7.4 | Chip landscape; Cerebras dropout paper |
| Thu 10 Sep | ~7–8 | 7.4 | Vijay Jain, Christian Pehle, Ameen Patel, MNIST launch, v4 ISA |
| Fri 11 Sep | ~5 | 2.1 | Apex demo, Dally emails, MNIST block, electricity + evolvability chats, Suhrud Kulkarni |
| Sat 12 Sep | ~4–5 | 4.3 | Esperanto, Jason Yosinski, AI Foundry, KV cache, backprop memory wall |
| Sun 13 Sep | ~4–5 | 3.5 | Esperanto/Cerebras, Natalia Vassilieva, reversible nets |
| Mon 14 Sep (wk 38), to 12:30 | ~0.5–1 | — | Check-ins, this request |
What it displaced, as the reports recorded it:
- Tue 8 Sep: the planning block; the correction note to the hosts; the GPU MODE ask to Mark Saroufim. Headaches at 16:18, 18:02 and 20:44 after a 16-day headache-free run (a weak association with "over-energized" days, ~1.4× lift, p ≈ 0.19).
- Wed 9 Sep: the Anthropic application and the competition both got zero minutes; the dropout paper took the day.
- Thu 10 Sep: Vijay Jain's reading list was half delivered when the Claude quota ran out ~14:30.
- Sat 12 Sep: the Anthropic block got ~1.5 of 4 planned hours.
- Sun 13 Sep: zero Anthropic minutes; "plan week" at 14:59 went back to the Esperanto doc — the third day hardware research displaced what you had chosen. The Claude weekly limit ran out, blocking Sunday's scheduled passes until 22:00.
- Money: Intel API-equivalent usage for week 37 was ~$1,091 (a list-price equivalent, not a bill), 63% of it multi-agent fan-out.
- Follow-through: Andy Zhang's PRs waiting 9–11 days; the uncommitted bundle, 1.22 µm calibration, chip landscape, cutoff and batching reports; the open loops listed in threads 2, 6, 7 and 8.
- Body (correlation only): week-37 WHOOP recovery averaged 47% vs 68% the week before; the night into today was 4 h 41 asleep with 24% recovery (6 h 20 combining both wrists).
- Week 38 ahead: 16 h of Anthropic blocks (Wed 16, Thu 17) and a draft to Roger Grosse by Fri 18 Sep; Thursday's 12:00 lunch splits the Thursday block.
The counterweight is real: the week-37 retrospective found the diversions produced finished work, not drift — Memory Wall, Dally Heuristics, a public competition, the reversible result, the pitch-128 model. The question is whether each was worth what it displaced, not whether diversions are bad.
7. Sources & gaps
What was read, what failed, what is still unverified
- Read: the live Tickertape; Google Docs (Interlude, Saroufim self-factcheck, Fable calibration, Esperanto, Sutro internal Log, MNIST docs, knols, sprint docs); claude.ai and ChatGPT shares; Claude Code and Codex logs on both Macs; meeting notes for the Interlude Show, Mark Saroufim, Lucas Cassiano, Vijay Jain, Christian Pehle, Ameen Patel, Hasan Unlu, Jason Yosinski, Natalia Vassilieva, Science Night; cloud check-ins; comms on the Intel (Telegram, Slack, WhatsApp, Signal, Gmail); GitHub repos and PRs; Chrome history on both Macs and the phone; Drive files viewed this week; Notability exports; the reports and briefs linked above.
- The second claude.ai account is
yaroslavvb2@gmail.com(Chrome Profile 2 on the M5). Its chats redirected to /new in the profile used, so they are not captured:1e1826ed"Backprop memory wall analysis and batch size trade-offs" (Fri 11 – Sun 13 Sep; likely the origin of the backprop-memory-wall note), the "espera" chat after 17:54 Sunday,595a67bd,b80dacdc(SPC members for backprop alternatives),f3f3c46f(dropout paper review),9b5ef9e4,8234ddab(influence functions) and a claude.ai project01a08246. Chats5a230605andcaa0ac16were recovered only partly, from screen text — a fallback you asked this morning to use sparingly. - Not captured: three Gemini chats (capacitance scaling; NVIDIA financing risk; original MNIST results); the source lists of four NotebookLM notebooks; the Co-Design video (M5 Downloads only, not watched); YouTube watch history (whether you watched the Ditzel talk is unknown); ChatGPT desktop use; the SPC Notion energy-forum notes and the GPU MODE "Energy cost of GPU operations" Discord thread; LinkedIn (Vijay Jain's message); iMessage.
- Found late by the completeness pass and folded in above: the Notability "sutro ideas" page (AT² scoring, backprop tutorial outline), your NVIDIA keynote annotations (root Tue 16 Dec 2025), the Engineering Architecture Bibliography (root Thu 30 Jul 2026), 17 SPC #energy-forum messages, Sutro Telegram decisions, your Thu 10 Sep X posts, and Alex Varga's outside use of the cost model.
- Data quality: the Mark Saroufim notes stop at 15:04 (~21 minutes untranscribed); speaker labels are unreliable throughout (content-based attribution); the Mac Wispr day-dir export dropped most dictations Wed–Sun, so the M5's live database was read instead; the Messenger and LinkedIn collectors are broken.
- Unverified or weakly sourced: vendor energy figures (no matched boundaries); Rabobank on-site gas costs (search excerpts); IEA and LBNL figures checked via snippets after 403s; Taalas SRAM size and context; ET-SoC-1 NoC link width; Numenta/Apical beyond Jayesh Iyer; whether you attended SPC energy forum #2; per-day hours (estimates).
- Intel-only, uncommitted: the reversible-net bundle, the 1.22 µm calibration, the chip-landscape, cutoff and batching reports.
- Housekeeping:
_service/service-chrome.mdstill says Profiles 1–3 went idle in June (Profile 2 is active). A mis-built agent command left a harmless/tmp/ap.jsonon the Intel. Scratchpad intermediates for this report contain pasted API keys and must not be published. Nothing was sent, edited or committed.
Total life satisfaction
This sprint was learning in its best form for you — a question you cared about, numbers that nerd-sniped you, and a stream of real people across the table — and on Sunday at 17:50 you said "I feel unusually good right now." The cost showed up where it always does: in the goal that has no person attached (Anthropic), in headaches on the most over-energized day, and in numbers that reached people before the fact-check did. Box it to Monday's room and one evening, attach one ask to one person, and the same curiosity that made this week enjoyable can also move the decisions it has been displacing.