Gyroscope report · Lighthouse / Sutro

Energy learning sprint — nine threads, what they produced, what they serve

Covers Thu 27 Aug – Mon 14 Sep 2026 (weeks 35–38), centred on Mon 7 – Mon 14 Sep 2026 (week 37). Generated Mon 14 Sep 2026 (week 38), 12:30–13:30 PDT, on the M5, from the live Tickertape, the linked Google Docs, claude.ai / ChatGPT shares, agent logs on both Macs, meeting notes, check-ins, comms, GitHub and the reports in ~/git0/gyroscope. Each gap was researched, then attacked by a separate verifier; only what survived is here.

1. Bottom line

This was one question pursued from four sides: is moving data, not arithmetic, what makes learning expensive — and can you build a scoring machine that rewards algorithms for moving fewer bytes? You fact-checked your own podcast claims the same day you made them (Tue 8 Sep 2026, week 37), turned the cost model into a public MNIST energy competition (Thu 10 Sep 2026, week 37), tested it against real chips (Esperanto, Cerebras) and against grid economics, and got the first measurement behind your no-backprop bet (reversible nets, Sun 13 Sep 2026, week 37). Your own week also shifted the thesis: the backprop tape wall is a capacity-and-dollars problem (~1.5–4.5% of the arithmetic energy at 7B width — derived, not measured), electricity is ~7% of an AI data centre's cost, and perf/W only decides purchases where power delivery is the binding limit — so "bytes and power delivery" is the framing that survives, not "joules". It serves the Lighthouse/Sutro goal strongly and the Anthropic / half-time-job decision and the runway hardly at all: no thread carries an ask, and hardware and paper research displaced the Anthropic block on Wed 9, Sat 12 and Sun 13 Sep. The decision it poses: box the energy work to Monday's Sutro room plus one evening so the Wed 16 / Thu 17 Sep (week 38) Anthropic blocks happen — or make it the job pitch by attaching one concrete ask to one of the people it produced.

2. The threads at a glance

# Thread Dates Main links Status Goal it serves
1 Podcast self-fact-check → memory-wall physics → Dally's heuristics Sun 23 Aug (wk 34) root; peak Tue 8 Sep (wk 37) Interlude doc · Fable fact-check · Memory Wall · Dally Heuristics Shipped; Dally Heuristics has 3 confirmed errors; correction note to hosts not sent Lighthouse (evidence for the cost model)
2 Why efficiency matters: grid power, electricity, the TCO paradox Thu 27 Aug (wk 35) – Sun 13 Sep (wk 37) Global electricity chat · PDF · GPU-hour note Open; article promised to Ameen Patel not seen sent Lighthouse "why now"; Resonance (SPC energy forum)
3 Dally grid model → MNIST energy competition Fri 28 Aug (wk 35) – Mon 14 Sep (wk 38) MNIST README · model 3 spec · redesign dossier Public since Thu 10 Sep; PR #70 conflicting; MNIST-medium release due tonight Lighthouse (first public learning benchmark)
4 Where bytes live: MoE, KV cache, FLOPs per byte, SRAM vs HBM Sat 7 Feb (wk 6) root; Tue 8 – Mon 14 Sep MoE chat · KV cache sizes · SRAM-vs-HBM cutoff Answered by agents; several spoken numbers wrong Lighthouse (which memory level an algorithm must live in)
5 Backprop vs the memory wall → reversible nets Sat 31 Jan (wk 5) bet; Sun 30 Aug – Mon 14 Sep Backprop brief · backprop-memorywall · batching report Reversible result uncommitted (Intel only) Lighthouse bet; the planned Anthropic artifact
6 Hardware landscape & co-design people Wed 26 Aug (wk 35) – Mon 14 Sep Chip landscape · reading page for Vijay Jain · SPC collaborators Several loops open in both directions Resonance
7 Esperanto / AI Foundry / Ainekko Fri 11 – Mon 14 Sep Esperanto doc · "espera" chat · AI Foundry research Outreach drafted, not sent; a wrong story is circulating Resonance; candidate real silicon
8 Cerebras wafer-scale dataflow (Natalia Vassilieva) Sun 23 Aug (wk 34) – Sun 13 Sep Natalia brief · notes · PR #7 Met; no follow-up sent Grid-model fidelity
9 Abstractions, evolvability, which hardware facts to freeze Thu 30 Jul (wk 31) root; Fri 11 – Sun 13 Sep Evolvability chat · CMOS survival · knol: engineering architecture Parked; conclusion lives only in chats Lighthouse (teachable core); competition rationale

Adjacent, not counted as energy threads: influence functions (the Anthropic project — doc, Sat 12 Sep) and the Gradient Dissent layer-dropout review of Cerebras's "Don't Drop Dropout" (Wed 9 – Thu 10 Sep), which touches threads 5 and 8.

3. Thread by thread

Thread 1 — Podcast self-fact-check → memory-wall physics → Dally's heuristics

Trigger. The Interlude Show recording, Tue 8 Sep 2026 (week 37) 09:00–10:14 at SPC, hosts Vatsal Bajaj and Jake Kang. On air you said "Maybe data centers were using 5% of the electricity." The thesis is older: Bill Dally lecture notes from Tue 16 Dec 2025 (week 51), a Dally chat started Sun 29 Mar 2026 (week 13), and the Science Night version on Sun 23 Aug 2026 (week 34) ~19:38, where moving a word was ~100× adding it.

Timeline
  • Thu 27 Aug 2026 (week 35) 13:52 — GPU-hour note (HBM pJ/bit, GPU latency ladder). Evening dinner report (after-math) corrects: register ~10×, DRAM ~6,400× an add (Horowitz 45 nm).
  • Mon 31 Aug 2026 (week 36) 10:30 — Backprop's Days Are Numbered brief (Dally 64,000×, Gholami gap).
  • Mon 7 Sep 2026 (week 37) 11:22–11:29 — Codex podcast brief warns that memory size alone does not give energy, and against the 420× claim.
  • Tue 8 Sep 2026 (week 37)
    • 09:00–10:14 recording (notes; the filename misdates it "08sep28").
    • 10:58–11:12 Gemini Notebook decks The Algorithmic Reset and Beyond Backpropagation, plus a 5-minute video.
    • 11:03 Fable fact-check chat starts (share); 11:17 Codex fact-check committed; 11:26 Fable page live; 11:35 posted in your #chat-yaroslav tagging both hosts (Fable page only).
    • 11:38 "08sep26 - interlude show" doc created; the capacitance passage is in by 13:11.
    • 13:04 Memory Wall and 13:44 Dally Heuristics published (unhashed URLs), both from Fable chat 5a230605 on the second claude.ai account.
    • 14:32–15:25 coffee with Mark Saroufim; 16:02 self-factcheck doc.
    • 16:37 you note that both Fable reports need fact-checking; 16:43 Intel Codex physical calibration (1.22 µm pitch, untracked); 17:08 "fable physical Dally calibration" doc — a partial fact-check.
    • 17:41 you re-teach wire capacitance to Lucas Cassiano.
  • By Sun 13 Sep 2026 (week 37) 16:43 — tailnet Memory Wall link and the chat link added to the Interlude doc.

What you learned.

Corrections to your own claims.

People. Vatsal Bajaj, Jake Kang, Mark Saroufim, Lucas Cassiano, Uliana Popov (you sent her the fact-check Tue 14:12). Sources: Bill Dally, Mark Horowitz, Amir Gholami.

Outputs. Fable fact-check · Codex fact-check · Memory Wall · Dally Heuristics · two decks and a video (Drive, 08sep26 day dir) · the recipe _request/request-claimresearch.md (written by the Codex agent in the 11:00 session) · the calibration doc.

Goal link (graded).

Filled in (survived verification).

Still open. The rest of Memory Wall's history and roadmap claims; whether the H100 carries HBM2e or (as Memory Wall says) HBM3; whether the episode has aired.

Next step. Send Vatsal Bajaj and Jake Kang a five-line correction note before the episode airs: NVLink not HBM; Horowitz 2014; name the source of each ratio; TPU was a forecast; "5%" is the US, not the world.


Thread 2 — Why efficiency matters: grid power, electricity economics and the TCO paradox

Trigger. Your on-air 5% estimate and the argument that grid capacity becomes the ceiling. Roots: Thu 27 Aug 2026 (week 35) 12:10 you told Nikita Kopotun "half of the cost of GPU runtime is energy", then caught it yourself by 14:36 with the GPU-hour note.

Timeline
  • Thu 27 Aug 2026 (week 35) 13:52 — GPU-hour note: electricity ~7% of TCO (Epoch 1 GW GB200 model). Mon 31 Aug (week 36): answered requests reconcile "half" vs 26% vs 7–12% (three denominators).
  • Tue 8 Sep 2026 (week 37) 09:51 on air; 11:38–13:45 grid notes in the Interlude doc; 12:33 you open the Introl article on PJM's 6 GW shortfall and WhatsApp it to Eric Munsing at 12:37.
  • Wed 9 Sep 14:30 — you create the Thu 17 Sep lunch invite on distributed AI boxes in homes. 17:00 SPC energy forum #2 (your attendance not established).
  • Thu 10 Sep 11:02 — side chat before the SPC all-hands: electricity is ~10% of TCO; 11:17 the energy forum is announced. 14:46:03 you reopen the Introl article and one second later tell Christian Pehle "In 2027, there's going to be a 6-gigawatt shortfall." 15:50 you tell Ameen Patel "6 gigawatt more chips than we have power to power them" and promise him the article (walk summary).
  • Fri 11 Sep 16:06–16:43 — "Global electricity usage and energy-efficient AI" chat, shared; 16:46 Introl reopened.
  • Sat 12 Sep 13:12 — global_electricity_and_efficient_ai.pdf (7 pp) in the day dir. 14:09 Signal to Jason Yosinski: "electricity is only 10% of TCO in datacenters".
  • Sun 13 Sep ~12:01 — you paste the same 10% line into claude.ai chat 5a230605; the partly captured reply says hyperscale is grid-power-limited, so perf/W buys capacity, not opex. ~21:51 at Science Night, Olia (surname unknown; identity from your 22:22 note) on methane from behind-the-meter gas-powered data centres.

What you learned.

Corrections to your own claims.

People. Ameen Patel, Christian Pehle, Nikita Kopotun (organises the SPC energy forum; accepted Thu's lunch), Eric Munsing, Austin Diamond, Olia, Jason Yosinski. Sources: Anders Andrae, Jens Malmodin.

Outputs. The chat share, the PDF, and the Thu 17 Sep 2026 (week 38) 12:00 lunch you organised (two invitees still pending).

Goal link (graded).

Filled in (survived verification).

Still open. Whether "1.4 TW total" meant installed capacity or queued generation; whether you attended energy forum #2; the paywalled SemiAnalysis model.

Next step. Before Thursday's lunch, send Ameen Patel the Introl link with one line: PJM's reserve shortfall, not chips waiting for power.


Thread 3 — Dally grid model → the MNIST energy competition (Sutro's scoring machine)

Trigger. Sutro planning sprint #2, Wed 2 Sep 2026 (week 36), floated "maybe MNIST or pattern recognition as an intermediate" goal; it became the goal through the Fri 4 Sep sanity check and Sutro #30 (Mon 7 Sep). The handwritten origin is your Notability page Notability.2026/Mine/sutro/sutro ideas.pdf (exported Tue 8 Sep 11:27): MNIST tiers at 3×3 / 9×9 / 28×28 and an "AT²" time-energy-area scoring sketch.

Timeline
  • Fri 28 Aug 2026 (week 35) 16:41–21:35 — NVML measurements on Modal A10G, B200, H100: the grid overstates the add/matmul gap ~7.5× (62.46× predicted vs 8.31× on B200).
  • Tue 1 Sep 2026 (week 36) 17:06–23:42 — first A100 runs (Intel Codex): 8192³ matmul FP32 15.76 J, FP16 1.43 J, INT8 0.904 J; a 1-bit AND+POPCOUNT kernel 0.12269 J vs grid 0.118953 J.
  • Wed 2 Sep — calibration against real silicon fails (24.4× die area, ~40× pitch); 12.82 J for MNIST-60k INT8 inference enters the "MNIST on grid model?" doc.
  • Thu 3 Sep 13:52–17:11 — M5 Codex settles c/160, 1 µm, 1 fJ/byte (posted to the Sutro Telegram 17:18).
  • Fri 4 Sep 16:40–17:06 — 9-agent sanity check: a 3×3 nearest-neighbour ceiling of 65.8%; recommends 8×8; not adopted.
  • Sun 6 Sep — Boris Ginsburg hike: add synchronisation to the cost; 16:51 multiprocessor Grid VM session.
  • Mon 7 Sep 2026 (week 37) 18:00 — Sutro #30: constants 1 µm / 1 fJ / 0.5 ps; A100 forward+backward is 2.85× forward in time, 2.74× in energy (the earlier "8×" was a BF16 artifact).
  • Tue 8 Sep 16:00–17:58 — redesign dossier (58 forks); 17:03 Telegram: "the task of the submitter will be to create an appropriate programming language".
  • Thu 10 Sep 16:48 Telegram: shift the IR choice to participants; 18:04–18:30 v4 tape ISA with a 50 fJ / 50 ps floor (PR #6 merged); 18:11 competition launched.
  • Fri 11 Sep 08:10 Telegram: "I'm finding that 48 is too low." 14:29 / 14:35 emails to Bill Dally ("What's a good way to add multicore?"; out until Tue 22 Sep) and Ronny Krashinsky (who had forwarded you to Dally on Sun 30 Aug 19:04 — not cold). 15:52 you flag the big single-core-grid vs A100 mismatch; 18:03–18:15 pitch-128 spatial computer (PR #7 merged).
  • Sat 12 Sep 04:06 answered requests: rank on grid energy; 07:34 PR #71 (Alex Varga) merged.
  • Mon 14 Sep 2026 (week 38) 04:09 — outstanding PRs: PR #70 conflicts; MNIST-medium release promised to Lucas Cassiano for tonight.

What you learned.

Corrections to your own claims.

People. Andy Zhang, Alex Varga, Lucas Cassiano, Mark Saroufim, Thomas Dybdahl Ahle, Boris Ginsburg, Cosmin Negruseri, Devrim Yasar, Bill Dally, Ronny Krashinsky, Ross Pantone.

Outputs. The competition (launched Thu 10 Sep 18:11), models 2 and 3, the multiprocessor page, the dossier, the Dally and Krashinsky emails.

Goal link (graded).

Filled in (survived verification).

Still open. Why PR #71's "no_slowdown" variant inverts (different CTA grids, inference batch 1 vs 30); re-scoring PRs #64–#66, #70, #71 under the spatial model.

Next step. Resolve PR #70's conflict, label its energy columns "single-core model", and release MNIST-medium at Sutro #31 tonight; on Tue 22 Sep send Bill Dally one constants table with the pitch-128 model.


Thread 4 — Where bytes live: MoE, KV cache, FLOPs per byte and the SRAM-vs-HBM cutoff

Trigger. The Boris Ginsburg hike, Sun 6 Sep 2026 (week 36): register 1 clock, SRAM ~10, HBM 500–1,000 (10:25), and, from Boris, that we are not limited by FLOPs (11:10, translated from Russian). On air Tue 8 Sep you said MoE lets you widen layers without paying the cost; at 11:58 you asked Claude about it. The "50 GB KV cache" line goes back to Mon 16 Feb 2026 (week 8).

Timeline
  • Tue 8 Sep 2026 (week 37) 12:02–12:16 two MoE chats (public, public2), 12:03 filed in knol: suboptimization; 14:41 at coffee, "1,000 FLOPs per HBM byte" to Mark Saroufim; 14:43 MoE is great for training, not inference.
  • Thu 10 Sep 12:49–12:50 at the Vijay Jain lunch: "10 gigabytes of SRAM is more than enough", "KV caches can be 50 gigabytes".
  • Sat 12 Sep 10:06 at breakfast with Jason Yosinski: a chip "can only do, like, 500 megabytes" (the next sentence names Groq); 10:22 paging from DRAM "like 1000 cycles" (a latency figure); fetching SRAM >10 mm away loses to HBM. 14:41–15:08 Claude Code traces the 50 GB story; 14:48 ChatGPT KV Cache Sizes; a second share at 15:08 holds the Taalas answer.
  • Sun 13 Sep 12:26 check-in asks for the cutoff; 18:44–20:10 Intel Codex sizes A100 memory levels; at 19:06 you invert the ratio to 300 bytes per FLOP; 19:16 you ask it to clarify sparse vs dense "15 vs 30".
  • Mon 14 Sep 2026 (week 38) 04:15 — SRAM-vs-HBM cutoff report.

What you learned.

Corrections to your own claims.

People. Boris Ginsburg, Mark Saroufim, Jason Yosinski, Vijay Jain, Hasan Unlu (both heard the KV story). Sources: Reiner Pope, Daria Soboleva (Cerebras; MoE impractical on wafer scale, per the Hasan brief).

Outputs. Two public MoE shares (Wed 9 Sep), the cutoff report (uncommitted, Intel), and a ridge-point answer that exists only in a Codex chat.

Goal link (graded).

Filled in (survived verification).

Still open. HC1's SRAM size and context limit; hot-die leakage; HBM active-standby power.

Next step. Write one short "numbers I say out loud" note — Groq 230 MB, Cerebras 44 GB, cutoff 67–136 mm (upper bound), A100 153/201 vs H100 ~300 FLOPs per byte — and send Jason Yosinski the leakage answer.


Thread 5 — Backprop vs the memory wall → activations → reversible networks (the no-backprop bet)

Trigger. Your public post on X, Sat 31 Jan 2026 (week 5) 08:49 PST: "In five years, most learning applications will not use backprop." Yann LeCun replied Wed 11 Feb 2026 (week 7) 12:58 PST with one word: "False." The clock runs to Fri 31 Jan 2031 (week 5). Reignited when the Interlude hosts asked about it (Tue 8 Sep), followed by your 11:16 request to go deeper into backprop's interdependencies.

Timeline
  • Sat 29 Aug 2026 (week 35) 12:42 — you retell the LeCun exchange to Massey Branscomb.
  • Sun 30 Aug 18:02 — you commit to a "wrong abstractions" post (paired with Anthropic prep); 18:42 a 389-word outline titled "Backprop days are numbered".
  • Mon 31 Aug 2026 (week 36) 10:30 — brief + research pack.
  • Sat 5 Sep — research pass: attention forward FLOPs are equal; Rabe & Staats (Dec 2021) came before FlashAttention.
  • Mon 7 Sep 2026 (week 37) — A100: forward+backward = 2.85× forward time, 2.74× energy (energy tracks operations).
  • Tue 8 Sep 11:03–11:30 — interdependencies chat.
  • Wed 9 – Thu 10 Sep — Gradient Dissent review of Cerebras's dropout paper; 15:20 Wed you link it to your question of whether backprop is needed; Thu 10:45 / 10:54 X posts on layer dropout; Mostafa Elhoushi (Cerebras) replied.
  • Fri 11 Sep — the MNIST system spec's motivation paragraph is the backprop-memory-wall argument (commit 303bd33, 18:37); 21:10 a second-account claude.ai chat "Backprop memory wall analysis and batch size trade-offs" opens — likely the origin of the report (inferred from matching Drive upload times).
  • Sat 12 Sep 07:48 report hosted on sutro-problems; 21:19 ChatGPT Backprop Memory Wall Analysis.
  • Sun 13 Sep 08:34 yaroslavvb/backprop-memorywall public. 18:06–20:58 Intel Codex runs reversible nets on MNIST-medium, steered partly from your phone; at 18:37 the trainers still stored activations; at 18:51 you tell Geek Club "It's basically a single forward pass twice as long"; 20:29 reconstruction backward validated.
  • Mon 14 Sep 2026 (week 38) 04:17 batching report; 11:34 Telegram to the Sutro Group: reversible nets "don't need to store activations". Bundle still uncommitted.

What you learned.

Corrections to your own claims.

People. Yann LeCun, Tim Salimans (Anthropic; the planned reader), Vatsal Bajaj, Massey Branscomb, Boris Ginsburg (backprop is not the bottleneck, in Feb 2026), Natalia Vassilieva (pushed back in minute seven), Suhrud Kulkarni (asked what replaces it; you said "I actually don't know."), Mostafa Elhoushi, Joel Hestness. Sources: Reiner Pope, Vijay Korthikanti.

Outputs. The brief, the public repo, the sutro-problems page, gradient-dissent, the batching report, and an uncommitted Intel bundle (docs/mnist-reversible-memory-tutorial.md + 12-page PDF, docs/mnist-reversible-a100-memory.md, cache-fit and A100-memory docs in ~/git/sutro-problems).

Goal link (graded).

Filled in (survived verification).

Still open. Replies under your X post; a finished "What's wrong with backprop" draft; the content of the second-account chat that likely wrote the report.

Next step. Commit the reversible bundle and say it honestly at Sutro tonight: flat 19.8 MiB workspace, 3.8% end to end until augmentation is streamed.


Thread 6 — Hardware landscape & co-design people (chip startups, photonics, analog, SPC)

Trigger. Devrim Yasar named Apex Compute on a chance street meeting Fri 4 Sep 2026 (week 36) 20:54–21:17 (with Maria Dubrovskaya) and at Sutro #30 offered a chip-friend intro (he described that friend as in Montréal, so "that was Hasan Unlu" is an inference). Root: your email to Thomas Dybdahl Ahle Wed 26 Aug 2026 (week 35), answered in 94 minutes with Valiant's work, the cell-probe model and compute-near-memory.

Timeline
  • Mon 7 Sep 2026 (week 37) 08:02 — call with Thomas Ahle (analog MNIST in SPICE, reversible nets).
  • Tue 8 Sep 10:00 — thomasnormal/spicenn2 v0; 18:34 Hasan Unlu brief session.
  • Wed 9 Sep 10:28–11:30 — AI-silicon landscape (uncommitted).
  • Thu 10 Sep 09:55 SPC collaborators; 11:18 Vijay Jain brief session; 12:04 lunch with Vijay Jain and Ping He (record); 13:32 Gemini Notebook deck The Physics of AI + video; 14:36 desk chat with Christian Pehle; 15:09 reading page published.
  • Fri 11 Sep 13:13–13:44 Apex FPGA demo (Hasan Unlu, Devrim Yasar); 14:43 you introduce Hasan to Nurcan Sönmez; 17:24–17:50 Suhrud Kulkarni and Sid Sethi (brief).
  • Sat 12 Sep 07:57 Buzz Cai brief; 09:56 email to Hasan about AI Foundry.
  • Sun 13 Sep 18:59 WhatsApp to Mitchell Nahmias: "Couple Optical computing people that are coming" (to Sutro).

What you learned.

Corrections to your own claims.

People. Hasan Unlu, Devrim Yasar, Vijay Jain, Ping He, Christian Pehle, Liam, Suhrud Kulkarni, Sid Sethi, Buzz Cai, Thomas Dybdahl Ahle, Mitchell Nahmias (Sphere Semi), Nurcan Sönmez, Bill Chang, Imbert Yuyen Wang, Seemandhar Jain, Maria Dubrovskaya.

Outputs. The reading page; the deck (Drive) and video (M5 Downloads only); the Hasan–Nurcan intro; the AI Foundry email; Vijay Jain added to the Sutro invite.

Goal link (graded).

Filled in (survived verification).

Still open. Awaiting them: Hasan Unlu, Nurcan Sönmez, Mitchell Nahmias. Owed by you: Suhrud Kulkarni's Slack "Hi!" (Fri 17:50), Vijay Jain's LinkedIn message (Sun 13 Sep 11:26, unread), Buzz Cai coffee today 13:00.

Next step. Reply to Vijay Jain's LinkedIn message with the reading page (not the deck) and invite Suhrud Kulkarni to Sutro tonight.


Thread 7 — Esperanto / AI Foundry / Ainekko: an open 7 nm manycore as a systolic-array case study

Trigger. Fri 11 Sep 2026 (week 37) 21:02 you asked an agent about Jason Yosinski's company (Apical); at 21:16 it offered an unconfirmed best guess that Apical merged with ex-Esperanto people. Saturday 08:56 you asked for Esperanto's trajectory, looked up Roman Shaposhnik, and met Jason at breakfast 10:04. Earliest root: you read Roman Shaposhnik's Mojo-vs-tinygrad tweet Sun 23 Aug 2026 (week 34).

Timeline
  • Tue 25 Aug 2026 (week 35) — Dave Ditzel's "Esperanto's Odyssey" at Cool Chips – Hot Takes (blog, Fri 4 Sep; slides).
  • Sat 12 Sep 2026 (week 37)
    • 09:48 the agent retracts the Anish Tondwalkar "unprogrammable" quote (it was about Tenstorrent); 09:50–09:56 you tell Geek Club and Hasan Unlu that Numenta "merged with the other half of Esperanto".
    • 09:59–11:29 breakfast (notes): maybe you can get an actual chip to optimise against.
    • 11:44 you paste the pre-correction Esperanto summary to Jason on Signal.
    • 11:53–12:05 AI Foundry deep research incl. Discord; trajectory note; network note.
    • 12:16–15:11 reading the AI Foundry Substack ("The Next Thousand Chips", including one long visit).
    • ~13:06 claude.ai chat 5a230605 produces et_platform_overview.md (13:55); 13:08 pitch session; 13:48 ChatGPT Almaz history.
    • 14:01–14:11 "12sep26 - research Esperanto" doc; 14:09 Signal to Jason on bandwidth and "10% of TCO".
  • Sun 13 Sep
    • By 12:00 the doc gains the network note, a Gemini session and the claude2 memory-wall link; 12:36 Signal to Jason: programmable cores only pay off on work that can't become a giant matmul or systolic array.
    • 16:49 Ditzel keynote page; by 17:17 the doc gains the keynote and both decks; 17:17–17:27 "Esperanto as systolic array" (shared 17:54); 17:20 you send Jason Ditzel's retrospective.
    • 17:40–17:50 walk: Signal to Andy about Ainekko's founders; 18:55 Geek Club: "Numenta, with a new name, is raising for a chip".
  • Mon 14 Sep 2026 (week 38) 04:17 "Roman Shapovalov" is Roman Shaposhnik, DM drafted, not sent; 08:27 you list "Numenta has a new sparse chip project" to Fatemeh among companies to consider.

What you learned.

Corrections to your own claims.

People. Jason Yosinski, Dave Ditzel, Tanya Dadasheva, Roman Shaposhnik, Dan Stolyarov, Bojan Bostjancic, Jayesh Iyer, Raj Khanna, Anish Tondwalkar, Christian Pehle, Hasan Unlu, Andy (Signal, surname unresolved), Allen Rush, Fatemeh.

Outputs. The Esperanto doc, the "espera" share, three reports (commit 83830a9), messages to Jason, Geek Club, Christian Pehle and Hasan, a drafted DM to Roman Shaposhnik.

Goal link (graded).

Filled in (survived verification).

Still open. ET-SoC-1 link width; whether the board pool is still open; what Ditzel said in the Q&A.

Next step. Send Geek Club (and Jason Yosinski) the one-line Esperanto/Numenta correction before sending anything new.


Thread 8 — Cerebras wafer-scale dataflow (Natalia Vassilieva)

Trigger. Your WhatsApp dataflow question to Natalia Vassilieva (Cerebras VP, Field CTO for ML), Sun 6 Sep 2026 (week 36) 18:16, carried as a blocker in the 7 Sep sweeps. Earlier root: Cerebras skepticism at Science Night, Sun 23 Aug 2026 (week 34). You re-commissioned that exchange on Sun 13 Sep 12:13 before meeting her.

Timeline
  • Sun 23 Aug 2026 (week 34) 20:08–20:23 — Science Night Cerebras exchange; someone (unattributed) asks for a Cerebras speaker.
  • Fri 4 Sep 2026 (week 36) — arXiv 2609.05275 "Don't Drop Dropout" (Cerebras). Sun 6 Sep: Boris Ginsburg describes 64 KB per core (the public figure is 48 kB).
  • Wed 9 – Thu 10 Sep 2026 (week 37) — gradient dissent prep, A100 depth-robustness runs, 13-slide review ("FLOPs, not bytes moved").
  • Fri 11 Sep 18:03 — you ask for Cerebras-inspired numbers, a processor every 128 nodes, tapes at the bottom → PR #7 (48 KiB per tile).
  • Sat 12 Sep 07:46 Telegram: a processor every 128 grid steps, Cerebras-like; 21:07 walk with Uliana Popov: ask the field CTO whether that model is right.
  • Sun 13 Sep 11:39–12:20 brief + Science Night feedback; ~12:13–12:39 claude2 chat on CS-6 heat and 100 W/cm²; 15:53–16:31 meeting (notes, Russian); 16:33 check-in: "didn't feel a strong connect"; 20:05 / 21:40 Science Night asks again for Cerebras contacts.
  • Mon 14 Sep 2026 (week 38) 10:02 — your feedback: the Natalia prep was too large (feedback).

What you learned (her description, checked against Cerebras's public papers and SDK).

Corrections to your own claims.

People. Natalia Vassilieva, Mostafa Elhoushi, Joel Hestness, Nolan Dey, Daria Soboleva, Liz Stein, Daisy Stanton, Subutai Ahmad, Boris Ginsburg, Uliana Popov.

Outputs. The brief and its feedback file, the gradient-dissent slides, the pitch-128 model.

Goal link (graded).

Filled in (survived verification).

Still open. Whether Claude's heat answer exists; entry-side task selection in public docs.

Next step. One short email to Natalia Vassilieva: ask for the paper links she offered, propose a follow-up with Mostafa Elhoushi on a hardware-fit learning rule, and pass on the Science Night speaker request.


Thread 9 — Abstractions, evolvability and which hardware assumptions to freeze

Trigger. Fri 11 Sep 2026 (week 37) 15:51: the grid-vs-A100 mismatch made you ask why keep an abstraction at all, if you can tune for the A100 directly. Roots go back further than the thread map said: your Engineering Architecture Bibliography (last edited Thu 30 Jul 2026, week 31 — Clark, Sangiovanni-Vincentelli, Doyle, Kirschner & Gerhart's "Evolvability"), a ChatGPT "evolution of evolvability" chat Tue 19 May 2026 (week 21), the suboptimization knol (Wed 28 Jan 2026, week 5), and "Wrong abstractions" as item 1 of your letter to Ali (Fri 22 Mar 2024, week 12).

Timeline
  • Sun 30 Aug 2026 (week 35) 17:51–18:02 — the wrong-abstractions argument (FlashAttention, NumPy) and a commitment to the post.
  • Fri 4 – Sat 5 Sep 2026 (week 36) — you ask whether linear algebra is the wrong abstraction; the research pass says the defensible target is the NumPy/PyTorch array interface.
  • Sun 6 Sep 18:08 — ChatGPT CMOS Survival Analysis.
  • Mon 7 Sep 2026 (week 37) — the week plan advises parking the CMOS-in-ten-years chapter; 15:20–17:50 you read Horowitz's MICRO 2023 keynote and CMOS 2.0 anyway; 18:00 at Sutro #30 you state the meta-skill thesis (evolution of evolvability).
  • Tue 8 Sep 19:18 check-in: the grid model is the part of the competition that doesn't change.
  • Fri 11 Sep 16:02 ChatGPT share; 16:03 Ben Recht's syllabus; 16:05–16:19 claude.ai evolvability chat; 20:28 Matni–Ames–Doyle (arXiv 2401.15185); 21:27 knol: engineering architecture created.
  • Sat 12 Sep 07:51–09:17 — 88-slide deck turned into a phone page.
  • Sun 13 Sep 12:00–12:02 — on Discord you wonder whether abstractions are needed at all, and note a limit to what compilers can do.

What you learned.

Corrections to your own claims.

People (sources). Ben Recht, John Doyle, Nikolai Matni, Herbert Simon, Marc Kirschner, John Gerhart, Leslie Valiant, Mark Horowitz.

Outputs. The evolvability share, the knol (still two links), the deck page and its build recipe.

Goal link (graded).

Filled in (survived verification) (confidence medium; a synthesis of AI chats plus your reports, not run).

Still open. Nothing ties this to the Sutro spec in writing.

Next step. Paste a five-line conclusion (freeze / parameterise / test) into "knol: engineering architecture".

4. How the threads connect

constants FLOPs per byte motivation rank test pitch-128 1,000→300 on-air 5% estimate suboptimization TCO paradox (resolved) intros real silicon mesh vs low DRAM bandwidth 1 · Memory-wall physics 4 · Where bytes live 2 · Grid & TCO economics 5 · Backprop → reversible 3 · Grid model → MNIST 9 · Abstractions 6 · Hardware people 8 · Cerebras dataflow 7 · Esperanto / Ainekko
Thread 3, the scoring machine, is the hub. Solid lines are what one thread fed into another; the dashed red line is the contradiction between your grid-ceiling motivation and your Esperanto verdict, resolved in thread 2. Scrolls sideways on a phone.

5. Connection to higher goals

Goal Threads Strength What's missing for the link to pay off
Lighthouse / Project Sutro — put learning on a physical footing; energy-efficient nanoGPT via MNIST; teach by 2044 3 (strong), 1, 5, 4, 9 (medium), 8 (medium), 2 (why-now), 7 (candidate silicon) Strongest One canonical constants table; the rank test under the spatial model; a correction pass on Dally Heuristics and Memory Wall; and a thesis sentence that survives your own week ("bytes and power delivery", not "joules")
Anthropic / half-time job decision — Anthropic as the gate, Vinci4D / Incept as fallback, $500K / $250K ask 5 (the planned post for Tim Salimans) Weakest A draft of that post. Influence functions got 1.5 of 4 h Saturday and 0 Sunday. Today's 17:00 call with Qingqing Mao (Incept Labs) has no prep from this sprint.
Runway (~three months, stated Fri 28 Aug 2026, week 35) none Weak No thread carries a revenue-linked ask. Apical and Cerebras (hardware-fit algorithms) are the two places where this work is the job pitch — neither has been asked.
Resonance / people 6, 7, 8, 2, 3 Medium Replies: Vijay Jain (LinkedIn), Suhrud Kulkarni, Natalia Vassilieva, Ameen Patel, Geek Club correction; and the Monday room reviewing Andy Zhang's waiting PRs

6. Cost side

Time. Your own hours on the energy/hardware sprint, Mon 7 – Mon 14 Sep 2026: about 35–40 h if Wednesday's dropout review counts, ~30–35 h if not. These are estimates from reports, meetings and agent logs — not the human-present measure in the protocol, and they overlap. Agent-active time on the ~30 core sessions was ~26 h (log timestamps, not hands-on time). For scale: your coaching floor is 5 focused h/week, and Brian Wang's budget for important-not-urgent work is 1 h/week rising toward 4.

Day Your hours (est.) Agent-active h What
Sun 6 Sep (wk 36) ~3–4 2.0 Boris Ginsburg hike; multiprocessor Grid VM
Mon 7 Sep (wk 37) ~4 3.0 Thomas Ahle call, podcast prep, A100 forward/backward, Sutro #30
Tue 8 Sep ~7–8 2.4 Podcast, fact-checks, Mark Saroufim, dossier, Lucas Cassiano
Wed 9 Sep ~1–2 (+4–7 dropout review) 7.4 Chip landscape; Cerebras dropout paper
Thu 10 Sep ~7–8 7.4 Vijay Jain, Christian Pehle, Ameen Patel, MNIST launch, v4 ISA
Fri 11 Sep ~5 2.1 Apex demo, Dally emails, MNIST block, electricity + evolvability chats, Suhrud Kulkarni
Sat 12 Sep ~4–5 4.3 Esperanto, Jason Yosinski, AI Foundry, KV cache, backprop memory wall
Sun 13 Sep ~4–5 3.5 Esperanto/Cerebras, Natalia Vassilieva, reversible nets
Mon 14 Sep (wk 38), to 12:30 ~0.5–1 — Check-ins, this request

What it displaced, as the reports recorded it:

The counterweight is real: the week-37 retrospective found the diversions produced finished work, not drift — Memory Wall, Dally Heuristics, a public competition, the reversible result, the pitch-128 model. The question is whether each was worth what it displaced, not whether diversions are bad.

7. Sources & gaps

What was read, what failed, what is still unverified
  • Read: the live Tickertape; Google Docs (Interlude, Saroufim self-factcheck, Fable calibration, Esperanto, Sutro internal Log, MNIST docs, knols, sprint docs); claude.ai and ChatGPT shares; Claude Code and Codex logs on both Macs; meeting notes for the Interlude Show, Mark Saroufim, Lucas Cassiano, Vijay Jain, Christian Pehle, Ameen Patel, Hasan Unlu, Jason Yosinski, Natalia Vassilieva, Science Night; cloud check-ins; comms on the Intel (Telegram, Slack, WhatsApp, Signal, Gmail); GitHub repos and PRs; Chrome history on both Macs and the phone; Drive files viewed this week; Notability exports; the reports and briefs linked above.
  • The second claude.ai account is yaroslavvb2@gmail.com (Chrome Profile 2 on the M5). Its chats redirected to /new in the profile used, so they are not captured: 1e1826ed "Backprop memory wall analysis and batch size trade-offs" (Fri 11 – Sun 13 Sep; likely the origin of the backprop-memory-wall note), the "espera" chat after 17:54 Sunday, 595a67bd, b80dacdc (SPC members for backprop alternatives), f3f3c46f (dropout paper review), 9b5ef9e4, 8234ddab (influence functions) and a claude.ai project 01a08246. Chats 5a230605 and caa0ac16 were recovered only partly, from screen text — a fallback you asked this morning to use sparingly.
  • Not captured: three Gemini chats (capacitance scaling; NVIDIA financing risk; original MNIST results); the source lists of four NotebookLM notebooks; the Co-Design video (M5 Downloads only, not watched); YouTube watch history (whether you watched the Ditzel talk is unknown); ChatGPT desktop use; the SPC Notion energy-forum notes and the GPU MODE "Energy cost of GPU operations" Discord thread; LinkedIn (Vijay Jain's message); iMessage.
  • Found late by the completeness pass and folded in above: the Notability "sutro ideas" page (AT² scoring, backprop tutorial outline), your NVIDIA keynote annotations (root Tue 16 Dec 2025), the Engineering Architecture Bibliography (root Thu 30 Jul 2026), 17 SPC #energy-forum messages, Sutro Telegram decisions, your Thu 10 Sep X posts, and Alex Varga's outside use of the cost model.
  • Data quality: the Mark Saroufim notes stop at 15:04 (~21 minutes untranscribed); speaker labels are unreliable throughout (content-based attribution); the Mac Wispr day-dir export dropped most dictations Wed–Sun, so the M5's live database was read instead; the Messenger and LinkedIn collectors are broken.
  • Unverified or weakly sourced: vendor energy figures (no matched boundaries); Rabobank on-site gas costs (search excerpts); IEA and LBNL figures checked via snippets after 403s; Taalas SRAM size and context; ET-SoC-1 NoC link width; Numenta/Apical beyond Jayesh Iyer; whether you attended SPC energy forum #2; per-day hours (estimates).
  • Intel-only, uncommitted: the reversible-net bundle, the 1.22 µm calibration, the chip-landscape, cutoff and batching reports.
  • Housekeeping: _service/service-chrome.md still says Profiles 1–3 went idle in June (Profile 2 is active). A mis-built agent command left a harmless /tmp/ap.json on the Intel. Scratchpad intermediates for this report contain pasted API keys and must not be published. Nothing was sent, edited or committed.

Total life satisfaction

This sprint was learning in its best form for you — a question you cared about, numbers that nerd-sniped you, and a stream of real people across the table — and on Sunday at 17:50 you said "I feel unusually good right now." The cost showed up where it always does: in the goal that has no person attached (Anthropic), in headaches on the most over-energized day, and in numbers that reached people before the fact-check did. Box it to Monday's room and one evening, attach one ask to one person, and the same curiosity that made this week enjoyable can also move the decisions it has been displacing.