Meeting brief · Sutro #31 · Mon 14 Sep 2026
Tonight 18:00–20:00, SPC, 380 Brannan. Covers Mon 7 – Mon 14 Sep. Three things to fix, in order: make the board usable, give Andy an answer, book an organizational sprint for tomorrow morning.
1 Bottom line
#30 agreed MNIST-as-transduction. Spec out Tue 8th, public Thu 10th, ranking moved to the pitch-128 multicore grid Fri 11th.
Every entry on the current spec is yours. PR #70 is conflicting, so
main's generator and evaluator still describe the old spec.
Planning sprint #2 stopped after five lines. Its one real sentence: “We need to listen to get more people involved.”
Recommendation for tonight: make the board work for a newcomer, then give Andy an answer and an owner role. Put a 90-minute organizational sprint on Tue 15 Sep so Wed 16 and Thu 17 stay Anthropic days.
2 Next ~100 minutes
The conflict is mnist/README.md only; the 5-line fix is in outstanding-prs.
Same commit: set accuracy_targets.json to match the README (still 60/98/98), and list
your 96.41% entry in the 8% and 12% bands. If only one gets done, do this one —
everything you show tonight depends on it.
One message. You promised him today.
About what you tested together in February.
Untracked in ~/git/sutro-problems. Label it honestly in the README:
activation workspace flat at 19.8 MiB, end-to-end saving 3.8%, because augmentation dominates.
3 Tonight, 18:00–20:00
The doc stub for #31 says only those three words. This fills them.
Moving a byte costs more than adding it: a 1 mm wire ≈ 100 adds, a DRAM read ≈ 64,000 (Dally, CACM 2022). The algorithms were designed for the opposite world. Goal: discover a learning algorithm with a small memory footprint.
MNIST public Thu 10th. Tiers: small 67%; medium 2/3/5/8/12% error bands; original 1%. Records: 4×4 matmul 1,316→675 (Juraj Selep); 16×16 66,300→63,819 (Alex Varga); MNIST-small −14% grid / −18% A100 energy (Alex Varga, PR #71).
Reversible nets keep activation memory flat — the whole run still saves only 3.8%. Say what the bottleneck turned out to be. Being honest here earns trust in the ruler.
Grid model or A100? In the one matched test (#71) the old single-core model got the A100 time ordering backwards, and pitch-128 has only 2 scored entries. Ask the room: what would make you trust the ruler enough to compete?
Thomas Dybdahl Ahle's spicenn2 shipped in 30 h; his video is Mon 21 Sep (#32). Cross-list as a track, or keep it separate? Its joules can't be compared with the A100 ladder.
Who wants to own something · what stopped you from submitting · which day should releases land on. These feed tomorrow's sprint.
Pair each newcomer with a runnable starter entry.
4 Decide in the room
matmul/?One ruling closes Andy's #49, #51, #52, #55, #56. Close #60 and #61 too — #61 would publish Telegram message IDs.
Pitch 1.0 / 1.2 / 1.22 µm and wire speed c/160 vs c/120 have all been used. Freeze one table; put the Manhattan-distance correction on the public design page.
Proposed: a spec or tooling version lands Monday before the meeting, never mid-week.
5 Say these out loud
6 Last week
7 People
| Person | Who they are | You owe |
|---|---|---|
| Andy Zhang | Most active collaborator; wrote the dally-eval Rust/GPU scorer. 8 open PRs. | The matmul/ ruling; approve dally-eval #1 and #2; reply to his 11:49 question |
| Lucas Cassiano | Wants to compete | The release — promised for today |
| Alex Varga | 16×16 record; MNIST-small PR #71 | Coffee you moved to “next week”, i.e. now. Book it |
| Hasan Unlu | Apex Compute (ex-Tesla Autopilot); arranged the FPGA demo | Invited verbally but not on the invite — add him |
| Juraj Selep | 4×4 matmul record (675) | Credit tonight |
| Thomas Dybdahl Ahle | spicenn2 analog MNIST; joined Telegram | Confirm the Mon 21 Sep video slot |
| Mark Saroufim | GPU MODE, megakernels — “I can try!” on a single-kernel MNIST | Check in; Boris Ginsburg ↔ Jared intro nudge |
| Jason Yosinski | Apical (formerly Numenta); possible real-chip access. Accepted | Correct the Esperanto story |
| Natalia Vassilieva | Cerebras; listed six ways model 3 differs from real hardware | Follow-up email (paper links, Mostafa Elhoushi) |
| Vijay Jain | Ex-PsiQuantum; asked to join, accepted | LinkedIn reply; reading page not sent |
| Suhrud Kulkarni | SPC; ex-TD Securities quant. Accepted | One bounded problem to own |
| Mitchell Nahmias | Sphere Semi; bringing optical-computing people | Ask who is coming |
| Bill Dally / Ronny Krashinsky | NVIDIA; multicore question sent | Resend Tue 22 Sep |
| Rohan Virani, Iván Bogatyy, Islam | Newcomers, added to invite/Telegram | Welcome by name |
| Cosmin Negruseri · Devrim Yasar · Christian Pehle | Competitive programmer · Lucky Robots · neuromorphic | Nothing outstanding |
“Accepted” on the calendar covers the whole recurring series, not tonight — twenty addresses show accepted, so it isn't a headcount.
8 The through-line
Promised a benchmark where the algorithm must be learned. Kept
Scorer and benchmark kept (scoring became the contestant's job). EqProp reading and Eric Frank's memo: no evidence
Spec kept, multicore decided, chip intro kept. Analog open, Andy's project not reviewed
The docs name this as the purpose of these evenings — and it's the test of sprint #1's “people love competitions”.
9 Direction
Sprint #1's cube still holds: move one axis at a time. Last week moved the problem axis a long way — sparse parity and matmul to MNIST. Move metric and process next, not problem.
Re-score #64–#66, #70 and #71 on pitch-128 and publish one frozen constants table. That work is also the substance of the Dally resend.
Make it work in 30 minutes, with a PR triage rule.
MNIST-small → medium → original, then sprint #1's target: energy-efficient training of nanoGPT.
Esperanto, Cerebras, optical — it produced people, but no competition entries.
The ruler isn't trusted yet at MNIST scale.
~3 months left. Keep Kateryna's split: Mon–Tue Sutro, Wed–Thu Anthropic, Fri reserve. Overflow goes to one job-shaped ask, not more research.
10 Next four weeks
| When | Milestone |
|---|---|
| Mon 14 Sep 18:00 week 38 · tonight |
#31 — board fixed, the three decisions, room input for the sprint |
| Tue 15 Sep 08:30–10:00 week 38 |
Organizational sprint proposed — the slot is free between “take second phone” and the 10:00 SymHub call |
| Wed 16 – Thu 17 Sep | Anthropic blocks. No Sutro work. |
| Fri 18 Sep 11:00 | Kateryna Peters #6, last paid session — report the Sutro/Anthropic split |
| Mon 21 Sep week 39 |
#32 — Thomas Ahle's analog video; proposed v1.1 release with frozen constants and re-scored entries |
| Tue 22 Sep | Resend to Bill Dally, with the re-scoring as evidence |
| Sun 27 Sep | End of your tooling freeze |
| Mon 28 Sep week 40 |
#33 — proposed first outside MNIST-medium entry: the success metric for the sprint |
| Mon 12 Oct week 42 |
proposed planning check — MNIST-original tier open, and decide whether the ruler is trusted enough for a text task |
11 Tomorrow, 08:30–10:00
Objective: an outsider can find, run, submit and get a review within one week, without messaging you.
| Min | Question | Output |
|---|---|---|
| 0–10 | Retrospective: what did sprint #1 produce, what stalled, why did #2 stop at five lines? | 3 bullets |
| 10–25 | Roles: who owns what besides you? | A roles table — ask each person, never assign |
| 25–40 | Review contract | “Every outside PR gets a human comment within 48 h; one ruling per category.” Turn branch protection on. |
| 40–55 | Release and meeting cadence | Release day; standing agenda; who takes notes and where |
| 55–70 | One channel of record | Announcements · design discussion · long-lived state. Retire or mirror the rest. |
| 70–85 | Timeline and success metrics | Outside entries per week; median time to first review |
| 85–90 | Send it | One paragraph to Telegram and Slack: the role asks, the release day, the review rule |
Spec/ruler: you · scorer & tooling: Andy Zhang · MNIST record steward: Alex Varga · analog track: Thomas Dybdahl Ahle · hardware-fidelity advisors: Natalia Vassilieva, Christian Pehle · newcomer onboarding: open, maybe Suhrud Kulkarni or Vijay Jain.
Success metric is an outside entry by Mon 28 Sep. If none arrives, change the process before changing the problem.
Separating organizers from participants is a split. Re-partition it at the October check.
Format: the one that worked in March — a written memo in a Google Doc with a fixed time box, plus tonight's three answers. Write into the “Sutro planning sprint #2” stub, or rename it.
12 Open loops
~/git/sutro-problems: reversible bundle,
ciresan_a100/, scoring-at-scale/, the four-port proposal13 Why it matters
The 16 Aug analysis of your check-ins found the March Sutro sprints were the most reliably satisfying work in the whole record. Tonight turns that into satisfaction that lasts if someone other than you gets on the board. The organizing is what protects Wednesday and Thursday for the job that funds all of it.
Rewritten for scanning from the brief “Sutro #31: last week's work, what to say tonight, and an organizational sprint” by @yaroslavvb, generated Mon 14 Sep 2026 ~15:20 PDT. Its own sources, all read that day: the Sutro Google Docs (top level, internal Log, #29, #30, planning sprints #1 and #2); recordings of #28 (24 Aug), #29 (31 Aug), #30 (7 Sep); cloud check-ins, Telegram/Slack/email threads and the calendar; live GitHub checked at 15:05; and the prior reports energy-learning-sprint, outstanding-prs, sutro-30-meeting and the week-37 weekly. Energy figures: Dally, CACM 2022. No numbers were added — items marked proposed are the original author's proposals.