Influence functions on the ET-SoC-1: one small step fits, none of the bill does

25 September 2026 · exploratory: estimates, not measurements · desk analysis built on numbers measured on the lab's cards (aifoundry2 and aifoundry3; the energy manual's catalogue, the gather/scatter rates and the host link also on aifoundry1's card 1); no card was run for this page · H100 and A100 figures are spec, published or estimated, as tagged · part of the ET-SoC-1 measurement reports

The question, from the author of the influence-functions report: which steps of large-scale influence-function work, the training-data attribution an interpretability team at a big lab runs, might this chip be uniquely suited to speed up? Perhaps finding the most influential training examples?

No step that matters to a big lab's influence bill suits the ET-SoC-1, and finding the most influential examples is not one either. Picking the top k is about 1.3×10⁵–1.5×10⁵ heap inserts per query even over 10¹⁰ candidates, free inside a GPU's scoring kernel. The cost lies in producing the scores: dense gradient work, on which one ET card is 21–35× slower than an H100 and spends 1.8–3.0× the energy on random data. Two narrow items survive, as exploratory research on the chip's handling of sparse and irregular work rather than as accelerators. S1: scoring a fixed shard of sketched training gradients held in the chip's scratchpads (72 MB usable of 80 MB), a few queries at a time, with the top k kept in the same pass. That spends 6–9× less board energy per pass than an H100 reading its HBM and 2.5–3.8× less than one holding the shard in its L2, and it is no faster. S2: an MNIST-sized influence pipeline run entirely on one chip, which I expect to lose to the author's measured A100. A third item was a measurement, since made (S3): the chip's gather, scatter and scatter-add rates, which every sparse idea here waited on. A scatter-add passes 10 G updates a second only while its buckets stay inside a shire.

How sure. The ET side scales one measured kernel: a 1024×4096 fp32 layer read from the scratchpads in 7.4 µs for 250 µJ of board energy (aifoundry3, one run), which the model below reproduces to within 9% above idle. The H100 side is spec, plus energies per byte measured on an A100; nobody has measured what an H100 spends on a 20 µs kernel, and that one number sets S1's lead. The verdict on value does not hang on either. The steps that survive are under a millionth of an amortized query; the one ET-shaped step with real size, scattering a gradient's sparse parts into a sketch, happens where the gradient is born, on the GPU. And S1's lead needs the card busy: at one query a second, its 25–36 W at rest cost 25–36 J per query, about 25,000–27,000 times the pass itself (the same card's idle on both sides). It breaks even with an H100 that is in the server anyway only above 2,900–20,000 queries a second, rates at which queries batch and the lead goes.

What would settle it. One new kernel on a lab card (section 5): score a static 65 MB int8 shard at 1 to 64 queries per pass with a fused top-k, next to the same scoring metered on a GPU. S1 is dead if one query costs 3 mJ or more of board energy, or if the GPU comes within 3×.

Share of an amortized query that S1 touches
about 10⁻⁹
the author's 8B model: 8.4 H100-hours per query, against about 23 µs of H100 time to score a shard
Dense gradient work, one ET card against an H100
21–35× slower
1.8–3.0× the energy on random data; this work is over 99% of the bill
S1, board energy per pass at one query
6–9× less
than an H100 reading HBM (2.5–3.8× if the shard is pinned in L2); no faster: 28 µs against 23 µs
Arrival rate S1 needs to pay for its idle card
2,900–20,000/s
queries a second, against an H100 already in the server; 390–5,000/s against the host CPU
Terms used on this page

The ET-SoC-1's compute cores are minions, small in-order RISC-V cores with a vector unit and a tensor unit, 32 to a shire; 1,024 of them run kernels. Each shire's SRAM holds a 2.5 MB scratchpad, software-managed memory any shire can address, of which about 2.25 MB is usable here: 72 MB on the chip. The tensor unit multiplies 16-row tiles (TensorFMA in fp32 and fp16, TensorIMA8A32 in int8) fed by TensorLoad; a hardware tree (TensorReduce) sums or takes the maximum across minions. More in the hub's glossary.

An influence function scores how much a training example moved a model's output on a query: the query's gradient, multiplied by the inverse of the curvature (an iHVP, here usually with EK-FAC, a Kronecker-factored approximation), dotted with the training example's gradient. A scan recomputes that gradient for each of N candidates; a sketch stores each one projected to k coordinates, so a query becomes k-dimensional dot products against an index. Recall then rerank: take the top k′ from the sketches, then score those exactly. Q is the number of queries scored in one pass over the data.

What the M/D/S/X/O/E evidence tags mean
How to read the numbers. Every figure carries a tag.
M
measured on an ET-SoC-1 lab card, in the linked reports. A measured statement that names no card holds on both aifoundry2 and aifoundry3 (the rule in the hub's §8); the gathers, scatters and packed atomics (E48) and the host link (E50) were measured on all three cards, aifoundry1's card 1 as well. A named card means that card only.
D
derived: arithmetic on measured numbers, nothing new measured.
S
vendor spec: datasheets and programmer's manuals.
X
published by others, measured on their hardware (for the GPUs, often an A100 standing in for an H100).
O
the influence-functions report's author: his measurements (the MNIST atlas, on an A100) or his 8B cost model.
E
estimate: an assumption of this page, never measured. Every conclusion below leans on at least one.

The numbers are computed by docs/reports/data/2026-09-25-influence-on-et/make_analysis.py, which reads the measured ones from the repository's data files and states the rest with their sources.

1. Where the cost of influence functions goes

The author's report splits the work into four phases paid on different schedules: fit the curvature once per checkpoint, precondition each query once (the iHVP), scan the candidates once per batch of queries, and validate once per method. At the shape of an 8B model with 10⁷ candidates O, one scan costs 702 H100-hours, because an unsketched scan is gradient recomputation: the dot products with 100 queries are 1.6% of it. Amortized, one query costs 8.4 H100-hours, 83% scan and 17% iterative iHVP. Sketching moves the scan into a one-off index build (1,150 H100-hours) and turns the lookup into a stream over the index; then the iHVP and, at big-lab scale, an exact rerank of the top candidates carry the bill.

What survives on the ET-SoC-1 sits at the cheap end: each step's cost in H100-hours, on a log scale, coloured by whether it suits the chip

Each dot is one step, costed per the unit in which it is paid (per query, per batch, per checkpoint), so the dots are not additive; the scale shows orders of magnitude. The two dashed lines are one amortized 8B query (8.4 H100-hours O) and a millionth of it. Tab to a dot, then use the arrow keys, for its size, what the ET-SoC-1 would do with it and why. Costs are the author's 8B model O except the 70B rerank E, the index stream (82 GB at the H100's 3.35 TB/s spec S), the prefilter (GPUSparse, 1.27 ms per query at batch 500 X), the sparse scatter (at most 1% of the scan's gradients E), S1 (the explorer's H100 pass E) and top-k (1 µs in a kernel epilogue E).

The same steps as a table
StepPaidH100-hoursOn the ET-SoC-1
Finding the most influential examples: why selection alone is worth nothing

Finding the most influential examples

Selecting the top k of N scores takes about k(1 + ln N/k) heap inserts: 1.3×10⁵–1.5×10⁵ per query for k = 10⁴ at N = 10⁹–10¹⁰ D. A GPU does that in the epilogue of the scoring kernel, or with a radix select for large k (FAISS, AIR top-k, RadiK X). The chip has a nice way to do it too: every minion keeps its own heap, and the running k-th-best threshold is shared chip-wide through the TensorReduce tree with a maximum instead of a sum, 2.3 µs for all 1,024 minions M. But there is nothing to win: the scores are what cost. In the two-stage pipeline Grosse described in April 2026 (secondhand, via the author's Bergson page O) and Bergson merged on 14 September, a sketch index gives the top k′ = 10³–10⁴, then each of those is re-scored with its exact gradient: at 70B that rerank is about 6 H100-hours per query E, over 99.9% of a query's energy. Selection is the one step of "find the most influential examples" that the chip is shaped for, and it is worth nothing on its own. It survives only folded into S1.

2. Explorer: an influence index against the chip

Set the size of a sketch index and how queries arrive. The capacity strip shows where the index fits; the chart shows the energy per query, board power included, on each device as the arrival rate changes. The ET card is charged its power at rest all the time, since it would be there only for this; the H100 and the host CPU are charged only for the pass, because they are in the server anyway (the H100 computes each query's gradient and iHVP, about 1.3 s of it at 8B O).

Index size
65.5 MB
16,000 codes × 4,096 × int8
ET-SoC-1, per query
25.1–36.3 J
at 1 query/s, idle included; 1 chip
H100 in the server, per query
8.3–9 mJ
the pass only: 23.1 µs from HBM
ET cheaper than the H100 above
2,900–4,500/s
queries a second (host CPU: 390–2,800/s)
Where the index fits: its size against each device's memory
Energy per query against how often queries arrive, board power included

Bands span the estimates' ranges: the ET card at rest draws 25.1–36.3 W M (aifoundry3 cool, aifoundry2 hot; the third card measured, aifoundry1's card 1, draws more at rest, 32–50 W at 56–81 °C, which would only raise the ET's band); the H100 60–90 W at rest and 105 pJ per byte from HBM or 38 from L2 E; the host CPU 40–200 W while it scores E. A band stops where its device is full at this Q; beyond it, more chips or a larger Q are needed. The shaded column is where the ET breaks even with the H100. Everything here is a model (section 6); only the ET's per-byte cost and bandwidth are measured.

3. What survived, ranked

Two candidates and a measurement survived a skeptical pass over two independent desk maps of the pipeline. None is a speed-up or a cut in a big lab's bill. They belong in the set's research and exploratory group, beside the sparse-compute measurements they build on: S1 as a hardware characterization ("energy per SRAM byte on an influence-shaped kernel"), S2 as a demonstrator with its expected loss stated, and a standing line that neither has value to a big lab: the steps that survive are under 10⁻⁶ of the bill, and every step over 1% of it is dense gradient work.

S1. Scoring a static shard held in SRAM, a few queries at a time, top-k fused in

The step
The recall stage of recall-then-rerank, or a fixed "hot" shard: the examples that are influential for many queries at once (the atlas found a handful of loud digits O; Grosse et al. found the top 1% of sequences carry 12–52% of the positive influence on a 22B model X), or a fine-tuning set small enough to fit. Codes of k coordinates in int8, fp16 or 1 bit sit in the 32 scratchpads; each minion scores its slice of about 16 codes against the Q queries, whose tiles it loads into its 3 KB L1 scratchpad a piece at a time (a 4,096-coordinate int8 query is 4 KB, so every minion reads each query once more per pass, which the model charges); it keeps the scores above the current threshold for an owner minion per query, and the threshold is shared through the TensorReduce tree.
Why this chip
Its one measured physical edge on this kind of work: a byte read by a tensor load from the shire's own scratchpad costs 4.2 pJ above idle at 2.46 TB/s M, against the 105 pJ from HBM and 38 pJ from L2 this page charges an H100 (section 6) E. The shard never leaves the chip, the maximum is combined in the tree for free, and every minion runs its own heap code.
Where it wins (all of these)
All five conditions, in full
  • The shard is static. A per-query shortlist has to be fetched first: 64 MB from the card's DRAM takes 0.84 ms and 30–39 mJ of board energy D, over PCIe 5.4 ms at the 11.8–11.9 GB/s the link was measured to move for a 64 MiB copy, or 8.4–12.0 ms as a program's staged copy M, 30–190× the 28 µs scoring pass (K1).
  • It fits 72 MB: at k = 4,096 that is 17,578 int8 codes, 8,789 in fp16 or 140,625 at 1 bit D; at the 3 million dimensions reported for one Anthropic index (secondhand O) a chip holds a dozen fp16 codes.
  • Few queries per pass, about 16 at most. The lead shrinks as Q grows: 3.1–4.2× at Q = 16 in int8. In this page's model, which lets the tensor unit reach its matmul-benchmark rate on fresh tiles, it is gone at Q ≈ 96–192 in int8 or 48–64 in fp16; the desk maps, charging the chip its full board power at a derived 39 TOP/s, put parity at Q ≈ 33–75 E. Neither rate is measured.
  • Each query's heap has one owner minion, so the state is Q × k × 8 B, at most about 10 MB; a heap on every minion would multiply it by 1,024.
  • The card is kept busy: at least 2,900–4,500 unbatched queries a second against an H100 already in the server that reads a 65 MB shard from HBM, 12,000–20,000 against one that pins a 37.5 MB shard in L2, or 390–5,000 against the host CPU E.
ET-SoC-1 against an H100 E
A 65.5 MB int8 shard, one query: the ET reads it in 28.4 µs for 1.01–1.32 mJ of board energy, idle included, bound by SRAM bandwidth; an H100 must read it from HBM (it exceeds the 37.5 MB this page lets it pin in L2) in 23.1 µs for 8.3–9.0 mJ. For a 37.5 MB shard the H100 keeps in L2: 0.60–0.79 mJ against 2.0–2.3 mJ, 2.6–3.9× E. No time win, and no capacity win: 72 MB per chip against 50 MB of L2 and 80 GB of HBM.
Weakest assumption
The three weakest assumptions, in full

That the card is busy. An interactive attribution tool runs at well under 10 queries a second, since each query costs about 1.3 H100-seconds of gradient and iHVP at 8B O; at 1–10 queries a second the card at rest costs 2.5–36 J per query, and the cheapest scorer is the host CPU that is already on. The second weakest is the H100's energy for a microsecond-scale kernel (How sure, above). The third is the cost of tensor ops with fewer than 16 rows (one per query): measured only in fp32 at batch 1, where the pass stayed load-bound. If every op costs a full 16 rows, one query takes 29.2 µs and 1.6–2.0 mJ, and the lead is 4.2–5.5×.

S1's lead shrinks as queries per pass grow: energy per query, log-log

Board energy per query for the 65.5 MB int8 shard and the 32 MB fp16 one, at the pass sizes actually run (this page's model). Where the bands cross is the parity Q quoted above; press Enter on a point to try that Q in the explorer above.

S2 in full: the MNIST-sized pipeline case, against the measured A100

S2. An MNIST-sized influence pipeline run entirely on one chip

The step
The author's atlas loop at research scale: per-example gradients of the 196–128–64–10 tanh MLP (34,122 parameters) over 60,000 training digits, contracted with 100 queries × 10 logits, top-10 per query, with K-FAC or iterative solves. Weights 136 KB, images 11.8 MB in uint8 at 14×14, queries 13.6 MB in fp32: everything fits the scratchpads, and a persistent kernel avoids the 0.56–0.57 ms an empty kernel takes to be launched from the host and waited for (0.10 ms each when launches queue back to back; E50, three cards) M.
ET-SoC-1 against the measured A100
The atlas's scan was measured at 0.231 s on an A100 O: 4.1×10¹² nominal FLOP, 17.8 TFLOP/s. The ET needs 0.43 s even at 100% of its fp32 peak as measured on aifoundry2, and 1.4–4.3 s at 10–30% D. Only in fp16 at 93% of its peak or more would it match the A100's time. On energy it wins if it sustains 30–48% of fp32 peak or 14–23% of fp16 peak (its board power on random operands, aifoundry2 at 80 °C M), against an A100 drawing 250–400 W E. Nobody has measured what fraction a persistent multi-layer kernel on 10–128-wide layers gets.

One sub-case has headroom: a single query through a sequential solver. The atlas's LiSSA took 14 ms a step and CG 268 ms an iteration on the A100 O, far above the arithmetic; a persistent kernel might be 10–100× faster there E. CUDA graphs, or batching the 100 queries per step, close most of that gap, and the author's own EK-FAC path answers a query in 6 ms. Weakest assumption: at least 1 TFLOP/s sustained in a persistent kernel on small tiles. The atlas's curvature is fp64, so fp32 factors on the chip would have to be checked against it. Value: a demonstrator, and a joule-per-phase answer to the author's open question on energy per attribution at MNIST scale; NVML can give the GPU side of that too.

S3 (a measurement, now made). Gather, scatter and scatter-add rates

No influence use survives on its own (K3 below), but every sparse or irregular idea on this chip waited on one primitive: the vector unit's gathers and scatters (fgw.ps, fscw.ps and their variants), which read or write eight elements at eight addresses per instruction. They were measured on 26 September on three cards (E48; the energy manual's §4.4 and §6), with both harts of all 1,024 minions and eight random words on eight lines of a 4 KB tile per instruction M. A word gathered from the L1 costs 12.8 pJ at 452 G/s over the chip, about what a scalar load costs; from the L2 354 pJ at 27.4 G/s and from the own scratchpad 368 pJ at the same rate, every element fetching its own 64 B line; from DRAM 9.8 nJ at 1.19 G/s. A word scattered into the L2 costs 732 pJ at 23.5 G/s.

Scatter-add, which the spec-derived estimates put anywhere from 10 to 300 G updates a second: a gather, an fadd.ps and a scatter on tables private to each hart run at 139 G updates/s (0.036 nJ each) when the table fits the hart's 512 B of L1, 18.2 G/s (0.85 nJ) in the L2, and 0.42 G/s (22 nJ) in DRAM. The packed atomic famoaddl.pi on a table a shire shares runs at 12.7 G/s (0.39 nJ), no faster than scalar amoaddl.w (12.8 G/s); famoaddg.pi on one table for the chip 2.8 G/s (1.7 nJ), against the spread global atomics' 1.92 G/s at 1.16 nJ. The condition this page set, 10 G updates a second, is met only by buckets that stay inside a shire (a hart's L1, its shire's L2 or scratchpad, or a table the shire shares); nothing that updates DRAM or the home L3 comes near it.

Scatter-add on the ET-SoC-1: updates a second against energy per update, by method and where the buckets live

Both harts of all 1,024 minions (E48, three passes on each of three cards M); the energy is card power above idle over the update rate. The dashed line is the 10 G updates a second this page asked for. The ring is the hot-line report's global atomics spread over 32 lines, the one rate this page had before S3. The energy manual draws the same rows in the same colours, with the hot line's contended line, lines of constant power and each card on its own (section 6).

4. What does not fit, and why

Everything that makes up the bill is dense arithmetic on gradients: one ET card does fp16 at 19.0 TFLOP/s for 60.8 W of board power on random operands (aifoundry2, 80 °C M), against an H100's 396 TFLOP/s at 40% of peak for 700 W O S, with no bf16, no fp64, no hardware divide or square root, 32 GB of DRAM at 76 GB/s and a host link that moves 12.5–12.6 GB/s to the card (E50 M). The candidates that looked ET-shaped die for other reasons.

Why each candidate (K1–K9) doesn't work, in full
CandidateWhy it dies
K1. Rerank a per-query shortlist in SRAMThe shortlist differs per query, so it must be fetched first: 64 MB is 0.84 ms and 30–39 mJ from the card's DRAM D, 5.4 ms over PCIe at the rate the link was measured to move for a 64 MiB copy (8.4–12.0 ms as a program's staged copy) M, 30–190× the 28 µs pass. An H100 reads it from HBM in about 21 µs. Only a static shard stays resident (S1).
K2. Many-query top-k with the hardware threshold, on its ownSelection is 10⁴–1.5×10⁵ inserts per query, free on a GPU; the state does not fit as first sized (1,000 queries × 10⁴ entries is 80 MB, over the 72 MB usable, and even 16 queries need 1.3 GB if every minion keeps its own heaps). Kept inside S1.
K3. Sparse-sketch accumulation, scatter-add recallUnder 1% of a gradient's cost, and a prefilter is about 10⁻⁶ of the bill. GPUs keep the buckets in L2, use the tensor-sketch FFT or factored projections (GPUSparse's 12.5 GB/s is for accumulators in HBM X). Measured since (S3 M): a scatter-add runs at 18.2–139 G updates/s while each hart's buckets stay in its shire's L2 or its own L1, and 12.7 G/s as packed atomics on a table a shire shares; a sketch the size of a real index lives in DRAM, where it runs at 0.42 G/s for 22 nJ an update, or at the home L3 (packed atomics on one table for the chip, 2.8 G/s), against the 10–19 G/s a posting stream from DRAM needs E. The recall use rests on an open bet that sparse codes keep the top k.
K4. Lexical or bitmap prefilter (BM25, TF-IDF)About 10⁻⁶ of the bill, and biased toward token overlap O. The host CPU's inverted index does it; on the chip it is DRAM-bound at 10⁷ documents (a 20-term query reads 25 MB of bitmaps: about 0.33 ms and 11–15 mJ E).
K5. ±1 and 1-bit cascadesNo lead at board level: ±1 in int8 from SRAM is 15–20 pJ per compare, idle included, against 18–19 pJ per bit for an H100 reading the same 65.5 MB as packed bits from HBM (one query, this page's model E). No popcount instruction (SWAR on the vector unit, about 1 pJ per bit E); 1-bit codes are shown to work for data selection (QLESS X), not for top-k recall. Kept as a storage format to try inside S1.
K6. Branchy pruned search over a DRAM shard (WAND, block-max, suffix arrays)32 GB per chip, and a random DRAM read costs 9.8 nJ per 4 B at 1.19 G/s over the chip (S3 M); real indexes are terabytes on SSD and served by CPUs (OLMoTrace X). Worth about 10⁻⁶.
K7. A sharded SRAM index, "29 µs at any N"10⁷ codes of 4,096 in int8 need 569 chips, 14–21 kW at rest D. A hot tier of the top 1% is 6 chips, and like S1 it pays for its idle chips only at 2,800–4,400 unbatched queries a second or more against an H100 E (try it in the explorer).
K8. Fine-tuning-scale attribution (LESS, QLESS)Queries come in batches: the whole 2.7×10⁵ × 8,192 score matrix at Q = 100 is 4.4×10¹⁴ operations, 0.2–0.4 s of H100 int8 E. Held resident it needs 31 chips in int8 or 4 at 1 bit.
K9. The dense steps: the scan, curvature statistics, eigendecompositions, the EK-FAC iHVP, iterative iHVPs, the exact rerank, batched scoring, scans from DRAM or flash, graph and IVF indexes, projections, zero-skip as a speed lever, mixture-of-experts gradients, validation retrains21–35× slower and 1.8–3.0× the energy on dense work D; eigendecompositions need fp64; the EK-FAC state (108 GB) would take about 1,500 chips of SRAM; streaming an index is bandwidth-bound and the chip's DRAM is 44× slower than HBM and 3–5× dearer per byte at board level; zeros save power, not time.

5. The first experiment on a lab card

One new kernel settles S1, in one to two days of work. Nothing below has been run.

Where the verdict would fall: this page's model for one query over the 65.5 MB shard, against the rule's pass and kill lines.

The first experiment's rule, and where this page's model puts S1: board energy against time for one query

The experiment plan, step by step
  1. Prerequisite, about an hour. The host link is timed (Over the PCIe link, E50, three cards: 12.5–12.6 GB/s to the card with the DMA alone, 5.2–7.8 GB/s as a program's staged copy, and an empty kernel 556–566 µs launched and waited for). What remains is the round trip of a persistent kernel waiting on a doorbell, and the board power of that waiting kernel. Poll a flag in the scratchpad, not a global atomic: 24 requesters on one line stop a shire (one hot line).
  2. The kernel, --test score in workloads/sparsity. Start from --test gemv --gemv-tree, which already reads a 1024×4096 fp32 matrix from the 32 scratchpads and reduces on chip (the 7.4 µs, 250 µJ anchor):
    • the shard: 65.5 MB of int8 codes (16,000 × 4,096) with TensorIMA8A32, then 8,000 fp16 codes with TensorFMA16A32, starting 256 KB into each scratchpad (offset 0 faults);
    • queries: Q = 1, 4, 16 and 64 as A tiles in the L1 scratchpads; keep A at 5 rows or more when B streams through TenB (erratum 1.29), or keep B in the L1 scratchpad as the gemv does;
    • top-k: each query's heap (k = 16–1,024) on one owner minion; every minion compares its scores with the current threshold (fltm.pi on int32 scores, fltm.ps on fp32) and sends the rare survivors to the owner; every T tiles, a chip-wide maximum of the thresholds through TensorReduce and TensorBroadcast; the host checks the final lists against its own top-k.
  3. Measure. Microseconds and board microjoules per pass, bracketed by idle as workloads/sparsity/run_energy.py does, each Q in its own timeout 10 process looping for about 8 s (the board meter takes a new value every 156–263 ms, depending on the card, and its rails average over about 1.1–1.2 s): some 3×10⁵ passes. On aifoundry3, pinned at 600 MHz, then aifoundry2 warm (68 °C or more, or its governor lifts the clock). From the passes and the measured idle, report energy per query at 1, 100 and 10⁴ arrivals a second.
  4. The GPU side, mandatory. The same scoring on an A100 or H100, the shard in HBM and pinned in L2, timed and metered through NVML's energy counter over loops of 10 s or more. S1's lead is that ratio.
  5. Verdict. Pass if one query over 65 MB takes 40 µs or less and 1.5 mJ or less of board energy, and the fused top-k adds at most 20. Kill S1 if one query costs 3 mJ or more, or the GPU comes within 3× per pass (the chart above).

Then, if S1 passes: S2 as atlasscan, a persistent kernel holding the atlas MLP, the digits and 100 query gradients; kill the speed story above 0.231 s and the energy story above the A100's metered joules, and time one LiSSA step against 14 ms. S3, independent of both, is done: the gather and scatter patterns were added to workloads/enercat and measured with the scatter-add on three cards on 26 September (E48; section 3).

6. Method and caveats

Method and caveats in full

7. Sources

Full source list