Influence functions on the ET-SoC-1: one small step fits, none of the bill does
The question, from the author of the influence-functions report: which steps of large-scale influence-function work, the training-data attribution an interpretability team at a big lab runs, might this chip be uniquely suited to speed up? Perhaps finding the most influential training examples?
No step that matters to a big lab's influence bill suits the ET-SoC-1, and finding the most influential examples is not one either. Picking the top k is about 1.3×10⁵–1.5×10⁵ heap inserts per query even over 10¹⁰ candidates, free inside a GPU's scoring kernel. The cost lies in producing the scores: dense gradient work, on which one ET card is 21–35× slower than an H100 and spends 1.8–3.0× the energy on random data. Two narrow items survive, as exploratory research on the chip's handling of sparse and irregular work rather than as accelerators. S1: scoring a fixed shard of sketched training gradients held in the chip's scratchpads (72 MB usable of 80 MB), a few queries at a time, with the top k kept in the same pass. That spends 6–9× less board energy per pass than an H100 reading its HBM and 2.5–3.8× less than one holding the shard in its L2, and it is no faster. S2: an MNIST-sized influence pipeline run entirely on one chip, which I expect to lose to the author's measured A100. A third item was a measurement, since made (S3): the chip's gather, scatter and scatter-add rates, which every sparse idea here waited on. A scatter-add passes 10 G updates a second only while its buckets stay inside a shire.
How sure. The ET side scales one measured kernel: a 1024×4096 fp32 layer read from the scratchpads in 7.4 µs for 250 µJ of board energy (aifoundry3, one run), which the model below reproduces to within 9% above idle. The H100 side is spec, plus energies per byte measured on an A100; nobody has measured what an H100 spends on a 20 µs kernel, and that one number sets S1's lead. The verdict on value does not hang on either. The steps that survive are under a millionth of an amortized query; the one ET-shaped step with real size, scattering a gradient's sparse parts into a sketch, happens where the gradient is born, on the GPU. And S1's lead needs the card busy: at one query a second, its 25–36 W at rest cost 25–36 J per query, about 25,000–27,000 times the pass itself (the same card's idle on both sides). It breaks even with an H100 that is in the server anyway only above 2,900–20,000 queries a second, rates at which queries batch and the lead goes.
What would settle it. One new kernel on a lab card (section 5): score a static 65 MB int8 shard at 1 to 64 queries per pass with a fused top-k, next to the same scoring metered on a GPU. S1 is dead if one query costs 3 mJ or more of board energy, or if the GPU comes within 3×.
Terms used on this page
The ET-SoC-1's compute cores are minions, small in-order RISC-V cores with a vector unit and a tensor unit, 32 to a shire; 1,024 of them run kernels. Each shire's SRAM holds a 2.5 MB scratchpad, software-managed memory any shire can address, of which about 2.25 MB is usable here: 72 MB on the chip. The tensor unit multiplies 16-row tiles (TensorFMA in fp32 and fp16, TensorIMA8A32 in int8) fed by TensorLoad; a hardware tree (TensorReduce) sums or takes the maximum across minions. More in the hub's glossary.
An influence function scores how much a training example moved a model's output on a query: the query's gradient, multiplied by the inverse of the curvature (an iHVP, here usually with EK-FAC, a Kronecker-factored approximation), dotted with the training example's gradient. A scan recomputes that gradient for each of N candidates; a sketch stores each one projected to k coordinates, so a query becomes k-dimensional dot products against an index. Recall then rerank: take the top k′ from the sketches, then score those exactly. Q is the number of queries scored in one pass over the data.
What the M/D/S/X/O/E evidence tags mean
- M
- measured on an ET-SoC-1 lab card, in the linked reports. A measured statement that names no card holds on both aifoundry2 and aifoundry3 (the rule in the hub's §8); the gathers, scatters and packed atomics (E48) and the host link (E50) were measured on all three cards, aifoundry1's card 1 as well. A named card means that card only.
- D
- derived: arithmetic on measured numbers, nothing new measured.
- S
- vendor spec: datasheets and programmer's manuals.
- X
- published by others, measured on their hardware (for the GPUs, often an A100 standing in for an H100).
- O
- the influence-functions report's author: his measurements (the MNIST atlas, on an A100) or his 8B cost model.
- E
- estimate: an assumption of this page, never measured. Every conclusion below leans on at least one.
The numbers are computed by docs/reports/data/2026-09-25-influence-on-et/make_analysis.py, which reads the
measured ones from the repository's data files and states the rest with their sources.
1. Where the cost of influence functions goes
The author's report splits the work into four phases paid on different schedules: fit the curvature once per checkpoint, precondition each query once (the iHVP), scan the candidates once per batch of queries, and validate once per method. At the shape of an 8B model with 10⁷ candidates O, one scan costs 702 H100-hours, because an unsketched scan is gradient recomputation: the dot products with 100 queries are 1.6% of it. Amortized, one query costs 8.4 H100-hours, 83% scan and 17% iterative iHVP. Sketching moves the scan into a one-off index build (1,150 H100-hours) and turns the lookup into a stream over the index; then the iHVP and, at big-lab scale, an exact rerank of the top candidates carry the bill.
Each dot is one step, costed per the unit in which it is paid (per query, per batch, per checkpoint), so the dots are not additive; the scale shows orders of magnitude. The two dashed lines are one amortized 8B query (8.4 H100-hours O) and a millionth of it. Tab to a dot, then use the arrow keys, for its size, what the ET-SoC-1 would do with it and why. Costs are the author's 8B model O except the 70B rerank E, the index stream (82 GB at the H100's 3.35 TB/s spec S), the prefilter (GPUSparse, 1.27 ms per query at batch 500 X), the sparse scatter (at most 1% of the scan's gradients E), S1 (the explorer's H100 pass E) and top-k (1 µs in a kernel epilogue E).
The same steps as a table
| Step | Paid | H100-hours | On the ET-SoC-1 |
|---|
Finding the most influential examples: why selection alone is worth nothing
Finding the most influential examples
Selecting the top k of N scores takes about k(1 + ln N/k) heap inserts: 1.3×10⁵–1.5×10⁵ per query for k = 10⁴ at N = 10⁹–10¹⁰ D. A GPU does that in the epilogue of the scoring kernel, or with a radix select for large k (FAISS, AIR top-k, RadiK X). The chip has a nice way to do it too: every minion keeps its own heap, and the running k-th-best threshold is shared chip-wide through the TensorReduce tree with a maximum instead of a sum, 2.3 µs for all 1,024 minions M. But there is nothing to win: the scores are what cost. In the two-stage pipeline Grosse described in April 2026 (secondhand, via the author's Bergson page O) and Bergson merged on 14 September, a sketch index gives the top k′ = 10³–10⁴, then each of those is re-scored with its exact gradient: at 70B that rerank is about 6 H100-hours per query E, over 99.9% of a query's energy. Selection is the one step of "find the most influential examples" that the chip is shaped for, and it is worth nothing on its own. It survives only folded into S1.
2. Explorer: an influence index against the chip
Set the size of a sketch index and how queries arrive. The capacity strip shows where the index fits; the chart shows the energy per query, board power included, on each device as the arrival rate changes. The ET card is charged its power at rest all the time, since it would be there only for this; the H100 and the host CPU are charged only for the pass, because they are in the server anyway (the H100 computes each query's gradient and iHVP, about 1.3 s of it at 8B O).
Bands span the estimates' ranges: the ET card at rest draws 25.1–36.3 W M (aifoundry3 cool, aifoundry2 hot; the third card measured, aifoundry1's card 1, draws more at rest, 32–50 W at 56–81 °C, which would only raise the ET's band); the H100 60–90 W at rest and 105 pJ per byte from HBM or 38 from L2 E; the host CPU 40–200 W while it scores E. A band stops where its device is full at this Q; beyond it, more chips or a larger Q are needed. The shaded column is where the ET breaks even with the H100. Everything here is a model (section 6); only the ET's per-byte cost and bandwidth are measured.
3. What survived, ranked
Two candidates and a measurement survived a skeptical pass over two independent desk maps of the pipeline. None is a speed-up or a cut in a big lab's bill. They belong in the set's research and exploratory group, beside the sparse-compute measurements they build on: S1 as a hardware characterization ("energy per SRAM byte on an influence-shaped kernel"), S2 as a demonstrator with its expected loss stated, and a standing line that neither has value to a big lab: the steps that survive are under 10⁻⁶ of the bill, and every step over 1% of it is dense gradient work.
S1. Scoring a static shard held in SRAM, a few queries at a time, top-k fused in
- The step
- The recall stage of recall-then-rerank, or a fixed "hot" shard: the examples that are influential for many queries at once (the atlas found a handful of loud digits O; Grosse et al. found the top 1% of sequences carry 12–52% of the positive influence on a 22B model X), or a fine-tuning set small enough to fit. Codes of k coordinates in int8, fp16 or 1 bit sit in the 32 scratchpads; each minion scores its slice of about 16 codes against the Q queries, whose tiles it loads into its 3 KB L1 scratchpad a piece at a time (a 4,096-coordinate int8 query is 4 KB, so every minion reads each query once more per pass, which the model charges); it keeps the scores above the current threshold for an owner minion per query, and the threshold is shared through the TensorReduce tree.
- Why this chip
- Its one measured physical edge on this kind of work: a byte read by a tensor load from the shire's own scratchpad costs 4.2 pJ above idle at 2.46 TB/s M, against the 105 pJ from HBM and 38 pJ from L2 this page charges an H100 (section 6) E. The shard never leaves the chip, the maximum is combined in the tree for free, and every minion runs its own heap code.
- Where it wins (all of these)
All five conditions, in full
- The shard is static. A per-query shortlist has to be fetched first: 64 MB from the card's DRAM takes 0.84 ms and 30–39 mJ of board energy D, over PCIe 5.4 ms at the 11.8–11.9 GB/s the link was measured to move for a 64 MiB copy, or 8.4–12.0 ms as a program's staged copy M, 30–190× the 28 µs scoring pass (K1).
- It fits 72 MB: at k = 4,096 that is 17,578 int8 codes, 8,789 in fp16 or 140,625 at 1 bit D; at the 3 million dimensions reported for one Anthropic index (secondhand O) a chip holds a dozen fp16 codes.
- Few queries per pass, about 16 at most. The lead shrinks as Q grows: 3.1–4.2× at Q = 16 in int8. In this page's model, which lets the tensor unit reach its matmul-benchmark rate on fresh tiles, it is gone at Q ≈ 96–192 in int8 or 48–64 in fp16; the desk maps, charging the chip its full board power at a derived 39 TOP/s, put parity at Q ≈ 33–75 E. Neither rate is measured.
- Each query's heap has one owner minion, so the state is Q × k × 8 B, at most about 10 MB; a heap on every minion would multiply it by 1,024.
- The card is kept busy: at least 2,900–4,500 unbatched queries a second against an H100 already in the server that reads a 65 MB shard from HBM, 12,000–20,000 against one that pins a 37.5 MB shard in L2, or 390–5,000 against the host CPU E.
- ET-SoC-1 against an H100 E
- A 65.5 MB int8 shard, one query: the ET reads it in 28.4 µs for 1.01–1.32 mJ of board energy, idle included, bound by SRAM bandwidth; an H100 must read it from HBM (it exceeds the 37.5 MB this page lets it pin in L2) in 23.1 µs for 8.3–9.0 mJ. For a 37.5 MB shard the H100 keeps in L2: 0.60–0.79 mJ against 2.0–2.3 mJ, 2.6–3.9× E. No time win, and no capacity win: 72 MB per chip against 50 MB of L2 and 80 GB of HBM.
- Weakest assumption
The three weakest assumptions, in full
That the card is busy. An interactive attribution tool runs at well under 10 queries a second, since each query costs about 1.3 H100-seconds of gradient and iHVP at 8B O; at 1–10 queries a second the card at rest costs 2.5–36 J per query, and the cheapest scorer is the host CPU that is already on. The second weakest is the H100's energy for a microsecond-scale kernel (How sure, above). The third is the cost of tensor ops with fewer than 16 rows (one per query): measured only in fp32 at batch 1, where the pass stayed load-bound. If every op costs a full 16 rows, one query takes 29.2 µs and 1.6–2.0 mJ, and the lead is 4.2–5.5×.
Board energy per query for the 65.5 MB int8 shard and the 32 MB fp16 one, at the pass sizes actually run (this page's model). Where the bands cross is the parity Q quoted above; press Enter on a point to try that Q in the explorer above.
S2 in full: the MNIST-sized pipeline case, against the measured A100
S2. An MNIST-sized influence pipeline run entirely on one chip
- The step
- The author's atlas loop at research scale: per-example gradients of the 196–128–64–10 tanh MLP (34,122 parameters) over 60,000 training digits, contracted with 100 queries × 10 logits, top-10 per query, with K-FAC or iterative solves. Weights 136 KB, images 11.8 MB in uint8 at 14×14, queries 13.6 MB in fp32: everything fits the scratchpads, and a persistent kernel avoids the 0.56–0.57 ms an empty kernel takes to be launched from the host and waited for (0.10 ms each when launches queue back to back; E50, three cards) M.
- ET-SoC-1 against the measured A100
- The atlas's scan was measured at 0.231 s on an A100 O: 4.1×10¹² nominal FLOP, 17.8 TFLOP/s. The ET needs 0.43 s even at 100% of its fp32 peak as measured on aifoundry2, and 1.4–4.3 s at 10–30% D. Only in fp16 at 93% of its peak or more would it match the A100's time. On energy it wins if it sustains 30–48% of fp32 peak or 14–23% of fp16 peak (its board power on random operands, aifoundry2 at 80 °C M), against an A100 drawing 250–400 W E. Nobody has measured what fraction a persistent multi-layer kernel on 10–128-wide layers gets.
One sub-case has headroom: a single query through a sequential solver. The atlas's LiSSA took 14 ms a step and CG 268 ms an iteration on the A100 O, far above the arithmetic; a persistent kernel might be 10–100× faster there E. CUDA graphs, or batching the 100 queries per step, close most of that gap, and the author's own EK-FAC path answers a query in 6 ms. Weakest assumption: at least 1 TFLOP/s sustained in a persistent kernel on small tiles. The atlas's curvature is fp64, so fp32 factors on the chip would have to be checked against it. Value: a demonstrator, and a joule-per-phase answer to the author's open question on energy per attribution at MNIST scale; NVML can give the GPU side of that too.
S3 (a measurement, now made). Gather, scatter and scatter-add rates
No influence use survives on its own (K3 below), but every sparse or irregular idea on this chip waited on one
primitive: the vector unit's gathers and scatters (fgw.ps, fscw.ps and their variants), which
read or write eight elements at eight addresses per instruction. They were measured on 26 September on three cards (E48;
the energy
manual's §4.4 and §6), with both harts of all 1,024 minions and eight random words on eight lines of a 4 KB tile per
instruction M. A word gathered from the L1 costs
12.8 pJ at
452 G/s over the chip, about what a scalar load
costs; from the L2 354 pJ at
27.4 G/s and from the own scratchpad
368 pJ at the same rate, every element fetching its own
64 B line; from DRAM 9.8 nJ at
1.19 G/s. A word scattered into the L2 costs
732 pJ at
23.5 G/s.
Scatter-add, which the spec-derived estimates put anywhere from 10 to 300 G updates a second: a gather, an
fadd.ps and a scatter on tables private to each hart run at
139 G updates/s
(0.036 nJ each) when the table fits the hart's 512 B of L1,
18.2 G/s
(0.85 nJ) in the L2, and
0.42 G/s
(22 nJ) in DRAM. The packed atomic famoaddl.pi on a table
a shire shares runs at 12.7 G/s
(0.39 nJ), no faster than scalar amoaddl.w
(12.8 G/s); famoaddg.pi on one table
for the chip 2.8 G/s
(1.7 nJ), against the spread global atomics'
1.92 G/s at
1.16 nJ. The condition this page set, 10 G updates a second, is met
only by buckets that stay inside a shire (a hart's L1, its shire's L2 or scratchpad, or a table the shire shares);
nothing that updates DRAM or the home L3 comes near it.
Both harts of all 1,024 minions (E48, three passes on each of three cards M); the energy is card power above idle over the update rate. The dashed line is the 10 G updates a second this page asked for. The ring is the hot-line report's global atomics spread over 32 lines, the one rate this page had before S3. The energy manual draws the same rows in the same colours, with the hot line's contended line, lines of constant power and each card on its own (section 6).
4. What does not fit, and why
Everything that makes up the bill is dense arithmetic on gradients: one ET card does fp16 at 19.0 TFLOP/s for 60.8 W of board power on random operands (aifoundry2, 80 °C M), against an H100's 396 TFLOP/s at 40% of peak for 700 W O S, with no bf16, no fp64, no hardware divide or square root, 32 GB of DRAM at 76 GB/s and a host link that moves 12.5–12.6 GB/s to the card (E50 M). The candidates that looked ET-shaped die for other reasons.
Why each candidate (K1–K9) doesn't work, in full
| Candidate | Why it dies |
|---|---|
| K1. Rerank a per-query shortlist in SRAM | The shortlist differs per query, so it must be fetched first: 64 MB is 0.84 ms and 30–39 mJ from the card's DRAM D, 5.4 ms over PCIe at the rate the link was measured to move for a 64 MiB copy (8.4–12.0 ms as a program's staged copy) M, 30–190× the 28 µs pass. An H100 reads it from HBM in about 21 µs. Only a static shard stays resident (S1). |
| K2. Many-query top-k with the hardware threshold, on its own | Selection is 10⁴–1.5×10⁵ inserts per query, free on a GPU; the state does not fit as first sized (1,000 queries × 10⁴ entries is 80 MB, over the 72 MB usable, and even 16 queries need 1.3 GB if every minion keeps its own heaps). Kept inside S1. |
| K3. Sparse-sketch accumulation, scatter-add recall | Under 1% of a gradient's cost, and a prefilter is about 10⁻⁶ of the bill. GPUs keep the buckets in L2, use the tensor-sketch FFT or factored projections (GPUSparse's 12.5 GB/s is for accumulators in HBM X). Measured since (S3 M): a scatter-add runs at 18.2–139 G updates/s while each hart's buckets stay in its shire's L2 or its own L1, and 12.7 G/s as packed atomics on a table a shire shares; a sketch the size of a real index lives in DRAM, where it runs at 0.42 G/s for 22 nJ an update, or at the home L3 (packed atomics on one table for the chip, 2.8 G/s), against the 10–19 G/s a posting stream from DRAM needs E. The recall use rests on an open bet that sparse codes keep the top k. |
| K4. Lexical or bitmap prefilter (BM25, TF-IDF) | About 10⁻⁶ of the bill, and biased toward token overlap O. The host CPU's inverted index does it; on the chip it is DRAM-bound at 10⁷ documents (a 20-term query reads 25 MB of bitmaps: about 0.33 ms and 11–15 mJ E). |
| K5. ±1 and 1-bit cascades | No lead at board level: ±1 in int8 from SRAM is 15–20 pJ per compare, idle included, against 18–19 pJ per bit for an H100 reading the same 65.5 MB as packed bits from HBM (one query, this page's model E). No popcount instruction (SWAR on the vector unit, about 1 pJ per bit E); 1-bit codes are shown to work for data selection (QLESS X), not for top-k recall. Kept as a storage format to try inside S1. |
| K6. Branchy pruned search over a DRAM shard (WAND, block-max, suffix arrays) | 32 GB per chip, and a random DRAM read costs 9.8 nJ per 4 B at 1.19 G/s over the chip (S3 M); real indexes are terabytes on SSD and served by CPUs (OLMoTrace X). Worth about 10⁻⁶. |
| K7. A sharded SRAM index, "29 µs at any N" | 10⁷ codes of 4,096 in int8 need 569 chips, 14–21 kW at rest D. A hot tier of the top 1% is 6 chips, and like S1 it pays for its idle chips only at 2,800–4,400 unbatched queries a second or more against an H100 E (try it in the explorer). |
| K8. Fine-tuning-scale attribution (LESS, QLESS) | Queries come in batches: the whole 2.7×10⁵ × 8,192 score matrix at Q = 100 is 4.4×10¹⁴ operations, 0.2–0.4 s of H100 int8 E. Held resident it needs 31 chips in int8 or 4 at 1 bit. |
| K9. The dense steps: the scan, curvature statistics, eigendecompositions, the EK-FAC iHVP, iterative iHVPs, the exact rerank, batched scoring, scans from DRAM or flash, graph and IVF indexes, projections, zero-skip as a speed lever, mixture-of-experts gradients, validation retrains | 21–35× slower and 1.8–3.0× the energy on dense work D; eigendecompositions need fp64; the EK-FAC state (108 GB) would take about 1,500 chips of SRAM; streaming an index is bandwidth-bound and the chip's DRAM is 44× slower than HBM and 3–5× dearer per byte at board level; zeros save power, not time. |
5. The first experiment on a lab card
One new kernel settles S1, in one to two days of work. Nothing below has been run.
Where the verdict would fall: this page's model for one query over the 65.5 MB shard, against the rule's pass and kill lines.
The experiment plan, step by step
- Prerequisite, about an hour. The host link is timed (Over the PCIe link, E50, three cards: 12.5–12.6 GB/s to the card with the DMA alone, 5.2–7.8 GB/s as a program's staged copy, and an empty kernel 556–566 µs launched and waited for). What remains is the round trip of a persistent kernel waiting on a doorbell, and the board power of that waiting kernel. Poll a flag in the scratchpad, not a global atomic: 24 requesters on one line stop a shire (one hot line).
- The kernel,
--test scoreinworkloads/sparsity. Start from--test gemv --gemv-tree, which already reads a 1024×4096 fp32 matrix from the 32 scratchpads and reduces on chip (the 7.4 µs, 250 µJ anchor):- the shard: 65.5 MB of int8 codes (16,000 × 4,096) with TensorIMA8A32, then 8,000 fp16 codes with TensorFMA16A32, starting 256 KB into each scratchpad (offset 0 faults);
- queries: Q = 1, 4, 16 and 64 as A tiles in the L1 scratchpads; keep A at 5 rows or more when B streams through TenB (erratum 1.29), or keep B in the L1 scratchpad as the gemv does;
- top-k: each query's heap (k = 16–1,024) on one owner minion; every minion compares its scores with the current threshold
(
fltm.pion int32 scores,fltm.pson fp32) and sends the rare survivors to the owner; every T tiles, a chip-wide maximum of the thresholds through TensorReduce and TensorBroadcast; the host checks the final lists against its own top-k.
- Measure. Microseconds and board microjoules per pass, bracketed by idle as
workloads/sparsity/run_energy.pydoes, each Q in its owntimeout 10process looping for about 8 s (the board meter takes a new value every 156–263 ms, depending on the card, and its rails average over about 1.1–1.2 s): some 3×10⁵ passes. On aifoundry3, pinned at 600 MHz, then aifoundry2 warm (68 °C or more, or its governor lifts the clock). From the passes and the measured idle, report energy per query at 1, 100 and 10⁴ arrivals a second. - The GPU side, mandatory. The same scoring on an A100 or H100, the shard in HBM and pinned in L2, timed and metered through NVML's energy counter over loops of 10 s or more. S1's lead is that ratio.
- Verdict. Pass if one query over 65 MB takes 40 µs or less and 1.5 mJ or less of board energy, and the fused top-k adds at most 20. Kill S1 if one query costs 3 mJ or more, or the GPU comes within 3× per pass (the chart above).
Then, if S1 passes: S2 as atlasscan, a persistent kernel holding the atlas MLP, the digits and 100 query
gradients; kill the speed story above 0.231 s and the energy story above the A100's metered joules, and time one LiSSA step
against 14 ms. S3, independent of both, is done: the gather and scatter patterns were added to workloads/enercat and
measured with the scatter-add on three cards on 26 September (E48; section 3).
6. Method and caveats
Method and caveats in full
- The ET model. One pass reads B = N·k·bytes of codes and does 2QNk operations (QNk bit compares at 1 bit). Chips = ⌈B / 72 MB⌉ in SRAM (⌈B / 32 GB⌉ in DRAM). A minion's 3 KB L1 scratchpad cannot hold a query, so every minion also reads the Q queries once for each tile of 16 codes it holds, from the shire's SRAM E: Bq = 1,024 · chips · ⌈N / (1,024 · chips · 16)⌉ · Q·k·bytes, which is 6% of B at Q = 1 for the 65.5 MB shard and as much as B at Q = 16. Time = the larger of (B + Bq)/chips ÷ bandwidth and operations/chips ÷ peak rate (from DRAM, the codes at DRAM bandwidth and the query tiles at SRAM bandwidth); energy above idle = the larger of the bytes × energy per byte and operations × energy per operation. Bandwidth and energy per byte are the energy manual's tensor loads on random data, three cards M (own scratchpad 2.46 TB/s and 4.21 pJ/B; DRAM 76 GB/s and 129 pJ/B); the peak rates and their power above the idle just before each workload come from the matmul benchmark on aifoundry2 M D (int8 0.43 pJ per operation, fp16 1.51, fp32 2.73); 1-bit codes use a SWAR estimate on the vector unit E. The larger-of rule for energy, rather than the sum, is chosen because it reproduces the one measured pass: the batch-1 fp32 layer (aifoundry3, one run) spent 68 µJ above idle M where the model gives 71 µJ. Its time, 7.4 µs for 16.8 MB at 600 MHz, is 2.3 TB/s, 7% under the model's bandwidth. That kernel splits each row across minions, so it is checked on its matrix bytes alone; the query tiles are this page's layout, not measured.
- The H100 model E. A shard of 37.5 MB or less is read from a pinned L2 at 4.3 TB/s (an A100's measured L2 rate as a floor X); larger ones from HBM at 3.0 TB/s (90% of the 3.35 TB/s spec S); 80 GB per GPU. Compute at 60% of the dense tensor-core peak S, at (700 − 75) W ÷ peak per operation above idle; at 1 bit, popc at 16 per clock per SM S, with an xor, a popc and an add per 32 bits each priced as one fp32 lane-operation at the TDP, 1.75 pJ per bit. 1.3 µs of launch per pass (CUDA graphs, measured on an H100 X). Energy = idle (60–90 W) × time + bytes × 105 pJ (HBM) or 38 pJ (L2) + operations × energy per operation. The per-byte figures are an A100's for the level alone (Antepara et al., SC'25), standing in for Hopper's newer process and HBM3. Charged the A100's whole path through L2 and L1 instead (155 and 50 pJ), the H100 would cost up to 1.4× more per pass and S1's lead would reach 12×; this page does not use that.
- The host CPU E: one socket at 200 GB/s and a few TOP/s, drawing 40–200 W more while it scores.
- Duty cycle. Energy per query at arrival rate λ is, for the ET, chips × Pidle ÷ λ + energy above idle per pass ÷ Q; for the H100 and the CPU, energy per pass ÷ Q. A device is full when λ exceeds Q ÷ time per pass.
- Dense work. "21–35× slower, 1.8–3.0× the energy" sets the chip's fp16 at 60–100% of its measured 19.0 TFLOP/s and 60.8 W of board power on random operands (aifoundry2, 80 °C) M E against an H100 at 40% of 989 TFLOP/s and 700 W, the author's cost model O.
- Measured since the first version: gathers, scatters, packed atomics and random DRAM reads (S3; E48, three passes on each of three cards, 26 September; the energy manual's §4.4 and §6), and the host link (E50, five runs on each of three cards, 27 September; Over the PCIe link): K1's copy over PCIe uses its rates.
- Not measured on the chip: tensor ops with fewer than 16 rows outside fp32 batch 1; int8 scoring with fresh tiles from the scratchpad; persistent multi-layer kernels; a polling kernel's power.
- Corrections. The two desk maps this page started from needed several; they are listed in the data README.
- Reproduce.
python3 docs/reports/data/2026-09-25-influence-on-et/make_analysis.py, thenpython3 scripts/build-report.py influence-on-et docs/reports/data/2026-09-25-influence-on-et/analysis.json docs/reports/2026-09-25-influence-on-et.html. - Versions. 25 September 2026: first version, desk analysis; no card was touched. 27 September: S3 measured (E48, three cards: section 3, K3, K6); the dense energy and S2 on random operands; the per-operation energies over each workload's own idle; the anchor's time at 600 MHz; the energy manual's three-card values; the corrections the two desk maps needed moved to the data README. Later on 27 September: the host link as measured (E50) in K1 (at the rate of a 64 MiB copy, the size nearest K1's), S2's host launch, section 4 and the first experiment's prerequisite; the idle-to-pass ratio taken on the same card (it had paired one card's idle with the other's pass: 19,000–36,000, now 25,000–27,000); the verdict chart in section 5, which marks that the model leaves out the fused top-k; S3's chart in the energy manual's colours. 28 September: the caveat that the H100's energies are an A100's per-byte figures is given once in "How sure" and once in section 6.
7. Sources
Full source list
- This set: the energy manual (§4.1 bytes by level, §3 per instruction, idle), sparse compute (546 cycles at any zero fraction, the batch-1 layer, the host launch), ridge points, memory hierarchy (and its A100 notes), matmul efficiency, on-chip communication (the 2.3 µs tree), on-chip relay (2.25 MB usable per scratchpad), one hot line (atomics); the energy manual's §4.4 and §6 (S3's gathers, scatters and scatter-add, E48); over the PCIe link (the host link and the launch path, E50).
- Data read by the producer:
docs/reports/data/2026-09-23-energy-manual/manual.json,docs/reports/data/2026-09-18-aifoundry2/results.json,docs/reports/data/2026-09-18-sparsity-aifoundry3/,docs/reports/data/2026-09-18-nocbench-aifoundry2/,docs/reports/data/2026-09-27-pcie/pcie.json. - The author's pages: Influence functions: Hessians, cost, sketching, and weak factoring (§3.1 the 8B price table, §4 sketching and retrieval); the MNIST influence atlas, its manifest's measured costs; the Bergson page (§9.4, the two-stage pipeline and the 3-million-dimension index, secondhand); the tutorial plan (chapters 11–13, the roofline and the "two machines"). The last three are private.
- Influence functions and attribution: Grosse et al., Studying Large Language Model Generalization with Influence Functions, 2023; Chang et al., TrackStar, ICLR 2025; Choe et al., LoGra, 2024; Hu et al., GraSS, NeurIPS 2025; Xia et al., LESS, 2024, and QLESS, 2025; Ran et al., RISE, 2026; Wang et al., ASTRA, 2025; Hu, Hu, Ma and Zhao, A Unified Theory of Random Projection for Influence Functions, 2026; Quirke et al., Bergson, 2026, and its repository; Liu et al., OLMoTrace, 2025.
- GPU kernels: FlashSketch, 2026; GPUSparse, 2026; Johnson, Douze and Jégou, Billion-scale similarity search with GPUs (FAISS); AIR top-k, SC'23; RadiK, 2025.
- GPU hardware: NVIDIA H100 (datasheet rates, 80 GB, 700 W); Antepara et al., Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Supercomputing, SC'25 (A100 energy per bit by level); CUDA C++ Programming Guide 12.4 (popc throughput); Luo et al., Benchmarking and Dissecting the Nvidia Hopper GPU Architecture (no binary tensor cores on Hopper).
8. Related reports
- Limits of observability, the hub: every report in the set, what each established, and the glossary.
- Sparse compute: what the silicon skips (power, not cycles), the batch-1 layer S1 is scaled from, and where the chip can meet an A100.
- Ridge points: the reuse each memory level demands, which decides when S1 turns from load-bound to compute-bound.
- Memory hierarchy: latency and bandwidth by level, and the A100 figures the H100 side leans on.
- The energy manual: every per-byte and per-operation energy used here, with bars from three cards, and S3's gather and scatter rates (§4.4).
- Influence functions: Hessians, cost, sketching, and weak factoring: the question's source, with the cost model this page prices against.