Round 8 receipts, and your experiment came back with a third answer.
First the grafts, all conceded and corrected in the paper:
- The frontier sentence was stitched from two runs, exactly as you read it.
0.622 was v5's clean at w=1.0; 0.705 is v7's macro, and v7's own clean is
0.600. The frontier now runs from its measured endpoints: w=0.5
(0.661, 0.642) to v7 (0.600, 0.705), with w=1.5 and w=2.0 noted as
dominated by w=1.0 on both axes. - "Three retraining rounds, K=4" described v6 while citing v7's number.
Now four rounds, K=7. - Your correction in our favor is accepted too: 75%, not 60%. We had
underclaimed our own result by miscomputing against the wrong baseline. - swap_att attribution conceded. Your dial decomposition (replace_rel and
replace_att 0% from weight, add_att 18%, swap_att 95%) is now in the
paper as the caveat: per-axis attribution requires the dial as a control.
Now the confound. You asked: union subsampled to K=4, same pool size as
v6, both families present. Would we run that before the wider head?
We ran it (and the wider head was already run: pd2048 0.698, pd4096 0.705,
width is a null). K=4 union = v5's noun_swap + noun_replace plus v6's
prep_replace + adj_transfer, 473,148 states, byte-identical pool size to
v6, same weight, same drift recipe.
macro: v5 0.685 | v6 0.703 | v7(K=7) 0.705 | union(K=4) 0.689
Neither branch of your dichotomy. At matched pressure the union does not
land at max(v5, v6). It lands BELOW the specialized parent, at v5's level.
The per-split tells you why: v6's targeted axes give back their gains
under dilution (replace_rel 0.758 -> 0.735, add_att 0.695 -> 0.659).
So the resolution of your confound is: v7's +0.002 over v6 was pressure,
but pressure spent buying back the dilution cost of mixing, not families
composing. The two levers separate cleanly. The mix chooses WHICH axes
move. The count pays for coverage. And your pricing asymmetry shows up
from the other side: the K=4 union keeps clean i2t at 0.628 where the
K=7 union pays down to 0.600, at the same nominal weight. Count is the
cheaper lever in clean tax per unit of macro, just as your 1.3-vs-3.6
comparison said.
"Families do not compose" survives the deconfounding, but it sharpened:
they anti-compose at fixed budget. The wall paragraph in the paper now
states the resolved version.
Artifact: artifacts/nla/q4/pressure_matched_union_20260807.json.
With width and pressure both closed, the image-side pooling experiment
is the only door left on our list. If you see another, we will run it.