# Hutter packet — **fractal_zip** submission (`bins/floor_b_rsslegal_staging`, RSS-legal packed-int4, 2026-09-16)

Submission name: **fractal_zip** (https://freement.cloud/fractal_zip/) — a
cmix-lex-transformer derivative developed under the fractal_zip project;
lineage and deltas disclosed below.

Mail **the program**, not a local archive9. They run `./cmix -e enwik9 archive9` on the prize machines.

<!-- TODO(user): review/edit all prose before mailing. Numbers below are measured/verified 2026-09-14/15. -->

## What to send

| File | Role |
| --- | --- |
| `cmix` | Self-extracting compressor — packed-int4 dense + static fvmem (UPX ELF + dict + article order + 6M q4 weights) |
| matching source zip | `cmix-lex-transformer` tree that built it (OSI) |
| `writeup.md` | Ivanov transformer writeup + our delta below |
| `SHA256SUMS` | packet sha256 |

Do **not** send `run/cmix` (Floor A lex). Do **not** send `floor_b_work`. Do **not** send a vocab mask.

## Packet identity

- `cmix` = **3,836,817** bytes, SHA256 `a9f2894d54f7eed9f6d7373248a830c8b97082e8e366b588f1151028719e3969`
- Construct dir: `bins/floor_b_rsslegal_staging` (SHA256SUMS + VERDICT.json alongside). Unpacked ELF `cmix_orig` (605,420 bytes) used only for prefix `-n` screens with the external `.tfwc2`; the fat `./cmix -e` path embeds everything.
- Supersedes the 2026-09-14 packet (`f957016d…`, 3,835,701 bytes) and the withdrawn 2026-09-15 build (`b47f498d…`, 3,836,609 bytes) after the RSS incidents below. Same seed/UPDATE_LIMIT, same embedded dict/order/tfweights (byte-identical caches), same model bits — only the fvmem store/eviction code changed (+1116 bytes vs the 09-14 stub).

## RSS incidents (2026-09-15/16) and fixes — disclosure

Two distinct defects in the fvmem compressed heap, found in sequence:

1. **Arena leak (found 2026-09-15, live full `-e`).** The 09-14 packet's
   bump-only arena allocator (`fvm_arena_alloc`) abandoned the old
   compressed chunk on every dirty re-eviction. The evict arm engaged at
   ~25% (RSS ≥ 9.0 GiB); the arena (cap ≈ 1.125× heap) filled in <3 h
   (evicts 0 → 528,996; store frozen at exactly 1,763,805,844 bytes),
   evictions stalled and RSS ran 8.6 → 20.1 GiB peak. Fix: per-block chunk
   ownership with in-place rewrite when the new compressed image fits the
   recorded capacity, else a size-bucketed free list; intern-table (dedup)
   chunks stay immutable and are bounded by the 4096-entry table.
2. **Eviction never released pages (found 2026-09-16, adversarial
   review).** `fvm_evict_one` only `mprotect(PROT_NONE)`'d the evicted
   block — which releases zero RSS on Linux. Every heap page ever touched
   stayed resident for the life of the process, so the pool cap bounded
   nothing physical and the 9 GiB arm could not reduce RSS. Fix:
   `madvise(MADV_DONTNEED)` immediately after the `PROT_NONE`; verified
   with `mincore` that evicted blocks drop to 0 resident pages.
   Correctness: every fault-in path fully rewrites the block from the
   compressed store (or memsets zero blocks) before it becomes
   accessible, so kernel zero-fill can never reach model state —
   torture-tested with an evict→release→refault→verify-contents cycle.

Model bits were never at risk from either defect — this code only stores
compressed copies of PPMD heap pages. Both fixes ship in source.zip under
`fractal_compute/`; torture suite (incl. ASan/UBSan) green.

## Expected score (full local -e in flight; measured S pending)

| Item | Bytes |
| --- | ---: |
| Live 1% bar (deepmix) | 107,203,662 |
| **Measured** archive9 (with fixed stub, full `-e` completed 2026-09-17 14:22 ET) | **97,617,144** |
| This packet C | 3,836,817 |
| **Measured S** | **101,453,961** — clears bar by 5,749,701 (≥300 KiB margin: GREEN) |

Measured archive9 SHA256 `5bf6021376b23ae3067110b994fbbfdbc560404d4143a41ced5ce2621f64d4d4`
(raw encode output `892d5617…`, 97,616,028 B with the old stub; converted via
`swap_archive9_stub.sh`, delta exactly +1116). Local encode wall 240,958.87 s
(66.93 h, heavily contended by the containment-proof reruns documented in
VERDICT.json's contention ledger; the time-legality projections below come from
the quiet solo retime, not this contended wall). Local peak RSS 20,129,460 KiB —
old leaky binary, disclosed above; not a property of the shipped packet.

S accounting: archive9 = decompressor stub ‖ dict.comp ‖ tfweights ‖ coded
body ‖ 16-byte header; the stub size is derived at decode, not stored, so
the local run's archive9 (old stub) is converted to the mailable one by
swapping in the fixed stub (`swap_archive9_stub.sh`): **S_new =
S_measured + 1116**. A committee `-e` run with this packet produces the same
bytes independently (body/dict/tfweights/header all bit-identical).

<!-- Measured S filled in 2026-09-17 14:30 ET. -->

## Time legality (quiet retime, solo core 11, 2026-09-14)

Rule: each of compress/decompress must satisfy **wall_hours × T < 70,000** (Geekbench 5
single-core T of the test machine). Measured 10M `-n` prefix, clock-normalized to each
machine's f_max, linear byte extrapolation:

- 1M `-n`: enc 330.68 s / dec 332.90 s (enc/dec ratio 1.007 clean).
- 10M `-n`: enc 3535.55 s @ 3.5621 GHz mean → prize-i7-normalized **42.12 h** encode
  (α=1.029 power fit: 47.36 h).
- 10M decode: first pass 3901.79 s @ 3.40 GHz (contaminated) → 44.40 h upper bound.
  Evening re-measure (also contaminated, desktop active, 3.14 GHz mean): 4344.43 s,
  DECODE_OK bit-identical, RSS 5.86 GiB → **45.66 h upper bound** → wall×T = 65,154 (i7)
  / 59,812 (Ryzen 7) — under 70,000 a fortiori even on contaminated data. Clean 1M
  enc/dec ratio is 1.007; both 10M decode figures are contention-inflated upper bounds.

| Machine | T | Raw budget h | Encode proj (wall×T) | Decode proj (wall×T) |
| --- | ---: | ---: | ---: | ---: |
| **Lenovo i7-1165G7** (named in rules) | ≈1427 | 49.05 | 42.12 h → **60,105** | 44.40 h → **63,359** |
| **AMD Ryzen 7** (named in rules) | 1310 | 53.44 | 42.12 h → **55,177** | 44.40 h → **58,164** |
| GCP c2-standard-4 Xeon (2024 fx2-cmix judging practice) | 1026 | 68.2 | → 43,215 | → 45,554 |
| Bowery 5900X (informational lab hedge only — NOT a rules machine) | 1648 | 42.48 | → 69,414 | → 73,171 |

All rows on the two machines named in the official rules (and the observed 2024 judging
machine) are **under 70,000 with 9–14% margin**. Encode also clears the lab's stricter
12% "mail bar" on the i7 outright (42.12 vs 43.17 h); the decode leg is within 1.24 h of
that bar on a contaminated measurement. The Bowery hedge row models a committee member's
presumed desktop and is not a pass/fail criterion.

<!-- TODO(user): decide how to phrase the time-margin discussion in the cover email. -->

## What this ELF is

- Published 6M transformer (Ivanov / fx2-cmix-transformer), vocab **205** baked (OOV extras ≤ 8 map into the train set).
- **fvmem** PPMD heap, statically linked, packed-int4 dense weights. No `ppm.temp`, no repo rpath. Resident pool 3 GiB (ceil 4 GiB), no periodic zstd-19 compact.
- No GrammarMatch, no KH_OBIAS, no toolkit mixer, no mask file, no PGO.

## Proofs in hand (all vs vendor control, bit-identical + decode OK, no ppm.temp)

For the RSS-legal packet (2026-09-15, `bins/floor_b_rsslegal_staging/proofs/`):

- 128k `-n`: **10,709** bit-identical + decode OK
- 1M `-n`: **80,576** bit-identical + decode OK
- 10M RELEASE-LOAD-BEARING containment under an ENFORCED cgroup-v2 ceiling
  (`systemd-run --user --scope -p MemoryMax=… -p MemorySwapMax=0`; no
  `ulimit -v` — fvmem legitimately reserves ~14 GiB VA, the bound is on
  resident memory). Test-only binary: arm 9.0→2.0 GiB, hysteresis
  8.5→1.75 GiB — the ONLY source difference — plus a small runtime
  `FVMEM_POOL_BYTES` to force maximal evict churn at prefix scale. The
  ceiling is deliberately set BELOW the binary's baseline + total touched
  heap, so the run can complete only if eviction physically returns pages.
  RESULTS: encode rc=0 bit-identical to the 10M control (**1,362,498**),
  peak 6,350,155,776 B < 6,442,450,944 B ceiling; decode rc=0, output ==
  prefix, peak 6,342,909,952 B < ceiling; zero OOM; 6.77 M evictions with
  **~423 GB cumulatively released** (MADV_DONTNEED counter). NEGATIVE
  controls: (a) the pre-release binary (identical except no MADV_DONTNEED,
  arena fix included) under a 6,350,176,256 B cage pinned exactly at the
  ceiling and was kernel-OOM-killed at 55% in 92 min (anon-rss
  6,187,764 kB) — a ceiling the release binary's ENTIRE 100% run fits
  under (peak 6,350,155,776 < 6,350,176,256); its unconstrained
  same-workload trajectory reaches 6,535,356,416 B. Three-series
  counterfactual exhibit (identical deterministic workload, release
  flattens at eviction onset while no-release climbs to OOM):
  `proofs/containment_release_10485760/EXHIBIT_counterfactual.tsv`.
  (b) the original bump-arena binary was OOM-killed at 17.6% in 16 min
  (anon-rss 6,800,144 kB, 2026-09-15). Full matrix in `proofs/` and
  VERDICT.json. Local unbounded-RAM observations were not used as
  containment evidence.
- `fvmem_torture` unit suite (incl. new reclamation churn test, ASan/UBSan
  build): all green. Steady-state arena growth decoupled from cumulative
  compressed output (old allocator: equal).
- Full enwik9 `-e` launched 2026-09-14 evening on this box (core 11 solo)
  with the OLD binary; its coded body/dict/tfweights are bit-identical to
  what this packet produces (identity proofs above), so archive9 is mailed
  with the fixed stub swapped in. Decode-verify runs against the swapped
  archive.

## Residual risk

Prefix `-n` identity is not a full `-e` proof until the local archive9 lands. Reorder pad uses `vec.size()` (enwik9 still 243425). RSS containment at full enwik9 scale is MEASURED only at 10M prefix scale (enforced ceiling, release-load-bearing, encode+decode, with negative controls); the enwik9-scale claim is an extrapolation from that measurement plus the allocator's structural bounds (pool cap enforced by real page release; live compressed store + bounded arena + fixed baseline), not a full-scale fixed-binary `-e` on this box — the incident run consumed that slot. Disclosed as such in the mail.
