feat(rebuild): shard the weekly rebuild and make each shard resumable #163

Merged
pat-s merged 2 commits from t3code/shard-weekly-rebuild into main 2026-08-12 08:30:30 +00:00
Owner

Problem

weekly-rebuild-missing runs one job per <os>-<arch> and walks that slot's list serially in a single R -q -e argument.
That was cheap while every source fallback was skipped as "already built".
Since bincraft #105/#106/#107 and #159 the gate works, and the lists are large: 8 917 source-served records on amd64/alpine324, 15 023 on amd64/resolute.

Pipeline 10910 (weekly_rebuild_missing:alpine-324-amd64) ran for two days, reached [8692/23885] cholera, and was killed there.

Two failures follow from that shape:

  • No parallelism. The work is embarrassingly parallel across packages; one job does all of it.
  • No resumability and no clean stopping point. The loop ends only by exhausting the list, so the only way to stop it is a kill. A restart re-walks from the first entry, paying a CRAN version resolution and an S3 HEAD per package before reaching new work. And a kill matches neither success nor failure, so the Purge CDN cache step never ran: the ~4 600 binaries 10910 did publish stayed hidden behind stale edge copies.

What this changes

Three shards per slot. Each of the 18 OS/ARCH rows gains SPLIT_INTO/SPLIT_INDEX, mirroring build-all-versions.yaml. Cron and manual routing are unchanged: both filters already match on ${OS}-${ARCH}, so they now match all three shards of a slot.

local/rebuild-missing.R replaces the ~1 500-character inline one-liner. The slice is interleaved rather than contiguous, because the list is alphabetical and cost clusters by name (Rcpp*, Bioc*, rstan*).

Resume by re-deriving state from the bucket. One s3_dir_info() listing gives ETags for the slot; a package is outstanding iff its object's ETag equals CRAN's published MD5sum, i.e. it is still byte-identical to CRAN's source. That is check_s3_root_package() evaluated in bulk. No progress file, no volume, no DB cursor, and correct when a sibling shard or a process-updates run completes something concurrently.

It reads ETags rather than the index's Built field the way packages-to-build.R does, because the index is no longer rewritten until the dependent pipeline runs and so cannot reflect the current run's progress.

Unknown always means "already a binary", never "rebuild it": a multipart ETag, an unreadable CRAN index or an empty listing can never mass-schedule work.

A 20 h wall-clock budget per shard. It exits 0, so the re-index and purge always fire and the remainder is picked up next run with no bookkeeping.

.crow/weekly-rebuild-reindex.yaml takes over re-indexing and the purge, with depends_on: [weekly-rebuild-missing] and runs_on: [success, failure]. Three shards writing one slot's PACKAGES concurrently would race: update_PACKAGES() lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work.

Verification

crow lint .crow/ passes on all 11 pipelines. prek run passes.

19 assertions in local/tests/test-rebuild-missing.R, 0 failures, covering the partition (disjoint, covering, deterministic, short lists, out-of-range index) and the outstanding filter (source ETag kept, binary ETag dropped, absent object kept, multipart and missing-from-CRAN treated as built).

One of those tests caught a real bug before it shipped: an empty ETag table indexed to zero length rather than to NA, which recycled the result away and reported "nothing to build" — the dangerous direction. Fixed with an explicit lookup().

The filter run against the live amd64/alpine324 index, using its MD5sum column as the ETag (established to match the objects):

index packages:               24343
outstanding (filter):          8950
no Built stamp:                8917
filter vs no-Built agreement:  8917 of 8917
outstanding but stamped Built:   33 (version drift vs CRAN)
shard sizes: 2984/2983/2983 (sum 8950, unique 8950)

It reproduces the source-served set exactly. The extra 33 are packages whose slot version differs from CRAN's current one, so no object exists at the CRAN version key: correctly outstanding.

Notes for review

  • The 20 h budget is a chosen default, exposed as REBUILD_BUDGET_HOURS in the pipeline.
  • depends_on is file-level, not row-level, so on a full cron run no slot is re-indexed until the slowest of all 54 jobs finishes. The budget bounds that at roughly a day.
  • An explicit cancel still skips the re-index. Recovery is to trigger weekly-rebuild-reindex on its own.
  • The purge runs on every re-index row rather than one designated slot: a cron fires only its own slot's row, so gating on a named slot would leave every other slot unpurged.
  • Out of scope: build-all-versions still cannot rebuild source fallbacks, because local/build-all.R:113-122 drops every version with any single_builds row, which is precisely the source-fallback set.

Design: specs/2026-08-12-shard-weekly-rebuild-design.md

## Problem `weekly-rebuild-missing` runs one job per `<os>-<arch>` and walks that slot's list serially in a single `R -q -e` argument. That was cheap while every source fallback was skipped as "already built". Since bincraft #105/#106/#107 and #159 the gate works, and the lists are large: 8 917 source-served records on `amd64/alpine324`, 15 023 on `amd64/resolute`. Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) ran for two days, reached `[8692/23885] cholera`, and was killed there. Two failures follow from that shape: - **No parallelism.** The work is embarrassingly parallel across packages; one job does all of it. - **No resumability and no clean stopping point.** The loop ends only by exhausting the list, so the only way to stop it is a kill. A restart re-walks from the first entry, paying a CRAN version resolution and an S3 `HEAD` per package before reaching new work. And a kill matches neither `success` nor `failure`, so the `Purge CDN cache` step never ran: the ~4 600 binaries 10910 did publish stayed hidden behind stale edge copies. ## What this changes **Three shards per slot.** Each of the 18 `OS`/`ARCH` rows gains `SPLIT_INTO`/`SPLIT_INDEX`, mirroring `build-all-versions.yaml`. Cron and manual routing are unchanged: both filters already match on `${OS}-${ARCH}`, so they now match all three shards of a slot. **`local/rebuild-missing.R`** replaces the ~1 500-character inline one-liner. The slice is interleaved rather than contiguous, because the list is alphabetical and cost clusters by name (`Rcpp*`, `Bioc*`, `rstan*`). **Resume by re-deriving state from the bucket.** One `s3_dir_info()` listing gives ETags for the slot; a package is outstanding iff its object's ETag equals CRAN's published `MD5sum`, i.e. it is still byte-identical to CRAN's source. That is `check_s3_root_package()` evaluated in bulk. No progress file, no volume, no DB cursor, and correct when a sibling shard or a `process-updates` run completes something concurrently. It reads ETags rather than the index's `Built` field the way `packages-to-build.R` does, because the index is no longer rewritten until the dependent pipeline runs and so cannot reflect the current run's progress. Unknown always means "already a binary", never "rebuild it": a multipart ETag, an unreadable CRAN index or an empty listing can never mass-schedule work. **A 20 h wall-clock budget** per shard. It exits 0, so the re-index and purge always fire and the remainder is picked up next run with no bookkeeping. **`.crow/weekly-rebuild-reindex.yaml`** takes over re-indexing and the purge, with `depends_on: [weekly-rebuild-missing]` and `runs_on: [success, failure]`. Three shards writing one slot's `PACKAGES` concurrently would race: `update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work. ## Verification `crow lint .crow/` passes on all 11 pipelines. `prek run` passes. 19 assertions in `local/tests/test-rebuild-missing.R`, 0 failures, covering the partition (disjoint, covering, deterministic, short lists, out-of-range index) and the outstanding filter (source ETag kept, binary ETag dropped, absent object kept, multipart and missing-from-CRAN treated as built). One of those tests caught a real bug before it shipped: an empty ETag table indexed to zero length rather than to `NA`, which recycled the result away and reported "nothing to build" — the dangerous direction. Fixed with an explicit `lookup()`. The filter run against the live `amd64/alpine324` index, using its `MD5sum` column as the ETag (established to match the objects): ``` index packages: 24343 outstanding (filter): 8950 no Built stamp: 8917 filter vs no-Built agreement: 8917 of 8917 outstanding but stamped Built: 33 (version drift vs CRAN) shard sizes: 2984/2983/2983 (sum 8950, unique 8950) ``` It reproduces the source-served set exactly. The extra 33 are packages whose slot version differs from CRAN's current one, so no object exists at the CRAN version key: correctly outstanding. ## Notes for review - The 20 h budget is a chosen default, exposed as `REBUILD_BUDGET_HOURS` in the pipeline. - `depends_on` is file-level, not row-level, so on a full cron run no slot is re-indexed until the slowest of all 54 jobs finishes. The budget bounds that at roughly a day. - An explicit cancel still skips the re-index. Recovery is to trigger `weekly-rebuild-reindex` on its own. - The purge runs on every re-index row rather than one designated slot: a cron fires only its own slot's row, so gating on a named slot would leave every other slot unpurged. - Out of scope: `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row, which is precisely the source-fallback set. Design: `specs/2026-08-12-shard-weekly-rebuild-design.md`
Split every OS/arch row of weekly-rebuild-missing into three shards and move
the build loop out of the inline R one-liner into local/rebuild-missing.R.

A shard re-derives its outstanding set from the bucket on every start: an
object whose ETag equals CRAN's published MD5sum is still a source fallback and
needs building. That is bincraft's check_s3_root_package() evaluated in bulk,
so a restart resumes rather than replaying thousands of per-package HEADs, and
it stays correct when a sibling shard or a process-updates run finishes work in
the meantime. No progress file, no volume, no database cursor.

Give each shard a 20h wall-clock budget so it exits cleanly instead of having
to be killed. A kill matches neither success nor failure, which is how pipeline
10910 skipped its CDN purge and left ~4600 rebuilt binaries behind stale edge
copies.

Move re-indexing and the purge into weekly-rebuild-reindex.yaml, which depends
on the rebuild and runs on success or failure. Three shards writing one slot's
PACKAGES concurrently would race: update_PACKAGES lists the live bucket, so an
early lister that uploads last publishes an index missing its siblings' work.
pat-s merged commit 4b7dc28cc8 into main 2026-08-12 08:30:30 +00:00
pat-s deleted branch t3code/shard-weekly-rebuild 2026-08-12 08:30:30 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
devxy/build-cran-binaries!163
No description provided.