From 4b7dc28cc83784fd350731b87f0d515df7c21299 Mon Sep 17 00:00:00 2001 From: pat-s Date: Wed, 12 Aug 2026 08:30:29 +0000 Subject: [PATCH] feat(rebuild): shard the weekly rebuild and make each shard resumable (#163) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Problem `weekly-rebuild-missing` runs one job per `-` and walks that slot's list serially in a single `R -q -e` argument. That was cheap while every source fallback was skipped as "already built". Since bincraft #105/#106/#107 and #159 the gate works, and the lists are large: 8 917 source-served records on `amd64/alpine324`, 15 023 on `amd64/resolute`. Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) ran for two days, reached `[8692/23885] cholera`, and was killed there. Two failures follow from that shape: - **No parallelism.** The work is embarrassingly parallel across packages; one job does all of it. - **No resumability and no clean stopping point.** The loop ends only by exhausting the list, so the only way to stop it is a kill. A restart re-walks from the first entry, paying a CRAN version resolution and an S3 `HEAD` per package before reaching new work. And a kill matches neither `success` nor `failure`, so the `Purge CDN cache` step never ran: the ~4 600 binaries 10910 did publish stayed hidden behind stale edge copies. ## What this changes **Three shards per slot.** Each of the 18 `OS`/`ARCH` rows gains `SPLIT_INTO`/`SPLIT_INDEX`, mirroring `build-all-versions.yaml`. Cron and manual routing are unchanged: both filters already match on `${OS}-${ARCH}`, so they now match all three shards of a slot. **`local/rebuild-missing.R`** replaces the ~1 500-character inline one-liner. The slice is interleaved rather than contiguous, because the list is alphabetical and cost clusters by name (`Rcpp*`, `Bioc*`, `rstan*`). **Resume by re-deriving state from the bucket.** One `s3_dir_info()` listing gives ETags for the slot; a package is outstanding iff its object's ETag equals CRAN's published `MD5sum`, i.e. it is still byte-identical to CRAN's source. That is `check_s3_root_package()` evaluated in bulk. No progress file, no volume, no DB cursor, and correct when a sibling shard or a `process-updates` run completes something concurrently. It reads ETags rather than the index's `Built` field the way `packages-to-build.R` does, because the index is no longer rewritten until the dependent pipeline runs and so cannot reflect the current run's progress. Unknown always means "already a binary", never "rebuild it": a multipart ETag, an unreadable CRAN index or an empty listing can never mass-schedule work. **A 20 h wall-clock budget** per shard. It exits 0, so the re-index and purge always fire and the remainder is picked up next run with no bookkeeping. **`.crow/weekly-rebuild-reindex.yaml`** takes over re-indexing and the purge, with `depends_on: [weekly-rebuild-missing]` and `runs_on: [success, failure]`. Three shards writing one slot's `PACKAGES` concurrently would race: `update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work. ## Verification `crow lint .crow/` passes on all 11 pipelines. `prek run` passes. 19 assertions in `local/tests/test-rebuild-missing.R`, 0 failures, covering the partition (disjoint, covering, deterministic, short lists, out-of-range index) and the outstanding filter (source ETag kept, binary ETag dropped, absent object kept, multipart and missing-from-CRAN treated as built). One of those tests caught a real bug before it shipped: an empty ETag table indexed to zero length rather than to `NA`, which recycled the result away and reported "nothing to build" — the dangerous direction. Fixed with an explicit `lookup()`. The filter run against the live `amd64/alpine324` index, using its `MD5sum` column as the ETag (established to match the objects): ``` index packages: 24343 outstanding (filter): 8950 no Built stamp: 8917 filter vs no-Built agreement: 8917 of 8917 outstanding but stamped Built: 33 (version drift vs CRAN) shard sizes: 2984/2983/2983 (sum 8950, unique 8950) ``` It reproduces the source-served set exactly. The extra 33 are packages whose slot version differs from CRAN's current one, so no object exists at the CRAN version key: correctly outstanding. ## Notes for review - The 20 h budget is a chosen default, exposed as `REBUILD_BUDGET_HOURS` in the pipeline. - `depends_on` is file-level, not row-level, so on a full cron run no slot is re-indexed until the slowest of all 54 jobs finishes. The budget bounds that at roughly a day. - An explicit cancel still skips the re-index. Recovery is to trigger `weekly-rebuild-reindex` on its own. - The purge runs on every re-index row rather than one designated slot: a cron fires only its own slot's row, so gating on a named slot would leave every other slot unpurged. - Out of scope: `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row, which is precisely the source-fallback set. Design: `specs/2026-08-12-shard-weekly-rebuild-design.md` Reviewed-on: https://git.devxy.io/devxy/build-cran-binaries/pulls/163 --- .crow/weekly-rebuild-missing.yaml | 306 +++++++++++++++--- .crow/weekly-rebuild-reindex.yaml | 193 +++++++++++ local/rebuild-missing-helpers.R | 87 +++++ local/rebuild-missing.R | 194 +++++++++++ local/tests/test-rebuild-missing.R | 94 ++++++ .../2026-08-12-shard-weekly-rebuild-design.md | 167 ++++++++++ 6 files changed, 1005 insertions(+), 36 deletions(-) create mode 100644 .crow/weekly-rebuild-reindex.yaml create mode 100644 local/rebuild-missing-helpers.R create mode 100644 local/rebuild-missing.R create mode 100644 local/tests/test-rebuild-missing.R create mode 100644 specs/2026-08-12-shard-weekly-rebuild-design.md diff --git a/.crow/weekly-rebuild-missing.yaml b/.crow/weekly-rebuild-missing.yaml index c86fe5a..5afb3c3 100644 --- a/.crow/weekly-rebuild-missing.yaml +++ b/.crow/weekly-rebuild-missing.yaml @@ -1,12 +1,21 @@ # Consolidated weekly-rebuild-missing pipeline (all platforms, both arches). -# One matrix row per OS/arch replaces the former per-platform files. +# Three matrix rows per OS/arch, one per shard of that slot's rebuild list. # Routing is preserved 1:1: # - cron: each existing `weekly-rebuild-missing--` cron fires only -# its matching matrix row (via the per-row `cron:` name filter). +# its matching matrix rows (via the per-row `cron:` name filter), +# which is now all three shards of that slot. # - manual: `weekly_rebuild_missing` dropdown, default "all" (matches the # previous bare manual trigger that ran every os/arch); pick a # single - to run just one. # Arch placement is handled by the group label (rpkgs-amd64, rpkgs-arm64). +# +# The shard picks up its own slice and re-derives what is still outstanding +# from the bucket, so a restart resumes rather than replaying; see +# local/rebuild-missing.R. +# +# Re-indexing and the CDN purge deliberately do NOT live here. Three shards +# writing one slot's PACKAGES concurrently would race, so they moved to +# .crow/weekly-rebuild-reindex.yaml, which depends on this pipeline. variables: # Gates this pipeline. A manual pipeline creation instantiates every file in # .crow/, and a declared default is applied even when the run never passed @@ -53,74 +62,326 @@ matrix: ARCH: amd64 R_VERSION: 4.5.3 IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-322 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-322 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: alpine-322 ARCH: arm64 R_VERSION: 4.5.3 IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-322 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-322 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: alpine-323 ARCH: amd64 R_VERSION: 4.5.3 IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-323 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-323 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: alpine-323 ARCH: arm64 R_VERSION: 4.5.3 IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-323 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-323 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: alpine-324 ARCH: amd64 R_VERSION: 4.5.3 IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-324 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-324 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: alpine-324 ARCH: arm64 R_VERSION: 4.5.3 IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: alpine-324 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: alpine-324 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-8 ARCH: amd64 R_VERSION: 4.4.3 IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-8 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-8 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-8 ARCH: arm64 R_VERSION: 4.4.3 IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-8 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-8 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:8 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-9 ARCH: amd64 R_VERSION: 4.4.3 IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-9 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-9 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-9 ARCH: arm64 R_VERSION: 4.4.3 IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-9 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-9 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:9 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-10 ARCH: amd64 R_VERSION: 4.5.3 IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-10 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-10 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: redhat-10 ARCH: arm64 R_VERSION: 4.5.3 IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: redhat-10 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: redhat-10 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: redhat:10 + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2204 ARCH: amd64 R_VERSION: 4.4.3 IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2204 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2204 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2204 ARCH: arm64 R_VERSION: 4.4.3 IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2204 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2204 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2404 ARCH: amd64 R_VERSION: 4.4.3 IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2404 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2404 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2404 ARCH: arm64 R_VERSION: 4.4.3 IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2404 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2404 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2604 ARCH: amd64 R_VERSION: 4.4.3 IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2604 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2604 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 3 - OS: ubuntu-2604 ARCH: arm64 R_VERSION: 4.4.3 IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 1 + - OS: ubuntu-2604 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 2 + - OS: ubuntu-2604 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + SPLIT_INTO: 3 + SPLIT_INDEX: 3 steps: - name: 'Rebuild missing binaries' @@ -153,6 +414,12 @@ steps: PLATFORM: ${OS} ARCH: ${ARCH} NCPUS: 2 + SPLIT_INTO: ${SPLIT_INTO} + SPLIT_INDEX: ${SPLIT_INDEX} + # Wall clock after which the shard stops cleanly instead of having to be + # killed. A kill matches neither `success` nor `failure`, so it would skip + # the dependent re-index and leave rebuilt binaries behind a stale edge. + REBUILD_BUDGET_HOURS: 20 commands: - git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git . - mkdir -p /mnt/cache/uvr/cache /mnt/cache/uvr/packages /mnt/cache/R-pkgs /mnt/cache/ccache /mnt/cache/packages @@ -162,18 +429,7 @@ steps: - XVFB=$(command -v xwfb-run 2>/dev/null || command -v xvfb-run); XVFB_ARGS=""; if command -v xwfb-run >/dev/null 2>&1; then dnf install -y -q weston 2>/dev/null; XVFB_ARGS="-c weston"; fi - UVR_R_BIN=/opt/R/$R_VERSION/bin/R local/uvr-install.sh httr2 - /opt/R/$R_VERSION/bin/R -q -e 'source("local/fetch-rebuild-packages-from-issue.R")' - - $XVFB $XVFB_ARGS -- /opt/R/$R_VERSION/bin/R -q -e "sink(stdout(), type = 'message'); options(crayon.enabled = TRUE, Ncpus = $NCPUS, future.globals.onReference = NULL); pkgs <- readLines('/tmp/rebuild_pkgs.txt'); if (length(pkgs) == 0) { cat('Nothing to rebuild\n'); q('no') }; excluded <- jsonlite::fromJSON('local/excluded-packages.json')[['package']]; pkgs <- setdiff(pkgs, excluded); cat(sprintf('Rebuilding %d packages\n', length(pkgs))); n <- length(pkgs); for (i in seq_along(pkgs)) { x <- pkgs[i]; cat(sprintf('[%d/%d] %s\n', i, n, x)); tryCatch(bincraft::build_binary_package(x, tag_limit = 1L, patches = 'local/patches', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'), metadata_db_host = 'r-binaries.devxy.io', metadata_db_name = 'build_metadata', metadata_db_table = 'single_builds', metadata_db_user = 'rpkgs', metadata_db_password = Sys.getenv('PGPASS'), metadata_db_sslmode = 'require', metadata_db_port = 15432, archive = TRUE, upload = TRUE, store_build_metadata = TRUE), error = function(e) cat(sprintf('ERROR building %s - %s\n', x, conditionMessage(e)))) }" 2>&1 - # A rebuild replaces objects in place, so the slot's index still advertises - # the old MD5 and, for anything that had been served from source, no Built - # stamp. Re-index here rather than waiting for the next process-updates - # run, or the rebuilt binaries stay invisible to clients until then. - # The codename is detected from the image's /etc/os-release. - - /opt/R/$R_VERSION/bin/R -q -e 'library(bincraft); upload_package_index(s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", s3_region = "eu-central-003", s3_bucket = "devxy-rpkgs-binaries", s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"))' - - | - for RBIN in /opt/R/[0-9]*/bin/R; do - RMINOR=$(basename "$(dirname "$(dirname "$RBIN")")" | cut -d. -f1-2) - /opt/R/$R_VERSION/bin/R -q -e "library(bincraft); upload_package_index(r_minor = '$RMINOR', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'))" || true - done + - $XVFB $XVFB_ARGS -n $SPLIT_INDEX -- /opt/R/$R_VERSION/bin/Rscript local/rebuild-missing.R $SPLIT_INTO $SPLIT_INDEX $REBUILD_BUDGET_HOURS 2>&1 backend_options: docker: resources: @@ -183,25 +439,3 @@ steps: limits: memory: 18Gi cpu: 3000m - - - name: Purge CDN cache - image: reg.devxy.io/docker.io/library/alpine:3.24 - environment: - OTEL_R_TRACES_EXPORTER: none - OTEL_R_LOGS_EXPORTER: none - OTEL_R_METRICS_EXPORTER: none - BUNNYNET_API_KEY: - from_secret: BUNNYNET_API_KEY - REPO_RO_TOKEN: - from_secret: REPO_RO_TOKEN - # All hostnames on the zone share this id, so one purge covers - # cran.devxy.io, cran.allianceswisspass.devxy.io and cran.rpkgs.com. - BUNNY_PULLZONE: '3857050' - commands: - - apk add --no-cache -q bash curl git - - git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git . - - bash scripts/purge_cdn_zone.sh "$BUNNYNET_API_KEY" "$BUNNY_PULLZONE" - # A rebuild that died part-way still replaced objects, and those are exactly - # the ones a stale edge would keep hiding, so purge either way. - when: - - status: [success, failure] diff --git a/.crow/weekly-rebuild-reindex.yaml b/.crow/weekly-rebuild-reindex.yaml new file mode 100644 index 0000000..224596b --- /dev/null +++ b/.crow/weekly-rebuild-reindex.yaml @@ -0,0 +1,193 @@ +# Re-index and purge after weekly-rebuild-missing. +# +# weekly-rebuild-missing runs three shards per slot. Each of them replaces +# objects in place, so the slot's index still advertises the old MD5 and, for +# anything that had been served from source, no Built stamp. Re-indexing from +# inside a shard would mean three concurrent `upload_package_index()` calls on +# one prefix: `cranlike::update_PACKAGES()` lists the live bucket, so an early +# lister that uploads last publishes an index missing its siblings' work. +# +# So it happens exactly once per slot, here, after every shard has finished. +# `runs_on: [success, failure]` keeps that true when a shard fails; only an +# explicit cancel skips it, and this pipeline can then be triggered on its own. + +variables: + # Mirrors the gate on weekly-rebuild-missing so a manual run re-indexes + # exactly the slots it rebuilt. A manual pipeline creation instantiates every + # file in .crow/, so the default must match no matrix row. + weekly_rebuild_missing: + description: "Manual run target: a specific -, 'all' for every os/arch, or 'none' to run nothing." + options: + - none + - all + - alpine-322-amd64 + - alpine-322-arm64 + - alpine-323-amd64 + - alpine-323-arm64 + - alpine-324-amd64 + - alpine-324-arm64 + - redhat-8-amd64 + - redhat-8-arm64 + - redhat-9-amd64 + - redhat-9-arm64 + - redhat-10-amd64 + - redhat-10-arm64 + - ubuntu-2204-amd64 + - ubuntu-2204-arm64 + - ubuntu-2404-amd64 + - ubuntu-2404-arm64 + - ubuntu-2604-amd64 + - ubuntu-2604-arm64 + default: none + +when: + - event: cron + cron: weekly-rebuild-missing-${OS}-${ARCH} + - event: manual + evaluate: 'weekly_rebuild_missing == "all" || weekly_rebuild_missing == "${OS}-${ARCH}"' + +depends_on: + - weekly-rebuild-missing + +runs_on: [success, failure] + +skip_clone: true + +labels: + group: rpkgs-${ARCH} + +matrix: + include: + - OS: alpine-322 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + - OS: alpine-322 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.22 + - OS: alpine-323 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + - OS: alpine-323 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.23 + - OS: alpine-324 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + - OS: alpine-324 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: alpine:3.24 + - OS: redhat-8 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:8 + - OS: redhat-8 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:8 + - OS: redhat-9 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: redhat:9 + - OS: redhat-9 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: redhat:9 + - OS: redhat-10 + ARCH: amd64 + R_VERSION: 4.5.3 + IMG: redhat:10 + - OS: redhat-10 + ARCH: arm64 + R_VERSION: 4.5.3 + IMG: redhat:10 + - OS: ubuntu-2204 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + - OS: ubuntu-2204 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:jammy + - OS: ubuntu-2404 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + - OS: ubuntu-2404 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:noble + - OS: ubuntu-2604 + ARCH: amd64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + - OS: ubuntu-2604 + ARCH: arm64 + R_VERSION: 4.4.3 + IMG: ubuntu:resolute + +steps: + - name: 'Re-index the slot' + image: reg.devxy.io/rpkgs/build-env-${IMG} + pull: true + environment: + OTEL_R_TRACES_EXPORTER: none + OTEL_R_LOGS_EXPORTER: none + OTEL_R_METRICS_EXPORTER: none + RED_HAT_DEV_PW: + from_secret: RED_HAT_DEV_PW + B2_S3_ACCESS_KEY: + from_secret: B2_S3_ACCESS_KEY + B2_S3_SECRET_KEY: + from_secret: B2_S3_SECRET_KEY + REPO_RO_TOKEN: + from_secret: REPO_RO_TOKEN + GIT_USER: pat-s + R_LIBS_USER: /mnt/cache/R-pkgs + R_VERSION: ${R_VERSION} + PLATFORM: ${OS} + ARCH: ${ARCH} + commands: + - git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git . + - mkdir -p /mnt/cache/R-pkgs + - rm -rf /mnt/cache/R-pkgs/00LOCK-* + - /opt/R/$R_VERSION/bin/Rscript local/install-bincraft.R + # The codename is detected from the image's /etc/os-release. + - /opt/R/$R_VERSION/bin/R -q -e 'library(bincraft); upload_package_index(s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", s3_region = "eu-central-003", s3_bucket = "devxy-rpkgs-binaries", s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"))' + - | + for RBIN in /opt/R/[0-9]*/bin/R; do + RMINOR=$(basename "$(dirname "$(dirname "$RBIN")")" | cut -d. -f1-2) + /opt/R/$R_VERSION/bin/R -q -e "library(bincraft); upload_package_index(r_minor = '$RMINOR', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'))" || true + done + + - name: Purge CDN cache + image: reg.devxy.io/docker.io/library/alpine:3.24 + environment: + OTEL_R_TRACES_EXPORTER: none + OTEL_R_LOGS_EXPORTER: none + OTEL_R_METRICS_EXPORTER: none + BUNNYNET_API_KEY: + from_secret: BUNNYNET_API_KEY + REPO_RO_TOKEN: + from_secret: REPO_RO_TOKEN + # All hostnames on the zone share this id, so one purge covers + # cran.devxy.io, cran.allianceswisspass.devxy.io and cran.rpkgs.com. + BUNNY_PULLZONE: '3857050' + commands: + - apk add --no-cache -q bash curl git + - git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git . + - bash scripts/purge_cdn_zone.sh "$BUNNYNET_API_KEY" "$BUNNY_PULLZONE" + # Runs on every row rather than on one designated slot: a cron fires only + # its own slot's row, so gating on a named slot would leave every other + # slot unpurged. A manual "all" run therefore purges the zone 18 times, + # which is a cheap API call and rare. + # + # Run it even when the re-index above failed: the objects were still + # replaced, and a stale edge is exactly what keeps them hidden. + when: + - status: [success, failure] diff --git a/local/rebuild-missing-helpers.R b/local/rebuild-missing-helpers.R new file mode 100644 index 0000000..16a588e --- /dev/null +++ b/local/rebuild-missing-helpers.R @@ -0,0 +1,87 @@ +# Pure helpers for local/rebuild-missing.R, kept separate so local/tests can +# source them without executing a rebuild. + +# Interleaved slice of the rebuild list. +# +# The list is alphabetical and build cost clusters by name (Rcpp*, Bioc*, +# rstan*), so contiguous thirds would be badly unbalanced. Interleaving also +# makes each shard's progress counter representative of the slot as a whole. +shard_slice <- function(pkgs, split_into, split_index) { + split_into <- as.integer(split_into) + split_index <- as.integer(split_index) + if (is.na(split_into) || is.na(split_index)) { + stop("shard_slice(): split_into and split_index must be integers") + } + if (split_into < 1L || split_index < 1L || split_index > split_into) { + stop(sprintf( + "shard_slice(): need 1 <= split_index <= split_into, got %s of %s", + split_index, + split_into + )) + } + # seq() errors on a descending range, which is what an empty list or a shard + # index past the end would produce. + if (length(pkgs) < split_index) { + return(pkgs[0L]) + } + pkgs[seq.int(split_index, length(pkgs), by = split_into)] +} + +# Packages that still need building, decided from the bucket rather than from +# remembered progress. +# +# This is bincraft's `check_s3_root_package()` evaluated in bulk: an object +# whose ETag equals CRAN's published MD5sum is byte-identical to CRAN's source, +# so the build that was supposed to replace it has not happened yet. +# +# `etag_by_file` named by `_.tar.gz`, values are unquoted ETags +# `cran_version` named by package +# `cran_md5` named by `_` +# +# Unknown always means "already a binary", never "rebuild it", so an unreadable +# CRAN index or a multipart ETag can never mass-schedule work. +outstanding_packages <- function(pkgs, etag_by_file, cran_version, cran_md5) { + if (length(pkgs) == 0L) { + return(pkgs) + } + + # An empty table indexes to zero length rather than to NA, which would + # recycle the whole result away and silently report "nothing to build". + lookup <- function(table, key) { + if (length(table) == 0L) { + return(rep(NA_character_, length(key))) + } + unname(as.character(table[key])) + } + + version <- lookup(cran_version, pkgs) + file <- sprintf("%s_%s.tar.gz", pkgs, version) + etag <- lookup(etag_by_file, file) + md5 <- lookup(cran_md5, paste(pkgs, version, sep = "_")) + + # No CRAN version means the package cannot be resolved to a tarball at all; + # leave it in and let bincraft report why. + unresolved <- is.na(version) + # No object at the key: never built, so it is outstanding by definition. + absent <- !unresolved & is.na(etag) + # A multipart upload carries a compound ETag rather than an MD5. + unknown <- !is.na(etag) & grepl("-", etag, fixed = TRUE) + + is_source <- !unresolved & + !is.na(etag) & + !unknown & + !is.na(md5) & + etag == md5 + + pkgs[unresolved | absent | is_source] +} + +parse_rebuild_args <- function(args) { + pos <- args[!startsWith(args, "--")] + budget <- as.numeric(pos[3L]) + list( + split_into = as.integer(pos[1L]), + split_index = as.integer(pos[2L]), + budget_hours = if (is.na(budget)) 20 else budget + ) +} diff --git a/local/rebuild-missing.R b/local/rebuild-missing.R new file mode 100644 index 0000000..369024e --- /dev/null +++ b/local/rebuild-missing.R @@ -0,0 +1,194 @@ +### Rebuild one shard of a slot's missing-binary list. +# +# Usage: Rscript local/rebuild-missing.R [budget_hours] +# +# The list itself comes from local/fetch-rebuild-packages-from-issue.R, which +# writes $REBUILD_PKG_LIST (default /tmp/rebuild_pkgs.txt). +# +# Two properties matter here and are the reason this is a script rather than an +# `R -q -e` argument in the pipeline: +# +# * it is restartable. The outstanding set is re-derived from the bucket on +# every start, so a shard that died resumes where it stopped without any +# progress file, and without replaying thousands of per-package HEADs. +# * it terminates. A wall-clock budget stops the loop cleanly instead of the +# run having to be killed, which is what previously skipped the re-index and +# CDN purge and left rebuilt binaries hidden behind stale edge copies. + +options(error = function() { + cat("ERROR:", geterrmessage(), "\n", file = stdout()) + traceback(2) + q(status = 1) +}) + +library(bincraft, quietly = TRUE) + +source(file.path("local", "rebuild-missing-helpers.R")) + +args <- parse_rebuild_args(commandArgs(trailingOnly = TRUE)) +if (is.na(args$split_into) || is.na(args$split_index)) { + stop("usage: rebuild-missing.R [budget_hours]") +} + +list_file <- Sys.getenv("REBUILD_PKG_LIST", "/tmp/rebuild_pkgs.txt") +pkgs <- if (file.exists(list_file)) readLines(list_file) else character(0) +pkgs <- pkgs[nzchar(pkgs)] +if (length(pkgs) == 0L) { + cat("Nothing to rebuild\n") + q("no") +} + +excluded <- jsonlite::fromJSON("local/excluded-packages.json")[["package"]] +pkgs <- setdiff(pkgs, excluded) + +mine <- shard_slice(pkgs, args$split_into, args$split_index) +cat(sprintf( + "Shard %s/%s: %s of %s listed packages\n", + args$split_index, + args$split_into, + length(mine), + length(pkgs) +)) + +### Resume: ask the bucket what is still outstanding + +codename <- bincraft::set_codename(NULL) +local_machine <- Sys.info()[["machine"]] +arch <- if (grepl("arm64|aarch64", local_machine)) "arm64" else "amd64" +slot_dir <- sprintf( + "devxy-rpkgs-binaries/%s/%s/latest/src/contrib", + arch, + codename +) + +s3fs::s3_file_system( + aws_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), + aws_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"), + endpoint = "https://s3.eu-central-003.backblazeb2.com", + region_name = "eu-central-003", + refresh = TRUE +) + +# One paginated listing instead of a HEAD per package. Not recursed: the +# rebuild passes no `is_r_minor_sensitive`, so it only ever targets the flat +# path, and the resume filter matches that scope deliberately. +info <- tryCatch(s3fs::s3_dir_info(slot_dir), error = function(e) NULL) +etag_by_file <- if (is.null(info) || nrow(info) == 0L) { + cat(sprintf( + "WARNING: could not list %s; building the whole shard\n", + slot_dir + )) + stats::setNames(character(), character()) +} else { + stats::setNames( + gsub('^"|"$', "", as.character(info$etag)), + basename(as.character(info$uri)) + ) +} + +cran <- tryCatch( + { + con <- gzcon(url( + "https://cloud.r-project.org/src/contrib/PACKAGES.gz", + open = "rb" + )) + on.exit(close(con), add = TRUE) + read.dcf(con, fields = c("Package", "Version", "MD5sum")) + }, + error = function(e) { + cat(sprintf( + "WARNING: could not read CRAN's index (%s)\n", + conditionMessage(e) + )) + NULL + } +) +cran_version <- stats::setNames(character(), character()) +cran_md5 <- stats::setNames(character(), character()) +if (!is.null(cran)) { + cran_version <- stats::setNames( + as.character(cran[, "Version"]), + as.character(cran[, "Package"]) + ) + keep <- !is.na(cran[, "MD5sum"]) + cran_md5 <- stats::setNames( + as.character(cran[keep, "MD5sum"]), + paste(cran[keep, "Package"], cran[keep, "Version"], sep = "_") + ) +} + +before <- length(mine) +mine <- outstanding_packages(mine, etag_by_file, cran_version, cran_md5) +cat(sprintf( + "Resume: %s of %s already carry a binary; %s outstanding\n", + before - length(mine), + before, + length(mine) +)) + +if (length(mine) == 0L) { + cat("Nothing outstanding for this shard\n") + q("no") +} + +### Build + +options( + crayon.enabled = TRUE, + Ncpus = as.integer(Sys.getenv("NCPUS", "2")), + future.globals.onReference = NULL +) + +started <- Sys.time() +n <- length(mine) +completed <- 0L +for (i in seq_along(mine)) { + elapsed <- as.numeric(difftime(Sys.time(), started, units = "hours")) + if (elapsed > args$budget_hours) { + cat(sprintf( + "Budget of %sh reached after %d/%d packages; stopping cleanly. The next run resumes from the bucket.\n", + args$budget_hours, + completed, + n + )) + break + } + + x <- mine[i] + cat(sprintf("[%d/%d] %s\n", i, n, x)) + tryCatch( + bincraft::build_binary_package( + x, + tag_limit = 1L, + patches = "local/patches", + s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", + s3_region = "eu-central-003", + s3_bucket = "devxy-rpkgs-binaries", + s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), + s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"), + metadata_db_host = "r-binaries.devxy.io", + metadata_db_name = "build_metadata", + metadata_db_table = "single_builds", + metadata_db_user = "rpkgs", + metadata_db_password = Sys.getenv("PGPASS"), + metadata_db_sslmode = "require", + metadata_db_port = 15432, + archive = TRUE, + upload = TRUE, + store_build_metadata = TRUE + ), + error = function(e) { + cat(sprintf("ERROR building %s - %s\n", x, conditionMessage(e))) + } + ) + completed <- completed + 1L +} + +cat(sprintf( + "Shard %s/%s finished: %d/%d packages processed in %.1fh\n", + args$split_index, + args$split_into, + completed, + n, + as.numeric(difftime(Sys.time(), started, units = "hours")) +)) diff --git a/local/tests/test-rebuild-missing.R b/local/tests/test-rebuild-missing.R new file mode 100644 index 0000000..a5e70b9 --- /dev/null +++ b/local/tests/test-rebuild-missing.R @@ -0,0 +1,94 @@ +source(file.path("..", "rebuild-missing-helpers.R")) + +test_that("shard_slice partitions the list without gaps or overlap", { + pkgs <- letters[1:10] + parts <- lapply(1:3, function(i) shard_slice(pkgs, 3, i)) + + expect_identical(parts[[1]], c("a", "d", "g", "j")) + expect_identical(parts[[2]], c("b", "e", "h")) + expect_identical(parts[[3]], c("c", "f", "i")) + + expect_identical(sort(unlist(parts)), sort(pkgs)) + expect_identical(anyDuplicated(unlist(parts)), 0L) +}) + +test_that("shard_slice is deterministic and survives short lists", { + expect_identical( + shard_slice(letters[1:10], 3, 2), + shard_slice(letters[1:10], 3, 2) + ) + expect_identical(shard_slice(character(0), 3, 1), character(0)) + # more shards than packages: the tail shards get nothing rather than erroring + expect_identical(shard_slice(c("a"), 3, 1), "a") + expect_identical(shard_slice(c("a"), 3, 2), character(0)) +}) + +test_that("shard_slice rejects an out-of-range index", { + expect_error(shard_slice(letters, 3, 4), "split_index") + expect_error(shard_slice(letters, 3, 0), "split_index") +}) + +test_that("outstanding_packages keeps source fallbacks and drops real binaries", { + cran_version <- c(httr = "1.4.8", R6 = "2.6.1", curl = "7.1.0") + cran_md5 <- c( + httr_1.4.8 = "8756015b94a9cff6f410ca4de8557f12", + R6_2.6.1 = "f01b1787f12797c29194d63c9afd5d70", + curl_7.1.0 = "8af2ccbf5d85dc18866f45f1f26f348d" + ) + etag <- c( + # byte-identical to CRAN: the build never happened + "httr_1.4.8.tar.gz" = "8756015b94a9cff6f410ca4de8557f12", + # a real binary was published + "R6_2.6.1.tar.gz" = "9d6087ee9adda3f0a3b8067cfc652c05" + # curl has no object at all + ) + + out <- outstanding_packages( + c("httr", "R6", "curl"), + etag, + cran_version, + cran_md5 + ) + expect_identical(out, c("httr", "curl")) +}) + +test_that("outstanding_packages treats unknowns as already built", { + cran_version <- c(a = "1.0", b = "1.0") + cran_md5 <- c(a_1.0 = "aaaa") + + # a multipart ETag carries no MD5, and `b` is missing from CRAN's index: + # neither may schedule a rebuild + etag <- c("a_1.0.tar.gz" = "abc-3", "b_1.0.tar.gz" = "bbbb") + + expect_identical( + outstanding_packages(c("a", "b"), etag, cran_version, cran_md5), + character(0) + ) +}) + +test_that("outstanding_packages keeps a package CRAN has no version for", { + out <- outstanding_packages( + "ghost", + c(), + c(other = "1.0"), + c(other_1.0 = "aaaa") + ) + expect_identical(out, "ghost") +}) + +test_that("outstanding_packages handles an empty list", { + expect_identical( + outstanding_packages(character(0), c(), c(), c()), + character(0) + ) +}) + +test_that("parse_rebuild_args defaults the budget", { + a <- parse_rebuild_args(c("3", "2")) + expect_identical(a$split_into, 3L) + expect_identical(a$split_index, 2L) + expect_identical(a$budget_hours, 20) + + b <- parse_rebuild_args(c("3", "2", "1.5")) + expect_identical(b$budget_hours, 1.5) +}) diff --git a/specs/2026-08-12-shard-weekly-rebuild-design.md b/specs/2026-08-12-shard-weekly-rebuild-design.md new file mode 100644 index 0000000..f4ccbaa --- /dev/null +++ b/specs/2026-08-12-shard-weekly-rebuild-design.md @@ -0,0 +1,167 @@ +# Design: Sharding and resuming the weekly rebuild + +Date: 2026-08-12 +Status: Approved (pending spec review) + +## Problem + +`weekly-rebuild-missing` runs one job per `-` and walks that slot's rebuild list serially in a single `R -q -e` invocation (`.crow/weekly-rebuild-missing.yaml:165`). +Until 2026-08-09 that was cheap, because every source fallback was skipped as "already built" and the list was effectively empty. +Since bincraft #105/#106/#107 and build-cran-binaries #159 the gate works, and the lists are now large. + +Share of records whose object is byte-identical to CRAN's source, measured against `cran.r-project.org` MD5s on 2026-08-12: + +| slot | records | source-served | share | +| ------------------ | ------: | ------------: | -----------: | +| `amd64/resolute` | 24 212 | 15 023 | 62.1% | +| `arm64/resolute` | 24 291 | 13 670 | 56.3% | +| `arm64/alpine324` | 24 328 | 9 514 | 39.2% | +| `amd64/alpine324` | 24 343 | 8 917 | 36.7% | +| `arm64/rhel10` | 24 695 | 5 384 | 21.9% | +| `amd64/rhel10` | 24 881 | 4 712 | 19.2% | +| 12 remaining slots | ~24 700 | 850 to 2 130 | 3.5% to 8.7% | + +A single serial job cannot absorb that. +Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) started on 2026-08-09, ran for roughly two days, reached `[8692/23885] cholera`, and was killed there. + +Two distinct failures follow from that shape. + +**No parallelism.** The work is embarrassingly parallel across packages, but one job does all of it. + +**No resumability, and no clean stopping point.** The loop has no terminating condition other than exhausting the list, so the only way to stop it is a kill. +A restarted run re-reads the same list and walks it from the first entry. +It skips completed packages via `check_s3_root_package()`, but that costs a CRAN version resolution and an S3 `HEAD` per package, thousands of times, before it reaches new work. +Worse, a kill is not a pipeline failure: the `Purge CDN cache` step is guarded by `when: status: [success, failure]` (`.crow/weekly-rebuild-missing.yaml:206-207`), and on 10910 it produced no output at all. +So the ~4 600 binaries that run did publish stayed hidden behind stale edge copies. + +## Goal + +Turn each slot's rebuild into bounded, parallel, restartable units, without introducing state that can disagree with the bucket. + +## Design + +### 1. Shard the matrix three ways + +Each of the 18 `OS`/`ARCH` rows in `.crow/weekly-rebuild-missing.yaml` gains `SPLIT_INTO: 3` and `SPLIT_INDEX: 1|2|3`, giving 54 rows. +This mirrors `.crow/build-all-versions.yaml:57-98`, which already shards its matrix four ways per arch. + +Routing needs no change. +The cron filter `cron: weekly-rebuild-missing-${OS}-${ARCH}` and the manual `evaluate: weekly_rebuild_missing == "${OS}-${ARCH}"` both match all three shards of a slot. +Placement stays on the `rpkgs-${ARCH}` group label, so shards queue against available capacity rather than oversubscribing it. + +### 2. Extract the loop into `local/rebuild-missing.R` + +The build is currently a single ~1 500-character `R -q -e` argument. +Shard arithmetic and resume logic do not belong in a YAML string, and none of it is testable there. +The loop moves to `local/rebuild-missing.R`, invoked as `Rscript local/rebuild-missing.R $SPLIT_INTO $SPLIT_INDEX`, mirroring `local/build-all.R`. +Its body is unchanged in substance: read `/tmp/rebuild_pkgs.txt`, subtract `local/excluded-packages.json`, loop with `tryCatch` around `bincraft::build_binary_package()`. + +The slice is **interleaved**, not contiguous: + +```r +# the list is alphabetical and build cost clusters by name (Rcpp*, Bioc*, +# rstan*), so contiguous thirds would be badly unbalanced +mine <- pkgs[seq(split_index, length(pkgs), by = split_into)] +``` + +`local/build-all.R:64` uses contiguous chunks via `cut()`. +That is fine there because its list is every CRAN package and version, so the chunks average out. +Here the list is a filtered backlog in which expensive families sit adjacent, so interleaving is the better default. +Interleaving also makes each shard's `[i/n]` progress representative of the slot as a whole. + +### 3. Resume by re-deriving state from the bucket + +Before the loop, the shard performs one `s3fs::s3_dir_info()` on `devxy-rpkgs-binaries///latest/src/contrib` and reads the `etag` column. +It fetches CRAN's `PACKAGES` once for the latest version and published `MD5sum` of every package. +A package is still outstanding if and only if the object at `_.tar.gz` has an ETag equal to CRAN's `MD5sum` for that version, which is the definition `check_s3_root_package()` already applies one package at a time. + +```r +# one paginated listing instead of ~2900 sequential HEAD requests per shard +info <- s3fs::s3_dir_info(slot_dir) +etag <- setNames(gsub('^"|"$', "", info$etag), basename(info$uri)) + +key <- sprintf("%s_%s.tar.gz", mine, cran_version[mine]) +# keep a package when no object exists yet, or when the object is still +# byte-identical to CRAN's source; drop it once a real binary is published +mine <- mine[is.na(etag[key]) | etag[key] == cran_md5[key]] +``` + +This is the whole resume mechanism. +There is no progress file, no volume, and no database cursor. +A restarted shard recomputes ground truth and continues where it stopped, and it is correct even when a sibling shard, a `process-updates` cron, or a manual `just rebuild` completed something in the meantime. + +Three properties make this the right source of truth: + +- **It is what the build itself checks.** Any other store can disagree with the bucket; this one cannot. +- **It is agent-independent.** `.crow/weekly-rebuild-missing.yaml` mounts no `volumes:`, unlike `.crow/build-all-versions.yaml:132-133`, so `/mnt/cache` is per-job and cannot carry progress anyway. +- **It costs one listing.** `cranlike`'s `s3` fork already does exactly this call against this bucket at ~24 000 objects, so the approach is proven at the required scale. + +It must read ETags rather than the slot index's `Built` field, which is how `local/packages-to-build.R:104-130` answers the same question. +Under this design the index is not rewritten until the dependent re-index pipeline runs (section 5), so mid-run it cannot reflect the current run's progress. + +Packages that genuinely fail to build re-publish their CRAN source, so they stay outstanding and would be retried on every restart. +That is already handled upstream: `bincraft::filter_packages_with_errors()` (`R/build_binaries.R:1018`, `:1143`) drops anything with `error_occurred = TRUE`, and `store_build_metadata = TRUE` is passed on every call. +No additional poison-pill filter is needed here. + +Only the flat `src/contrib` path is considered. +The rebuild call passes no `is_r_minor_sensitive`, so it defaults to `FALSE` and only ever targets the flat path; the resume filter matches that scope deliberately. + +### 4. Give each shard a wall-clock budget + +`local/rebuild-missing.R` takes a budget, defaulting to 20 hours, and breaks out of the loop once it is exceeded: + +```r +# exit cleanly rather than being killed, so the dependent re-index still runs +if (difftime(Sys.time(), started, units = "hours") > budget_hours) { + cat(sprintf("Budget of %sh reached after %d/%d packages; stopping cleanly\n", budget_hours, i, n)) + break +} +``` + +It exits 0 and reports how much of the slice it covered. +Every run then has a terminating condition, the re-index and purge always fire, and the remainder is picked up by the next run with no bookkeeping, because section 3 recomputes the outstanding set from scratch. + +### 5. Move the re-index and purge into `.crow/weekly-rebuild-reindex.yaml` + +Three shards per slot means three concurrent `upload_package_index()` calls on the same S3 prefix. +`cranlike::update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work. +The re-index steps (`.crow/weekly-rebuild-missing.yaml:171-176`) and the purge step (`:187-207`) therefore leave that file entirely. + +The new file carries: + +```yaml +depends_on: + - weekly-rebuild-missing +runs_on: [success, failure] +``` + +`runs_on: [success, failure]` validates as a workflow-level key under `crow lint`, so a failing shard no longer withholds the re-index. +The file uses the same 18-row matrix and the same `when:` gating as `weekly-rebuild-missing`, so it only re-indexes slots that actually ran. +Each row re-indexes the flat slot and every per-minor slot. +`scripts/purge_cdn_zone.sh` runs once on a single row, because all hostnames share pull zone `3857050` and 18 identical zone purges would be waste. + +## Failure behaviour + +| case | today | after | +| -------------------------- | ----------------------------------- | ----------------------------------------------------- | +| one package errors | `tryCatch` logs, loop continues | unchanged | +| a shard fails outright | purge runs, re-index does not | re-index and purge run via `runs_on` | +| a shard exceeds its budget | cannot happen, runs until killed | exits 0, re-index and purge run | +| a shard is killed | nothing runs | still nothing; trigger the re-index pipeline alone | +| a shard restarts | re-walks the list, HEAD per package | one listing, resumes at the first outstanding package | + +The known cost of `depends_on` being file-level rather than row-level: on the weekly cron no slot is re-indexed until the slowest of all 54 jobs finishes. +The 20-hour budget bounds that at roughly one day. + +## Out of scope + +- `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row for the platform and arch, which is precisely the source-fallback set. That is a separate change. +- Bunny Perma-Cache eviction. `scripts/purge_cdn_zone.sh` purges the regular edge cache only; see the note in `CLAUDE.md` and issue history. +- The audit that produces the rebuild list is unchanged. + +## Verification + +- `crow lint .crow/` passes for both pipeline files. +- `local/rebuild-missing.R` gets unit coverage in `local/tests/` for the two pure pieces: the interleaved slice (disjoint, covering, deterministic) and the outstanding-set filter (source-served ETag kept, binary ETag dropped, absent object kept). +- A single-slot manual run of `alpine-324-amd64` shard 1 confirms the listing shortcut against the live bucket, and that the reported outstanding count is close to the 8 917 measured above divided by three. +- Restarting that shard mid-run confirms it resumes rather than replaying, by comparing the outstanding count it reports on the second start.