Compare commits

...
Sign in to create a new pull request.
Author SHA1 Message Date
faa678c0c7
docs(rebuild): record the matrix size against the Crow permutation cap 2026-08-12 09:18:48 +00:00
c072b780b2
feat(rebuild): shard the weekly rebuild and make each shard resumable
Split every OS/arch row of weekly-rebuild-missing into three shards and move
the build loop out of the inline R one-liner into local/rebuild-missing.R.

A shard re-derives its outstanding set from the bucket on every start: an
object whose ETag equals CRAN's published MD5sum is still a source fallback and
needs building. That is bincraft's check_s3_root_package() evaluated in bulk,
so a restart resumes rather than replaying thousands of per-package HEADs, and
it stays correct when a sibling shard or a process-updates run finishes work in
the meantime. No progress file, no volume, no database cursor.

Give each shard a 20h wall-clock budget so it exits cleanly instead of having
to be killed. A kill matches neither success nor failure, which is how pipeline
10910 skipped its CDN purge and left ~4600 rebuilt binaries behind stale edge
copies.

Move re-indexing and the purge into weekly-rebuild-reindex.yaml, which depends
on the rebuild and runs on success or failure. Three shards writing one slot's
PACKAGES concurrently would race: update_PACKAGES lists the live bucket, so an
early lister that uploads last publishes an index missing its siblings' work.
2026-08-12 08:27:45 +00:00
02e3966523
docs(rebuild): add design for sharding and resuming the weekly rebuild 2026-08-12 08:17:49 +00:00
6 changed files with 1011 additions and 36 deletions

View file

@ -1,12 +1,27 @@
# Consolidated weekly-rebuild-missing pipeline (all platforms, both arches). # Consolidated weekly-rebuild-missing pipeline (all platforms, both arches).
# One matrix row per OS/arch replaces the former per-platform files. # Three matrix rows per OS/arch, one per shard of that slot's rebuild list.
# Routing is preserved 1:1: # Routing is preserved 1:1:
# - cron: each existing `weekly-rebuild-missing-<os>-<arch>` cron fires only # - cron: each existing `weekly-rebuild-missing-<os>-<arch>` cron fires only
# its matching matrix row (via the per-row `cron:` name filter). # its matching matrix rows (via the per-row `cron:` name filter),
# which is now all three shards of that slot.
# - manual: `weekly_rebuild_missing` dropdown, default "all" (matches the # - manual: `weekly_rebuild_missing` dropdown, default "all" (matches the
# previous bare manual trigger that ran every os/arch); pick a # previous bare manual trigger that ran every os/arch); pick a
# single <os>-<arch> to run just one. # single <os>-<arch> to run just one.
# Arch placement is handled by the group label (rpkgs-amd64, rpkgs-arm64). # Arch placement is handled by the group label (rpkgs-amd64, rpkgs-arm64).
#
# 9 OS versions x 2 arches x 3 shards = 54 rows. Crow counts the *declared*
# matrix against CROW_MAX_MATRIX_SIZE before any `when:` gate is applied, so a
# single-slot manual run expands all 54 too. The server default is 50 and was
# raised for this; `crow lint` does not check the limit, so adding an OS
# version here is only caught when a pipeline is triggered.
#
# The shard picks up its own slice and re-derives what is still outstanding
# from the bucket, so a restart resumes rather than replaying; see
# local/rebuild-missing.R.
#
# Re-indexing and the CDN purge deliberately do NOT live here. Three shards
# writing one slot's PACKAGES concurrently would race, so they moved to
# .crow/weekly-rebuild-reindex.yaml, which depends on this pipeline.
variables: variables:
# Gates this pipeline. A manual pipeline creation instantiates every file in # Gates this pipeline. A manual pipeline creation instantiates every file in
# .crow/, and a declared default is applied even when the run never passed # .crow/, and a declared default is applied even when the run never passed
@ -53,74 +68,326 @@ matrix:
ARCH: amd64 ARCH: amd64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.22 IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-322
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-322
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: alpine-322 - OS: alpine-322
ARCH: arm64 ARCH: arm64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.22 IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-322
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-322
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.22
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: alpine-323 - OS: alpine-323
ARCH: amd64 ARCH: amd64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.23 IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-323
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-323
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: alpine-323 - OS: alpine-323
ARCH: arm64 ARCH: arm64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.23 IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-323
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-323
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.23
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: alpine-324 - OS: alpine-324
ARCH: amd64 ARCH: amd64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.24 IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-324
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-324
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: alpine-324 - OS: alpine-324
ARCH: arm64 ARCH: arm64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: alpine:3.24 IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: alpine-324
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: alpine-324
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.24
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-8 - OS: redhat-8
ARCH: amd64 ARCH: amd64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: redhat:8 IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-8
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-8
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-8 - OS: redhat-8
ARCH: arm64 ARCH: arm64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: redhat:8 IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-8
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-8
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:8
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-9 - OS: redhat-9
ARCH: amd64 ARCH: amd64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: redhat:9 IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-9
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-9
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-9 - OS: redhat-9
ARCH: arm64 ARCH: arm64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: redhat:9 IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-9
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-9
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:9
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-10 - OS: redhat-10
ARCH: amd64 ARCH: amd64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: redhat:10 IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-10
ARCH: amd64
R_VERSION: 4.5.3
IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-10
ARCH: amd64
R_VERSION: 4.5.3
IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: redhat-10 - OS: redhat-10
ARCH: arm64 ARCH: arm64
R_VERSION: 4.5.3 R_VERSION: 4.5.3
IMG: redhat:10 IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: redhat-10
ARCH: arm64
R_VERSION: 4.5.3
IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: redhat-10
ARCH: arm64
R_VERSION: 4.5.3
IMG: redhat:10
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2204 - OS: ubuntu-2204
ARCH: amd64 ARCH: amd64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:jammy IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2204
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2204
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2204 - OS: ubuntu-2204
ARCH: arm64 ARCH: arm64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:jammy IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2204
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2204
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2404 - OS: ubuntu-2404
ARCH: amd64 ARCH: amd64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:noble IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2404
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2404
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2404 - OS: ubuntu-2404
ARCH: arm64 ARCH: arm64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:noble IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2404
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2404
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:noble
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2604 - OS: ubuntu-2604
ARCH: amd64 ARCH: amd64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:resolute IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2604
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2604
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 3
- OS: ubuntu-2604 - OS: ubuntu-2604
ARCH: arm64 ARCH: arm64
R_VERSION: 4.4.3 R_VERSION: 4.4.3
IMG: ubuntu:resolute IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 1
- OS: ubuntu-2604
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 2
- OS: ubuntu-2604
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
SPLIT_INTO: 3
SPLIT_INDEX: 3
steps: steps:
- name: 'Rebuild missing binaries' - name: 'Rebuild missing binaries'
@ -153,6 +420,12 @@ steps:
PLATFORM: ${OS} PLATFORM: ${OS}
ARCH: ${ARCH} ARCH: ${ARCH}
NCPUS: 2 NCPUS: 2
SPLIT_INTO: ${SPLIT_INTO}
SPLIT_INDEX: ${SPLIT_INDEX}
# Wall clock after which the shard stops cleanly instead of having to be
# killed. A kill matches neither `success` nor `failure`, so it would skip
# the dependent re-index and leave rebuilt binaries behind a stale edge.
REBUILD_BUDGET_HOURS: 20
commands: commands:
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git . - git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- mkdir -p /mnt/cache/uvr/cache /mnt/cache/uvr/packages /mnt/cache/R-pkgs /mnt/cache/ccache /mnt/cache/packages - mkdir -p /mnt/cache/uvr/cache /mnt/cache/uvr/packages /mnt/cache/R-pkgs /mnt/cache/ccache /mnt/cache/packages
@ -162,18 +435,7 @@ steps:
- XVFB=$(command -v xwfb-run 2>/dev/null || command -v xvfb-run); XVFB_ARGS=""; if command -v xwfb-run >/dev/null 2>&1; then dnf install -y -q weston 2>/dev/null; XVFB_ARGS="-c weston"; fi - XVFB=$(command -v xwfb-run 2>/dev/null || command -v xvfb-run); XVFB_ARGS=""; if command -v xwfb-run >/dev/null 2>&1; then dnf install -y -q weston 2>/dev/null; XVFB_ARGS="-c weston"; fi
- UVR_R_BIN=/opt/R/$R_VERSION/bin/R local/uvr-install.sh httr2 - UVR_R_BIN=/opt/R/$R_VERSION/bin/R local/uvr-install.sh httr2
- /opt/R/$R_VERSION/bin/R -q -e 'source("local/fetch-rebuild-packages-from-issue.R")' - /opt/R/$R_VERSION/bin/R -q -e 'source("local/fetch-rebuild-packages-from-issue.R")'
- $XVFB $XVFB_ARGS -- /opt/R/$R_VERSION/bin/R -q -e "sink(stdout(), type = 'message'); options(crayon.enabled = TRUE, Ncpus = $NCPUS, future.globals.onReference = NULL); pkgs <- readLines('/tmp/rebuild_pkgs.txt'); if (length(pkgs) == 0) { cat('Nothing to rebuild\n'); q('no') }; excluded <- jsonlite::fromJSON('local/excluded-packages.json')[['package']]; pkgs <- setdiff(pkgs, excluded); cat(sprintf('Rebuilding %d packages\n', length(pkgs))); n <- length(pkgs); for (i in seq_along(pkgs)) { x <- pkgs[i]; cat(sprintf('[%d/%d] %s\n', i, n, x)); tryCatch(bincraft::build_binary_package(x, tag_limit = 1L, patches = 'local/patches', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'), metadata_db_host = 'r-binaries.devxy.io', metadata_db_name = 'build_metadata', metadata_db_table = 'single_builds', metadata_db_user = 'rpkgs', metadata_db_password = Sys.getenv('PGPASS'), metadata_db_sslmode = 'require', metadata_db_port = 15432, archive = TRUE, upload = TRUE, store_build_metadata = TRUE), error = function(e) cat(sprintf('ERROR building %s - %s\n', x, conditionMessage(e)))) }" 2>&1 - $XVFB $XVFB_ARGS -n $SPLIT_INDEX -- /opt/R/$R_VERSION/bin/Rscript local/rebuild-missing.R $SPLIT_INTO $SPLIT_INDEX $REBUILD_BUDGET_HOURS 2>&1
# A rebuild replaces objects in place, so the slot's index still advertises
# the old MD5 and, for anything that had been served from source, no Built
# stamp. Re-index here rather than waiting for the next process-updates
# run, or the rebuilt binaries stay invisible to clients until then.
# The codename is detected from the image's /etc/os-release.
- /opt/R/$R_VERSION/bin/R -q -e 'library(bincraft); upload_package_index(s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", s3_region = "eu-central-003", s3_bucket = "devxy-rpkgs-binaries", s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"))'
- |
for RBIN in /opt/R/[0-9]*/bin/R; do
RMINOR=$(basename "$(dirname "$(dirname "$RBIN")")" | cut -d. -f1-2)
/opt/R/$R_VERSION/bin/R -q -e "library(bincraft); upload_package_index(r_minor = '$RMINOR', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'))" || true
done
backend_options: backend_options:
docker: docker:
resources: resources:
@ -183,25 +445,3 @@ steps:
limits: limits:
memory: 18Gi memory: 18Gi
cpu: 3000m cpu: 3000m
- name: Purge CDN cache
image: reg.devxy.io/docker.io/library/alpine:3.24
environment:
OTEL_R_TRACES_EXPORTER: none
OTEL_R_LOGS_EXPORTER: none
OTEL_R_METRICS_EXPORTER: none
BUNNYNET_API_KEY:
from_secret: BUNNYNET_API_KEY
REPO_RO_TOKEN:
from_secret: REPO_RO_TOKEN
# All hostnames on the zone share this id, so one purge covers
# cran.devxy.io, cran.allianceswisspass.devxy.io and cran.rpkgs.com.
BUNNY_PULLZONE: '3857050'
commands:
- apk add --no-cache -q bash curl git
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- bash scripts/purge_cdn_zone.sh "$BUNNYNET_API_KEY" "$BUNNY_PULLZONE"
# A rebuild that died part-way still replaced objects, and those are exactly
# the ones a stale edge would keep hiding, so purge either way.
when:
- status: [success, failure]

View file

@ -0,0 +1,193 @@
# Re-index and purge after weekly-rebuild-missing.
#
# weekly-rebuild-missing runs three shards per slot. Each of them replaces
# objects in place, so the slot's index still advertises the old MD5 and, for
# anything that had been served from source, no Built stamp. Re-indexing from
# inside a shard would mean three concurrent `upload_package_index()` calls on
# one prefix: `cranlike::update_PACKAGES()` lists the live bucket, so an early
# lister that uploads last publishes an index missing its siblings' work.
#
# So it happens exactly once per slot, here, after every shard has finished.
# `runs_on: [success, failure]` keeps that true when a shard fails; only an
# explicit cancel skips it, and this pipeline can then be triggered on its own.
variables:
# Mirrors the gate on weekly-rebuild-missing so a manual run re-indexes
# exactly the slots it rebuilt. A manual pipeline creation instantiates every
# file in .crow/, so the default must match no matrix row.
weekly_rebuild_missing:
description: "Manual run target: a specific <os>-<arch>, 'all' for every os/arch, or 'none' to run nothing."
options:
- none
- all
- alpine-322-amd64
- alpine-322-arm64
- alpine-323-amd64
- alpine-323-arm64
- alpine-324-amd64
- alpine-324-arm64
- redhat-8-amd64
- redhat-8-arm64
- redhat-9-amd64
- redhat-9-arm64
- redhat-10-amd64
- redhat-10-arm64
- ubuntu-2204-amd64
- ubuntu-2204-arm64
- ubuntu-2404-amd64
- ubuntu-2404-arm64
- ubuntu-2604-amd64
- ubuntu-2604-arm64
default: none
when:
- event: cron
cron: weekly-rebuild-missing-${OS}-${ARCH}
- event: manual
evaluate: 'weekly_rebuild_missing == "all" || weekly_rebuild_missing == "${OS}-${ARCH}"'
depends_on:
- weekly-rebuild-missing
runs_on: [success, failure]
skip_clone: true
labels:
group: rpkgs-${ARCH}
matrix:
include:
- OS: alpine-322
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.22
- OS: alpine-322
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.22
- OS: alpine-323
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.23
- OS: alpine-323
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.23
- OS: alpine-324
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.24
- OS: alpine-324
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.24
- OS: redhat-8
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:8
- OS: redhat-8
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:8
- OS: redhat-9
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:9
- OS: redhat-9
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:9
- OS: redhat-10
ARCH: amd64
R_VERSION: 4.5.3
IMG: redhat:10
- OS: redhat-10
ARCH: arm64
R_VERSION: 4.5.3
IMG: redhat:10
- OS: ubuntu-2204
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
- OS: ubuntu-2204
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
- OS: ubuntu-2404
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:noble
- OS: ubuntu-2404
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:noble
- OS: ubuntu-2604
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
- OS: ubuntu-2604
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
steps:
- name: 'Re-index the slot'
image: reg.devxy.io/rpkgs/build-env-${IMG}
pull: true
environment:
OTEL_R_TRACES_EXPORTER: none
OTEL_R_LOGS_EXPORTER: none
OTEL_R_METRICS_EXPORTER: none
RED_HAT_DEV_PW:
from_secret: RED_HAT_DEV_PW
B2_S3_ACCESS_KEY:
from_secret: B2_S3_ACCESS_KEY
B2_S3_SECRET_KEY:
from_secret: B2_S3_SECRET_KEY
REPO_RO_TOKEN:
from_secret: REPO_RO_TOKEN
GIT_USER: pat-s
R_LIBS_USER: /mnt/cache/R-pkgs
R_VERSION: ${R_VERSION}
PLATFORM: ${OS}
ARCH: ${ARCH}
commands:
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- mkdir -p /mnt/cache/R-pkgs
- rm -rf /mnt/cache/R-pkgs/00LOCK-*
- /opt/R/$R_VERSION/bin/Rscript local/install-bincraft.R
# The codename is detected from the image's /etc/os-release.
- /opt/R/$R_VERSION/bin/R -q -e 'library(bincraft); upload_package_index(s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", s3_region = "eu-central-003", s3_bucket = "devxy-rpkgs-binaries", s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"))'
- |
for RBIN in /opt/R/[0-9]*/bin/R; do
RMINOR=$(basename "$(dirname "$(dirname "$RBIN")")" | cut -d. -f1-2)
/opt/R/$R_VERSION/bin/R -q -e "library(bincraft); upload_package_index(r_minor = '$RMINOR', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'))" || true
done
- name: Purge CDN cache
image: reg.devxy.io/docker.io/library/alpine:3.24
environment:
OTEL_R_TRACES_EXPORTER: none
OTEL_R_LOGS_EXPORTER: none
OTEL_R_METRICS_EXPORTER: none
BUNNYNET_API_KEY:
from_secret: BUNNYNET_API_KEY
REPO_RO_TOKEN:
from_secret: REPO_RO_TOKEN
# All hostnames on the zone share this id, so one purge covers
# cran.devxy.io, cran.allianceswisspass.devxy.io and cran.rpkgs.com.
BUNNY_PULLZONE: '3857050'
commands:
- apk add --no-cache -q bash curl git
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- bash scripts/purge_cdn_zone.sh "$BUNNYNET_API_KEY" "$BUNNY_PULLZONE"
# Runs on every row rather than on one designated slot: a cron fires only
# its own slot's row, so gating on a named slot would leave every other
# slot unpurged. A manual "all" run therefore purges the zone 18 times,
# which is a cheap API call and rare.
#
# Run it even when the re-index above failed: the objects were still
# replaced, and a stale edge is exactly what keeps them hidden.
when:
- status: [success, failure]

View file

@ -0,0 +1,87 @@
# Pure helpers for local/rebuild-missing.R, kept separate so local/tests can
# source them without executing a rebuild.
# Interleaved slice of the rebuild list.
#
# The list is alphabetical and build cost clusters by name (Rcpp*, Bioc*,
# rstan*), so contiguous thirds would be badly unbalanced. Interleaving also
# makes each shard's progress counter representative of the slot as a whole.
shard_slice <- function(pkgs, split_into, split_index) {
split_into <- as.integer(split_into)
split_index <- as.integer(split_index)
if (is.na(split_into) || is.na(split_index)) {
stop("shard_slice(): split_into and split_index must be integers")
}
if (split_into < 1L || split_index < 1L || split_index > split_into) {
stop(sprintf(
"shard_slice(): need 1 <= split_index <= split_into, got %s of %s",
split_index,
split_into
))
}
# seq() errors on a descending range, which is what an empty list or a shard
# index past the end would produce.
if (length(pkgs) < split_index) {
return(pkgs[0L])
}
pkgs[seq.int(split_index, length(pkgs), by = split_into)]
}
# Packages that still need building, decided from the bucket rather than from
# remembered progress.
#
# This is bincraft's `check_s3_root_package()` evaluated in bulk: an object
# whose ETag equals CRAN's published MD5sum is byte-identical to CRAN's source,
# so the build that was supposed to replace it has not happened yet.
#
# `etag_by_file` named by `<pkg>_<ver>.tar.gz`, values are unquoted ETags
# `cran_version` named by package
# `cran_md5` named by `<pkg>_<ver>`
#
# Unknown always means "already a binary", never "rebuild it", so an unreadable
# CRAN index or a multipart ETag can never mass-schedule work.
outstanding_packages <- function(pkgs, etag_by_file, cran_version, cran_md5) {
if (length(pkgs) == 0L) {
return(pkgs)
}
# An empty table indexes to zero length rather than to NA, which would
# recycle the whole result away and silently report "nothing to build".
lookup <- function(table, key) {
if (length(table) == 0L) {
return(rep(NA_character_, length(key)))
}
unname(as.character(table[key]))
}
version <- lookup(cran_version, pkgs)
file <- sprintf("%s_%s.tar.gz", pkgs, version)
etag <- lookup(etag_by_file, file)
md5 <- lookup(cran_md5, paste(pkgs, version, sep = "_"))
# No CRAN version means the package cannot be resolved to a tarball at all;
# leave it in and let bincraft report why.
unresolved <- is.na(version)
# No object at the key: never built, so it is outstanding by definition.
absent <- !unresolved & is.na(etag)
# A multipart upload carries a compound ETag rather than an MD5.
unknown <- !is.na(etag) & grepl("-", etag, fixed = TRUE)
is_source <- !unresolved &
!is.na(etag) &
!unknown &
!is.na(md5) &
etag == md5
pkgs[unresolved | absent | is_source]
}
parse_rebuild_args <- function(args) {
pos <- args[!startsWith(args, "--")]
budget <- as.numeric(pos[3L])
list(
split_into = as.integer(pos[1L]),
split_index = as.integer(pos[2L]),
budget_hours = if (is.na(budget)) 20 else budget
)
}

194
local/rebuild-missing.R Normal file
View file

@ -0,0 +1,194 @@
### Rebuild one shard of a slot's missing-binary list.
#
# Usage: Rscript local/rebuild-missing.R <split_into> <split_index> [budget_hours]
#
# The list itself comes from local/fetch-rebuild-packages-from-issue.R, which
# writes $REBUILD_PKG_LIST (default /tmp/rebuild_pkgs.txt).
#
# Two properties matter here and are the reason this is a script rather than an
# `R -q -e` argument in the pipeline:
#
# * it is restartable. The outstanding set is re-derived from the bucket on
# every start, so a shard that died resumes where it stopped without any
# progress file, and without replaying thousands of per-package HEADs.
# * it terminates. A wall-clock budget stops the loop cleanly instead of the
# run having to be killed, which is what previously skipped the re-index and
# CDN purge and left rebuilt binaries hidden behind stale edge copies.
options(error = function() {
cat("ERROR:", geterrmessage(), "\n", file = stdout())
traceback(2)
q(status = 1)
})
library(bincraft, quietly = TRUE)
source(file.path("local", "rebuild-missing-helpers.R"))
args <- parse_rebuild_args(commandArgs(trailingOnly = TRUE))
if (is.na(args$split_into) || is.na(args$split_index)) {
stop("usage: rebuild-missing.R <split_into> <split_index> [budget_hours]")
}
list_file <- Sys.getenv("REBUILD_PKG_LIST", "/tmp/rebuild_pkgs.txt")
pkgs <- if (file.exists(list_file)) readLines(list_file) else character(0)
pkgs <- pkgs[nzchar(pkgs)]
if (length(pkgs) == 0L) {
cat("Nothing to rebuild\n")
q("no")
}
excluded <- jsonlite::fromJSON("local/excluded-packages.json")[["package"]]
pkgs <- setdiff(pkgs, excluded)
mine <- shard_slice(pkgs, args$split_into, args$split_index)
cat(sprintf(
"Shard %s/%s: %s of %s listed packages\n",
args$split_index,
args$split_into,
length(mine),
length(pkgs)
))
### Resume: ask the bucket what is still outstanding
codename <- bincraft::set_codename(NULL)
local_machine <- Sys.info()[["machine"]]
arch <- if (grepl("arm64|aarch64", local_machine)) "arm64" else "amd64"
slot_dir <- sprintf(
"devxy-rpkgs-binaries/%s/%s/latest/src/contrib",
arch,
codename
)
s3fs::s3_file_system(
aws_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"),
aws_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"),
endpoint = "https://s3.eu-central-003.backblazeb2.com",
region_name = "eu-central-003",
refresh = TRUE
)
# One paginated listing instead of a HEAD per package. Not recursed: the
# rebuild passes no `is_r_minor_sensitive`, so it only ever targets the flat
# path, and the resume filter matches that scope deliberately.
info <- tryCatch(s3fs::s3_dir_info(slot_dir), error = function(e) NULL)
etag_by_file <- if (is.null(info) || nrow(info) == 0L) {
cat(sprintf(
"WARNING: could not list %s; building the whole shard\n",
slot_dir
))
stats::setNames(character(), character())
} else {
stats::setNames(
gsub('^"|"$', "", as.character(info$etag)),
basename(as.character(info$uri))
)
}
cran <- tryCatch(
{
con <- gzcon(url(
"https://cloud.r-project.org/src/contrib/PACKAGES.gz",
open = "rb"
))
on.exit(close(con), add = TRUE)
read.dcf(con, fields = c("Package", "Version", "MD5sum"))
},
error = function(e) {
cat(sprintf(
"WARNING: could not read CRAN's index (%s)\n",
conditionMessage(e)
))
NULL
}
)
cran_version <- stats::setNames(character(), character())
cran_md5 <- stats::setNames(character(), character())
if (!is.null(cran)) {
cran_version <- stats::setNames(
as.character(cran[, "Version"]),
as.character(cran[, "Package"])
)
keep <- !is.na(cran[, "MD5sum"])
cran_md5 <- stats::setNames(
as.character(cran[keep, "MD5sum"]),
paste(cran[keep, "Package"], cran[keep, "Version"], sep = "_")
)
}
before <- length(mine)
mine <- outstanding_packages(mine, etag_by_file, cran_version, cran_md5)
cat(sprintf(
"Resume: %s of %s already carry a binary; %s outstanding\n",
before - length(mine),
before,
length(mine)
))
if (length(mine) == 0L) {
cat("Nothing outstanding for this shard\n")
q("no")
}
### Build
options(
crayon.enabled = TRUE,
Ncpus = as.integer(Sys.getenv("NCPUS", "2")),
future.globals.onReference = NULL
)
started <- Sys.time()
n <- length(mine)
completed <- 0L
for (i in seq_along(mine)) {
elapsed <- as.numeric(difftime(Sys.time(), started, units = "hours"))
if (elapsed > args$budget_hours) {
cat(sprintf(
"Budget of %sh reached after %d/%d packages; stopping cleanly. The next run resumes from the bucket.\n",
args$budget_hours,
completed,
n
))
break
}
x <- mine[i]
cat(sprintf("[%d/%d] %s\n", i, n, x))
tryCatch(
bincraft::build_binary_package(
x,
tag_limit = 1L,
patches = "local/patches",
s3_endpoint = "https://s3.eu-central-003.backblazeb2.com",
s3_region = "eu-central-003",
s3_bucket = "devxy-rpkgs-binaries",
s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"),
s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"),
metadata_db_host = "r-binaries.devxy.io",
metadata_db_name = "build_metadata",
metadata_db_table = "single_builds",
metadata_db_user = "rpkgs",
metadata_db_password = Sys.getenv("PGPASS"),
metadata_db_sslmode = "require",
metadata_db_port = 15432,
archive = TRUE,
upload = TRUE,
store_build_metadata = TRUE
),
error = function(e) {
cat(sprintf("ERROR building %s - %s\n", x, conditionMessage(e)))
}
)
completed <- completed + 1L
}
cat(sprintf(
"Shard %s/%s finished: %d/%d packages processed in %.1fh\n",
args$split_index,
args$split_into,
completed,
n,
as.numeric(difftime(Sys.time(), started, units = "hours"))
))

View file

@ -0,0 +1,94 @@
source(file.path("..", "rebuild-missing-helpers.R"))
test_that("shard_slice partitions the list without gaps or overlap", {
pkgs <- letters[1:10]
parts <- lapply(1:3, function(i) shard_slice(pkgs, 3, i))
expect_identical(parts[[1]], c("a", "d", "g", "j"))
expect_identical(parts[[2]], c("b", "e", "h"))
expect_identical(parts[[3]], c("c", "f", "i"))
expect_identical(sort(unlist(parts)), sort(pkgs))
expect_identical(anyDuplicated(unlist(parts)), 0L)
})
test_that("shard_slice is deterministic and survives short lists", {
expect_identical(
shard_slice(letters[1:10], 3, 2),
shard_slice(letters[1:10], 3, 2)
)
expect_identical(shard_slice(character(0), 3, 1), character(0))
# more shards than packages: the tail shards get nothing rather than erroring
expect_identical(shard_slice(c("a"), 3, 1), "a")
expect_identical(shard_slice(c("a"), 3, 2), character(0))
})
test_that("shard_slice rejects an out-of-range index", {
expect_error(shard_slice(letters, 3, 4), "split_index")
expect_error(shard_slice(letters, 3, 0), "split_index")
})
test_that("outstanding_packages keeps source fallbacks and drops real binaries", {
cran_version <- c(httr = "1.4.8", R6 = "2.6.1", curl = "7.1.0")
cran_md5 <- c(
httr_1.4.8 = "8756015b94a9cff6f410ca4de8557f12",
R6_2.6.1 = "f01b1787f12797c29194d63c9afd5d70",
curl_7.1.0 = "8af2ccbf5d85dc18866f45f1f26f348d"
)
etag <- c(
# byte-identical to CRAN: the build never happened
"httr_1.4.8.tar.gz" = "8756015b94a9cff6f410ca4de8557f12",
# a real binary was published
"R6_2.6.1.tar.gz" = "9d6087ee9adda3f0a3b8067cfc652c05"
# curl has no object at all
)
out <- outstanding_packages(
c("httr", "R6", "curl"),
etag,
cran_version,
cran_md5
)
expect_identical(out, c("httr", "curl"))
})
test_that("outstanding_packages treats unknowns as already built", {
cran_version <- c(a = "1.0", b = "1.0")
cran_md5 <- c(a_1.0 = "aaaa")
# a multipart ETag carries no MD5, and `b` is missing from CRAN's index:
# neither may schedule a rebuild
etag <- c("a_1.0.tar.gz" = "abc-3", "b_1.0.tar.gz" = "bbbb")
expect_identical(
outstanding_packages(c("a", "b"), etag, cran_version, cran_md5),
character(0)
)
})
test_that("outstanding_packages keeps a package CRAN has no version for", {
out <- outstanding_packages(
"ghost",
c(),
c(other = "1.0"),
c(other_1.0 = "aaaa")
)
expect_identical(out, "ghost")
})
test_that("outstanding_packages handles an empty list", {
expect_identical(
outstanding_packages(character(0), c(), c(), c()),
character(0)
)
})
test_that("parse_rebuild_args defaults the budget", {
a <- parse_rebuild_args(c("3", "2"))
expect_identical(a$split_into, 3L)
expect_identical(a$split_index, 2L)
expect_identical(a$budget_hours, 20)
b <- parse_rebuild_args(c("3", "2", "1.5"))
expect_identical(b$budget_hours, 1.5)
})

View file

@ -0,0 +1,167 @@
# Design: Sharding and resuming the weekly rebuild
Date: 2026-08-12
Status: Approved (pending spec review)
## Problem
`weekly-rebuild-missing` runs one job per `<os>-<arch>` and walks that slot's rebuild list serially in a single `R -q -e` invocation (`.crow/weekly-rebuild-missing.yaml:165`).
Until 2026-08-09 that was cheap, because every source fallback was skipped as "already built" and the list was effectively empty.
Since bincraft #105/#106/#107 and build-cran-binaries #159 the gate works, and the lists are now large.
Share of records whose object is byte-identical to CRAN's source, measured against `cran.r-project.org` MD5s on 2026-08-12:
| slot | records | source-served | share |
| ------------------ | ------: | ------------: | -----------: |
| `amd64/resolute` | 24 212 | 15 023 | 62.1% |
| `arm64/resolute` | 24 291 | 13 670 | 56.3% |
| `arm64/alpine324` | 24 328 | 9 514 | 39.2% |
| `amd64/alpine324` | 24 343 | 8 917 | 36.7% |
| `arm64/rhel10` | 24 695 | 5 384 | 21.9% |
| `amd64/rhel10` | 24 881 | 4 712 | 19.2% |
| 12 remaining slots | ~24 700 | 850 to 2 130 | 3.5% to 8.7% |
A single serial job cannot absorb that.
Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) started on 2026-08-09, ran for roughly two days, reached `[8692/23885] cholera`, and was killed there.
Two distinct failures follow from that shape.
**No parallelism.** The work is embarrassingly parallel across packages, but one job does all of it.
**No resumability, and no clean stopping point.** The loop has no terminating condition other than exhausting the list, so the only way to stop it is a kill.
A restarted run re-reads the same list and walks it from the first entry.
It skips completed packages via `check_s3_root_package()`, but that costs a CRAN version resolution and an S3 `HEAD` per package, thousands of times, before it reaches new work.
Worse, a kill is not a pipeline failure: the `Purge CDN cache` step is guarded by `when: status: [success, failure]` (`.crow/weekly-rebuild-missing.yaml:206-207`), and on 10910 it produced no output at all.
So the ~4 600 binaries that run did publish stayed hidden behind stale edge copies.
## Goal
Turn each slot's rebuild into bounded, parallel, restartable units, without introducing state that can disagree with the bucket.
## Design
### 1. Shard the matrix three ways
Each of the 18 `OS`/`ARCH` rows in `.crow/weekly-rebuild-missing.yaml` gains `SPLIT_INTO: 3` and `SPLIT_INDEX: 1|2|3`, giving 54 rows.
This mirrors `.crow/build-all-versions.yaml:57-98`, which already shards its matrix four ways per arch.
Routing needs no change.
The cron filter `cron: weekly-rebuild-missing-${OS}-${ARCH}` and the manual `evaluate: weekly_rebuild_missing == "${OS}-${ARCH}"` both match all three shards of a slot.
Placement stays on the `rpkgs-${ARCH}` group label, so shards queue against available capacity rather than oversubscribing it.
### 2. Extract the loop into `local/rebuild-missing.R`
The build is currently a single ~1 500-character `R -q -e` argument.
Shard arithmetic and resume logic do not belong in a YAML string, and none of it is testable there.
The loop moves to `local/rebuild-missing.R`, invoked as `Rscript local/rebuild-missing.R $SPLIT_INTO $SPLIT_INDEX`, mirroring `local/build-all.R`.
Its body is unchanged in substance: read `/tmp/rebuild_pkgs.txt`, subtract `local/excluded-packages.json`, loop with `tryCatch` around `bincraft::build_binary_package()`.
The slice is **interleaved**, not contiguous:
```r
# the list is alphabetical and build cost clusters by name (Rcpp*, Bioc*,
# rstan*), so contiguous thirds would be badly unbalanced
mine <- pkgs[seq(split_index, length(pkgs), by = split_into)]
```
`local/build-all.R:64` uses contiguous chunks via `cut()`.
That is fine there because its list is every CRAN package and version, so the chunks average out.
Here the list is a filtered backlog in which expensive families sit adjacent, so interleaving is the better default.
Interleaving also makes each shard's `[i/n]` progress representative of the slot as a whole.
### 3. Resume by re-deriving state from the bucket
Before the loop, the shard performs one `s3fs::s3_dir_info()` on `devxy-rpkgs-binaries/<arch>/<codename>/latest/src/contrib` and reads the `etag` column.
It fetches CRAN's `PACKAGES` once for the latest version and published `MD5sum` of every package.
A package is still outstanding if and only if the object at `<pkg>_<version>.tar.gz` has an ETag equal to CRAN's `MD5sum` for that version, which is the definition `check_s3_root_package()` already applies one package at a time.
```r
# one paginated listing instead of ~2900 sequential HEAD requests per shard
info <- s3fs::s3_dir_info(slot_dir)
etag <- setNames(gsub('^"|"$', "", info$etag), basename(info$uri))
key <- sprintf("%s_%s.tar.gz", mine, cran_version[mine])
# keep a package when no object exists yet, or when the object is still
# byte-identical to CRAN's source; drop it once a real binary is published
mine <- mine[is.na(etag[key]) | etag[key] == cran_md5[key]]
```
This is the whole resume mechanism.
There is no progress file, no volume, and no database cursor.
A restarted shard recomputes ground truth and continues where it stopped, and it is correct even when a sibling shard, a `process-updates` cron, or a manual `just rebuild` completed something in the meantime.
Three properties make this the right source of truth:
- **It is what the build itself checks.** Any other store can disagree with the bucket; this one cannot.
- **It is agent-independent.** `.crow/weekly-rebuild-missing.yaml` mounts no `volumes:`, unlike `.crow/build-all-versions.yaml:132-133`, so `/mnt/cache` is per-job and cannot carry progress anyway.
- **It costs one listing.** `cranlike`'s `s3` fork already does exactly this call against this bucket at ~24 000 objects, so the approach is proven at the required scale.
It must read ETags rather than the slot index's `Built` field, which is how `local/packages-to-build.R:104-130` answers the same question.
Under this design the index is not rewritten until the dependent re-index pipeline runs (section 5), so mid-run it cannot reflect the current run's progress.
Packages that genuinely fail to build re-publish their CRAN source, so they stay outstanding and would be retried on every restart.
That is already handled upstream: `bincraft::filter_packages_with_errors()` (`R/build_binaries.R:1018`, `:1143`) drops anything with `error_occurred = TRUE`, and `store_build_metadata = TRUE` is passed on every call.
No additional poison-pill filter is needed here.
Only the flat `src/contrib` path is considered.
The rebuild call passes no `is_r_minor_sensitive`, so it defaults to `FALSE` and only ever targets the flat path; the resume filter matches that scope deliberately.
### 4. Give each shard a wall-clock budget
`local/rebuild-missing.R` takes a budget, defaulting to 20 hours, and breaks out of the loop once it is exceeded:
```r
# exit cleanly rather than being killed, so the dependent re-index still runs
if (difftime(Sys.time(), started, units = "hours") > budget_hours) {
cat(sprintf("Budget of %sh reached after %d/%d packages; stopping cleanly\n", budget_hours, i, n))
break
}
```
It exits 0 and reports how much of the slice it covered.
Every run then has a terminating condition, the re-index and purge always fire, and the remainder is picked up by the next run with no bookkeeping, because section 3 recomputes the outstanding set from scratch.
### 5. Move the re-index and purge into `.crow/weekly-rebuild-reindex.yaml`
Three shards per slot means three concurrent `upload_package_index()` calls on the same S3 prefix.
`cranlike::update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work.
The re-index steps (`.crow/weekly-rebuild-missing.yaml:171-176`) and the purge step (`:187-207`) therefore leave that file entirely.
The new file carries:
```yaml
depends_on:
- weekly-rebuild-missing
runs_on: [success, failure]
```
`runs_on: [success, failure]` validates as a workflow-level key under `crow lint`, so a failing shard no longer withholds the re-index.
The file uses the same 18-row matrix and the same `when:` gating as `weekly-rebuild-missing`, so it only re-indexes slots that actually ran.
Each row re-indexes the flat slot and every per-minor slot.
`scripts/purge_cdn_zone.sh` runs once on a single row, because all hostnames share pull zone `3857050` and 18 identical zone purges would be waste.
## Failure behaviour
| case | today | after |
| -------------------------- | ----------------------------------- | ----------------------------------------------------- |
| one package errors | `tryCatch` logs, loop continues | unchanged |
| a shard fails outright | purge runs, re-index does not | re-index and purge run via `runs_on` |
| a shard exceeds its budget | cannot happen, runs until killed | exits 0, re-index and purge run |
| a shard is killed | nothing runs | still nothing; trigger the re-index pipeline alone |
| a shard restarts | re-walks the list, HEAD per package | one listing, resumes at the first outstanding package |
The known cost of `depends_on` being file-level rather than row-level: on the weekly cron no slot is re-indexed until the slowest of all 54 jobs finishes.
The 20-hour budget bounds that at roughly one day.
## Out of scope
- `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row for the platform and arch, which is precisely the source-fallback set. That is a separate change.
- Bunny Perma-Cache eviction. `scripts/purge_cdn_zone.sh` purges the regular edge cache only; see the note in `CLAUDE.md` and issue history.
- The audit that produces the rebuild list is unchanged.
## Verification
- `crow lint .crow/` passes for both pipeline files.
- `local/rebuild-missing.R` gets unit coverage in `local/tests/` for the two pure pieces: the interleaved slice (disjoint, covering, deterministic) and the outstanding-set filter (source-served ETag kept, binary ETag dropped, absent object kept).
- A single-slot manual run of `alpine-324-amd64` shard 1 confirms the listing shortcut against the live bucket, and that the reported outstanding count is close to the 8 917 measured above divided by three.
- Restarting that shard mid-run confirms it resumes rather than replaying, by comparing the outstanding count it reports on the second start.