feat(rebuild): shard the weekly rebuild and make each shard resumable (#163)
Some checks failed
ci/crow/cron/process-updates/7 Pipeline was successful
ci/crow/cron/process-updates/8 Pipeline was successful
ci/crow/cron/process-updates/9 Pipeline was successful
ci/crow/cron/process-updates/3 Pipeline was successful
ci/crow/cron/process-updates/10 Pipeline was successful
ci/crow/cron/process-updates/4 Pipeline was successful
ci/crow/cron/weekly-audit-missing/6 Pipeline was successful
ci/crow/cron/weekly-audit-missing/5 Pipeline was successful
ci/crow/cron/process-updates/13 Pipeline was successful
ci/crow/cron/process-updates/14 Pipeline was successful
ci/crow/cron/process-updates/15 Pipeline was successful
ci/crow/cron/process-updates/16 Pipeline failed
ci/crow/cron/process-updates/18 Pipeline was successful
ci/crow/cron/process-updates/11 Pipeline was successful
ci/crow/cron/process-updates/17 Pipeline was successful
ci/crow/cron/process-updates/12 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/18 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/16 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/17 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/6 Pipeline was successful
ci/crow/cron/process-updates/6 Pipeline was successful
ci/crow/cron/process-updates/5 Pipeline was successful
ci/crow/cron/process-updates/1 Pipeline was successful
ci/crow/cron/process-updates/2 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/51 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/13 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/14 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/15 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/5 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/49 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/17 Pipeline failed
ci/crow/manual/weekly-rebuild-missing/50 Pipeline was successful

## Problem

`weekly-rebuild-missing` runs one job per `<os>-<arch>` and walks that slot's list serially in a single `R -q -e` argument.
That was cheap while every source fallback was skipped as "already built".
Since bincraft #105/#106/#107 and #159 the gate works, and the lists are large: 8 917 source-served records on `amd64/alpine324`, 15 023 on `amd64/resolute`.

Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) ran for two days, reached `[8692/23885] cholera`, and was killed there.

Two failures follow from that shape:

- **No parallelism.** The work is embarrassingly parallel across packages; one job does all of it.
- **No resumability and no clean stopping point.** The loop ends only by exhausting the list, so the only way to stop it is a kill. A restart re-walks from the first entry, paying a CRAN version resolution and an S3 `HEAD` per package before reaching new work. And a kill matches neither `success` nor `failure`, so the `Purge CDN cache` step never ran: the ~4 600 binaries 10910 did publish stayed hidden behind stale edge copies.

## What this changes

**Three shards per slot.** Each of the 18 `OS`/`ARCH` rows gains `SPLIT_INTO`/`SPLIT_INDEX`, mirroring `build-all-versions.yaml`. Cron and manual routing are unchanged: both filters already match on `${OS}-${ARCH}`, so they now match all three shards of a slot.

**`local/rebuild-missing.R`** replaces the ~1 500-character inline one-liner. The slice is interleaved rather than contiguous, because the list is alphabetical and cost clusters by name (`Rcpp*`, `Bioc*`, `rstan*`).

**Resume by re-deriving state from the bucket.** One `s3_dir_info()` listing gives ETags for the slot; a package is outstanding iff its object's ETag equals CRAN's published `MD5sum`, i.e. it is still byte-identical to CRAN's source. That is `check_s3_root_package()` evaluated in bulk. No progress file, no volume, no DB cursor, and correct when a sibling shard or a `process-updates` run completes something concurrently.

It reads ETags rather than the index's `Built` field the way `packages-to-build.R` does, because the index is no longer rewritten until the dependent pipeline runs and so cannot reflect the current run's progress.

Unknown always means "already a binary", never "rebuild it": a multipart ETag, an unreadable CRAN index or an empty listing can never mass-schedule work.

**A 20 h wall-clock budget** per shard. It exits 0, so the re-index and purge always fire and the remainder is picked up next run with no bookkeeping.

**`.crow/weekly-rebuild-reindex.yaml`** takes over re-indexing and the purge, with `depends_on: [weekly-rebuild-missing]` and `runs_on: [success, failure]`. Three shards writing one slot's `PACKAGES` concurrently would race: `update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work.

## Verification

`crow lint .crow/` passes on all 11 pipelines. `prek run` passes.

19 assertions in `local/tests/test-rebuild-missing.R`, 0 failures, covering the partition (disjoint, covering, deterministic, short lists, out-of-range index) and the outstanding filter (source ETag kept, binary ETag dropped, absent object kept, multipart and missing-from-CRAN treated as built).

One of those tests caught a real bug before it shipped: an empty ETag table indexed to zero length rather than to `NA`, which recycled the result away and reported "nothing to build" — the dangerous direction. Fixed with an explicit `lookup()`.

The filter run against the live `amd64/alpine324` index, using its `MD5sum` column as the ETag (established to match the objects):

```
index packages:               24343
outstanding (filter):          8950
no Built stamp:                8917
filter vs no-Built agreement:  8917 of 8917
outstanding but stamped Built:   33 (version drift vs CRAN)
shard sizes: 2984/2983/2983 (sum 8950, unique 8950)
```

It reproduces the source-served set exactly. The extra 33 are packages whose slot version differs from CRAN's current one, so no object exists at the CRAN version key: correctly outstanding.

## Notes for review

- The 20 h budget is a chosen default, exposed as `REBUILD_BUDGET_HOURS` in the pipeline.
- `depends_on` is file-level, not row-level, so on a full cron run no slot is re-indexed until the slowest of all 54 jobs finishes. The budget bounds that at roughly a day.
- An explicit cancel still skips the re-index. Recovery is to trigger `weekly-rebuild-reindex` on its own.
- The purge runs on every re-index row rather than one designated slot: a cron fires only its own slot's row, so gating on a named slot would leave every other slot unpurged.
- Out of scope: `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row, which is precisely the source-fallback set.

Design: `specs/2026-08-12-shard-weekly-rebuild-design.md`
Reviewed-on: #163
This commit is contained in:
Patrick Schratz 2026-08-12 08:30:29 +00:00 committed by Patrick Schratz
commit 4b7dc28cc8

View file

@ -0,0 +1,193 @@
# Re-index and purge after weekly-rebuild-missing.
#
# weekly-rebuild-missing runs three shards per slot. Each of them replaces
# objects in place, so the slot's index still advertises the old MD5 and, for
# anything that had been served from source, no Built stamp. Re-indexing from
# inside a shard would mean three concurrent `upload_package_index()` calls on
# one prefix: `cranlike::update_PACKAGES()` lists the live bucket, so an early
# lister that uploads last publishes an index missing its siblings' work.
#
# So it happens exactly once per slot, here, after every shard has finished.
# `runs_on: [success, failure]` keeps that true when a shard fails; only an
# explicit cancel skips it, and this pipeline can then be triggered on its own.
variables:
# Mirrors the gate on weekly-rebuild-missing so a manual run re-indexes
# exactly the slots it rebuilt. A manual pipeline creation instantiates every
# file in .crow/, so the default must match no matrix row.
weekly_rebuild_missing:
description: "Manual run target: a specific <os>-<arch>, 'all' for every os/arch, or 'none' to run nothing."
options:
- none
- all
- alpine-322-amd64
- alpine-322-arm64
- alpine-323-amd64
- alpine-323-arm64
- alpine-324-amd64
- alpine-324-arm64
- redhat-8-amd64
- redhat-8-arm64
- redhat-9-amd64
- redhat-9-arm64
- redhat-10-amd64
- redhat-10-arm64
- ubuntu-2204-amd64
- ubuntu-2204-arm64
- ubuntu-2404-amd64
- ubuntu-2404-arm64
- ubuntu-2604-amd64
- ubuntu-2604-arm64
default: none
when:
- event: cron
cron: weekly-rebuild-missing-${OS}-${ARCH}
- event: manual
evaluate: 'weekly_rebuild_missing == "all" || weekly_rebuild_missing == "${OS}-${ARCH}"'
depends_on:
- weekly-rebuild-missing
runs_on: [success, failure]
skip_clone: true
labels:
group: rpkgs-${ARCH}
matrix:
include:
- OS: alpine-322
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.22
- OS: alpine-322
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.22
- OS: alpine-323
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.23
- OS: alpine-323
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.23
- OS: alpine-324
ARCH: amd64
R_VERSION: 4.5.3
IMG: alpine:3.24
- OS: alpine-324
ARCH: arm64
R_VERSION: 4.5.3
IMG: alpine:3.24
- OS: redhat-8
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:8
- OS: redhat-8
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:8
- OS: redhat-9
ARCH: amd64
R_VERSION: 4.4.3
IMG: redhat:9
- OS: redhat-9
ARCH: arm64
R_VERSION: 4.4.3
IMG: redhat:9
- OS: redhat-10
ARCH: amd64
R_VERSION: 4.5.3
IMG: redhat:10
- OS: redhat-10
ARCH: arm64
R_VERSION: 4.5.3
IMG: redhat:10
- OS: ubuntu-2204
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
- OS: ubuntu-2204
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:jammy
- OS: ubuntu-2404
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:noble
- OS: ubuntu-2404
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:noble
- OS: ubuntu-2604
ARCH: amd64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
- OS: ubuntu-2604
ARCH: arm64
R_VERSION: 4.4.3
IMG: ubuntu:resolute
steps:
- name: 'Re-index the slot'
image: reg.devxy.io/rpkgs/build-env-${IMG}
pull: true
environment:
OTEL_R_TRACES_EXPORTER: none
OTEL_R_LOGS_EXPORTER: none
OTEL_R_METRICS_EXPORTER: none
RED_HAT_DEV_PW:
from_secret: RED_HAT_DEV_PW
B2_S3_ACCESS_KEY:
from_secret: B2_S3_ACCESS_KEY
B2_S3_SECRET_KEY:
from_secret: B2_S3_SECRET_KEY
REPO_RO_TOKEN:
from_secret: REPO_RO_TOKEN
GIT_USER: pat-s
R_LIBS_USER: /mnt/cache/R-pkgs
R_VERSION: ${R_VERSION}
PLATFORM: ${OS}
ARCH: ${ARCH}
commands:
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- mkdir -p /mnt/cache/R-pkgs
- rm -rf /mnt/cache/R-pkgs/00LOCK-*
- /opt/R/$R_VERSION/bin/Rscript local/install-bincraft.R
# The codename is detected from the image's /etc/os-release.
- /opt/R/$R_VERSION/bin/R -q -e 'library(bincraft); upload_package_index(s3_endpoint = "https://s3.eu-central-003.backblazeb2.com", s3_region = "eu-central-003", s3_bucket = "devxy-rpkgs-binaries", s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"), s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"))'
- |
for RBIN in /opt/R/[0-9]*/bin/R; do
RMINOR=$(basename "$(dirname "$(dirname "$RBIN")")" | cut -d. -f1-2)
/opt/R/$R_VERSION/bin/R -q -e "library(bincraft); upload_package_index(r_minor = '$RMINOR', s3_endpoint = 'https://s3.eu-central-003.backblazeb2.com', s3_region = 'eu-central-003', s3_bucket = 'devxy-rpkgs-binaries', s3_access_key_id = Sys.getenv('B2_S3_ACCESS_KEY'), s3_secret_access_key = Sys.getenv('B2_S3_SECRET_KEY'))" || true
done
- name: Purge CDN cache
image: reg.devxy.io/docker.io/library/alpine:3.24
environment:
OTEL_R_TRACES_EXPORTER: none
OTEL_R_LOGS_EXPORTER: none
OTEL_R_METRICS_EXPORTER: none
BUNNYNET_API_KEY:
from_secret: BUNNYNET_API_KEY
REPO_RO_TOKEN:
from_secret: REPO_RO_TOKEN
# All hostnames on the zone share this id, so one purge covers
# cran.devxy.io, cran.allianceswisspass.devxy.io and cran.rpkgs.com.
BUNNY_PULLZONE: '3857050'
commands:
- apk add --no-cache -q bash curl git
- git clone -q https://pat-s:$$REPO_RO_TOKEN@git.devxy.io/devxy/build-cran-binaries.git .
- bash scripts/purge_cdn_zone.sh "$BUNNYNET_API_KEY" "$BUNNY_PULLZONE"
# Runs on every row rather than on one designated slot: a cron fires only
# its own slot's row, so gating on a named slot would leave every other
# slot unpurged. A manual "all" run therefore purges the zone 18 times,
# which is a cheap API call and rare.
#
# Run it even when the re-index above failed: the objects were still
# replaced, and a stale edge is exactly what keeps them hidden.
when:
- status: [success, failure]