feat(rebuild): shard the weekly rebuild and make each shard resumable (#163)
Some checks failed
ci/crow/cron/process-updates/7 Pipeline was successful
ci/crow/cron/process-updates/8 Pipeline was successful
ci/crow/cron/process-updates/9 Pipeline was successful
ci/crow/cron/process-updates/3 Pipeline was successful
ci/crow/cron/process-updates/10 Pipeline was successful
ci/crow/cron/process-updates/4 Pipeline was successful
ci/crow/cron/weekly-audit-missing/6 Pipeline was successful
ci/crow/cron/weekly-audit-missing/5 Pipeline was successful
ci/crow/cron/process-updates/13 Pipeline was successful
ci/crow/cron/process-updates/14 Pipeline was successful
ci/crow/cron/process-updates/15 Pipeline was successful
ci/crow/cron/process-updates/16 Pipeline failed
ci/crow/cron/process-updates/18 Pipeline was successful
ci/crow/cron/process-updates/11 Pipeline was successful
ci/crow/cron/process-updates/17 Pipeline was successful
ci/crow/cron/process-updates/12 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/18 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/16 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/17 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/6 Pipeline was successful
ci/crow/cron/process-updates/6 Pipeline was successful
ci/crow/cron/process-updates/5 Pipeline was successful
ci/crow/cron/process-updates/1 Pipeline was successful
ci/crow/cron/process-updates/2 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/51 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/13 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/14 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/15 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/5 Pipeline was successful
ci/crow/manual/weekly-rebuild-missing/49 Pipeline was successful
ci/crow/manual/weekly-rebuild-reindex/17 Pipeline failed
ci/crow/manual/weekly-rebuild-missing/50 Pipeline was successful

## Problem

`weekly-rebuild-missing` runs one job per `<os>-<arch>` and walks that slot's list serially in a single `R -q -e` argument.
That was cheap while every source fallback was skipped as "already built".
Since bincraft #105/#106/#107 and #159 the gate works, and the lists are large: 8 917 source-served records on `amd64/alpine324`, 15 023 on `amd64/resolute`.

Pipeline 10910 (`weekly_rebuild_missing:alpine-324-amd64`) ran for two days, reached `[8692/23885] cholera`, and was killed there.

Two failures follow from that shape:

- **No parallelism.** The work is embarrassingly parallel across packages; one job does all of it.
- **No resumability and no clean stopping point.** The loop ends only by exhausting the list, so the only way to stop it is a kill. A restart re-walks from the first entry, paying a CRAN version resolution and an S3 `HEAD` per package before reaching new work. And a kill matches neither `success` nor `failure`, so the `Purge CDN cache` step never ran: the ~4 600 binaries 10910 did publish stayed hidden behind stale edge copies.

## What this changes

**Three shards per slot.** Each of the 18 `OS`/`ARCH` rows gains `SPLIT_INTO`/`SPLIT_INDEX`, mirroring `build-all-versions.yaml`. Cron and manual routing are unchanged: both filters already match on `${OS}-${ARCH}`, so they now match all three shards of a slot.

**`local/rebuild-missing.R`** replaces the ~1 500-character inline one-liner. The slice is interleaved rather than contiguous, because the list is alphabetical and cost clusters by name (`Rcpp*`, `Bioc*`, `rstan*`).

**Resume by re-deriving state from the bucket.** One `s3_dir_info()` listing gives ETags for the slot; a package is outstanding iff its object's ETag equals CRAN's published `MD5sum`, i.e. it is still byte-identical to CRAN's source. That is `check_s3_root_package()` evaluated in bulk. No progress file, no volume, no DB cursor, and correct when a sibling shard or a `process-updates` run completes something concurrently.

It reads ETags rather than the index's `Built` field the way `packages-to-build.R` does, because the index is no longer rewritten until the dependent pipeline runs and so cannot reflect the current run's progress.

Unknown always means "already a binary", never "rebuild it": a multipart ETag, an unreadable CRAN index or an empty listing can never mass-schedule work.

**A 20 h wall-clock budget** per shard. It exits 0, so the re-index and purge always fire and the remainder is picked up next run with no bookkeeping.

**`.crow/weekly-rebuild-reindex.yaml`** takes over re-indexing and the purge, with `depends_on: [weekly-rebuild-missing]` and `runs_on: [success, failure]`. Three shards writing one slot's `PACKAGES` concurrently would race: `update_PACKAGES()` lists the live bucket, so an early lister that uploads last publishes an index missing its siblings' work.

## Verification

`crow lint .crow/` passes on all 11 pipelines. `prek run` passes.

19 assertions in `local/tests/test-rebuild-missing.R`, 0 failures, covering the partition (disjoint, covering, deterministic, short lists, out-of-range index) and the outstanding filter (source ETag kept, binary ETag dropped, absent object kept, multipart and missing-from-CRAN treated as built).

One of those tests caught a real bug before it shipped: an empty ETag table indexed to zero length rather than to `NA`, which recycled the result away and reported "nothing to build" — the dangerous direction. Fixed with an explicit `lookup()`.

The filter run against the live `amd64/alpine324` index, using its `MD5sum` column as the ETag (established to match the objects):

```
index packages:               24343
outstanding (filter):          8950
no Built stamp:                8917
filter vs no-Built agreement:  8917 of 8917
outstanding but stamped Built:   33 (version drift vs CRAN)
shard sizes: 2984/2983/2983 (sum 8950, unique 8950)
```

It reproduces the source-served set exactly. The extra 33 are packages whose slot version differs from CRAN's current one, so no object exists at the CRAN version key: correctly outstanding.

## Notes for review

- The 20 h budget is a chosen default, exposed as `REBUILD_BUDGET_HOURS` in the pipeline.
- `depends_on` is file-level, not row-level, so on a full cron run no slot is re-indexed until the slowest of all 54 jobs finishes. The budget bounds that at roughly a day.
- An explicit cancel still skips the re-index. Recovery is to trigger `weekly-rebuild-reindex` on its own.
- The purge runs on every re-index row rather than one designated slot: a cron fires only its own slot's row, so gating on a named slot would leave every other slot unpurged.
- Out of scope: `build-all-versions` still cannot rebuild source fallbacks, because `local/build-all.R:113-122` drops every version with any `single_builds` row, which is precisely the source-fallback set.

Design: `specs/2026-08-12-shard-weekly-rebuild-design.md`
Reviewed-on: #163
This commit is contained in:
Patrick Schratz 2026-08-12 08:30:29 +00:00 committed by Patrick Schratz
commit 4b7dc28cc8
3 changed files with 1005 additions and 36 deletions

View file

@ -0,0 +1,87 @@
# Pure helpers for local/rebuild-missing.R, kept separate so local/tests can
# source them without executing a rebuild.
# Interleaved slice of the rebuild list.
#
# The list is alphabetical and build cost clusters by name (Rcpp*, Bioc*,
# rstan*), so contiguous thirds would be badly unbalanced. Interleaving also
# makes each shard's progress counter representative of the slot as a whole.
shard_slice <- function(pkgs, split_into, split_index) {
split_into <- as.integer(split_into)
split_index <- as.integer(split_index)
if (is.na(split_into) || is.na(split_index)) {
stop("shard_slice(): split_into and split_index must be integers")
}
if (split_into < 1L || split_index < 1L || split_index > split_into) {
stop(sprintf(
"shard_slice(): need 1 <= split_index <= split_into, got %s of %s",
split_index,
split_into
))
}
# seq() errors on a descending range, which is what an empty list or a shard
# index past the end would produce.
if (length(pkgs) < split_index) {
return(pkgs[0L])
}
pkgs[seq.int(split_index, length(pkgs), by = split_into)]
}
# Packages that still need building, decided from the bucket rather than from
# remembered progress.
#
# This is bincraft's `check_s3_root_package()` evaluated in bulk: an object
# whose ETag equals CRAN's published MD5sum is byte-identical to CRAN's source,
# so the build that was supposed to replace it has not happened yet.
#
# `etag_by_file` named by `<pkg>_<ver>.tar.gz`, values are unquoted ETags
# `cran_version` named by package
# `cran_md5` named by `<pkg>_<ver>`
#
# Unknown always means "already a binary", never "rebuild it", so an unreadable
# CRAN index or a multipart ETag can never mass-schedule work.
outstanding_packages <- function(pkgs, etag_by_file, cran_version, cran_md5) {
if (length(pkgs) == 0L) {
return(pkgs)
}
# An empty table indexes to zero length rather than to NA, which would
# recycle the whole result away and silently report "nothing to build".
lookup <- function(table, key) {
if (length(table) == 0L) {
return(rep(NA_character_, length(key)))
}
unname(as.character(table[key]))
}
version <- lookup(cran_version, pkgs)
file <- sprintf("%s_%s.tar.gz", pkgs, version)
etag <- lookup(etag_by_file, file)
md5 <- lookup(cran_md5, paste(pkgs, version, sep = "_"))
# No CRAN version means the package cannot be resolved to a tarball at all;
# leave it in and let bincraft report why.
unresolved <- is.na(version)
# No object at the key: never built, so it is outstanding by definition.
absent <- !unresolved & is.na(etag)
# A multipart upload carries a compound ETag rather than an MD5.
unknown <- !is.na(etag) & grepl("-", etag, fixed = TRUE)
is_source <- !unresolved &
!is.na(etag) &
!unknown &
!is.na(md5) &
etag == md5
pkgs[unresolved | absent | is_source]
}
parse_rebuild_args <- function(args) {
pos <- args[!startsWith(args, "--")]
budget <- as.numeric(pos[3L])
list(
split_into = as.integer(pos[1L]),
split_index = as.integer(pos[2L]),
budget_hours = if (is.na(budget)) 20 else budget
)
}

194
local/rebuild-missing.R Normal file
View file

@ -0,0 +1,194 @@
### Rebuild one shard of a slot's missing-binary list.
#
# Usage: Rscript local/rebuild-missing.R <split_into> <split_index> [budget_hours]
#
# The list itself comes from local/fetch-rebuild-packages-from-issue.R, which
# writes $REBUILD_PKG_LIST (default /tmp/rebuild_pkgs.txt).
#
# Two properties matter here and are the reason this is a script rather than an
# `R -q -e` argument in the pipeline:
#
# * it is restartable. The outstanding set is re-derived from the bucket on
# every start, so a shard that died resumes where it stopped without any
# progress file, and without replaying thousands of per-package HEADs.
# * it terminates. A wall-clock budget stops the loop cleanly instead of the
# run having to be killed, which is what previously skipped the re-index and
# CDN purge and left rebuilt binaries hidden behind stale edge copies.
options(error = function() {
cat("ERROR:", geterrmessage(), "\n", file = stdout())
traceback(2)
q(status = 1)
})
library(bincraft, quietly = TRUE)
source(file.path("local", "rebuild-missing-helpers.R"))
args <- parse_rebuild_args(commandArgs(trailingOnly = TRUE))
if (is.na(args$split_into) || is.na(args$split_index)) {
stop("usage: rebuild-missing.R <split_into> <split_index> [budget_hours]")
}
list_file <- Sys.getenv("REBUILD_PKG_LIST", "/tmp/rebuild_pkgs.txt")
pkgs <- if (file.exists(list_file)) readLines(list_file) else character(0)
pkgs <- pkgs[nzchar(pkgs)]
if (length(pkgs) == 0L) {
cat("Nothing to rebuild\n")
q("no")
}
excluded <- jsonlite::fromJSON("local/excluded-packages.json")[["package"]]
pkgs <- setdiff(pkgs, excluded)
mine <- shard_slice(pkgs, args$split_into, args$split_index)
cat(sprintf(
"Shard %s/%s: %s of %s listed packages\n",
args$split_index,
args$split_into,
length(mine),
length(pkgs)
))
### Resume: ask the bucket what is still outstanding
codename <- bincraft::set_codename(NULL)
local_machine <- Sys.info()[["machine"]]
arch <- if (grepl("arm64|aarch64", local_machine)) "arm64" else "amd64"
slot_dir <- sprintf(
"devxy-rpkgs-binaries/%s/%s/latest/src/contrib",
arch,
codename
)
s3fs::s3_file_system(
aws_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"),
aws_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"),
endpoint = "https://s3.eu-central-003.backblazeb2.com",
region_name = "eu-central-003",
refresh = TRUE
)
# One paginated listing instead of a HEAD per package. Not recursed: the
# rebuild passes no `is_r_minor_sensitive`, so it only ever targets the flat
# path, and the resume filter matches that scope deliberately.
info <- tryCatch(s3fs::s3_dir_info(slot_dir), error = function(e) NULL)
etag_by_file <- if (is.null(info) || nrow(info) == 0L) {
cat(sprintf(
"WARNING: could not list %s; building the whole shard\n",
slot_dir
))
stats::setNames(character(), character())
} else {
stats::setNames(
gsub('^"|"$', "", as.character(info$etag)),
basename(as.character(info$uri))
)
}
cran <- tryCatch(
{
con <- gzcon(url(
"https://cloud.r-project.org/src/contrib/PACKAGES.gz",
open = "rb"
))
on.exit(close(con), add = TRUE)
read.dcf(con, fields = c("Package", "Version", "MD5sum"))
},
error = function(e) {
cat(sprintf(
"WARNING: could not read CRAN's index (%s)\n",
conditionMessage(e)
))
NULL
}
)
cran_version <- stats::setNames(character(), character())
cran_md5 <- stats::setNames(character(), character())
if (!is.null(cran)) {
cran_version <- stats::setNames(
as.character(cran[, "Version"]),
as.character(cran[, "Package"])
)
keep <- !is.na(cran[, "MD5sum"])
cran_md5 <- stats::setNames(
as.character(cran[keep, "MD5sum"]),
paste(cran[keep, "Package"], cran[keep, "Version"], sep = "_")
)
}
before <- length(mine)
mine <- outstanding_packages(mine, etag_by_file, cran_version, cran_md5)
cat(sprintf(
"Resume: %s of %s already carry a binary; %s outstanding\n",
before - length(mine),
before,
length(mine)
))
if (length(mine) == 0L) {
cat("Nothing outstanding for this shard\n")
q("no")
}
### Build
options(
crayon.enabled = TRUE,
Ncpus = as.integer(Sys.getenv("NCPUS", "2")),
future.globals.onReference = NULL
)
started <- Sys.time()
n <- length(mine)
completed <- 0L
for (i in seq_along(mine)) {
elapsed <- as.numeric(difftime(Sys.time(), started, units = "hours"))
if (elapsed > args$budget_hours) {
cat(sprintf(
"Budget of %sh reached after %d/%d packages; stopping cleanly. The next run resumes from the bucket.\n",
args$budget_hours,
completed,
n
))
break
}
x <- mine[i]
cat(sprintf("[%d/%d] %s\n", i, n, x))
tryCatch(
bincraft::build_binary_package(
x,
tag_limit = 1L,
patches = "local/patches",
s3_endpoint = "https://s3.eu-central-003.backblazeb2.com",
s3_region = "eu-central-003",
s3_bucket = "devxy-rpkgs-binaries",
s3_access_key_id = Sys.getenv("B2_S3_ACCESS_KEY"),
s3_secret_access_key = Sys.getenv("B2_S3_SECRET_KEY"),
metadata_db_host = "r-binaries.devxy.io",
metadata_db_name = "build_metadata",
metadata_db_table = "single_builds",
metadata_db_user = "rpkgs",
metadata_db_password = Sys.getenv("PGPASS"),
metadata_db_sslmode = "require",
metadata_db_port = 15432,
archive = TRUE,
upload = TRUE,
store_build_metadata = TRUE
),
error = function(e) {
cat(sprintf("ERROR building %s - %s\n", x, conditionMessage(e)))
}
)
completed <- completed + 1L
}
cat(sprintf(
"Shard %s/%s finished: %d/%d packages processed in %.1fh\n",
args$split_index,
args$split_into,
completed,
n,
as.numeric(difftime(Sys.time(), started, units = "hours"))
))

View file

@ -0,0 +1,94 @@
source(file.path("..", "rebuild-missing-helpers.R"))
test_that("shard_slice partitions the list without gaps or overlap", {
pkgs <- letters[1:10]
parts <- lapply(1:3, function(i) shard_slice(pkgs, 3, i))
expect_identical(parts[[1]], c("a", "d", "g", "j"))
expect_identical(parts[[2]], c("b", "e", "h"))
expect_identical(parts[[3]], c("c", "f", "i"))
expect_identical(sort(unlist(parts)), sort(pkgs))
expect_identical(anyDuplicated(unlist(parts)), 0L)
})
test_that("shard_slice is deterministic and survives short lists", {
expect_identical(
shard_slice(letters[1:10], 3, 2),
shard_slice(letters[1:10], 3, 2)
)
expect_identical(shard_slice(character(0), 3, 1), character(0))
# more shards than packages: the tail shards get nothing rather than erroring
expect_identical(shard_slice(c("a"), 3, 1), "a")
expect_identical(shard_slice(c("a"), 3, 2), character(0))
})
test_that("shard_slice rejects an out-of-range index", {
expect_error(shard_slice(letters, 3, 4), "split_index")
expect_error(shard_slice(letters, 3, 0), "split_index")
})
test_that("outstanding_packages keeps source fallbacks and drops real binaries", {
cran_version <- c(httr = "1.4.8", R6 = "2.6.1", curl = "7.1.0")
cran_md5 <- c(
httr_1.4.8 = "8756015b94a9cff6f410ca4de8557f12",
R6_2.6.1 = "f01b1787f12797c29194d63c9afd5d70",
curl_7.1.0 = "8af2ccbf5d85dc18866f45f1f26f348d"
)
etag <- c(
# byte-identical to CRAN: the build never happened
"httr_1.4.8.tar.gz" = "8756015b94a9cff6f410ca4de8557f12",
# a real binary was published
"R6_2.6.1.tar.gz" = "9d6087ee9adda3f0a3b8067cfc652c05"
# curl has no object at all
)
out <- outstanding_packages(
c("httr", "R6", "curl"),
etag,
cran_version,
cran_md5
)
expect_identical(out, c("httr", "curl"))
})
test_that("outstanding_packages treats unknowns as already built", {
cran_version <- c(a = "1.0", b = "1.0")
cran_md5 <- c(a_1.0 = "aaaa")
# a multipart ETag carries no MD5, and `b` is missing from CRAN's index:
# neither may schedule a rebuild
etag <- c("a_1.0.tar.gz" = "abc-3", "b_1.0.tar.gz" = "bbbb")
expect_identical(
outstanding_packages(c("a", "b"), etag, cran_version, cran_md5),
character(0)
)
})
test_that("outstanding_packages keeps a package CRAN has no version for", {
out <- outstanding_packages(
"ghost",
c(),
c(other = "1.0"),
c(other_1.0 = "aaaa")
)
expect_identical(out, "ghost")
})
test_that("outstanding_packages handles an empty list", {
expect_identical(
outstanding_packages(character(0), c(), c(), c()),
character(0)
)
})
test_that("parse_rebuild_args defaults the budget", {
a <- parse_rebuild_args(c("3", "2"))
expect_identical(a$split_into, 3L)
expect_identical(a$split_index, 2L)
expect_identical(a$budget_hours, 20)
b <- parse_rebuild_args(c("3", "2", "1.5"))
expect_identical(b$budget_hours, 1.5)
})