fix(rebuild): re-index and purge the CDN after a rebuild #160

Merged
pat-s merged 2 commits from fix/rebuild-reindex-and-purge into main 2026-08-10 06:30:36 +00:00
Owner

Problem

The rebuild now works — AATtools 0.0.3 was detected as a source fallback, built, and published:

ℹ `upload_single_binary()`: Replacing the CRAN source published for AATtools 0.0.3 … with the binary.
✔ Successfully uploaded package AATtools with tag 0.0.3.

But clients still get the source, and will for about a year:

$ curl -sI .../src/contrib/AATtools_0.0.3.tar.gz
etag: "ea8127d953ca6a2f118ea49441772af6"   # CRAN's source MD5
cdn-cache: HIT
cdn-cachedat: 08/09/2026 16:43:32          # predates the 18:12 upload

Two causes, both specific to a rebuild:

  1. The slot is never re-indexed. weekly-rebuild-missing has no upload_package_index step, so the index keeps the old MD5 and — for anything that had been served from source — no Built stamp. This one self-heals at the next process-updates run.
  2. The tarball URL is never purged. A normal update publishes new packages at new URLs, so purge_cdn_cache.sh only needs the five index files. A rebuild replaces an object in place, and the zone caches tarballs for cache_expiration_time = 31919000 (~370 days). This does not self-heal.

Nothing about a stale package looks wrong from the outside, which is what makes it worth fixing rather than documenting.

What this changes

Re-index at the end of a rebuild, flat and per-minor, mirroring the tail of process-updates. The codename is detected from the image's /etc/os-release (as local/packages-to-build.R already does) rather than adding OS_ID to all 18 matrix rows.

Purge the zone afterwards, via a new scripts/purge_cdn_zone.sh. One call to POST /pullzone/{id}/purgeCache covers every replaced object, and all three hostnames — cran.devxy.io, cran.allianceswisspass.devxy.io, cran.rpkgs.com — share pull zone 3857050, confirmed from the cdn-pullzone response header.

Purging per URL was the alternative and is worse here: ~13.5k rate-limited calls per arch, where a single missed call leaves a package silently stale. The cost of the zone purge is a cold cache for everything else, which is why it stays out of the daily update path — purge_cdn_cache.sh is untouched.

The purge runs on failure too (when: status: [success, failure]): a rebuild that died part-way still replaced objects, and those are exactly the ones a stale edge keeps hiding.

Verification

crow lint .crow/ reports all ten configs valid; bash -n on the new script passes; prek hooks pass.

Not yet exercised against Bunny — it needs BUNNYNET_API_KEY, which is a CI secret. The failure mode is explicit rather than silent: any status other than 200/204 prints the response body and exits non-zero.

## Problem The rebuild now works — `AATtools 0.0.3` was detected as a source fallback, built, and published: ``` ℹ `upload_single_binary()`: Replacing the CRAN source published for AATtools 0.0.3 … with the binary. ✔ Successfully uploaded package AATtools with tag 0.0.3. ``` But clients still get the source, and will for about a year: ``` $ curl -sI .../src/contrib/AATtools_0.0.3.tar.gz etag: "ea8127d953ca6a2f118ea49441772af6" # CRAN's source MD5 cdn-cache: HIT cdn-cachedat: 08/09/2026 16:43:32 # predates the 18:12 upload ``` Two causes, both specific to a rebuild: 1. **The slot is never re-indexed.** `weekly-rebuild-missing` has no `upload_package_index` step, so the index keeps the old MD5 and — for anything that had been served from source — no `Built` stamp. This one self-heals at the next `process-updates` run. 2. **The tarball URL is never purged.** A normal update publishes new packages at *new* URLs, so `purge_cdn_cache.sh` only needs the five index files. A rebuild replaces an object *in place*, and the zone caches tarballs for `cache_expiration_time = 31919000` (~370 days). This does not self-heal. Nothing about a stale package looks wrong from the outside, which is what makes it worth fixing rather than documenting. ## What this changes **Re-index at the end of a rebuild**, flat and per-minor, mirroring the tail of `process-updates`. The codename is detected from the image's `/etc/os-release` (as `local/packages-to-build.R` already does) rather than adding `OS_ID` to all 18 matrix rows. **Purge the zone afterwards**, via a new `scripts/purge_cdn_zone.sh`. One call to `POST /pullzone/{id}/purgeCache` covers every replaced object, and all three hostnames — `cran.devxy.io`, `cran.allianceswisspass.devxy.io`, `cran.rpkgs.com` — share pull zone `3857050`, confirmed from the `cdn-pullzone` response header. Purging per URL was the alternative and is worse here: ~13.5k rate-limited calls per arch, where a single missed call leaves a package silently stale. The cost of the zone purge is a cold cache for everything else, which is why it stays out of the daily update path — `purge_cdn_cache.sh` is untouched. The purge runs on failure too (`when: status: [success, failure]`): a rebuild that died part-way still replaced objects, and those are exactly the ones a stale edge keeps hiding. ## Verification `crow lint .crow/` reports all ten configs valid; `bash -n` on the new script passes; prek hooks pass. Not yet exercised against Bunny — it needs `BUNNYNET_API_KEY`, which is a CI secret. The failure mode is explicit rather than silent: any status other than 200/204 prints the response body and exits non-zero.
The cache handed to build_binary_package() as s3_package_cache was the raw
bucket listing, and the build list subtracted the same listing. Neither could
tell a binary from a package whose build failed and was published as its CRAN
source, so every source fallback read as "already built" and was skipped for
good - which is how alpine324 accumulated ~13.5k of them.

- drop objects the slot's index reports as served from source, which bincraft
  marks by leaving the Built stamp off
- build s3_dt from the filtered listing too, since it is subtracted from the
  build list and would otherwise exclude the packages that need building
- log how many objects were dropped

Archived objects have no index record and are kept: unknown means binary, never
"rebuild it". A slot last indexed by a bincraft that predates the Built change
stamps everything, so its cache is unchanged from before.
A rebuild replaces an object in place: a package whose build failed was
published as its CRAN source, and the rebuilt binary takes exactly the same
URL. Two things then hide the result from clients.

The slot's index still advertises the old MD5 and, for anything served from
source, no Built stamp, because weekly-rebuild-missing never re-indexed. And
the pull zone caches tarballs for ~370 days, while purge_cdn_cache.sh only
purges the five index files, so the edge keeps serving the source tarball for
up to a year with nothing about it looking wrong.

Observed after rebuilding AATtools 0.0.3: the pipeline reported a successful
upload while the edge still served the CRAN source, etag ea8127... and no Meta/.

- re-index the slot at the end of a rebuild, flat and per-minor, detecting the
  codename from the image rather than adding OS_ID to 18 matrix rows
- add scripts/purge_cdn_zone.sh and call it afterwards. One zone purge covers
  every replaced object and all three hostnames, which share pull zone 3857050;
  purging per URL would be ~13.5k rate-limited calls per arch where one missed
  call leaves a silently stale package
- purge on failure too, since a rebuild that died part-way still replaced
  objects and those are exactly the ones a stale edge keeps hiding
pat-s merged commit 01b8ab43df into main 2026-08-10 06:30:36 +00:00
pat-s deleted branch fix/rebuild-reindex-and-purge 2026-08-10 06:30:36 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
devxy/build-cran-binaries!160
No description provided.