Skip to content

fix(ci): ride out transient network flakes instead of failing the run - #27

Merged
mhenrixon merged 1 commit into
dashfrom
fix/ci-flake-resilience
Jul 26, 2026
Merged

mhenrixon merged 1 commit into
dashfrom
fix/ci-flake-resilience

Conversation

@mhenrixon

Copy link
Copy Markdown
Collaborator

Summary

The three red CI runs on dash (merges of #22, #23, #24) were all infrastructure, none were test regressions:

Run Symptom Root cause
#24 merge bundle exit 17 in Install Ruby (rails_edge, Ruby 4.0) rubygems.org SSL reset mid-resolution — every job resolves from scratch since Gemfile.lock is removed
#23 merge every MainTest failed at docker compose build Docker Hub dial tcp i/o timeout pulling base images; build_images_once had no retry, so each test re-hit the outage
#22 merge steps ended with conclusion: null runner died mid-job (nothing to fix repo-side)

Changes:

  • test/integration/integration_test.rb — build_images_once retries docker compose build up to 3 attempts with 15s/30s backoff, matching the compose_up_with_retry idiom right above it
  • .github/workflows/ci.yml — BUNDLE_RETRY: "6" on the tests job; guard the always() disk-usage step so du /mnt/docker no longer paints a second misleading failure when an earlier step is what actually failed

Test plan

  • CI green on this PR (actionlint + zizmor validate the workflow edits)
  • Next transient Hub blip during integration build shows a "retrying" line instead of failing the suite
  • An early job failure no longer flags Check disk usage as a second failed step

Deviations & judgment calls

  • Branched off dash, not main — deliberate exception to the feature-branch rule: all touched files are fork-owned and heavily diverged from upstream (ci.yml matrix, integration harness), so a main-rooted branch would be pure conflict noise and this change is not upstream-PR-able anyway.
  • Retries won't save a sustained outage — the feat(app): expose kamal-proxy rollout as kamal app rollout deploy/set/stop #23-merge Hub outage lasted 4+ minutes; 3 attempts with backoff rides out blips, not incidents. Re-run stays the answer for those, but a blip no longer redlines the suite.
  • No registry mirror for host builds — the compose stack's hub-cache can't serve the host daemon's base-image pulls without a chicken-and-egg bootstrap; not worth the complexity for the observed failure rate.
  • Nothing fixed for the conclusion: null run — that's a GitHub runner death; no repo-side change can prevent it.

## Summary

The last three red CI runs on dash were all infrastructure, not code:

- rubygems.org SSL reset during bundle install (rails_edge, Ruby 4.0) -
  bundler exit 17. Every matrix job resolves from scratch because
  Gemfile.lock is removed, so one blip kills the job. Give bundler
  BUNDLE_RETRY=6 headroom.
- Docker Hub i/o timeouts during `docker compose build` failed every
  MainTest in a run: build_images_once had no retry (unlike
  compose_up_with_retry right above it) and $IMAGES_BUILT never got set,
  so each test re-hit the outage. Retry the build up to 3 times with
  15s/30s backoff.
- The always() disk-usage step ran `du /mnt/docker` even when the Docker
  configure step was skipped, adding a second, misleading red step to
  every early failure. Guard on the directory existing.

## Verification

- [x] ruby -c and rubocop clean on the harness change
- [x] ci.yml edits are env + shell-guard only; actionlint/zizmor run in CI
@mhenrixon mhenrixon self-assigned this Jul 26, 2026
@mhenrixon mhenrixon added the bug Something isn't working label Jul 26, 2026
@mhenrixon
mhenrixon merged commit 46dc295 into dash Jul 26, 2026
17 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant