Repository navigation
Conversation
|
This looks good. I'll take some time to look over the changeset. I'll switch some of my projects over to the develop branch and start testing it. Should we invite others to do the same, with an issue or a discussion? Can you describe how this changes the maintenance routine, if at all? How do we add new toolchain versions when it becomes periodically necessary? Does the CI work the same way, discovering/reporting compatibility etc? |
|
Maybe we can tag a v2.rc1 or v2.beta1 to advertise the new branch rather than suggesting to use |
|
I see now @minhqdao made a tag already
|
Yes, absolutely. I'll create an issue. And the maintenance routine is an excellent point, too. I'll have some thoughts on that and on how to improve it, too, then I'll add a dedicated section. The CI works a little bit differently from before to support a larger number of jobs. In short, the job limit of 256 only applies to a single file – split it into one file per compiler and we can definitely test all permutations. I think the CI is self-explanatory. You can start reading things from there. And the entry point of the Action is always For maintainers, the biggest difference might be the absence of a unified compatibility matrix. I found the old one to grow a bit too crowded already, and with added compilers, platforms, and versions, it just made more sense to move to a "one table per compiler" design. But if you are now missing any maintainer features of the old design, then we should think about how to get those features back. And yes, the tag is
|
|
Tried with your new beta but still the same issue :s https://github.com/certik/fplotlib/actions/runs/35761151145/job/106865586419?pr=24#step:3:109 |
|
Sometimes we get bad responses from APT. It lasts for some time and then stops. I don't know if retries will be of much use unless you set a much higher retry count and make the wait time unreasonably long. IMO it's better for the action to fail relatively quickly rather than waiting a while for transient network errors to resolve, but maybe others feel differently. |
|
Let's try again with We could probably do without retries and timeouts if our test matrices were much smaller. But with this many jobs, it would be very hard to get CI green on the first attempt. The retries have already saved us a lot of manual reruns. |
|
The point of retries/timeouts is UX, right? It's nice that they make tests likelier to pass, but that just reflects the fact that they give the action a bit of fault tolerance. It's a design decision to be made mainly from a user perspective, rather than a developer's, I mean. |
|
I'd put retries/timeouts in the resilience/robustness rather than the UX department. There have been issues with GitHub Actions lately that persisted for days, and I think having those in place helps. To address your main concern: I usually find it easy to distinguish a transient network hiccup from a deterministic error. I don't think a retry will easily mask a misconfigured setup. |
|
isn't resilience part of the user experience? 🙂 just to clarify I agree retries are a good idea. I think they'd be a good idea even ignoring the tests. I meant that retries may not help much with long 3rd-party service outages. they will definitely help with short hiccups. I'm not worried a network error will mask a misconfiguration. If something is down for say 3 hours, I'd just rather get a failure quickly and rerun the job later, than have it keep retrying with longer and longer delay until the service is back. |
|
Regarding the retries/timeouts: Matches what I found on the Sept 15-16 ifort macOS DNS failures (getaddrinfo ENOTFOUND registrationcenter-download.intel.com) - the outage lasted hours, not minutes (a job succeeded ~15:40 UTC, a later one was still failing at 17:56-17:58 UTC the same day, and it recurred again the next day), which is also why the ~6 min DNS-wait loop got reverted (c99ad8f). Each failed attempt itself is fast though, only ~2 min for the whole tool-cache + curl cascade, so short retries barely cost anything even when they can't outlast the outage. I tried a longer retry window plus an exit-code-based classifier for network-vs-real errors on the ifort workflow, and dropped both - no reasonable retry window covers an hours-long outage, and pattern-matching on the error can't see a hang killed by its own timeout, which is a blind spot a human reading the log doesn't have. +1 on failing fast for long outages instead of an ever-growing backoff. The short, per-attempt retries already living inside the installers (#274) feel like a separate, still-worth-keeping layer either way. |
|
I'm also fine with calling it user experience :) Initially I had zero retries and added them only when real hiccups/rate-limiting happened and saw that a retry would actually help. If I could confirm that the matrix finished more reliably, I'd keep that retry, otherwise I'd remove it again (just as @djkees identified with c99ad8f). Up to some point, I'd say they were indeed evidence-based, but I might have lost track of them a little bit. So maybe this is a good time to revisit all the retries and timeouts that we have and record the "maximum delay until complete failure" for each case. I think it will still be difficult to distinguish hiccups/rate-limiting from a longer network outage, but I think it'll be good to have an overview and then see if we can reduce some. I'll create an issue and we can continue that conversation there. |
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* Add retry to installer downloads/installs that had none Mirrors the retry-with-backoff patterns already used by ifx/win32.ts and gfortran/darwin.ts, applied to ifort/win32.ts, aocc/debian.ts, and flang/darwin.ts's download and brew-install calls, plus lfortran/debian.ts's conda create (already had a shared retry helper, just wasn't using it). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(flang): require explicit destination, match gfortran's brew flags Addresses review: destination is no longer optional in downloadToolWithRetry (removes the dead cleanup guard), and brewInstallWithRetry now passes --skip-post-install like gfortran's does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This is a full rewrite (
bashscript --> TypeScript) that addresses the majority of the open issues. Additionally it now:Test with:
flangandarmflang.gfortran,nvfortran,flang, andarmflangonubuntu-arm.flangonwindows-arm.ucrt64andclang64on Windows.This will be the README after the merge: https://github.com/fortran-lang/setup-fortran/tree/develop
Transferred from https://github.com/minhqdao/setup-fortran
Breaking changes:
ifxon macOS now fails instead of silently redirecting toifort.ifxversions are now addressed by compiler version instead of their oneAPI release number (e.g.2022.2.0instead of2022.1).Closes #11, #12, #32, #98, #104, #125, #126, #129, #131, #246, #248, #249
Thanks to @wpbonelli for maintaining
setup-fortranthroughout the years and inviting its replacement, and to @awvwgk initially authoring this project and for the green light.