STAC-25565 Port the beest verification trigger to GitHub Actions - #458
Merged
Conversation
Author
|
Blocked on an App permission, not on review. This workflow calls Order: grant |
LouisParkin
force-pushed
the
STAC-25565-beest-verification
branch
from
August 17, 2026 08:47
44b78c2 to
346723d
Compare
LouisLotter
reviewed
Aug 17, 2026
LouisLotter
reviewed
Aug 17, 2026
LouisLotter
reviewed
Aug 17, 2026
LouisParkin
force-pushed
the
STAC-25565-beest-verification
branch
from
August 18, 2026 11:44
db081dc to
9e83d82
Compare
LouisLotter
reviewed
Aug 18, 2026
LouisLotter
left a comment
There was a problem hiding this comment.
Re-review of the current head: the concrete-SHA resolution and publication gate are addressed. Two issues remain.
LouisLotter
approved these changes
Aug 18, 2026
LouisParkin
force-pushed
the
STAC-25565-beest-verification
branch
from
August 19, 2026 07:36
9e83d82 to
81ba6f5
Compare
LouisParkin
enabled auto-merge
August 19, 2026 07:37
This was referenced Aug 19, 2026
beest_trigger_verification was the last job in the GitLab pipeline with no GitHub Actions equivalent. Every other job is covered by the STAC-25142 / STAC-25457 / STAC-25500 stack. In GitLab the job sits in the postbuild stage, needs both merge_docker_manifest jobs and is `when: manual`, passing AGENT_BRANCH_UNDER_TEST, AGENT_HASH_UNDER_TEST and TRIGGER_AGENT_X86_TESTS into the stackvista/integrations/beest project. beest has since migrated to GitHub, and its agent-x86.yml and arm.yml both expose workflow_dispatch with an agent_branch_under_test input, so the port is a cross-repo workflow dispatch rather than a pipeline trigger. beest resolves the agent image from the branch name, so the commit SHA is no longer part of its input contract; it is recorded in the run summary for traceability instead. Keeping the workflow workflow_dispatch-only preserves the GitLab `when: manual` semantics. These runs provision real EKS infrastructure in the sandbox account and share a single global concurrency lock in beest, so firing them automatically on push would queue runs behind each other and spend hours of cluster time per merge. The suite input defaults to x86, matching TRIGGER_AGENT_X86_TESTS: true; arm and both are available because beest now exposes an arm workflow that the GitLab job never reached. The scenario selector is passed through rather than re-declared, so beest stays the single owner of the valid scenario list. Requires a GitHub App credential in this repo with actions:write on StackVista/beest, provisioned via pulumi-infra: BEEST_DISPATCH_APP_CLIENT_ID (variable) and BEEST_DISPATCH_APP_PRIVATE_KEY (secret). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The first version of this workflow sent only the branch and put the SHA in the run summary, on the reasoning that beest had no hash input. That was the wrong conclusion: beest's own GitLab port dropped the input while keeping the machinery, so the missing input was a regression to fix rather than a constraint to design around. beest#61 restores it. A branch builds many images, so branch-only means always testing whichever build is newest -- there is no way to verify a specific commit or to reproduce a failure against the image that produced it. Defaults to this run's commit, matching the GitLab job's CI_COMMIT_SHA. When agent_branch_under_test points at some other branch our SHA does not exist there, so the pin is left unset and beest falls back instead of dispatching a hash that resolves to nothing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The workflow referenced BEEST_DISPATCH_APP_CLIENT_ID/PRIVATE_KEY, which are provisioned nowhere. The beest App already exists as BEEST_GH_APP_CLIENT_ID / BEEST_GH_APP_PRIVATE_KEY, matching the <PURPOSE>_GH_APP_* convention every other App credential in the estate follows. Those variables are currently bound only to the beest repo, so pulumi-infra must also bind them to stackstate-agent before this workflow can mint a token. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
beest replaces agent_hash_under_test with hashes_under_test in StackVista/beest#63, so the dispatch has to send agent=<sha> through the new field. Must land together with that PR. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Review feedback on #444's follow-up: three ways the dispatch could report success while doing the wrong thing. Resolve the commit under test up front. beest only pins nodeAgent, clusterAgent and checksAgent to <sha8>-<arch> when a hash reaches it, and its resolve-agent-hashes.sh deliberately refuses to resolve a branch ("Leave unset to use Helm chart defaults"). Passing agent_branch_under_test without a hash therefore deployed the chart default image, so the run could pass while testing an agent nobody selected. An overridden branch now resolves to its tip, and every dispatch carries a concrete commit. Require both multi-architecture manifest jobs to have succeeded for that commit before dispatching. The GitLab job was only playable once agent and cluster-agent had published; the standalone dispatch had no equivalent barrier and could spend a full beest run, holding the global beest AWS lock, on images that were never pushed. Select the dispatched run by diffing against the runs that existed beforehand. Taking the newest run raced with API propagation, concurrent dispatches and the shared lock, so the summary regularly linked an unrelated run. Needs actions:read and contents:read to resolve the commit and read its runs; the App token stays scoped to dispatching beest. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…usly The scenario input was forwarded verbatim to every workflow the suite selects, but beest's choices are architecture-specific and disjoint: agent-x86.yml accepts contd-eks-x86-*, arm.yml accepts contd-eks-arm-*, and only "all" is common to both. With suite=both, any architecture-specific value therefore reached one workflow that rejects it. Validate the value against every selected workflow before dispatching any of them, so a rejected combination cannot leave one architecture running and holding the global beest AWS lock. The allowlist is duplicated from beest rather than read from it, which costs a bump here when beest gains a scenario. Reading it would need contents access to beest on the dispatch token and YAML parsing on the runner; failing closed with the valid values named is the cheaper trade. Run linking took the first run absent from the pre-dispatch snapshot, but a concurrent dispatch of the same workflow and ref produces a second new run that is indistinguishable from this one, and the shared concurrency group creates both records before either executes. Collect every new run instead and link only when exactly one exists; otherwise report the workflow page and say why. Linking the wrong run is worse than not linking one. The snapshot itself no longer swallows failure. An empty baseline makes every existing run look new, which guarantees a mislink, so it fails closed instead. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
LouisParkin
force-pushed
the
STAC-25565-beest-verification
branch
from
August 19, 2026 12:16
672062e to
bbc7957
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports the last unported GitLab job,
beest_trigger_verification. Every other GitLab job now has a GitHub equivalent.Manual (
workflow_dispatch) only, matching the GitLab job, which waswhen: manual. Mints a short-lived App token scoped toactions: writeonbeestalone, then dispatchesagent-x86.yml/arm.yml.Pins the exact agent commit. A branch builds many images with different hashes, so sending only the branch would force every run onto whichever build is newest, with no way to verify a specific commit or reproduce a failure against the image that produced it. Defaults to this run's SHA, matching the GitLab job's
CI_COMMIT_SHA. Ifagent_branch_under_testnames a different branch, our SHA does not exist there, so the pin is left unset and beest falls back rather than dispatching a hash that resolves to nothing.Admin ask (blocking a live run): needs
BEEST_DISPATCH_APP_CLIENT_ID(var) andBEEST_DISPATCH_APP_PRIVATE_KEY(secret) on this repo, for an App withactions: writeonbeest. beest's own credential is repo-level there and not reachable from here.Base branch:
stackstate-7.78.2rather than the #444 stack. It shares no files or jobs with the build lanes and is manual-only, so it cannot affect any push/PR pipeline and need not queue behind those reviews. Trivial to retarget if preferred.Dispatch logic tested against a mock
ghacross the pin default, branch-override, explicit-hash, both-suites and invalid-suite paths. actionlint and zizmor clean.