Skip to content

Continuous integration

Both clients need the same two things from your CI, and nothing else:

  1. A run id that every shard of one run agrees on. It is what makes eight parallel jobs join a single pipeline and split one allocation between them instead of eight jobs each claiming the whole suite.
  2. Which shard this job is — N of M, one-based.

On GitLab both are detected automatically. Everywhere else you supply at least one of them.

CI Run id Branch Shard
GitLab CI CI_PIPELINE_ID CI_COMMIT_REF_NAME CI_NODE_INDEX / CI_NODE_TOTAL
GitHub Actions GITHUB_RUN_ID + GITHUB_RUN_ATTEMPT GITHUB_HEAD_REF, GITHUB_REF_NAME pass --shard=N/M
Anything else set TESTFLAKE_RUN_ID set TESTFLAKE_BRANCH pass --shard=N/M

TESTFLAKE_RUN_ID and TESTFLAKE_BRANCH always win when set, so you can override the detected values on GitLab and GitHub too.

Two details that catch people out on CI specifically:

  • Shard numbers are one-based. GitLab’s CI_NODE_INDEX already is. CircleCI’s CIRCLE_NODE_INDEX is zero-based and is not read by either client — add one yourself, as the CircleCI tab below does.
  • The run id must survive a retry. Re-running a single failed job has to produce the same id as its siblings, or that job starts a pipeline of its own. This is why the GitHub id includes GITHUB_RUN_ATTEMPT: a full re-run is a new run, while retrying one job of an existing run is not.

Each tab is a complete, minimal pipeline. TESTFLAKE_KEY comes from your CI’s secret store in every one of them — it is a credential, so do not commit it.

parallel: sets CI_NODE_INDEX and CI_NODE_TOTAL, and every parallel job of one pipeline shares CI_PIPELINE_ID. Nothing has to be passed — this is the only CI where the wrapper needs no arguments at all.

Add TESTFLAKE_KEY under Settings → CI/CD → Variables, masked.

.gitlab-ci.yml
playwright:
image: mcr.microsoft.com/playwright:v1.62.1-noble
parallel: 8
variables:
TESTFLAKE_HOST: https://app.testflake.com
script:
- npm ci
- npx testflake npx playwright test
  1. Run the pipeline twice. The first run has no timing data and falls back to an alphabetical split, so a single run tells you nothing about balance.

  2. Compare the shard durations of the second run. They should finish within a few percent of each other. If one shard is still much slower, it is probably a single test longer than the average shard — nothing can split below one test.

  3. Check stderr for [testflake] warnings. Every failure path says so there, and every one of them means the suite ran unsharded.

Variable Default Purpose
TESTFLAKE_KEY — Attributes runs to a project. Treat as a credential.
TESTFLAKE_HOST — API base URL.
TESTFLAKE_RUN_ID detected on GitLab and GitHub Ties the shards of one run together.
TESTFLAKE_BRANCH detected on GitLab and GitHub Branch the run belongs to.

Your build still passes. Any failure — an unreachable host, a non-2xx response, a malformed body — makes the client warn on stderr and run the entire suite, unsharded, exactly as if the wrapper were not there. You get a slow build, never a red one, and the exit code is always your test runner’s own.

That is worth knowing before you debug a pipeline that suddenly got slower: a run taking eight times as long usually means eight shards each ran everything.