Parallel CI Jobs With Isolated API State

Each parallel CI job needs its own isolated API state to avoid flaky tests.

Senior Staff Writer · · 10 min read
Cover illustration for “Parallel CI Jobs With Isolated API State”
CI Pipeline Testing · October 10, 2026 · 10 min read · 2,255 words

Splitting a test suite across parallel jobs only cuts wall-clock time if each job owns its own state. When jobs share state instead, parallelism trades slow, predictable failures for fast, nondeterministic ones that cost far more engineering time to track down. The specific failure mode looks like this: tests that pass every time when run one after another start deleting each other's records, consuming each other's queued jobs, or stepping on the same external side effects the moment they run at the same time. Nobody touched the code, yet the suite starts failing. Picture a GitHub Actions matrix job where every shard talks to the same live Stripe sandbox account: one shard creates a customer, another lists and deletes customers, a third asserts that a customer still exists. Each of those three tests can be correct on its own and still produce a flaky suite when they run together, because none of them was written with the others' timing in mind. The common response, serializing every test that touches an API, makes the flakiness go away but also removes the entire reason parallelism was introduced: the suite becomes reliable again by becoming slow again.

What "isolated state" means across the different dependencies a CI job touches

Isolation has to hold across every stateful dependency a job touches, and a single gap breaks the guarantee for the whole job, no matter how well the rest of the pipeline is built. For a database, isolation can take several forms, each with its own cost and its own way of failing. A separate named database on a shared server works well for most API and integration shards and comes with moderate setup overhead. Schema-level separation runs fast but adds migration complexity, which makes it a reasonable fit for a PostgreSQL monolith where spinning up a fresh container for every shard would be too slow. A dangerous middle ground occurs often in practice: a team namespaces test data at the row level, giving each shard its own usernames or record prefixes, while still running migrations and truncations against one shared schema. Unique usernames do not stop a TRUNCATE in one shard from wiping out another shard's data.

Databases are also not the only stateful dependency a job touches. Async workers are often forgotten: if the web service under test points at a per-job database but an email worker or a webhook processor still points at a shared queue or a shared database, isolation is broken even though the main service looks isolated. The worker needs the same DATABASE_URL and the same queue namespace as the job that owns it, or it will quietly act on the wrong job's data. Third-party APIs break this model. An external service has no equivalent of "spin up a fresh container": a team cannot provision a new Stripe account for every CI shard. Isolation has to be built on the client side, inside the pipeline.

Why third-party API state is the hardest dependency to isolate in a parallel pipeline

Diagram: Three Bad Options for Third-Party API Isolation — and Why Each Fails. Visualizes: Show three parallel options a team faces when isolating external API calls (e.g.

External APIs bring the isolation problem in from outside the pipeline's control. Rate limits apply across every parallel runner using the same sandbox credentials, and providers offer nothing like "create a fresh database per job." Without a simulator standing between the test suite and the real API, a team is left with three options, and none of them is good. Every shard can hit the live sandbox directly, which risks exhausting rate limits and adds flakiness from network latency and from state the provider shares across shards that have no awareness of each other. All API-touching jobs can be serialized, which removes the flakiness but also removes the speed gain parallelism was supposed to deliver. Or the team can switch to stateless mocks, which solves the rate-limit problem but creates a different correctness problem, covered next.

The constraint is structural, independent of tuning. A CircleCI tutorial on parallelism notes that running too many parallel containers can exhaust shared resources, rate limits among them, and no amount of workflow configuration changes that ceiling. Stripe's own documentation says as much directly: it recommends running tests that validate Stripe API responses infrequently, specifically to avoid hitting rate limits. That is a vendor telling its own customers that its test environment is not built to be hammered by a parallel CI suite. The fix cannot live in the provider's sandbox, because the provider has no reason to build per-job isolation into a shared sandbox account. It has to live in how the CI pipeline itself is architected.

Why stateless mocks fail multi-step API workflows in parallel jobs

A stateless mock is a routing table, not a running system. It matches an incoming request to a response that was written in advance, with no memory of what request came before it and no way to check that the state implied by one response actually matches what a later response claims. The gap is concrete: a mock can hand back a hardcoded customer object on step one and a hardcoded subscription object on step three, but it has no way to confirm that the customer created in step one still exists by the time step three runs. Whatever consistency exists there was built by whoever configured the mock, not verified by anything resembling the real API's behavior.

That gap matters most on exactly the workflows parallel CI exists to speed up: create an order, capture payment, fire a webhook, update fulfillment state; create a channel, post a message, add a reaction, retrieve the thread. Any flow where the output of one step feeds the input of the next is a flow a stateless mock cannot check. Hand-maintained stubs also decay. Every time the real API changes, the stub has to be updated by hand, and the divergence it causes typically surfaces in production. A test that passes against a stale mock only proves the mock agrees with itself. The common guidance across the mocking ecosystem, to keep mocks stateless by default and reserve stateful infrastructure for complex multi-step workflows, points at the exact class of test parallel CI jobs are built to run faster.

How a stateful simulator gives every parallel job its own isolated API world

A stateful simulator running as a sidecar or service container inside each CI job gives that job a full, private copy of the API's state. Create a customer in job A, and it exists only inside job A's simulator. Delete it in job B, and job A never notices, because there is nothing connecting the two. The isolation here is structural rather than procedural: each job's simulator is its own process with its own memory or disk-backed state, so there is no shared channel through which one job's actions could leak into another's. It is the same guarantee a team gets from giving each shard its own database container, applied to the API layer.

State inside a single job's simulator behaves the way state behaves against the real API: add a record, list it, update it, delete it, and the simulator's internal world updates at each step, so a test can check every transition the same way it would against production. A fan-out structure fits this model cleanly. A single build job compiles and caches the shared artifacts, and each downstream test job starts its own simulator instance, runs in parallel with the others, and tears everything down when it finishes. This is the same pattern the LibreChat CI refactor applied to its databases and build artifacts, just extended to cover the API layer as well.

jobs:
  test-api:
    runs-on: ubuntu-latest
    services:
      api-simulator:
        image: your-org/api-simulator:latest
        ports:
          - 4010:4010
    env:
      API_BASE_URL:
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npm test

Each job in the matrix gets its own api-simulator service container, bound to its own job, with the test suite pointed at localhost as the provider's endpoint. Nothing about that configuration needs to change as the matrix grows, since every job brings its own simulator with it.

Continuous verification against the real API and CI reliability over time

A simulator carries the same risk a hand-maintained mock carries if nobody checks it against the real API: it drifts quietly, and a passing test gives a false sense of safety until the drift appears in production. The way to prevent that is a verification process, not a cleverer architecture. Before every release, a probe suite runs against both the simulator and the live API, compares the two sets of responses, and blocks the release if fidelity drops below a set threshold. Drift gets caught at that gate, before it reaches a developer's test run, instead of showing up later as a production incident.

Publishing the result of that process, a specific probe count and a named fidelity score, turns "we think it's accurate" into a claim that can actually be checked and that changes version to version. The Stood open-source project (on GitHub as ma-za-kpe/stood) makes the same argument in its own design: it keeps fake implementations clearly separate from its HTTP simulator and labels synthetic outcomes explicitly, surfacing contract drift. A fidelity score, tracked over time, does for a simulator what a test suite does for application code: it gives a concrete, repeatable reason to trust it, rather than asking the team to take the trust on faith.

Structuring a GitHub Actions workflow for parallel API-touching jobs with per-job simulator instances

Diagram: Fan-Out: One Build, Four Isolated Simulator Jobs. Visualizes: Illustrate the fan-out CI workflow shape: a single build job compiles and caches artifacts, then four parallel test-api shard jobs (shard 1, 2, 3, 4) each spin up their own…

The fan-out shape described earlier translates into a workflow with one build job and a matrix of test jobs that depend on it through needs: build, each running its own simulator as a service container. A few details decide whether that structure actually holds up under concurrent runs.

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npm run build
      - uses: actions/cache@v4
        with:
          path: node_modules
          key: node-modules-backend-${{ hashFiles('package-lock.json') }}

  test-api:
    needs: build
    runs-on: ubuntu-latest
    strategy:
      matrix:
        shard: [1, 2, 3, 4]
    services:
      api-simulator:
        image: your-org/api-simulator:latest
        ports:
          - 4010:4010
    env:
      API_BASE_URL:
      MONGO_URI: ${{ secrets.MONGO_URI }}
      OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx jest --shard=${{ matrix.shard }}/4 --maxWorkers=50%
      - name: Teardown
        if: always()
        run: npm run teardown

Persistent identifiers, database names, simulator ports, artifact names, need the run ID and shard identity baked into them, so two pipeline runs happening at the same time on the same runner pool don't collide. This is the same naming discipline that applies to isolating databases across parallel jobs, extended to the API layer. Secrets scoping is a correctness issue here, not only a security one: credentials such as MONGO_URI, OPENAI_API_KEY, JWT_SECRET, CREDS_KEY, CREDS_IV, and the BAN_* variables belong in the job-level env of the specific job that needs them, not in a workflow-level env block, which is the change the LibreChat refactor made. A job that inherits live API credentials by accident will talk to the real API instead of the simulator, defeating the entire architecture without producing an obvious error.

Cache keys need the same discipline. A key like node-modules-backend-${{ hashFiles('package-lock.json') }}, scoped to its own workflow, stops one workflow's cached node_modules from being restored into a different workflow's job, which can otherwise cause ABI mismatches or stale dependencies that are hard to trace back to a cache. Worker counts need a ceiling too: setting Jest's maxWorkers to something like '50%' keeps one job from saturating the runner's CPU and quietly cancelling out the speed gain parallelism was supposed to provide, a change the LibreChat refactor applied across all six of its Jest configurations. CircleCI users get at the same goal with the parallelism key and circleci tests split, using timing-based splitting, which relies on historical per-file duration data and produces more evenly loaded shards than splitting by filename or filesize. Finally, teardown has to run under if: always(), so simulator instances and named databases get stopped and dropped even when a job fails partway through. Skipping that step lets state accumulate on persistent runners until it starts interfering with later runs.

Fault injection as a first-class concern in a parallel API testing architecture

Testing how a system handles API failures, rate limits, 5xx errors, timeouts, malformed responses, only works if the failure can be triggered on demand and in a predictable way. That requires owning the API surface the test talks to, not just subscribing to it. A live sandbox gives no control over when it fails. A stateless mock can hand back a hardcoded error body, but it cannot reproduce the way a real outage actually unfolds: requests succeeding, then rate-limit errors starting to appear, then the service recovering over time.

Fault injection operates on two separate layers. At the transport layer sit problems like latency, packet loss, connection refusal, and timeouts, which tools such as Toxiproxy, built by Shopify, handle by acting as a TCP proxy sitting between the test and the simulator. At the semantic layer sit faults transport tools cannot see at all: a structurally valid HTTP 200 response carrying malformed or truncated content, a failure mode that matters in particular for AI agents parsing API output. A proxy that only manipulates packets never touches that layer, because the connection itself is healthy.

A parallel architecture with isolated simulator state makes this kind of testing practical. One shard runs the happy path, a second runs the same scenario with latency injected, a third runs a sequence of 429 responses, all three running at the same time, with each job's isolated simulator state guaranteeing that the fault injected in one shard has no way of reaching the others.

More in CI Pipeline Testing