Flaky Integration Tests in CI Root Causes and Fixes

Timing issues and external APIs cause most CI flakiness, but need different fixes.

Staff Writer · · 11 min read
Cover illustration for “Flaky Integration Tests in CI Root Causes and Fixes”
CI Pipeline Testing · October 8, 2026 · 11 min read · 2,477 words

A flaky test is one that produces both passing and failing results on the same code, the same commit, with nothing changed in the application between runs. The defining trait is non-determinism, not some hardware glitch that occasionally trips a wire. The instability lives in timing, in shared state, or in the environment surrounding the test, hiding somewhere specific rather than striking at random, which tells you where to look.

Integration and end-to-end tests carry far more of this risk than unit tests do, because they touch browsers, networks, shared databases, and outside services, every one of which is a place where non-determinism can enter. A unit test controls its own inputs and outputs and rarely misbehaves. An integration test hands control to systems it does not own, and each of those systems has its own clock, its own load, and its own failure modes.

The cost appears in two separate ways. One is visible: engineers spend hours re-running builds and chasing down failures that have nothing to do with the code they wrote. The other is quieter but more damaging: confidence in CI erodes, and once a team learns to expect red builds that mean nothing, a real regression can hide inside the noise and ship anyway. Fixing flakiness starts with naming its causes precisely, because "flaky" is a symptom, not a diagnosis. The next step is building out that list.

The five root causes of flaky integration tests, ranked by frequency

Diagram: Five Root Causes of Flaky Tests, Ranked by Frequency. Visualizes: Show five root causes of flaky integration tests ranked from most to least frequent, as a vertical ranked list with a clear visual break after rank 1.

The largest category by volume is async wait and timing assumptions: tests that check a result before an asynchronous operation has actually finished, with the outcome depending on how loaded the CI runner happens to be at that moment. Sauce Labs' 2026 guide points to hard-coded sleep statements as a particularly bad version of this problem, since a fixed delay is either too short, and the test fails, or too long, and the test wastes time even when it passes.

Test order dependency comes next. A test passes on its own but fails when it runs after another test has left something behind: a seeded database row, a cached session, a cookie nobody cleaned up. This problem tends to stay invisible until a team adds parallelism to its pipeline, at which point tests that used to run in a fixed sequence start running in whatever order the scheduler picks, and the hidden dependency appears.

Resource leaks and shared state make up a related but distinct category: file handles, open ports, in-memory caches, or fixtures that persist across test runs and cause intermittent collisions between tests that should have nothing to do with each other.

Network and infrastructure variability covers DNS flaps, latency from third-party APIs, slow container cold starts, and CI runners competing with each other for CPU and memory. External API calls live inside this category, though, as later sections argue, they behave differently enough to deserve separate treatment.

Selector fragility and DOM drift round out the list: tests anchored to CSS classes or XPath expressions that break the moment a front-end team refactors the UI. Academic taxonomies from a decade ago underweight this category, but it has grown substantially with the rise of modern front-end frameworks.

The ranking carries a practical consequence. Fixing the ten flakiest tests by failure frequency typically clears somewhere between 60 and 70 percent of flakiness by volume. Most flaky-test tooling gets built and tuned for timing problems, since that is where the biggest, fastest win lives. Network and API variability, as a result, gets less engineering attention than its structural risk actually warrants.

External API dependencies and structural non-determinism

Network and API variability reads, at first glance, like just another row in that taxonomy, one failure category among five. It behaves differently the moment a team tries to fix it, and that difference is the reason it deserves its own argument rather than a shared bullet point with DNS flaps and cold starts.

External API flakiness is a control problem. The test has no authority over the thing causing the failure. A live third-party call can fail or time out for reasons that have nothing to do with the code under test: a rate limit kicks in, the provider's API has a brief outage, a sandbox environment goes unstable, or a credential expires in the middle of a run. PayPal's access tokens illustrate the mechanism concretely. PayPal's own documentation states that access tokens have a life of 15 minutes or eight hours depending on the scopes associated. Any CI run long enough, or unlucky enough in its timing, can have a token expire mid-execution. Harness's 2026 guide names environmental factors like this, including external API rate limits, as a documented and recurring cause of intermittent CI failure, and none of it has anything to do with whether the application code is correct.

Compare that to the other four root causes. Timing flakiness gets fixed by changing how a test waits for a result. Test order and shared-state flakiness gets fixed by isolating test data between runs. Selector flakiness gets fixed by changing the selector strategy. In every one of those cases, the fix lives inside the test itself or inside the test's own setup and teardown. External API flakiness does not offer that option: changing the test changes nothing about whether PayPal's token expires or whether a sandbox has a bad five minutes. The only lever available is changing what the test calls. That is why the standard advice to "fix the root cause" points somewhere entirely different for API dependencies than it does for any other category on the list.

The limits of retries and quarantine for API flakiness

Retries and quarantine are good, sensible responses to most flaky tests, and nothing here argues otherwise. For external API dependencies specifically, both tools manage the symptom while leaving the underlying cause fully intact. The failure rate does not go down no matter how many times the team applies them.

Look at what a retry actually does. It hides the failure signal, since a passing rerun simply discards the original failure and the build goes green as if nothing happened. Reruns inflate CI wall-clock time and cost in ways that rarely appear on anyone's dashboard. Worse, it can produce false-green builds: if a real regression happens to pass on its retry attempt, the regression ships, and the team finds out from a customer instead of from CI.

Harness's 2026 guide traces where this habit ends up. Developers learn that a red build doesn't necessarily mean something is broken, so they start clicking rerun on reflex. From there it is a short step to merging through red entirely, and from there a shorter step still to writing fewer new tests, because "tests are flaky anyway" has become the team's working assumption. At that point flakiness has stopped being a productivity drain and turned into a quality problem, because the thing CI exists to catch is no longer being caught.

Quarantine is the better tactical move of the two. Flag the flaky test, move it into a non-blocking suite so it stops holding up releases, and keep running it anyway so it still generates data. Quarantine should work as a queue with a deadline, not a drawer where tests go to be forgotten. But for API-driven flakiness, quarantine still leaves the actual cause standing: the next CI run still dials out to a live API, and the rate at which that call fails stays exactly where it was. Quarantine buys time. It does not remove the dependency that is causing the problem. That raises the real question: what does removing a live dependency actually look like in practice?

Stateless Mocks as a Replacement for Live API Calls

The obvious answer, and the one most developers reach for first, is mocking. A stateless mock gets rid of the live call, but it trades one failure mode for another, because it cannot represent what actually happens across a sequence of API calls, only what a single call looks like in isolation.

A stateless mock is, at bottom, a configuration file that maps a request pattern to a canned response. It has no memory of anything that happened before it. It cannot tell you whether a record created in step one of a workflow still exists when step three of that same workflow tries to read it back, because it was never built to track that kind of continuity.

Consider a concrete failure pattern around Stripe's PaymentIntent object. A team mocks a POST to /v1/payment_intents and asserts that the response comes back 200, the test goes green, and the pull request merges. But the handler being tested never actually stored the PaymentIntent ID anywhere, so when the next step in the webhook path tries to read that value, it comes back undefined. The mock proved that the request shape matched what was expected. It proved nothing about whether the workflow that depends on that data actually functions, and the first real checkout against staging exposes the gap.

Sauce Labs' guide names external dependencies as a root cause of flakiness and recommends mocking as a fix, and that recommendation holds only as long as the mock actually reflects real API behavior across a full sequence of calls, something a stateless mock cannot guarantee once a workflow has more than one step. A second, slower problem compounds the first: drift. A hand-maintained mock is a snapshot of what the real API looked like on the day someone wrote it. Every change the API provider makes afterward requires someone to go update the stub by hand, and most teams find out that update never happened when the integration breaks in production, not when a test fails in CI.

Stripe's own documentation is a useful marker of how seriously this problem is taken even by companies that run the APIs in question. Stripe recommends using separate general sandboxes for local development and for CI rather than relying on the shared test-mode sandbox, a piece of guidance that quietly concedes that even Stripe's own live sandbox introduces variability teams need to isolate against. But sandbox-first still means live network calls, with every bit of fragility that implies. None of this rules out mocking as a concept. It raises a narrower question: what would it take for a mock to actually remember what happened in the step before it?

What stateful simulators do differently

A stateful simulator maintains memory across calls, answering each request in the context of what came before it. It solves external-API flakiness at the infrastructure level, replacing a live, uncontrolled dependency with a locally hosted stand-in that behaves the same way every time it runs.

The architectural difference from a stateless mock is specific. Where a stateless mock maps one request pattern to one canned response, a stateful simulator holds a working model of the API's world. Create a resource, list it, update it, delete it, and the simulator's internal state updates the same way a real API's internal state would, so a value written in step one of a test is still there when step three of that test goes looking for it.

That property removes the entire set of external causes that retries and quarantine could only work around: rate limit timeouts, a credential expiring mid-run the way PayPal's tokens can, sandbox outages, shared test-mode data from other teams contaminating a run, and the general unpredictability of a network nobody on the engineering team controls.

A simulator also opens a door that a live API keeps shut: fault injection. A simulator that can inject latency, 429 rate-limit responses, and 503 errors on demand lets a team actually test its resilience code, timeout handling, retry logic, backoff behavior, without risking real damage or tripping an actual rate limit against a production-adjacent system. That kind of testing is close to impossible to do safely against a live third-party API.

None of this holds up on its own without upkeep. A simulator that isn't checked regularly against the real API it stands in for will drift the same way a hand-maintained mock drifts, just more slowly, and it will hand developers the same false confidence in the meantime. Verification against the real API before every release keeps a simulator's behavior honest and prevents it from quietly falling out of sync.

Simulator-Based Testing in a CI Pipeline

A stateful simulator is one piece of a layered fix, not a replacement for the other four. Each root cause in the taxonomy needs its own specific intervention, and the simulator is built to handle exactly one of them, the external-dependency category, without reaching into the others or requiring them to change.

Timing issues still get fixed with explicit conditional waits. Test order dependency and shared state still get fixed with isolated test data and proper teardown between cases, so no test leaves a database record behind for the next one to trip over. Selector fragility still gets fixed with stable test IDs rather than CSS classes or XPath paths that break on a UI refactor. External API dependencies are the one category where the fix is structural: a stateful simulator running locally or in CI, standing in for every live third-party call the integration suite would otherwise make.

Where the simulator runs matters for how consistent this fix actually is. It can run locally on a developer's laptop during development, offline and without needing real credentials, and it can run in CI as a sidecar or a containerized service, using the same binary in both places. Local development and CI are the two contexts where this kind of flakiness occurs most often, and running the same simulator in both closes the gap between "works on my machine" and "fails in the pipeline.

Stripe's own guidance recommends isolated sandboxes for each environment, separate for local work, separate for CI, separate again for staging, with the option to scale isolated sandboxes out across teams. A stateful simulator satisfies that same isolation requirement without keeping the live network dependency that a sandbox-based approach still carries.

Quarantine and simulator testing complement each other directly. A test quarantined because of API-driven flakiness can move back into the blocking suite once it runs against a simulator instead of a live endpoint, because the source of the non-determinism has actually been removed for good.

The same infrastructure extends to a newer testing problem: evaluating AI agents that call external APIs through tool use. An agent needs a stateful, repeatable environment to check whether it picked the right endpoint, extracted the right parameters, and handled an error response correctly, and a live API cannot offer the repeatability that meaningful agent evaluation requires. The same architectural shift that fixes flaky integration tests, trading a live, uncontrolled dependency for a stateful, locally controlled one, turns out to be the same shift agent testing needs for exactly the same reason.

More in CI Pipeline Testing