Simulating Stripe Payment Failures and Decline Codes in Tests

Test your payment failures across full transaction sequences, not just single API calls.

Contributing Editor · · 9 min read
Cover illustration for “Simulating Stripe Payment Failures and Decline Codes in Tests”
Fault Injection Testing · October 7, 2026 · 9 min read · 2,110 words

Knowing the card number that triggers insufficient_funds is not the same as having proof that an application handles insufficient_funds correctly. That is the gap this piece sets out to close. Stripe's sandbox is a full parallel environment: it has its own API keys and its own objects, and a test-mode toggle in the shared dashboard reaches it. Stripe's testing documentation is explicit that nothing in this environment touches real money or a real bank, and that test cards exist to trigger specific, reproducible outcomes on demand. A card number, though, functions as a credential rather than a test: it tells Stripe which outcome to produce, and says nothing about whether the surrounding application handles that outcome correctly across the sequence of calls that precede and follow it. If a team runs 4242 4242 4242 4242 through a signup form once and calls the integration validated, it has only confirmed that the happy path exists. It has not confirmed the decline path, the retry path, the webhook handler that reacts to the decline, or the state the system is left in after each of those steps resolves.

What Stripe's decline surface covers

Stripe's testing documentation organizes its simulation surface into several distinct categories, and you need to see what that surface actually contains before you can argue about what it leaves out. Issuer declines come with specific decline_code values attached to specific card numbers: 4000 0000 0000 0002 returns generic_decline, 4000 0000 0000 9995 returns insufficient_funds, 4000 0000 0000 9987 returns lost_card, 4000 0000 0000 9979 returns stolen_card, 4000 0000 0000 0069 returns expired_card, 4000 0000 0000 0127 returns incorrect_cvc, 4000 0000 0000 0119 returns processing_error, and 4000 0000 0000 6975 returns card_velocity_exceeded.

A second category covers 3D Secure and strong customer authentication. 4000 0000 0000 3220 requires 3DS2 and succeeds when authentication completes. 4000 0000 0000 3063 requires 3DS2 for off-session payments unless the card is set up for future use. 4000 0025 0000 3155 requires authentication unless the card is set up for off-session payments. 4000 0027 6000 3184 requires authentication on every payment, with no exceptions. 4000 0084 0000 1629 requires authentication and then declines regardless.

Fraud and dispute simulation makes up a third category. 4100 0000 0000 0019 is always blocked by Radar as the highest risk tier. 4000 0000 0000 0259 is disputed as fraudulent after it succeeds. 4000 0000 0000 2685 succeeds and is then disputed as not received. 4000 0000 0000 5423 succeeds and receives an early fraud warning. 4000 0000 0000 9235 carries elevated risk and may be queued for manual review. Bank transfers have their own failure tokens: ACH test accounts can trigger account closed, Radar-blocked, no account found, and insufficient funds outcomes on US bank accounts. Stripe Terminal testing largely parallels the online flow: you get a built-in simulated card reader or a set of physical test cards.

Every one of these card numbers produces one specific, deterministic outcome, and the card number is the only field that decides what happens: expiry date, CVC, and postal code all accept any value in the correct format. That determinism makes the list useful as a reference. Nothing about a card number encodes what happens when several of these calls are chained together in the order a real transaction actually produces them, and that is where its usefulness ends.

Why decline codes matter inside a sequence of calls

A decline code is one state transition inside a multi-step flow: the application's behavior across that whole sequence, not just at the instant the decline fires, is what makes an integration work or fail. Consider the canonical recovery flow built around 4000 0000 0000 9995. A customer's card declines with insufficient_funds. The application has to notify the customer, hold the order in a "payment required" state, accept a new payment method, and retry the charge, all without creating a duplicate. If a test checks only that the card returns insufficient_funds, it has verified none of those four subsequent steps.

Subscription cards make the same point sharper. 4000 0000 0000 0341 attaches successfully as a payment method but fails on the first actual charge, so the application ends up with a subscriber who has never once paid successfully. That is a distinct state from a subscriber whose payment method worked for three months and then started failing, and it demands different dunning logic. A stateless test, one that only inspects the response of a single API call, cannot tell these two subscribers apart, because the difference only exists across time and across calls.

Dispute cards extend the same argument past the point of charge success, because 4000 0000 0000 0259 succeeds immediately and then, asynchronously, arrives as a dispute. The application has to track that charge as "under dispute" and manage an evidence submission flow that did not exist at the moment the charge was authorized. The charge object at the moment of charge and the charge object at the moment of dispute are not the same object, and a test asserting only on the first has missed the actual failure mode the card was built to simulate.

3D Secure cards complicate the picture one layer further. They insert an authentication step between card entry and charge, and 4000 0084 0000 1629 requires that authentication to succeed and then declines the charge anyway. This is a real customer scenario, and it depends on the application never treating a successful authentication as equivalent to a successful payment. A reasonable objection at this point is that Stripe's test mode already tracks state, so none of this should require extra effort. It does track state, but it tracks Stripe's state: the PaymentIntent object, the dispute object, the subscription object. The application's own state machine decides what a customer sees, what gets retried, and what gets logged as resolved, and a card-number lookup table has no way to exercise it.

Webhooks as the part of failure simulation most teams skip

Webhooks are the mechanism by which Stripe communicates a state change back to an application after the fact, and they carry the entire asynchronous half of the dispute and subscription scenarios just described. payment_intent.payment_failed fires on insufficient funds, an expired voucher, or a closed account, and Stripe's documentation recommends notifying the customer and requesting a new payment method in response. charge.dispute.created fires the moment a dispute opens, at which point funds have already been pulled and an evidence submission window has already started. customer.subscription.updated and customer.subscription.deleted carry the subscription state changes driven by recurring payment failures, including the delinquent-subscriber scenario raised by 4000 0000 0000 0341. You cannot reach any of these events by sending a single payment request and inspecting the response. They arrive on their own schedule, and an integration that never tests its webhook handler has never tested the half of the flow where the actual business consequences land.

Three failure modes in webhook signature verification sit entirely outside the choice of card, and they need separate attention. The first concerns body parsing: if a web framework parses the incoming JSON before the signature-verification handler runs, the raw bytes Stripe originally signed no longer match the re-serialized bytes being hashed, so the raw request body has to be captured before any middleware touches it. The second concerns the string comparison used to check the signature, which has to run in constant time to avoid leaking timing information that could help an attacker brute-force it. The third concerns timestamp tolerance: accepting a signature with any timestamp at all opens the door to replay attacks, and a five-minute tolerance window is the standard mitigation. A broken webhook handler that silently fails a dispute event typically goes unnoticed until a bank actually pulls funds from a live account, not when you run a test. Stripe's dispute test cards make that scenario fully rehearsable in advance, but only for teams that put the webhook handler inside the scope of the test. The open question this raises is what kind of environment could plausibly hold all of this state, across payments, webhooks, and retries, at once.

Idempotency under decline-and-retry is a correctness requirement

Idempotency keys are not a hygiene detail to clean up once an integration is otherwise working. In a decline-and-retry flow, omitting them means a network timeout or an ordinary retry can create a second charge for the same order, a correctness failure. The mechanism Stripe provides for this is straightforward: a unique idempotency key, typically a UUID, has to be generated for each distinct payment attempt and included in every payment creation request. Stripe returns the identical response for any duplicate request that carries the same key within a 24-hour window under API v1, and that window extends under API v2.

If a test sends one POST /v1/payment_intents call and inspects the response, it cannot tell whether the application generates a fresh idempotency key on the next attempt, which is correct, or reuses the original key, which risks a duplicate charge. The scenario that actually exercises this logic requires several steps in sequence: a decline, followed by the customer updating their payment method, followed by the application retrying with a new key, followed by a check that exactly one charge exists at the end. Running that scenario requires an environment capable of remembering that the first attempt happened.

What kind of test environment can hold payment state

The environment you choose for failure simulation decides which of these flows you can actually reach, and most lightweight options trade coverage for speed, so the stateful scenarios go untested. Stateless mocks, which return a hardcoded response for each card number, are fast and simple, but each call is independent of the last, so they cannot represent a decline followed by a retry or a subscriber who moves from delinquent to current, because the mock has no memory of what happened in the prior call. Recorded cassettes that replay a captured sequence of real calls solve part of this by encoding one specific sequence deterministically, but they break the moment the actual sequence diverges, whether that means a new decline code or a different retry path, and they offer no way to inject a fault on demand.

Live Stripe test mode tracks genuine Stripe-side state, so it returns the correct decline_code for every card in the list above. It also requires live credentials, hits rate limits under sustained load, and is unreachable offline. Stripe's own testing documentation warns against using test environments for load testing because those rate limits apply, so test mode fits poorly into any suite you run repeatedly in continuous integration. Stateful simulators close this gap by remembering state across calls: a PaymentIntent gets created, declined, updated with a new payment method, and retried, and the simulated world updates accordingly at each step. You can run them locally or in CI without credentials, without rate limits, and without a network dependency, and they support fault injection on demand, not just the fixed set of outcomes a card number encodes.

Stripe's own documentation contains a quiet acknowledgment of exactly this problem. Its guidance states that "card-level state from one test can affect later tests that use the same card," and that "when a test depends on card-level state, use a different test card for each independent end-to-end scenario." That is Stripe itself confirming that state persists across calls within its own sandbox, and that tests have to be designed around that persistence rather than around the assumption that each call is isolated.

Fault injection as a first-class part of payment resilience testing

Testing that a payment fails with insufficient_funds establishes one thing: that the application reacts correctly when Stripe tells it, in an orderly fashion, that a card lacks funds. It establishes nothing about how the application behaves when a payment request times out, comes back with a 429, or fails with a 503, because none of those outcomes correspond to any card number. They are infrastructure-level failures, and you can only reach them through deliberate fault injection, not a decline simulation. Latency injection is the clearest example: configuring a simulator or proxy to delay a response on command makes it possible to confirm that an orchestrator's timeout logic and its circuit breaker actually fire when they are supposed to, rather than assuming they do because no one has ever forced the slow path to occur. A stateful test environment capable of tracking a PaymentIntent through decline, retry, and dispute is also the kind of environment that can be asked to delay, corrupt, or drop a response on demand, and that combination, state plus fault injection, is what separates a card-number lookup table from an actual test of payment resilience.

Sources

  1. Test card numbers
  2. Test Stripe Terminal
  3. Stripe decline codes
  4. Idempotent requests
  5. Automated testing
  6. Error handling