AI Automation

n8n workflows: retries, dead letters, safe re-runs

n8n workflows: retries, dead letters, safe re-runs

Opening answer (BLUF)

Production n8n workflows fail in three places we see over and over: a downstream API returns HTTP 429, a CRM or finance write succeeds and then the node still errors, or a later node fails after earlier side effects already landed. n8n gives you node-level Retry On Fail, an On Error policy (stop, continue, or continue on an error output), a dedicated Error Workflow that starts with the Error Trigger, and an Executions retry that can replay a failed run with original data or with the currently saved workflow [1][2][18]. Those controls stay safe only if every mutating write is idempotent, because HTTP POST is not idempotent by default and a retry can create a second invoice, ticket, or contact [10]. Our team designs n8n workflows so a failed execution can be replayed without duplicating records.

Why production runs break after a side effect already landed

A webhook or schedule fires, the workflow enriches a record, posts to a CRM, then posts to accounting. If accounting times out after the CRM create already succeeded, a naive replay creates a second CRM record. If the HTTP Request node hits a rate limit, n8n surfaces HTTP 429 ("The service is receiving too many requests from you") [3]. The 2012 HTTP status-code RFC that defined 429 says the user sent too many requests in a given amount of time, and the response may include Retry-After [4].

That is an operations problem. NIST's 2024 Cybersecurity Framework 2.0 treats DETECT, RESPOND, and RECOVER as concurrent functions: find anomalies quickly, contain the effects, and restore operations when something fails [5]. NIST SP 800-61 Revision 3 (April 2025) says the same for incident handling: prepare, reduce impact, and improve detection, response, and recovery [6]. For automation, that means classifying errors and recovering with a replay that does not double-write.

We design around three classes:

  • Transient transport failures (timeouts, 429, 502, 503) belong on Retry On Fail or a paced Wait loop [3].
  • Permanent request failures (400, 401, 403, schema mismatches) should stop the item and go to an error branch. Retrying them only burns quota.
  • Partial success, where the remote system accepted the write and the node still threw, is the class that duplicates records if you retry blindly.

Node-level retry and backoff for HTTP 429s

Every node Settings tab includes Retry On Fail. When it is on, n8n reruns that node after a failure until it succeeds or it exhausts the configured tries [2]. For rate-limited APIs, official n8n guidance is to set Wait Between Tries (ms) above the service's request interval. If the API allows one request per second, start at 1000 ms [3]. The HTTP Request node also documents Max Tries, Wait Between Tries, and a Batching option (Items per Batch and Batch Interval) so you do not spray a full item list at the API in one burst [3].

Retry is the wrong tool for a bad payload. A 400 from a CRM will fail every attempt. Use Retry On Fail on nodes that talk to flaky HTTP services, not on nodes that validate your own data. For cool-downs longer than a few seconds, the Wait node pauses the execution and, for waits of 65 seconds or more, offloads execution data to the database [16]. Pair Loop Over Items with Wait when you must send one record, pause, then send the next [3].

On HTTP Request nodes that write to SaaS APIs, we typically enable Retry On Fail (two to four extra tries, 1000 to 5000 ms wait) and Batching when many items arrive. Leave On Error at Stop Workflow for money-moving nodes; use Continue (using error output) for bulk jobs that should keep going. If the API sends Retry-After on a 429, honor it in a Wait node instead of hammering Max Tries [4].

Continue, stop, or branch: choosing On Error

Retry decides whether the node tries again. On Error decides what happens if it still fails. n8n documents three options [2]:

  • Stop Workflow (default): the execution halts. This marks the run failed and, if you configured one, starts the Error Workflow [1]. Use it on writes you cannot skip: invoice create, payment capture, inventory decrement.
  • Continue: the workflow proceeds and passes the last valid data. Use this only for optional enrichment. Silent continue is how rows go missing.
  • Continue (using error output): the node grows a second output. Successful items leave the main output; failed items leave the error output so you can log, notify, or park them. This is the pattern when one bad item in a loop should not kill the other eight.

Always Output Data returns an empty item even when the node produced nothing. n8n warns against turning it on for IF nodes, because an empty item can loop forever [2]. If a business rule should fail the whole run (duplicate invoice number, amount over a threshold), add Stop And Error so the Error Workflow fires [1][15].

Error workflows as the dead-letter path

For each production workflow, set Error Workflow in Workflow Settings to a handler that starts with the Error Trigger [1][13]. One handler can serve many workflows. The Error Trigger receives execution id, URL, last node executed, error message and stack, workflow id and name, and, when the run is itself a retry, `execution.retryOf` [1][12]. If the trigger node of the main workflow is what failed, n8n sends a thinner `execution` object and more data under `trigger` [12].

Three product rules matter. You do not have to publish the error workflow; if a workflow contains the Error Trigger, n8n uses that workflow as its own error workflow by default [12]. You cannot test the Error Trigger with a manual Execute Workflow click; it runs when an automatic workflow errors [12]. Runs of a workflow set as an error workflow do not count toward the execution quota (manual runs and sub-workflow runs also do not) [14].

That error workflow is your dead-letter queue. We have it post to Slack or email with the execution URL, write a row to a data table (workflow name, node, error, execution id, timestamp), and, only for a short allow-list of transient errors, call the n8n API retry endpoint. Do not auto-retry every failure. A 400 mapping error will fail again and loop.

On self-hosted instances that use the durable scheduler (available from n8n 2.36.0), scheduled runs that keep failing are reclaimed a limited number of times. `N8N_SCHEDULER_MAX_ATTEMPTS` defaults to 5; after that, n8n dead-letters the run [8]. Failed scheduler history is kept longer than successful history (default seven days versus one day) [8]. That is the platform-level dead letter. Your Error Workflow is the business-level one.

Turn on Save failed production executions. Without saved executions you cannot retry, debug, or attach an execution URL in the alert [13]. Save execution progress persists each node's data so a later retry can resume from where it stopped; n8n notes that this may increase latency [13].

Idempotency keys so a replay does not duplicate records

Retries are only safe if the write is idempotent. RFC 9110 (June 2022) defines an idempotent method as one whose intended effect on the server of multiple identical requests is the same as a single request. PUT, DELETE, and the safe methods are idempotent; POST is not [10]. CRM "create contact" and finance "create invoice" are POST operations. If the node succeeds on the server and then n8n records a failure, the next retry is a second POST.

AWS Well-Architected reliability guidance states the same rule: a mutating API should accept an idempotency token so multiple identical requests have the same effect as one [9]. Stripe's API implements that pattern with an Idempotency-Key on POST, storing the first status and body for that key and returning the same result on later requests. Keys can be V4 UUIDs, last at least 24 hours, and must not be personal identifiers [11].

In n8n workflows we generate the key in a Code or Edit Fields node before the write, then send it as the header or body field the destination API documents. A stable key looks like `{workflowId}:{sourceRecordId}:{operation}`. Do not generate a fresh UUID on every retry. If the key changes, the server treats the retry as a new operation [9][11]. When the destination has no idempotency header, we search by external id or invoice number, create only if missing, otherwise patch. For internal data tables, upsert on a unique business key rather than insert. Pass the same business id to every downstream write. AWS notes that in event-driven chains, every consumer must honor the token or a duplicate message still double-applies [9].

How to re-run a failed execution without doubling writes

n8n stores each run on the Executions list. Filter by Failed. From there you have two recovery paths.

Debug in editor. On a failed execution, Debug in editor copies the run data into the current canvas and pins it on the first node so you can inspect the payload, fix the mapping, and re-run with that data. Successful executions offer Copy to editor instead. This is available on n8n Cloud and on self-hosted Registered Community, Business, and Enterprise [7]. Confirm that failed production executions are being saved, or the list will be empty [7][13].

Retry from the Executions list. Open the failed row, use Refresh, then choose Retry with currently saved workflow after you fixed the node, or Retry with original workflow when the failure was transient (429, timeout) and the canvas did not need a change [18].

The public API exposes the same choice: `POST /executions/{executionId}/retry` with `loadWorkflow: true` to retry against the currently saved workflow, or omit it to replay the version stored with the original run. The new execution has `mode: retry` and `retryOf` set to the original id [17]. The Error Trigger payload includes `execution.retryOf` when the failing run was already a retry, which is how your dead-letter handler can refuse to retry a retry [1].

Before you click Retry: identify the last green write node and assume that side effect landed; confirm mutating nodes send an idempotency key or upsert (if they do not, replay only the failed tail); for 429s, wait out Retry-After [3][4]; use currently saved workflow after a fix, original after a blip [18]; if the new run fails the same way, park it. Pinning data in the editor is for diagnosis. Production recovery should go through Executions retry so you keep `retryOf` lineage.

Practical takeaways

  • Put Retry On Fail on HTTP nodes that face rate limits and transient 5xx responses. Set Wait Between Tries at or above the API interval, and add Batching or Loop Over Items plus Wait for bulk sends [2][3].
  • Leave On Error at Stop Workflow for money and CRM writes. Use Continue (using error output) when a batch should park bad items and keep the rest [2].
  • Attach one Error Workflow (Error Trigger first) to every published workflow. Alert, log, and retry only documented transient errors. It does not fire on manual runs [1][12][14].
  • Save failed production executions so you can retry and debug [13].
  • Make every create operation idempotent: native Idempotency-Key where the API supports it, otherwise upsert on a business key [9][10][11].
  • Re-run from Executions with currently saved workflow after a fix, or original workflow after a blip. Check `retryOf` so you do not retry a retry [18][17].
  • On self-hosted durable scheduler setups, a scheduled run is dead-lettered after `N8N_SCHEDULER_MAX_ATTEMPTS` (default 5) [8].

How we can help

Have more questions or want to get in touch? Our team in Charlotte, with work across Raleigh, Asheville, and Philadelphia, builds production n8n workflows with retry policy, error workflows, and idempotent writes so a failed run can be replayed without duplicating CRM or finance records. See how we approach AI business tools, then contact Idea Forge Studios, call (980) 322-4500, or email [email protected].

Citations

  1. n8n Docs, "Handle errors gracefully" (2026)
  2. n8n Docs, "Work with nodes" (2026)
  3. n8n Docs, "Handle rate limits" (2026)
  4. IETF RFC 6585, "Additional HTTP Status Codes, Section 4: 429 Too Many Requests" (2012-04)
  5. NIST, "The NIST Cybersecurity Framework (CSF) 2.0" (2024-02-26)
  6. NIST CSRC, "SP 800-61 Rev. 3, Incident Response Recommendations and Considerations for Cybersecurity Risk Management" (2025-04)
  7. n8n Docs, "Debug executions" (2026)
  8. n8n Docs, "Scheduler" (2026)
  9. AWS Well-Architected, "REL04-BP04 Make mutating operations idempotent" (2026)
  10. IETF RFC 9110, "HTTP Semantics, Section 9.2.2: Idempotent Methods" (2022-06)
  11. Stripe, "Idempotent requests" (2026)
  12. n8n Docs, "Error Trigger" (2026)
  13. n8n Docs, "Configure workflow settings" (2026)
  14. n8n Docs, "Understand executions" (2026)
  15. n8n Docs, "Stop And Error" (2026)
  16. n8n Docs, "Wait" (2026)
  17. n8n Docs, "Executions API: Retry an execution" (2026)
  18. n8n Docs, "View executions for a single workflow" (2026)
Our Strongest Offering

Forge Your Next Website

Forged Sites are custom-built, static-first websites with a full AI content engine on board — no CMS to log into, no plugins to break, no builder to fight.

  • Working target: WCAG 2.2 AA
  • During work hours, an account manager still reviews material changes
  • DraftDash auto-drafted blogs keep your content engine running
  • Ethel AI-powered forms filter spam and capture genuine leads