Errors and retries in n8n: a workflow that doesn't stay silent when it fails
Sooner or later someone else's API will slow down, return an error or be unreachable. The question is whether you find out right away or a week later. We learn three layers of protection — node settings (retry and behaviour on error), the Stop and Error node and a separate error workflow with an Error Trigger — and how to keep your executions long enough to investigate them.
$json.error.message and $json.node.name do not exist; the real ones are $json.execution.error.message and $json.execution.lastNodeExecuted. (4) $http.get in the Code node example does not exist — removed; the new example was run in Node.js. (5) The retry limits were missing: Max Tries is 2–5 (default 3, includes the first attempt), Wait Between Tries is 0–5000 ms (default 1000). (6) The Error Trigger cannot be tested with a manual run — it fires only when an automatically started workflow fails. (7) Unverified claims were removed ("Concurrency in Workflow Settings", "unpin test data", a log folder path). Added: the Stop and Error node, "Continue (using error output)", "Never Error" on HTTP Request, exactly what the Error Trigger receives (and when there is no execution.url), the settings for saving executions and the variables that delete old executions (14 days and 10,000 by default), a warning about retrying writes. Removed: a personal recipient email, the example with a specific Postgres database and a specific chat — the exercise now runs with a channel of your choice.
01What you'll learn
- What n8n does by default on an error and which three node settings change that.
- How to switch on retry (Retry On Fail) and what its limits are.
- How to choose between "stop", "continue" and "continue through an error output" (On Error).
- How to fail a workflow on purpose with Stop and Error when the data is bad.
- How to build an error workflow with an Error Trigger and attach it to other workflows.
- How long executions are kept and how to configure their deletion.
02Before you start
- You have done Lessons 1 and 2 — you have a running n8n 2.x installation 🔒 local and you know the Webhook and HTTP Request nodes.
- For the exercise: access to a notification channel of your choice — an email node or a chat node 🌐 global. Log-in details live in a credential, not in the node.
- Permission in n8n to create workflows and change their settings.
03Steps
-
What happens by default — and where the three layers are
When a node fails and nothing is configured, n8n stops the whole execution and marks it as failed. Nobody is told — unless you prepared for it. So we think in three layers:
1 · Node settings→2 · Checks and Stop and Error→3 · Error workflowWhy three? First you let the node try again (the network blinks sometimes). Then, if the data is bad, you decide — and fail the workflow with a clear message. Finally, whatever is left uncaught must reach a human.
-
Layer 1a: Retry On Fail — a second chance without code
Open the node → Settings tab → switch on Retry On Fail. Max Tries and Wait Between Tries appear. The documentation only says "the node reruns until it succeeds"; we took the exact limits from n8n's source code:
Field Default Limits Note Max Tries 3 2 to 5 this is all attempts, including the first (3 = first + 2 retries) Wait Between Tries 1000 ms 0 to 5000 ms pause between attempts, in milliseconds n8n · the node's Settings tabRetry On Fail: on Max Tries: 3 Wait Between Tries: 1000 # ms On Error: Stop Workflow⚠️Retrying a write can create a duplicateIf the node creates something (for example a POST that makes a record), a retry after a lost reply can create a second record. Retry reads and writes that survive repetition. This is good practice, not a claim from the documentation.For a fixed pause of up to 5 seconds this is enough. If you want a longer or growing pause — see step 7.
-
Layer 1b: On Error — what happens when the attempts run out
The same tab has an On Error setting with three values:
Value What it does When Stop Workflow halts the whole execution; no following node runs a critical step — better not to go on Continue moves to the next node despite the error, with the last valid data a side step whose failure does not matter Continue (using error output) continues, but sends the error information through a separate output of the node when you want to handle the error inside the same workflow With the third value the node gets a second output — connect it to its own chain (a record, a message, a fallback path). The other two are simpler, but "Continue" hides the problem: the following nodes work with data that may be stale, so use it only when that is safe.
✅Always Output Data — not for errorsThis setting does one thing: if the node returns nothing, it returns an empty item so the chain after it does not stop. Don't expect it to handle "errors". Be careful with IF nodes — with it you can get an endless loop (a warning from the documentation). -
Layer 2a: HTTP Request and the response code
By default HTTP Request fails when the API returns a code outside 2xx. If you want to decide yourself what happens (for example, take another path on a 404), choose Response → Never Error in the node's options and put an IF on the response code after it (in Lesson 2 you saw how to include the full response).
n8n · HTTP Request → IFHTTP Request: Options → Response → Include Response Headers and Status → Never Error IF: Value 1: {{ $json.statusCode }} Operation: Number → is equal to Value 2: 200 # true → continue normally # false → error path (message, record, Stop and Error)Some APIs return 200 with an error inside the body. Then the check is on the body, for example whether the field
errorexists — look at the response of your specific API. -
Layer 2b: Stop and Error — you fail the workflow on purpose
Sometimes nothing is "broken", but the data is unusable — a required field is missing, an amount is negative. The Stop and Error node forces the execution to fail with your message and so triggers the error workflow. It has two operations:
Error Type What you provide Error Message the message text — clear, without secrets Error Object a JSON object with the properties of the error you want to throw Put it on the "bad" branch of an IF node: every validation that fails then ends with a call for help instead of a silent skip.
-
Layer 3: an error workflow with the Error Trigger
This is the safety net under everything else. You make a separate workflow whose first node is the Error Trigger, and point to it from the others.
n8n · creating it# 1) New workflow → first node: Error Trigger # 2) Name it, for example: Error Handler → Save # 3) In the workflow you want to watch: # Options (⋯) → Settings → Error workflow → pick "Error Handler" → Save # One error workflow can serve many workflows.What the Error Trigger receives — one item of this shape:
JSON · Error Trigger data (example)[{ "execution": { "id": "231", "url": "https://n8n.example.com/execution/231", "retryOf": "34", "error": { "message": "Example Error Message", "stack": "Stacktrace" }, "lastNodeExecuted": "Node With Error", "mode": "manual" }, "workflow": { "id": "1", "name": "Example Workflow" } }]expressions in the error workflow{{ $json.workflow.name }} // which workflow failed {{ $json.execution.lastNodeExecuted }} // on which node {{ $json.execution.error.message }} // what the message is {{ $json.execution.url }} // link to the execution (not always present) {{ $json.execution.id }}✅Rules for the Error Trigger (from the documentation)The error workflow does not have to be published. A workflow that contains an Error Trigger uses itself as its error workflow by default. You cannot test it with a manual run — it fires only when an automatically started workflow fails (step 9).execution.idandexecution.urlexist only if the execution was saved to the database; they are missing when the error is in the trigger of the main workflow.retryOfappears only when a failed execution is run again.If the error is in the trigger of the main workflow (for example it cannot attach), the data is different: instead of
executionthere is atriggerblock witherrorandmode: "trigger", andworkflowstays. Write your expression so it survives both kinds — for example{{ $json.execution?.error?.message ?? $json.trigger?.error?.message }}⚠️ (the expression follows the documented fields; we did not run it). -
Retry with a growing pause in a Code node
The built-in Retry On Fail goes up to 5 attempts and a 5-second pause. If you need a growing pause (1 s, 2 s, 4 s…) or different logic on each attempt, write a loop in a Code node. Here is an example that we ran in Node.js: after two failures the third attempt succeeds; on a permanent failure it throws a clear error after all attempts.
JavaScript · retry with exponential backoffconst maxAttempts = 4; // all attempts, including the first const baseDelay = 1000; // ms; the pauses are 1000, 2000, 4000 async function retryWithBackoff(fn) { let lastErr; for (let attempt = 1; attempt <= maxAttempts; attempt++) { try { return await fn(attempt); } catch (err) { lastErr = err; if (attempt === maxAttempts) break; const delay = baseDelay * 2 ** (attempt - 1); await new Promise(r => setTimeout(r, delay)); } } throw new Error(`Failed after ${maxAttempts} attempts: ${lastErr.message}`); } // Your logic that may fail: const result = await retryWithBackoff(async (attempt) => { // ... a call that throws on failure return { ok: true, attempt }; }); return [{ json: result }];⚠️Not checked inside the Code node itselfThe code was run in Node.js (with shorter pauses), not inside n8n. The documentation says the Code node supports promises, but says nothing aboutsetTimeoutin its environment. So for simple cases use the built-in Retry On Fail, and for a longer pause use the Wait node. The old example called$http.get— there is no such thing; for requests use the HTTP Request node. -
Alert and log: how not to lose the error
In an error workflow you usually want two things: notify a person and keep a trace.
Error Trigger→Edit Fields · compose the text→Notification node→Write to a table (optional)alert textSubject: n8n error: {{ $json.workflow.name }} Body: Workflow: {{ $json.workflow.name }} Node: {{ $json.execution.lastNodeExecuted }} Error: {{ $json.execution.error.message }} Link: {{ $json.execution.url }}The notification node is your choice (email, chat, something else) — the log-in details are in a credential. A write to a table (your own table or database) gives you a "list of broken things for manual review" — old executions get deleted (see below), while the row in the table stays.
✅The alert carries the text, not only the linkExecutions are deleted by age and by count (next table). If the alert is only a link, after a while it leads nowhere. So put names and the message in it. -
How long executions are kept
In each workflow's Workflow Settings there are Save failed production executions, Save successful production executions, Save manual executions and Save execution progress. At the installation level (self-hosted) environment variables control what is kept and for how long:
Variable Default What it does EXECUTIONS_DATA_PRUNEtrue deletes old executions on a rolling basis EXECUTIONS_DATA_MAX_AGE336 (hours = 14 days) age after which an execution is deleted EXECUTIONS_DATA_PRUNE_MAX_COUNT10000 maximum executions kept in the database; 0 = no limit EXECUTIONS_DATA_SAVE_ON_ERRORall whether data is saved on error (all / none) EXECUTIONS_DATA_SAVE_ON_SUCCESSall the same on success (all / none) EXECUTIONS_DATA_SAVE_ON_PROGRESSfalse saves progress after every node EXECUTIONS_TIMEOUT·EXECUTIONS_TIMEOUT_MAX-1 (off) · 3600 total time in seconds for a workflow · upper limit for a single workflow's setting N8N_WORKFLOW_AUTODEACTIVATION_ENABLEDfalse unpublishes a workflow after repeated crashed executions; the number is in …_MAX_LAST_EXECUTIONS(3)example · keep failures for 30 daysEXECUTIONS_DATA_PRUNE=true EXECUTIONS_DATA_MAX_AGE=720 # 30 days in hours EXECUTIONS_DATA_PRUNE_MAX_COUNT=10000 EXECUTIONS_DATA_SAVE_ON_ERROR=allMatch them to your own volume: if a workflow runs thousands of times a day, the count limit will delete executions much sooner than 14 days. The variables apply to self-hosted n8n; on the cloud version ⚠️ the settings are different and we did not check them.
-
Exercise: a workflow that fails on purpose, and an alert
You build two workflows: one that fails on demand, and one that catches the error and sends an alert.
n8n · the "Error Handler" workflowError Trigger → Edit Fields: msg = Workflow "{{ $json.workflow.name }}" failed on node "{{ $json.execution.lastNodeExecuted }}": {{ $json.execution.error.message }} → Notification node (your channel): text = {{ $json.msg }}n8n · the "Fail Demo" workflowWebhook: POST · Path: fail-demo · Authentication: Header Auth Respond: Immediately → Stop and Error: Error Type: Error Message Error Message: Demonstration error (exercise 8) Workflow Settings → Error workflow: "Error Handler"Publish "Fail Demo" (errors from a manually run workflow do not reach the Error Trigger) and call it with the production URL, as in Lesson 2:
bash · trigger itKEY="<your-generated-key>" curl -i -X POST https://<your-n8n-host>/webhook/fail-demo \ -H "X-Api-Key: $KEY"In Executions you should see two records: the failed "Fail Demo" and the successful "Error Handler", and an alert should arrive on your chosen channel. Afterwards remove the Stop and Error node or stop the workflow — otherwise every call will raise an alert.
-
When something doesn't work
❌The error workflow does not startCheck that the failing workflow was started automatically (a published trigger), not by a manual run, and that its Workflow Settings point to the error workflow.❌Noexecution.urlThe execution was not saved to the database (the "Save failed production executions" setting), or the error is in the trigger of the main workflow. Then useworkflow.nameand the error text.❌Retry doesn't helpMax Tries goes up to 5 at most and the pause up to 5 seconds. For longer waiting use the Wait node or your own loop (step 7).❌The workflow "continues" but the data is wrongMost likely On Error: Continue is chosen. The following nodes get the last valid data. Switch to "Continue (using error output)" and handle the error on a separate branch.💡The question for every step"If this fails at 3 a.m. — who will know and what will they see?" If the answer is "nobody", add an error workflow.
04Check
Checklist
- Nodes that call an unstable service have Retry On Fail (for example 3 tries, 1000 ms); writes are not retried blindly.
- For every important node, On Error is chosen deliberately.
- There is an error workflow with an Error Trigger and it is selected in the Workflow Settings of every working workflow.
- The alert contains the workflow name, the node name and the message, not only a link.
- You provoked an error on purpose (Stop and Error) and got an alert; you also tried the success path.
- Failed executions are kept as long as you need to investigate them.
- There are no keys or personal data in the alerts or in your screenshots.
Quiz
1. What values can Max Tries take in Retry On Fail?
2. Where is the name of the node on which the workflow failed in the Error Trigger data?
3. Which node fails the execution on purpose, with your message, and triggers the error workflow?
4. Why should the alert carry the error text and not only a link to the execution?
05What's next
06Sources
- n8n: working with nodes — node settings (Always Output Data, Execute Once, Retry On Fail, On Error).
- n8n: Handle errors gracefully — error workflow, data, Stop and Error.
- n8n: the Error Trigger node · the Stop And Error node.
- n8n: workflow settings — Error Workflow, saving executions.
- n8n: execution environment variables — pruning old executions, timeouts.
- n8n on GitHub: node execution — the limits of Retry On Fail (Max Tries 2–5, Wait 0–5000 ms).
- n8n: using the Code node.