Skip to content
PD
n8n

Error Handling in n8n: Retries, Error Workflows and Reliable Automation

How to make an n8n workflow durable: node-level error settings, retries and timeouts, a dedicated error workflow for alerts, idempotency under re-runs, and why automation usually fails silently.

All articles in the guide n8n · 18

A workflow built in an evening works. A workflow still working six months later without your attention is a different engineering object. The difference is almost entirely error handling.

The real danger is the silent failure

Automation rarely breaks loudly. Usually one of two things happens: the workflow deactivates after a run of failures and simply stops firing, or it runs empty - the trigger fires, no data arrives, zero output, no error.

Both share one property: nobody finds out. Data stops flowing, a business process stalls, and it surfaces a week later, usually from outside - a customer or an accountant.

So the first thing to do in production is not to improve the logic but to make failure visible.

The error workflow

n8n lets you designate a workflow that runs automatically when another one fails. It receives what failed, where, and with which error.

This is the cheapest reliability available: one alerting workflow bound to every production automation. A minimally useful message contains the workflow name, the node and the error text - otherwise you get “something broke” and still have to go digging.

One limitation worth understanding: an error workflow catches failures, not empty results. A workflow that ran and did nothing is formally a success. For that you need a separate check - a daily “how many records were processed” with an alert when the answer is zero.

Node-level settings

Every node has error behaviour options. Three actually get used:

Retry on fail. Attempt count and a pause between attempts. This covers most external flakiness: a network blip, a 502, a response that took longer than usual.

Continue on fail. The node does not kill the execution but passes the error onward as data for you to handle. Useful in batch work: one bad record out of a hundred should not stop the other 99.

Timeout. A cap on how long a response may take. Without it a slow API holds the execution indefinitely - and on a frequent schedule that also piles up concurrent runs.

What to retry and what not to

The split is simple and almost always the same.

Worth retrying: network failures, timeouts, 5xx responses, rate limits (with a longer pause). These are states that pass on a second attempt.

Not worth retrying: authorisation errors, 404s, invalid data, an exhausted quota. A retry will not fix them - it only burns time, and with rate limits it makes things worse.

The practical corollary: a retry without a limit is a hidden infinite loop. Always cap the attempts.

Idempotency

Retries and re-runs raise a question people skip at build time: what happens if this step executes twice?

For reading data, nothing. For creating an order, charging a card or sending an email, quite a lot. A workflow that creates a duplicate record on retry is more dangerous than one that simply failed.

The defence is making the action safe to repeat: check whether the record exists before creating it, use an operation id the service recognises as a duplicate, or mark processed items in your own store.

This is not theoretical: with retries enabled, and every time you manually re-run a failed execution, repetition happens regularly.

Partial processing

A separate nuisance: the execution died halfway - half the items processed, half not. Restarting from the beginning reprocesses the first half.

Two practical rules follow. First, record progress as you go, not at the end: write the processed flag immediately after each item rather than after the whole loop. Second, on a re-run, filter out what is already done before starting. Both require somewhere to keep state: a database, a sheet, a flag on the record itself.

The production minimum

Before automation becomes part of a real process:

  1. An error workflow with alerts bound to every production workflow.
  2. Retries configured where failures are transient.
  3. Timeouts set on external calls.
  4. Re-execution is safe for anything that creates or sends.
  5. An alert on empty results, not only on failures.

The fifth is the one most often skipped - and the one that catches the most expensive failures.

FAQ

How do I find out that an n8n workflow failed?

Configure a dedicated error workflow - it runs automatically when any workflow bound to it fails, and receives the error details. Send a Telegram or email alert from inside it. Without one you learn about failures only when you open the executions list, or when a customer tells you.

What should I do when an external API fails intermittently?

Enable node-level retries with a few attempts and a pause between them - most network blips and 5xx responses pass on the second try. Do not retry authorisation errors or 4xx: they will not turn into successes and only waste time.

More on this topic

Done for you

I will build the automation in n8n or in code

Leads, sheets, CRM and Telegram connected, so nobody moves data by hand again.

from $300 · 3 to 7 days

Similar caseBAS Script License Issuing Automated on MakeA Make scenario that turns one Telegram message into a full licence handover: generated login and password, a licence for the requested term, FingerprintSwitcher Business enabled, and a row written to Google Sheets.

"Thanks to Pavel, the task is done. Always reachable, gave me detailed instructions and a guide, I will come back and I recommend him to everyone."

MarkBorisov · KworkTranslated from Russian