Skip to main content

Building Reliable Flows

One failure ends a whole run. That single fact drives most of what follows: there is no per-step error branch, no dead-letter queue and no automatic replay. A flow that matters has to be designed so that ending halfway through is survivable. This page covers what the runner actually does. For the step types themselves see Steps; for what can start a run see Triggers.

What a failure does

The first step to error wins, and everything after it is a consequence:
  1. The error is recorded against the run, and the run’s status becomes error.
  2. The run is cancelled. No further step starts.
  3. Steps already in flight are abandoned — they are not waited for and their results are discarded.
  4. The results of steps that had finished are kept and returned with the failure.
  5. The failed step itself stores nothing, so there is no result for a later step to inspect.
Because branches run in parallel, “steps that had finished” can include work on a branch unrelated to the one that failed. A run that dies at step 7 may already have written to storage at step 3.
Design writes to be safe to repeat, and put them as late as you can. The runner will not undo a write made before the failure, and rerunning the flow makes it a second time.

Recovering a run

Open Monitoring, select the run, and three controls appear depending on its status: Retry is possible because the failure is captured with everything needed to carry on: the results of every completed step, and the remainder of the graph from the failed step onward. It is the right control after a transient failure — an API that was down, a timeout, a bad credential you have since fixed.
If you have changed the flow since it failed, Retry will not pick the change up — it deliberately replays the configuration version the run started with. Publish the fix and use Rerun, which starts again on the current configuration.

Retries

Retries are configured per step, under Advanced Settings, and are off by default.
integer
Total attempts, not additional ones. 1 or 0 means no retry. Values above 5 are capped at 5.
integer
Milliseconds to wait between attempts. It is a constant wait — there is no exponential backoff and no jitter.
array
Response codes that must not trigger a retry.
integer
Per-step timeout in milliseconds.
Only four step types retry: Function, HTTP Request, Integration and nested flow. Setting attempts on a Map, Eval, Rules or Condition step does nothing — those either work or they do not.
“Ignore Response Codes” does not ignore the error. A code in that list is logged, the retry is skipped, and the step still fails the run. There is no setting that lets a flow continue past a 4xx or 5xx from an HTTP Request step. If a non-2xx response is a normal outcome for your integration, call it through a Function step and branch on the body instead.
A retry re-runs the step from the beginning, including its side effect. A step that publishes a message, writes a key or posts to an API will do it again.
Make repeated work harmless rather than trying to prevent it:
  • Write under a deterministic key — the order ID, not a timestamp — so a repeat overwrites rather than duplicating.
  • Read before you write when a duplicate would be visible to a person.
  • Keep the retried step small. A step that does three things cannot repeat one of them.

Timeouts

A Function step and a nested-flow step time out after 20 seconds unless you set timeout. An HTTP Request step falls back to 30 seconds, and takes its bound from the payload’s own timeout if that carries one, otherwise from timeout under Advanced Settings. timeout is a millisecond count everywhere it appears. Set it on anything that leaves the platform and is slower than the default.

Failures that do not look like failures

Three things go wrong quietly. All three leave the run in status success. A function that returns an error object. Most built-in functions catch their own errors and return {"error": "…"}. The invocation succeeded, so the step succeeded, and the next step reads an error object where it expected data. Branch on $.STEP_ID.error or $.STEP_ID.status_code when the result matters. See Functions. A JSONPath expression that matches nothing. It resolves to [], not null and not an error. A field name typo therefore produces an empty array that flows on downstream. Test with Array.isArray(x) && x.length === 0, not for a missing key. A rule that produces nothing. A misspelled operator is now caught before the step runs and fails it outright, naming the operator. Several other mistakes are not: a calculation that produces NaN, a condition ladder where no cell matches and there is no default, and an outcome of null all drop the fact from the result with no error. The flow continues with the fact simply absent, which usually shows up much later as a condition taking the wrong branch. See when a rule produces nothing.

Controlling how often a flow runs

The gear icon in the flow editor toolbar opens Flow options, where Run Type changes how concurrent triggers are handled. Debounce On and Order On are JSONPath expressions resolved against the trigger, so $.trigger.customer_id gives you one queue per customer. Both times are in nanoseconds: five seconds is 5000000000.
A debounced or ordered trigger returns straight away with a message rather than a tracking ID, because the run has not started yet. Anything expecting a synchronous result — an HTTP Request trigger answering its caller — should not use either.

Secrets

Put SECRET::<name>:: anywhere in a node’s configuration and the platform substitutes the value once, before the first step runs. That is the right way to reach a credential from a flow, and it is better than a step that fetches one, because the value never becomes a step result.
  • A name containing a . is refused. Account-shared secrets have plain names.
  • SECRET::user.<KEY>:: resolves against your own identity for a personal secret.
  • Step results are stored and shown in Monitoring. A credential that passes through a step’s result is visible to anyone who can read the run.

The limits worth knowing


Before you publish

  • Test the draft. The editor’s Test control runs the draft with a payload you type. It does not exercise the trigger.
  • Then publish. Editing changes the draft only. A flow that behaves correctly under test keeps running its old published version until you publish.
  • Check a trigger survived the publish. A trigger belongs to the flow version it was created on — see each trigger page.
  • Watch the first live runs in Monitoring. Step results are recorded per run, which is the only place a silent failure becomes visible.

Steps Overview

How a run moves through the graph

Triggers Overview

Which triggers can lose an event

Condition Step

Branching on a result before you use it

Built-in Functions

The function catalogue