Resume failed workflow runs from checkpoints with Postgres and Slack

Go to Workflow
0 views
Built by Oneclick AI Squad Oneclick AI Squad
Created on October 02, 2026

Description

Quick overview
This workflow implements a checkpoint-and-resume system for a multi-stage pipeline, storing stage outputs in Postgres and sending failure/recovery alerts to Slack, with manual, webhook, and scheduled triggers to start runs and automatically resume failed or stalled executions.

How it works
Starts a new run via Manual Trigger, receives a POST request via webhook, or runs every 10 minutes on a schedule to sweep for resumable runs.
Ensures the required Postgres tables for runs, checkpoints, and run events exist.
For scheduled sweeps, queries Postgres for failed runs whose retry time has passed or running runs whose lease expired, then calls this workflow’s webhook to resume each run.
For manual/webhook runs, loads the run state and completed checkpoints from Postgres, restores prior stage outputs into context, and determines the first stage that still needs to run.
Acquires a lease-based lock in Postgres to ensure only one execution owns the run, and exits if the run is already completed or currently locked by another execution.
Executes the pipeline stages in order (validate order, reserve inventory, charge payment, create shipment, send confirmation), saving a Postgres checkpoint and event after each successful stage.
On stage failure, retries the stage inline with exponential backoff when allowed, otherwise marks the run as failed or dead-lettered in Postgres and sends a Slack alert, and on completion marks the run completed and optionally posts a Slack recovery notification when resuming from checkpoints.

Setup
Add Postgres credentials for the database where the workflow can create and update the workflow_runs, workflow_checkpoints, and workflow_run_events tables.
Add Slack credentials and set the target Slack channel in the configuration (and update the Slack nodes if you want a different channel behavior).
Update the selfWebhookUrl value in the configuration to this workflow’s production webhook URL ending in /webhook/checkpoint-run so the sweeper can trigger resumes.
If you want crash notifications, enable the Error Trigger node and set this workflow as the Error workflow in your n8n instance settings.

Nodes Used (4)

Code
n8n-nodes-base.code
HTTP Request
n8n-nodes-base.httpRequest
Postgres
n8n-nodes-base.postgres
Slack
n8n-nodes-base.slack