BlogEngineering

Self-healing data pipelines: a complete walkthrough of ETL that fixes itself

· By Rhombus AI Engineering Team · 9 min read

Overview

Data pipelines rarely fail when you build them. They fail at 2 a.m., on a column that changed type, a join key with trailing spaces, or a file three times bigger than last month. Someone gets paged, opens the logs, finds the one-line fix, and reruns the job.

We wanted the pipeline to do that part itself.

A self-healing data pipeline is a pipeline that finds its own errors, works out the cause, applies a fix and checks that the fix worked, without a person stepping in. If it can't fix the problem within a set number of tries, it stops and tells you why.

Self-healing is not the same as retrying

Most pipeline tools already have retries and alerts. Neither one fixes anything.

ApproachWhat happens when a step failsWho fixes it
RetryRuns the same step again, unchangedNobody. It only helps if the failure was temporary
Monitoring and alertsTells someone the step failedA person
Self-healingReads the error, changes the code or plan, reruns it and checks the resultThe pipeline, within set limits, then a person if it can't

A retry is useful when a network call times out. It does nothing for a join on columns of different types, because that step will fail the same way every time.

Why data pipelines break, and why an AI model alone doesn't fix them

Pipelines break in three different places, and each one needs a different kind of fix:

  • In the design. The plan itself is wrong: a join is missing its key, a parameter has the wrong type, or a step points at an input that doesn't exist.
  • In custom code. Some logic doesn't fit a standard step, like "split this free-text address into street, suburb and postcode". That code has to be written, and new code has bugs.
  • At scale. The pipeline works on a sample and then fails on the full dataset in a distributed engine like Spark. These bugs only show up on the real job.

The obvious idea is to point a large language model at the error and let it rewrite the code. On its own, that doesn't work well. A single model writing ETL code is like a junior engineer working alone: fast and often right, but nothing checks its work before it ships. It can also "fix" an error by quietly dropping the column that caused it.

So we made the model one member of a team, and surrounded it with checks that never get tired.

How a self-healing data pipeline works

In Rhombus AI, self-healing comes from three healing loops, one for each place a pipeline can break. Each loop has its own check, its own way of fixing things and a hard limit on attempts.

A lead agent with a team behind it

User
“Join claims to members and flag duplicates”
Rhombus AI agent team
Rhombo · lead agent
plans · reasons · decides
Subagents
do focused research and checks in their own context
Skills
proven playbooks it follows for each kind of task
Code sandbox
tests the custom code it writes, safely
Pipeline compiler
checks every change before it can be saved
saved
Automatic pipeline repair
fixes failed jobs
Execution
runs the pipeline
Pipeline DAG
the saved, validated pipeline
live flowfailure caughtfixed and verified
One lead agent, with specialists and checks behind it. The compiler is the only path to a saved pipeline; failed runs flow back through automatic repair.
MemberWhat it doesWhy it matters
Rhombo (lead agent)Understands the request, inspects the data, plans the pipeline, decides what to changeThe reasoning happens in one place, with the whole conversation in view
SubagentsShort-lived agents Rhombo spawns in a separate context for focused research or validationDeep dives don't flood the lead agent's working memory
SkillsPlaybooks for building pipelines and for analysis, loaded when the task needs themThe same proven workflow runs every time
Pipeline compilerTurns proposed actions into a validated pipeline changeThe compiler catches invalid pipeline plans before they're saved
Code sandboxRuns generated transform code under static and runtime guardrailsCustom logic is safe to execute on real data
Automatic pipeline repairDiagnoses failed big-data jobs, patches the script and resubmitsScale-only failures heal without a person in the loop

The Rhombus AI agent team.

Six stages, the way a senior data engineer works

  1. Understand. Rhombo works out whether you want a reusable pipeline or a one-off answer. If the request is unclear ("remove the bad rows"), it asks you a multiple-choice question instead of guessing.
  2. Discover. It finds the right sources in your project and profiles them: types, nulls, distinct counts and sample rows. Large files are queried where they live instead of being loaded into memory.
  3. Design. Rhombo designs each step of the pipeline, drawing on a library of 200 transformer types, such as joins, filters, aggregations, pivots, deduplication, outlier handling, missing-value imputation, fuzzy merges, date parsing and type conversion. It reads each one's exact spec before using it, so it works from the spec rather than from memory.
  4. Compile. The agent turns the design into a small, typed plan. A compiler checks it before anything is saved.
  5. Run. The pipeline runs, whether the input is a small file or a very large dataset.
  6. Verify. Rhombo previews the output, checks it against your request and explains in plain language what changed. You watch its reasoning and progress live in the chat.

For a request like "join claims to members and flag duplicates", the plan has two steps. Each one names a transformer, where its data comes from and its settings.

Sources
claims
members
Step 1 · merge
attach member details
left join on member_id
Step 2 · remove_duplicate
keep one row per claim
unique on claim_id, keep first
✓ Pipeline saved
Compiler checks
step types, required settings
inputs connect, graph valid
The typed plan for "join claims to members and flag duplicates". Because every step is a known type with known settings, the compiler can check the whole plan before anything runs.

The three healing loops

Loop 1 · Design time
Catches a bad plan
fail
Agent plans
typed actions
Compiler checks
against what exists
✓ Pipeline saved
Fail: error with context goes back
Same check, every time
Loop 2 · Custom code
Catches bad code
fail
Agent writes code
custom transform
Sandbox checks
code, run, output
✓ Node ready
Fail: error + past attempts
Hard limit on attempts
Loop 3 · Run time
Catches scale-only bugs
fail
Job runs at scale
on the full dataset
Job checked
script + error kept
✓ Output delivered
Fail: known fix or AI patch
Hard limit on retries
Three healing loops catch errors at design, code and run time. Each has a check, a fix and a limit.

Each loop sends its error straight back to the agent with full context, and each one stops at a set limit.

Loop 1: the compiler catches design mistakes. This loop exists to stop hallucination. Before anything is saved, a deterministic compiler verifies the whole plan against what actually exists in your project — the model cannot make up a step, a column or a source that isn't there. When part of the plan doesn't hold up, the agent gets back an error with enough context, hints and semantic meaning to understand exactly what went wrong, and revises the plan. Because the compiler is deterministic, the same mistake is caught the same way every time. It never depends on the model noticing its own error.

Loop 2: the sandbox catches bad code. When logic doesn't fit a standard step, Rhombus writes Python and runs it through three layers of checks:

  • Before it runs: the generated code is checked for unsafe operations and limited to the intended data transformation — nothing else can run.
  • While it runs: it executes in a sandbox on a limited sample of the data.
  • After it runs: the output must be a well-formed table of the expected shape, with no input columns quietly dropped.

When a check fails, the error and the history of earlier failed attempts go into the next try. The model learns from the whole trail of mistakes, not just the latest one. There is a hard limit on attempts; if the code still isn't right, the loop stops and reports why.

Loop 3: automatic pipeline repair catches failures that only happen at scale. When a job fails on the full dataset, Rhombus stores the failing script and the error, then:

  • Tries known fixes first. A library of rule-based repairs handles the most common failure patterns, like a key column dropped before it is used, or code that behaves differently at scale.
  • If none apply, asks the model for a targeted patch and checks that the patched script compiles before running it.
  • Resubmits the job, and if the failure still can't be fixed, hands you a clear diagnosis.

Known fixes run first because they are instant and certain. The model handles the long tail.

Watch it live

Self-Healing ETL Pipelines · watch on YouTube

Putting it together

One request, end to end, in 11 steps. The compiler catches a design mistake at step 4, and a known rule repairs a failure at scale at step 8. The user only sees the finished output and a plain-language summary.

failure caughtfixed and verified
User
Rhombo
Compiler
Executor
Healing loops
Join claims to members, flag duplicates
1
Discover and profile sources
2
Healing loop 1 · compiler
Proposed plan
3
Error: the join is missing its key
4
Revised plan · join key added
5
Valid. Pipeline saved
6
Run the pipeline
7
Healing loop 3 · automatic repair
Job fails on the full dataset
8
Known fix applied, job resubmitted
9
Output ready
10
Preview and plain-language summary
11
One request, end to end, in 11 steps. Two failures caught and fixed before the output is ready.

The result is a pipeline that was designed by reasoning, validated by a compiler, tested in a sandbox and repaired at scale. You can read, edit and schedule it like any other pipeline on the canvas.

Try it on your own data

If you have a pipeline that breaks every few weeks for the same boring reasons, bring it to Rhombus AI and describe what it should do in plain English. Start at rhombusai.com.

← All posts
© 2026 Rhombus AI®. All rights reserved.