Self-healing data pipelines: a complete walkthrough of ETL that fixes itself
· By Rhombus AI Engineering Team · 9 min read
Overview
Data pipelines rarely fail when you build them. They fail at 2 a.m., on a column that changed type, a join key with trailing spaces, or a file three times bigger than last month. Someone gets paged, opens the logs, finds the one-line fix, and reruns the job.
We wanted the pipeline to do that part itself.
A self-healing data pipeline is a pipeline that finds its own errors, works out the cause, applies a fix and checks that the fix worked, without a person stepping in. If it can't fix the problem within a set number of tries, it stops and tells you why.
Self-healing is not the same as retrying
Most pipeline tools already have retries and alerts. Neither one fixes anything.
| Approach | What happens when a step fails | Who fixes it |
|---|---|---|
| Retry | Runs the same step again, unchanged | Nobody. It only helps if the failure was temporary |
| Monitoring and alerts | Tells someone the step failed | A person |
| Self-healing | Reads the error, changes the code or plan, reruns it and checks the result | The pipeline, within set limits, then a person if it can't |
A retry is useful when a network call times out. It does nothing for a join on columns of different types, because that step will fail the same way every time.
Why data pipelines break, and why an AI model alone doesn't fix them
Pipelines break in three different places, and each one needs a different kind of fix:
- In the design. The plan itself is wrong: a join is missing its key, a parameter has the wrong type, or a step points at an input that doesn't exist.
- In custom code. Some logic doesn't fit a standard step, like "split this free-text address into street, suburb and postcode". That code has to be written, and new code has bugs.
- At scale. The pipeline works on a sample and then fails on the full dataset in a distributed engine like Spark. These bugs only show up on the real job.
The obvious idea is to point a large language model at the error and let it rewrite the code. On its own, that doesn't work well. A single model writing ETL code is like a junior engineer working alone: fast and often right, but nothing checks its work before it ships. It can also "fix" an error by quietly dropping the column that caused it.
So we made the model one member of a team, and surrounded it with checks that never get tired.
How a self-healing data pipeline works
In Rhombus AI, self-healing comes from three healing loops, one for each place a pipeline can break. Each loop has its own check, its own way of fixing things and a hard limit on attempts.
A lead agent with a team behind it
| Member | What it does | Why it matters |
|---|---|---|
| Rhombo (lead agent) | Understands the request, inspects the data, plans the pipeline, decides what to change | The reasoning happens in one place, with the whole conversation in view |
| Subagents | Short-lived agents Rhombo spawns in a separate context for focused research or validation | Deep dives don't flood the lead agent's working memory |
| Skills | Playbooks for building pipelines and for analysis, loaded when the task needs them | The same proven workflow runs every time |
| Pipeline compiler | Turns proposed actions into a validated pipeline change | The compiler catches invalid pipeline plans before they're saved |
| Code sandbox | Runs generated transform code under static and runtime guardrails | Custom logic is safe to execute on real data |
| Automatic pipeline repair | Diagnoses failed big-data jobs, patches the script and resubmits | Scale-only failures heal without a person in the loop |
The Rhombus AI agent team.
Six stages, the way a senior data engineer works
- Understand. Rhombo works out whether you want a reusable pipeline or a one-off answer. If the request is unclear ("remove the bad rows"), it asks you a multiple-choice question instead of guessing.
- Discover. It finds the right sources in your project and profiles them: types, nulls, distinct counts and sample rows. Large files are queried where they live instead of being loaded into memory.
- Design. Rhombo designs each step of the pipeline, drawing on a library of 200 transformer types, such as joins, filters, aggregations, pivots, deduplication, outlier handling, missing-value imputation, fuzzy merges, date parsing and type conversion. It reads each one's exact spec before using it, so it works from the spec rather than from memory.
- Compile. The agent turns the design into a small, typed plan. A compiler checks it before anything is saved.
- Run. The pipeline runs, whether the input is a small file or a very large dataset.
- Verify. Rhombo previews the output, checks it against your request and explains in plain language what changed. You watch its reasoning and progress live in the chat.
For a request like "join claims to members and flag duplicates", the plan has two steps. Each one names a transformer, where its data comes from and its settings.
left join on member_id
unique on claim_id, keep first
inputs connect, graph valid
The three healing loops
Each loop sends its error straight back to the agent with full context, and each one stops at a set limit.
Loop 1: the compiler catches design mistakes. This loop exists to stop hallucination. Before anything is saved, a deterministic compiler verifies the whole plan against what actually exists in your project — the model cannot make up a step, a column or a source that isn't there. When part of the plan doesn't hold up, the agent gets back an error with enough context, hints and semantic meaning to understand exactly what went wrong, and revises the plan. Because the compiler is deterministic, the same mistake is caught the same way every time. It never depends on the model noticing its own error.
Loop 2: the sandbox catches bad code. When logic doesn't fit a standard step, Rhombus writes Python and runs it through three layers of checks:
- Before it runs: the generated code is checked for unsafe operations and limited to the intended data transformation — nothing else can run.
- While it runs: it executes in a sandbox on a limited sample of the data.
- After it runs: the output must be a well-formed table of the expected shape, with no input columns quietly dropped.
When a check fails, the error and the history of earlier failed attempts go into the next try. The model learns from the whole trail of mistakes, not just the latest one. There is a hard limit on attempts; if the code still isn't right, the loop stops and reports why.
Loop 3: automatic pipeline repair catches failures that only happen at scale. When a job fails on the full dataset, Rhombus stores the failing script and the error, then:
- Tries known fixes first. A library of rule-based repairs handles the most common failure patterns, like a key column dropped before it is used, or code that behaves differently at scale.
- If none apply, asks the model for a targeted patch and checks that the patched script compiles before running it.
- Resubmits the job, and if the failure still can't be fixed, hands you a clear diagnosis.
Known fixes run first because they are instant and certain. The model handles the long tail.
Watch it live
Self-Healing ETL Pipelines · watch on YouTube
Putting it together
One request, end to end, in 11 steps. The compiler catches a design mistake at step 4, and a known rule repairs a failure at scale at step 8. The user only sees the finished output and a plain-language summary.
The result is a pipeline that was designed by reasoning, validated by a compiler, tested in a sandbox and repaired at scale. You can read, edit and schedule it like any other pipeline on the canvas.
Try it on your own data
If you have a pipeline that breaks every few weeks for the same boring reasons, bring it to Rhombus AI and describe what it should do in plain English. Start at rhombusai.com.

