A document pipeline takes PDFs, drawings, and reports, pulls out text and metadata, and makes them searchable. When it crashes, someone notices. When it quietly skips a page, drops a field, or gives up on a file without saying so, nobody notices until a person asks a question and gets a confident, incomplete answer.
Most of my reliability work has been about that second kind of failure. These are the patterns I keep coming back to.
Give failures a small, honest vocabulary
“Something went wrong” is not a state. Every failed document should land in one of a few named states, and the most important split is simple: can this fix itself, or does a person need to act?
A rate limit from an AI provider is retryable. A missing API key is not, and retrying it a hundred times just hides the real problem behind noise. If the code that records a failure cannot tell those apart, the system will eventually retry the unfixable and give up on the fixable.
ingested everything extracted
provisional usable, but something was missing
retrying temporary problem, will try again
failed a person needs to act (with a reason)Show the reason where people look
A failure that only exists in a server log is invisible to the people who depend on the data. The interface should say which documents are incomplete and why, in plain words. A warning badge with an empty tooltip is worse than no badge: it tells people something is wrong and then refuses to explain.
One subtle trap: if health reasons are only ever appended, an old problem can keep showing after it has been fixed. Each check should own its reasons and replace them on every run, so the badge always reflects the current state.
Retries have to survive a restart
Retry logic usually works fine until the server restarts. Then work that was waiting for a retry is simply forgotten, or it comes back with its attempt counter reset to one, so the retry limit never kicks in.
Persist the retry state with the document: how many attempts, why it is waiting, and when to try next. On startup, look for anything that was mid-retry and pick it back up. Then add a check for documents stuck in limbo, neither finished nor failed, because that is exactly where silent losses hide.
If a document can enter a state, there should be a way to find every document in that state. “Stuck” is a state too.
Watch the boring code paths
Some of the worst bugs I have found were not in the clever parts. They were in plumbing: an update function with an allow-list of fields that silently dropped a new one, or a log message that said a step succeeded and failed at the same time. Plumbing like that deserves small, focused tests, because nobody reads it closely once it works.
Characterize before you fix
When a number looks wrong, my first instinct used to be to change the code. Now I write a test that captures what the code does today, wrong answer included. That test proves the bug exists, shows how big it is, and tells me exactly when my fix changes behavior and when it does not.
What I would build first
- Name the failure states and split retryable from permanent.
- Store a human-readable reason with every failure.
- Surface incomplete documents in the interface, not just the logs.
- Persist retry state and resume it on startup.
- Add a query for anything stuck between states.
The result
A pipeline that fails loudly is less impressive in a demo, because it admits what it could not do. It is far more useful in production, because the people relying on it know which answers they can trust.
This article describes general engineering principles from my experience and independent study. It does not disclose proprietary architecture, source code, customer information, or confidential details from kW Engineering.