Skip to content
NakodaAI

AI FIELD MANUALDEVELOPERS, OPERATIONS

It's 2am, the payment webhook is silently dropping events, and the stack trace alone isn't telling you why

A concrete retrieval-augmented workflow for using an LLM to debug a production incident against your actual codebase and logs - not against its training data's vague idea of what your framework does.

Last reviewed 1 September 2026

THE PROBLEM

An on-call backend engineer gets paged: a payment provider's webhook handler is returning 200 but roughly 8% of events never reach the order-update code path. The stack trace is clean - there is no exception, which is worse than an exception, because nothing is obviously broken.

Pasting the handler function into ChatGPT and asking "why would this silently drop events" gets a plausible-sounding answer built from the model's general knowledge of webhook handlers in whatever framework it recognizes - not from this specific codebase's actual middleware stack, retry logic, or the three other services that touch this queue before the handler runs. A generically-plausible bug is not the actual bug, and debugging with a wrong hypothesis costs more time than debugging with no hypothesis.

The fix is to retrieve the actual, current, relevant code and logs into the model's context before asking it to reason - a retrieval-augmented generation (RAG) pattern, run manually, for a debugging session rather than a chatbot.

THE APPROACH

Manually construct the retrieval step instead of relying on the model's memory of "how webhooks usually work": pull the exact handler code, the middleware it passes through, the queue/retry config, and a sample of the actual failing log lines, and put all of it in context together before asking for a hypothesis.

This is the same retrieve-then-generate pattern production RAG systems automate with an embedding index - here it's done by hand for one incident, because setting up a vector index is not the right tool for a single 2am debugging session, but the underlying discipline (ground the model in retrieved, current, specific material before it reasons) is identical.

WHY IT WORKS

An LLM's training data has a knowledge cutoff and no idea what your team's code looked like an hour ago. Retrieval closes that gap by putting the actual current state of the system in the prompt, so the model reasons over what's really there instead of a statistically likely guess at what's usually there.

Narrowing the retrieved set matters as much as including it: dumping the entire repository into context buries the relevant fifty lines in thousands of irrelevant ones and measurably degrades the model's ability to find the real signal (a well-documented "lost in the middle" effect in long-context reasoning). The discipline here is retrieving *precisely* the handler, its direct dependencies, and matching log lines - not everything that might be related.

STEP BY STEP

  1. 01

    Failure signal

    Real failing requests pulled from logs/observability - not a hypothetical.

  2. 02

    Code retrieval

    Exact handler, middleware chain and retry config - the real call path, not a remembered one.

  3. 03

    Contrast sample

    Matched failing vs. succeeding log entries, retrieved side by side.

  4. 04

    Grounded reasoning

    Model reasons only over what was retrieved, ranked by fit to the actual failure rate.

  5. 05

    Verify, then fix

    Hypothesis checked against the evidence before any production change.

  1. Pull the failure signal first

    Query the logging/observability tool for the actual failing requests - timestamps, payload shape, any partial trace - before opening an editor. You need real failing examples, not a hypothetical one.

  2. Retrieve the exact code path

    Find and copy the webhook handler function, every middleware it passes through in order, and the queue or retry configuration it depends on. Use your IDE's "find references" or a code-search tool (Cursor, GitHub Copilot Chat with codebase context) to be sure you have the real call chain, not the one you remember.

  3. Retrieve a representative failing sample and a representative succeeding sample

    Pull three or four failing log entries and three or four succeeding ones with the same shape. The contrast between the two sets is often where the actual signal is - a field present in one set and silently absent in the other.

  4. Assemble one prompt with all of it in context

    Give the model the handler code, the middleware chain, the retry config, and both log samples together, and ask it for the three most likely causes ranked by how well each explains why *some* events succeed and others don't - not just "what could be wrong".

  5. Test the top hypothesis against the evidence, not in production

    Before touching production code, check whether the top hypothesis actually explains the specific failing/succeeding contrast you retrieved. A hypothesis that would explain 100% of events failing is wrong on its face if the real failure rate is 8%.

  6. Fix, then write the incident note yourself

    Once verified, write the postmortem note in your own words. The model helped generate the hypothesis; the record of what actually happened and why is a human accountability document, not a delegated one.

TOOLS

LIMITATIONS

  • This is a manual, one-incident version of RAG. If your team debugs against the same large codebase weekly, a real indexed retrieval setup (embeddings over the repo, refreshed on commit) pays for itself; doing this by hand every time does not scale past occasional use.

  • A silent, intermittent bug (8% failure, no exception) is exactly the case where the retrieved context can still be incomplete - a race condition or a dependency on external service timing may not show up in any log line you thought to retrieve. If the top-ranked hypothesis doesn't survive contact with a reproduction attempt, retrieve more context rather than trusting the model's confidence.

  • The model can be confidently wrong even with good context - it will rank a plausible-but-incorrect cause highly if the plausible cause pattern-matches something common in its training data. Verification against your specific evidence, not the model's stated confidence, is what catches this.

EXAMPLE

A Node.js payment webhook handler is silently dropping ~8% of events with no exception thrown.

  1. Sentry query surfaces 40 failing requests over 6 hours out of roughly 500 total - no exception attached to any of them.

  2. Retrieved: the handler function, an async rate-limiting middleware it passes through, and the queue retry config.

  3. Retrieved: 4 failing log entries and 4 succeeding ones with matching payload shape.

  4. Contrast: every failing entry's payload was marginally larger, close to a request-body size limit set in the rate-limiting middleware that silently truncates rather than rejects.

  5. Claude ranked "payload near the middleware's body-size limit, truncated not rejected" as the top hypothesis, given the contrast.

  6. Verified against three more real failing payloads - all were within 200 bytes of the limit. Fix: reject-and-retry instead of silent truncate.

Root cause found in one session by grounding the model in the actual middleware config and a real failing/succeeding contrast, instead of a generic "check your webhook handler" answer that would have missed the middleware entirely.

RELATED

Nakoda editorial · last reviewed

This entry describes a workflow Nakoda recommends - it is not a claim about how any named tool behaves in every case, and it is not paid placement. Spotted something out of date? Tell us.

MORE FOR THIS AUDIENCE