THE PROBLEM
An on-call backend engineer gets paged: a payment provider's webhook handler is returning 200 but roughly 8% of events never reach the order-update code path. The stack trace is clean - there is no exception, which is worse than an exception, because nothing is obviously broken.
Pasting the handler function into ChatGPT and asking "why would this silently drop events" gets a plausible-sounding answer built from the model's general knowledge of webhook handlers in whatever framework it recognizes - not from this specific codebase's actual middleware stack, retry logic, or the three other services that touch this queue before the handler runs. A generically-plausible bug is not the actual bug, and debugging with a wrong hypothesis costs more time than debugging with no hypothesis.
The fix is to retrieve the actual, current, relevant code and logs into the model's context before asking it to reason - a retrieval-augmented generation (RAG) pattern, run manually, for a debugging session rather than a chatbot.