Single Pass vs. Graph Retry
AI Experiment
A head-to-head comparison of a single-pass RAG pipeline against a bounded retry graph on the same AI Engineering Wiki benchmark, using the same retrieval, prompts, and judge model for both.
Question
Does adding a bounded grade-and-retry graph around generation measurably improve answer quality over a single-pass pipeline, and at what latency cost?
Approach
Ran the same AI Engineering Wiki Benchmark through two configurations: a single-pass pipeline (retrieve, generate, done) and a LangGraph-based bounded retry graph (retrieve, generate, grade the answer, retry generation up to a limit if the grade is weak). Both used the same retrieval settings, prompt, and LLM-judge scoring so the only variable was the presence of the retry loop.
Results — AI Engineering Wiki Benchmark
| Metric | Baseline | Variant |
|---|---|---|
| Hit@K | 100% | 100% |
| Generation quality | 0.93 | 0.97 |
| Grounding | 0.95 | 0.98 |
| Outcome | 0.97 | 0.98 |
| Avg latency | 49.4s | 51.5s |
| Avg retry iterations | n/a | 0.0 |
Key Finding
The retry graph scored modestly higher on generation quality, grounding, and outcome than the single-pass pipeline, while matching it on Hit@K, and it did so with retries rarely triggering at all (0.0 average iterations), for only about 2 seconds of added latency.
What We Learned
On this benchmark, a bounded retry graph is a low-risk upgrade over a single-pass pipeline: it never made things worse, delivered a small but real quality improvement, and the retry loop itself was mostly insurance, it rarely had to fire to produce the better result. The graph's value here came from the grading step forcing a better first-pass generation, not from actual retries.