N
pending
LettuceDetect v2: Span-Level Hallucination Detection for Agentic AI
Grounded / Real
Inflated / Uruttu
Original Content
𝗧𝗼𝗱𝗮𝘆 𝘄𝗲 𝗮𝗿𝗲 𝗿𝗲𝗹𝗲𝗮𝘀𝗶𝗻𝗴 𝗟𝗲𝘁𝘁𝘂𝗰𝗲𝗗𝗲𝘁𝗲𝗰𝘁 𝘃𝟮: 𝗦𝗢𝗧𝗔 𝗵𝗮𝗹𝗹𝘂𝗰𝗶𝗻𝗮𝘁𝗶𝗼𝗻 𝗱𝗲𝘁𝗲𝗰𝘁𝗶𝗼𝗻 𝗳𝗼𝗿 𝗰𝗼𝗱𝗲, 𝘁𝗼𝗼𝗹 𝗼𝘂𝘁𝗽𝘂𝘁, 𝗮𝗻𝗱 𝗮𝗴𝗲𝗻𝘁𝗶𝗰 𝘄𝗼𝗿𝗸𝗳𝗹𝗼𝘄𝘀 — 𝗶𝗻 𝗰𝗼𝗹𝗹𝗮𝗯𝗼𝗿𝗮𝘁𝗶𝗼𝗻 𝘄𝗶𝘁𝗵 𝘁𝗵𝗲 [**vLLM**](https://www.linkedin.com/company/vllm-project?trk=public_post-text) 𝗦𝗲𝗺𝗮𝗻𝘁𝗶𝗰 𝗥𝗼𝘂𝘁𝗲𝗿 𝘁𝗲𝗮𝗺. In agentic workflows, hallucinations are rarely just text. Agents ground their answers in 𝘀𝗼𝘂𝗿𝗰𝗲 𝗰𝗼𝗱𝗲, 𝘁𝗼𝗼𝗹 𝗼𝘂𝘁𝗽𝘂𝘁, 𝗺𝗮𝗿𝗸𝗱𝗼𝘄𝗻, 𝘁𝗮𝗯𝗹𝗲𝘀. Existing hallucination datasets and detectors were built for document QA and 𝗻𝗼𝗻𝗲 𝘄𝗲𝗿𝗲 𝗯𝘂𝗶𝗹𝘁 𝗳𝗼𝗿 𝘁𝗵𝗶𝘀: on code-agent answers prior detectors reach 𝟬.𝟭𝟳 𝘀𝗽𝗮𝗻-𝗙𝟭, and even 550B-class LLM judges reach at most 𝟬.𝟮𝟮. So we built the 𝗳𝗶𝗿𝘀𝘁 𝘀𝗽𝗮𝗻-𝗹𝗲𝘃𝗲𝗹 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸 𝗮𝗻𝗱 𝗺𝗼𝗱𝗲𝗹𝘀 for hallucination detection over code, tool output, and structured documents, with multilingual coverage built in. The results: 🥇 𝗦𝗢𝗧𝗔 𝗼𝗻 𝗼𝘂𝗿 𝗻𝗲𝘄 𝗰𝗼𝗱𝗲/𝗮𝗴𝗲𝗻𝘁𝗶𝗰 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸: 𝟬.𝟲𝟬 span-F1 on code-agent answers, 𝟬.𝟲𝟴𝟵 on the full test set — \~3× the best prior detector 🌍 𝗦𝗢𝗧𝗔 𝗼𝗻 𝗣𝘀𝗶𝗹𝗼𝗤𝗔: best reported English IoU (𝟬.𝟳𝟮𝟰), with 14-language coverage 💪 𝘀𝘁𝗿𝗼𝗻𝗴 𝗼𝗻 𝗥𝗔𝗚𝗧𝗿𝘂𝘁𝗵: 𝟴𝟭.𝟴 example-F1 from a 2B model, so specializing on code gives strong general RAG performance The detectors take the request, the context, and the answer, and return the 𝗲𝘅𝗮𝗰𝘁 𝗵𝗮𝗹𝗹𝘂𝗰𝗶𝗻𝗮𝘁𝗲𝗱 𝘀𝗽𝗮𝗻𝘀, located to the character, 𝘁𝘆𝗽𝗲𝗱 (contradiction / fabricated reference / unsupported addition × 13 subcategories), each with an explanation. We are releasing: 🧪 a 𝘂𝗻𝗶𝗳𝗶𝗲𝗱 𝘀𝗽𝗮𝗻-𝗹𝗲𝘃𝗲𝗹 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸: 𝟳𝟰,𝟮𝟴𝟱 newly constructed examples across SWE-bench coding-agent traces, developer tool output, and structured documents (papers, READMEs, Wikipedia) — 𝟭𝟰𝟱𝗞+ training examples with RAGTruth and 14-language PsiloQA folded in, all with exact character labels 🤖 𝗹𝗲𝘁𝘁𝘂𝗰𝗲𝗱𝗲𝗰𝘁-𝘃𝟮-𝗾𝘄𝗲𝗻-𝟮𝗯 — SOTA generative detector, typed spans + explanations in a single pass, 𝟯𝟮𝗞-𝘁𝗼𝗸𝗲𝗻 context ⚡ 𝗹𝗲𝘁𝘁𝘂𝗰𝗲𝗱𝗲𝗰𝘁-𝘃𝟮-𝗺𝗺𝗯𝗲𝗿𝘁-𝗯𝗮𝘀𝗲 — 𝟯𝟬𝟳𝗠 multilingual encoder for high-throughput pipelines 🏷️ a 𝘁𝗮𝘅𝗼𝗻𝗼𝗺𝘆 𝘁𝘆𝗽𝗶𝗻𝗴 𝗵𝗲𝗮𝗱 (𝟬.𝟴𝟮 category accuracy on gold spans) that types any binary span detector 💻 the full 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲 construction, training, and evaluation pipeline This comes from our new paper and collaboration with [**Bowei He**](https://hk.linkedin.com/in/bowei-he-8a9450199?trk=public_post-text), [**Xunzhuo Liu**](https://sg.linkedin.com/in/bitliu?trk=public_post-text) and [**Huamin Chen**](https://www.linkedin.com/in/huaminchen?trk=public_post-text) from the [**vLLM**](https://www.linkedin.com/company/vllm-project?trk=public_post-text) 𝗦𝗲𝗺𝗮𝗻𝘁𝗶𝗰 𝗥𝗼𝘂𝘁𝗲𝗿 team. The models are integrated into [**vLLM**](https://www.linkedin.com/company/vllm-project?trk=public_post-text) SR with span-level verification inside the serving stack. More on that in the next days. 🚀 Everything is open source under permissive licenses. Paper / HF links in the comments. If you like it, a star means a lot to us ⭐
Validated Content
Confirmed Accurate
- LettuceDetect v2 is a real open-source hallucination detection project focused on RAG, code, tool output, and agentic workflows.
-
The project released two v2 models:
-
lettucedect-v2-qwen-2b(generative detector) -
lettucedect-v2-mmbert-base(encoder detector)
-
- The benchmark includes code-agent traces, tool output, and structured documents, and contains roughly 74,285 newly constructed examples.
- The reported 0.602 span-F1 on code-agent answers and 0.689 span-F1 on the unified test set match the published results.
- The detector identifies exact hallucinated spans and can classify them into categories and subcategories.
- The project is open source and releases models, datasets, and training pipelines.
- The work was conducted in collaboration with members of the vLLM Semantic Router team.
Mostly Accurate
-
"SOTA hallucination detection"
The published results show state-of-the-art performance on the authors' benchmark and strong results on PsiloQA and RAGTruth. However, "SOTA" depends on the benchmark, metric, and evaluation setup used. It should not be interpreted as best across every hallucination detection benchmark.
-
"3× the best prior detector"
The published benchmark shows substantial improvement over previously evaluated detectors and LLM judges on code-agent tasks. The 3× framing is broadly supported by the reported numbers.
Promotional / Research-Marketing Language
- "SOTA hallucination detection"
- "None were built for this"
- "Strong general RAG performance"
- "Everything is open source"
- Star request and launch-announcement framing
These are common research-launch statements and not independent facts.
Missing Context
- Most performance claims come from benchmarks created and released by the same research team.
- Independent third-party replication appears limited so far.
- Hallucination detection remains highly benchmark-dependent.
- Strong benchmark performance does not automatically guarantee equivalent production performance in every agentic workflow.
- Span-level detection is substantially harder than answer-level detection, making comparisons with other systems difficult unless the evaluation protocol is identical.