N
pending
RuleChef: Let the LLM Write Rules, Not Answer Every Query
Grounded / Real
Inflated / Uruttu
Original Content
Validated Content
The specific benchmark numbers (78.7 vs 74.8 F1, the feedback-loop improvement figures) are self-reported claims from the same release and haven't been independently verified yet — presented with full confidence as established results rather than "our own measured numbers, pending outside scrutiny."