| name | reference-free-evaluation-of-reasoning-in-open-end |
| description | Skill generated from arXiv paper 2607.19678: Reference-Free Evaluation of Reasoning in Open-Ended Question Answering |
| metadata | {"arxiv":{"id":"2607.19678","title":"Reference-Free Evaluation of Reasoning in Open-Ended Question Answering","authors":["Guneet Singh Kohli","Yuxiang Zhou","Michael Sejr Schlichtkrull","Gregory E Dean","Maria Liakata"],"published":"2026-07-22","categories":["cs.CL","cs.AI","cs.LG"],"url":"https://arxiv.org/abs/2607.19678","utility":1}} |
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
arXiv: 2607.19678
Published: 2026-07-22
Authors: Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata
Categories: cs.CL, cs.AI, cs.LG
Utility: 1.00
Key Innovation
AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then as...
Potential Application
This paper presents advancements that could be applied to enhance agent capabilities in the areas of cs.CL, cs.AI, cs.LG.
References