YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
Loading...
Date
Editor
Advisor
Volume
Issue
Journal
Series Titel
Book Title
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Publisher
Kerrville, TX : Association for Computational Linguistics (ACL)
Supplementary Material
Other Versions
Link to publishers' Version
Abstract
Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. We release multidisciplinary scienceQ&A datasets, including adversarial variants, with evaluation scores from multiple LLMs. Independent of proprietary models and human feedback, our approach enables scalable, cost-free evaluation. By advancing reliable LLM-as-a-judge models, this work supports AI alignment and fosters robust, transparent evaluation essential for scientific inquiry.
Description
Keywords
Keywords GND
Conference
63rd Annual Meeting of the Association for Computational Linguistics, 2025.07.27-08.01, Vienna, Austria
Publication Type
BookPart
Version
publishedVersion
Collections
License
CC BY 4.0 International
