TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
Journal
Proceedings of the Annual Meeting of the Association for Computational Linguistics
Journal Volume
1
Start Page
2934
End Page
2952
ISSN
0736587X
ISBN (of the container)
979-889176251-0
ISBN
[9798891762510]
Date Issued
2025
Author(s)
Abstract
The LLM-as-a-judge paradigm uses large language models (LLMs) for automated text evaluation, which assigns a score to the text based on some scoring rubrics. Existing methods for LLM-as-a-judge use cross-entropy (CE) loss for fine-tuning, which neglects the numeric nature of score prediction. Recent work addresses numerical prediction limitations of LLM fine-tuning through regression-aware fine-tuning, which, however, does not consider chain-of-thought (CoT) reasoning for score prediction. In this paper, we introduce TRACT (Two-stage Regression-Aware fine-tuning with CoT), a method combining CoT reasoning with regression-aware training. The training objective of TRACT combines the CE loss for learning the CoT reasoning and the regression-aware loss for the score prediction. TRACT consists of two stages: first, a seed LLM is fine-tuned to generate CoTs; next, we retrain the seed LLM using the CoTs generated by the LLM trained in stage 1. Experiments across four LLM-as-a-judge datasets and two LLMs show that TRACT significantly outperforms existing methods. Extensive ablation studies validate the importance of each component in TRACT.
Event(s)
63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
Publisher
Association for Computational Linguistics (ACL)
Type
conference paper
