Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap

ArXi:2508.04149v2 Announce Type: replace-cross Aligning large language models (LLMs) with human preferences is a critical challenge in AI research. While methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) are widely used, they often rely on large, costly preference datasets. The current work lacks methods for high-quality data selection specifically for preference data. In this work, we