Confidence-Aware Alignment Makes Reasoning LLMs More Reliable

ArXi:2605.07353v1 Announce Type: new Large reasoning models often reach correct answers through flawed intermediate steps, creating a gap between final accuracy and reasoning reliability. Existing alignment strategies address this with external verifiers or massive sampling, limiting scalability. In this work, we