SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy

ArXi:2604.02423v1 Announce Type: new Large language models exhibit sycophancy: the tendency to shift outputs toward user-expressed stances, regardless of correctness or consistency. While prior work has studied this issue and its impacts, rigorous computational linguistic metrics are needed to identify when models are being sycophantic. Here, we