AI RESEARCH

On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration

arXiv CS.AI

ArXi:2603.22061v1 Announce Type: cross Inasmuch as the removal of refusal behavior from instruction-tuned language models by directional abliteration requires the extraction of refusal-mediating directions from the residual stream activation space, and inasmuch as the construction of the contrast baseline against which harmful prompt activations are compared has been treated in the existing literature as an implementation detail rather than a methodological concern, the present work investigates whether a topically matched contrast baseline yields superior refusal directions.