A Theoretical Game of Attacks via Compositional Skills

ArXi:2605.01034v1 Announce Type: new As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we