CR

Colin Raffel

Machine-learning researcher

Many Builders reference roster

Open machine-learning research and evaluation of language-model robustness under adversarial pressure.

Many Builders reference → Safety & alignment coverage →

Writing and research by Colin Raffel

1 supplied byline match

These articles, papers and essays carry Colin Raffel in the source-supplied author field. Verify the definitive byline and text at the original publisher.

arXiv

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models — open the original publisher

By Malikeh Ehghaghi, Boglárka Ecsedi, Marsha Chechik, Colin Raffel

Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequently, ASR at a fixed budget can obscure the true effort required to jailbreak a model, thereby making it hard to determine whether an attack's cost justifies its payoff to the attacker. We propose a co

Research Safety & alignment

News and discussion

name mentions, not authorship

These source records mention Colin Raffel in the title or summary. They are kept separate because a mention does not mean the item was written by them.

OpenAlex

Crosslingual Generalization through Multitask Finetuning — open the original publisher

Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, Colin Raffel. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.

Research