YB
Yoshua Bengio
Professor & Founder, LawZero
University of Montreal / Mila / LawZero
Turing Award laureate who now leads international AI-risk assessment and founded a nonprofit for 'safe-by-design' AI.
Latest in the feed
Language models recognize dropout and Gaussian noise applied to their activations
We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely, we either (a) mask activations, simulating dropout, or (b) add Gaussian noise to them, at a target sentence. We then ask a multiple-choice question such as "Which of the previous sentences was perturbed?" or "Which of the two perturbations was applied?". We test models from the Llama, Olmo, and Qwen families, with siz
Managing extreme AI risks amid rapid progress
Preparation requires technical research and development, as well as adaptive, proactive governance.