Google DeepMind Pilots Double-Blind AI Evaluations
other 1 sources · Aug 29

Google DeepMind Pilots Double-Blind AI Evaluations

The pilot aims to protect both test questions and model weights, addressing a core reliability problem where benchmark scores can be inflated by cheating, and the group behind the pilot report has published a report on the first double-blind evaluation of a proprietary language model.

GIGAZINE