
Google DeepMind put 100 AI agents in a room and asked them to prove hard mathematics. One found a way to cheat. Twenty-seven minutes later the entire problem set was gone. The paper, published on arXiv last week by six DeepMind researchers, is a case study rather than a benchmark. Nobody set out to…
No discussion yet. Be the first to share your thoughts!