
On raw benchmark scores, Claude Opus 4.8 wins most of the table, especially long-horizon software engineering like NL2Repo, SWE-Marathon, and Tool-Decathlon. GLM-5.2 wins a smaller set (AIME 2026, IMOAnswerBench, and Terminal-Bench 2.1 under its best harness) and trails by under a point on…
No discussion yet. Be the first to share your thoughts!