
On the headline coding benchmark, SWE-bench Verified, Claude Opus 4.8 and GPT-5.5 are effectively tied at the top, with Gemini 3.1 Pro a clear step behind on that specific test but ahead on raw context window and often on price.
No discussion yet. Be the first to share your thoughts!