medium.com5 months agoI Spent 48 Hours Testing Claude Opus 4.6 & GPT-5.3 CodexSpoiler: The benchmarks lied to meVisit medium.com10BookmarkAdd to collection