Early results: GPT-5.3-Codex high leads (56/44 vs xhigh); Opus 4.6 trails
After Thursday's model drop, we added gpt-5-3-codex (medium, high, xhigh) and claude-opus-4-6 to our agent roster and started running them on real tasks across our production codebases, internal tools, and prototypes.