What kind of challenges are you giving arena.ai? When I tried Grok (latest at the time available in Cursor) vs Opus 4.5, Grok was much faster to respond, and hillariously terrible at wrting functional code - but confident in its responses.
I tried your link, got a Grok 4.2 v Grok 4.3 pairing, they did O.K. 4.2 better than 4.3 at making a webserver that queries and re-serves NEXRAD data, but… they didn’t get too far down the feature list (history of rainfall graphs) before they fell apart.
In my experience (I use arena.ai battle mode as a main chatbot), higher versions of grok on a good day are at the level of Claude Opus 4.X
Huh, interesting. I wonder why it’s so infrequently used then. Maybe people are afraid of using an AI that referred to itself as “mechahitler”.
Maybe that.
What kind of challenges are you giving arena.ai? When I tried Grok (latest at the time available in Cursor) vs Opus 4.5, Grok was much faster to respond, and hillariously terrible at wrting functional code - but confident in its responses.
Usual everyday tasks. For coding it is: Claude>Chatgpt>Chinese open source>Gemini>Grok>Everything else.
I tried your link, got a Grok 4.2 v Grok 4.3 pairing, they did O.K. 4.2 better than 4.3 at making a webserver that queries and re-serves NEXRAD data, but… they didn’t get too far down the feature list (history of rainfall graphs) before they fell apart.