Looking at the current leaderboard / paper, it seems like a number of open source / open weight models were evaluated (eg. deepseek-chat-20250301, Codestral-22B-v0.1, DeepSeek-Coder-V2, Llama-3.1, Mixtral-v3). It would be interesting to see how some other / newer models (eg. gpt-oss) rank in this benchmark too.
See Also
Looking at the current leaderboard / paper, it seems like a number of open source / open weight models were evaluated (eg.
deepseek-chat-20250301,Codestral-22B-v0.1,DeepSeek-Coder-V2,Llama-3.1,Mixtral-v3). It would be interesting to see how some other / newer models (eg. gpt-oss) rank in this benchmark too.See Also