A LLM benchmark that gave a hard programming tests to gpt 5.6, but for much more languages than common benchmarks
He's right that python/js tend to score better on most tasks. OpenAI is very bad at less popular languages. Especially J, so that is a very bad skew in his results.
No replies yet