OpenAI points out that model evaluations measure not only the models themselves but also API settings, the test environment, and prompt decisions. Citing its experiments, the company advises API developers to use the Responses API instead of the older Chat Completions API to maximize performance, and to enable reasoning retention and compaction; it also offers public games for testing.
OpenAI: evaluations measure multiple factors, recommends using Responses API and settings
OpenAI points out that model evaluations measure not only the models themselves but also API settings, the test environment, and prompt decisions.



