Tools

OpenAI: evaluations measure multiple factors, recommends using Responses API and settings

OpenAI points out that model evaluations measure not only the models themselves but also API settings, the test environment, and prompt decisions.

OpenAI: evaluations measure multiple factors, recommends using Responses API and settings

OpenAI points out that model evaluations measure not only the models themselves but also API settings, the test environment, and prompt decisions. Citing its experiments, the company advises API developers to use the Responses API instead of the older Chat Completions API to maximize performance, and to enable reasoning retention and compaction; it also offers public games for testing.