Research

AI-generated text

Criticism of Bridgemind AI benchmarks' lack of methodology

A user publicly criticized Bridgemind AI's published benchmark post, saying it does not detail how the tests were carried out (e.g., harness, tasks, number of runs, identification of day-to-day variance, APIs, traces).

Criticism of Bridgemind AI benchmarks' lack of methodology

A user publicly criticized Bridgemind AI's published benchmark post, saying it does not detail how the tests were carried out (e.g., harness, tasks, number of runs, identification of day-to-day variance, APIs, traces). They specifically highlighted the choice of +/-10% variance and the equal weighting of tokens with costs, which could undermine reproducibility and fair comparison.