A user publicly criticized Bridgemind AI's published benchmark post, saying it does not detail how the tests were carried out (e.g., harness, tasks, number of runs, identification of day-to-day variance, APIs, traces). They specifically highlighted the choice of +/-10% variance and the equal weighting of tokens with costs, which could undermine reproducibility and fair comparison.
AI-generated text
Criticism of Bridgemind AI benchmarks' lack of methodology
A user publicly criticized Bridgemind AI's published benchmark post, saying it does not detail how the tests were carried out (e.g., harness, tasks, number of runs, identification of day-to-day variance, APIs, traces).



