because human engineers are not reproducible, so having an accurate measure of the skills of one of them has a much lower utility return (those benchmarks are very expensive to run, and making them for humans will be at least as much expensive)
because human engineers are not reproducible, so having an accurate measure of the skills of one of them has a much lower utility return (those benchmarks are very expensive to run, and making them for humans will be at least as much expensive)
Are you saying LLMs are reproducible? I seem to get at least a slightly different answer every time I ask the same question
two istances of the same model, with low temperature, are infinitely more similar than any couple of human minds