>If the computer can't do it better than a human being, then what's the point?
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!"
This is what the tech industry has become?
Less of a failure is still failure.
That is the software industry. Our product is less bugged than our competitor’s.