I wonder if this is partly because “pelican riding a bicycle” has become a kind of benchmark prompt by now. If so, could the models actually be getting better at the benchmark rather than getting better at following the prompt?
I wonder if this is partly because “pelican riding a bicycle” has become a kind of benchmark prompt by now. If so, could the models actually be getting better at the benchmark rather than getting better at following the prompt?