Like I said in the edit, when people want specific formatting they ask for well known formats: Markdown, XML, JSON
I don't even need to debate if the benchmark is useful, it doesn't pass a sniff test: GPT-5.4 is not worse than Gemini 2.5 Flash in any way that matters to most users. In your benchmark it's meaningfully worse.