I noticed some AI tells, but found overall the article not too bad. It did seem to waffle at times though.
> Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.
The "does not rescue it". No human would write like that.
> I'll explain what that means on a table you already know.
No I don't already know that table.
Also
> and claims 40x on graph reachability
Is really hard to parse.
The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...
> The "does not rescue it". No human would write like that.
This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?
They’re not just trained on human prose. They’re sent to RLHF, and also their language changes as a result of RL on verifiable rewards.
Getting it to write well is really hard because there’s no real way to verify whether it’s good prose or not. You and I can tell, but we can’t write a verifier that codifies our judgment.
Maybe they’ll find a way to improve this, but for now it’s certainly one of the harder problems to solve for LLMs.
Part of it is that I think they also have poor theory of mind, which I imagine is also a hard thing to train it to do.
Why is it hard? Ask it to write professionally in mid-twentieth century style English, and without resorting to the clickbait style of writing.
In any event, other LLMs may not automatically have the problem, and don't even require such a prompt. This is a Claude problem.
I don’t think you realize how long ago the middle of the 20th century was.
I don't think you realize the prompt actually works. The word "style" does it. I guess you like clickbait too much.
For autoregressive models (practically all hosted ones), it's because of the nature of next-token prediction. LLMs lock themselves into a particular sentence structure ahead of time and have to guess at the rest of the sentence. Samplers have no insight into the LLM's "intent" aside from the probability of each next token, and the LLM has no insight into its previous "intent" that resulted in a given probability in the first place. I don't think this is possible to solve with more training, I think a fundamental architectural shift will be needed, like more research into diffusion language models.
[dead]
> The "does not rescue it". No human would write like that.
On the contrary. It is unnatural for a native speaker, which I don't think the author is. For someone that speaks English as a second language, it is not uncommon to use expressions literally translated from their first language, which may be understandable but weird for native speakers.
Hilarious that even DuckDB's article has many of LLM tells as well