> this is NOT something most people want to do

I've always felt that "proper" and responsible LLM use for search [0] would be two boxes: You can describe what you want in the first, and it'll propose search terms in the second, and then those get executed normally.

Yes, the "average user" [1] might not usually care about the second box... Until they need to because the query/results are wrong.

Showing them in tandem means:

1. Users are at least capable of learning through exposure.

2. Users may realize a key term can be added which the model could never have guessed.

3. Users may recognize a term in there that doesn't make sense, allowing them to detect a translation error.

4. If good search-terms leads to a bad outcome, it is possible for someone to report and diagnose it, rather than a fully black-box mystery.

_____

[0] Not just for websites, but also things like internal business software, or SQL queries.

[1] The average that might not exist. ( https://www.thestar.com/news/insight/when-u-s-air-force-disc... ) There are some features that everybody needs, just at different times.

In the words of Edsger Dijkstra, "Projects promoting programming in 'natural language' are intrinsically doomed to fail." I think for something as bespoke as a search engine, a DSL that accepts quotes, logical operators, and specific terms definitely allows for far greater specificity than natural language can (at least in an equivalent amount of text).

I think the two-tiered input-output approach you propose makes sense. Allowing users to inspect and mutate lower-level languages allows the user to make modifications as needed. I think it's very much in the spirit of free software.

But then again, for text-based search specifically (search engines, notes, etc.), I think there is some value in querying the LLM directly, as it is able to fuzzy search by inspecting its weights. This results in a lesser degree of specificity, which allows for more false positives, but can maybe capture similar words (i.e., synonyms / typos / tenses) or higher-level semantic concepts.

Maybe they're just two different search algorithms, and the user should be able to choose between them.

Thanks for linking the article, I found it very interesting.

Your described thinking pattern does not apply to 95%+ users, and might add confusion leading to retention loss (e.g. 2 text boxes? Wtf, two boxes? What do i type in the second one? Whatever i’ll switch to the usual). Every project with at least 100K non-tech users I’ve worked on, has shown that sort of thinking just doesn’t translate.

And regarding “average doesn’t exist” - that’s true. But no company does a/b testing to land on average. I’d assume a good 85%-kinda pass rate for these type of experiments.