> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient
Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
Pareto frontier the Pareto frontier and then with a little Pareto frontier you might get to the Pareto frontier
P^2
Its just an insufferable AI way to say best set of options given
Waaay older than AI. Basic econ 101,it is.
Though it being used constantly in AI-related topics from ~2024 onwards is very much a function of LLM output, I would argue.
Oh for sure. As has em-dashes -- a glorious tidbit primarily used by writers in the past to circumvent parenthetical clauses. Also the incredibly abused "load-bearing."
Meanwhile, I just want people to think about the production possibilities frontier.
Not really. The Pareto frontier is itself a distribution over all optimisation paths and dominance is a real and useful property we test when researching evolutionary methods. It’s not chosen to be insufferable but to try to be precise.
> Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly)
What does this mean?