Compression in this modern day and age is so slop.
Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and it's an optimal compressor. It's completely vacuous though. You can always hide the "heavy work" in your mappings and descriptions. It's ironic that a kolomogorov complexity minimizer is so loaded that it's vacuous.
This is pretty much why I roll my eyes at this point at all the compression is intelligence memes.
I wonder when intervention and causality will hit the mainstream. These tools were designed specifically to counteract purely predictive theories. But your average compression dude will hold tight to their paradigms and slogans, not realize their internal contradictions (that their own field has brought up), and then whenever a new paradigm suddenly becomes visible and mainstream, they'll latch onto that. It's not principled at all.
And to be clear - I do think intelligence is some amount of compression, and I am well aware of formal results such as the arithmetic decoding theoretical and empricial result. Just annoyed. It's literally no different than the whole Bayesianism meme. If you're not actually practicing that type of intelligence as a basis, then you don't get to go around beating the drum about how it's the ultimate reality. You're just spouting dogma to feel like part of an in-group.
I think that the ratio of work done by priors and search is the interesting question at this point. We aren't quite at the point where reading off the most likely hypothesis decompressed from a transformer is sufficient, but it's a lot closer than I originally suspected when we were ~solving chess. I think AlphaGo was kind of the watershed moment that prior-guided search was so much better than either alone.
I think the adversarial policies against Go AIs directly show the gap between intelligence and compression/priors.
> I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them
Then why don't you ? and did you mean all or those that halt ? I presume you have a way of separating those.
I’m guessing your downvotes were for tone? You’re correct though regarding Solomonoff induction, as the choice of reference universal partial recursive function gives drastically different results for predictions based on finite data (even with access to a halting oracle). Asymptotically, any choice eventually converges to the same predictions, but that’s no help when there are infinitely many choices for U and no obvious natural prior over universal functions. And I don’t find the argument that our natural environment “implements some choice of U” particularly convincing. There’s definitely an open mystery there.