There was this post a few days ago https://news.ycombinator.com/item?id=49797323
It had this to say in the linked post:
This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:
gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200
MENENIUS:
'Though all at once canq
MARCIUS:
Pray now, nocamest thou to a morsel.
LARTIUS:
Hence, and
I' the end admire, where G
again; and after it ag .
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.
I don't know how this is related, but it reminds me of how I have always believed compression to be the ultimate sign of intelligence. If you can reduce something while keeping comprehension, you are finding more abstract symbols to represent the information of the original source.
> I don't know how this is related
https://en.wikipedia.org/wiki/Hutter_Prize
How can things be compressed without losing information or structure?
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
comments would be omitted.
``` th #1 - numbers are just comments ng #2, this has a space at the end information #3 compress #4 letter #5 this has a space at the start and #6 this has a space at the start and end sentence #7 this has a space at the start
How can 1ings be 4ed without losi23 or structure?
Like for text, what would 1at involve? How do you 4 a stri2or multi-line stri2wi1out losi236hopefully structure (paragraphs, would it be like replaci2periods61e followi2space wi1 just sticki21e starti2capitalized5 of 1e followi2word to 1e previous7's last56when it de4es 1eres some kind of note 1at converts 1at back into 1e. First5of the next7 ```
I'm on a phone, so I may have mistakes here, but I'm pretty sure that's shorter than your original text, in bytes, by about (9+18+20+21+12+12+16=108), minus the dictionary size of 51 -- so, 57 bytes shorter, but still containing your full text. With predistributed compressor binaries and a lot of analysis, you can even predistribute a global dictionary for common sequences, and simply specify "xyz0", x, y, and z being 24-bit numbers, or whatever bit size can index into your full reference dictionary, and 0 meaning end-of-file-dictonary. then, assuming the byte sequences in your text above are common enough to be in the 24-bit indexed dictionary, that initial dictionary could be just 22 bytes (21 and a terminator) -- so, 86 bytes shorter than the original, but still containing your original message unaltered. ..assuming i didn't make mistakes in my hand-compression.
You analyze the frequency of combinations of bytes, then replace those with high frequency with pointers to a single instance.
Deduplication ^
Other than what the others mentioned about finding more efficient representations, you can also compress by pre-agreeing on some common terminology.
In many ways, we are communicating using a compressed channel (words) since we both have pre-agreed on the meaning of these words.
That and also that agreements can evolve with time and also within a discussion, so a basic level of agreement is necessary, but complete consensus about the meaning of all words is unnecessary and often unproductive to communications held in good faith.
Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
The initial content is the most efficient representation but not for persistent digital storage and that's an important distinction because what we call sparse or inneficient is actually quite efficient, but only if you consider human consumption as the optimization target.