1) yes, but with more than channels than 3 (and no correspondence to color, its "pixel space" in the sense of position, not value)

1D grid example for clarity (obviously meant to be done in 2D though)

So embed -> [N dim, N dim, N dim ... , N dim] instead of embed -> [RGB, RGB, RGB ... , RGB]

2) / [followup 1)] Stride of 1, so we place a window at each pixel (so lots of overlapping)

Also I wouldn't phrase it as "come up with a better image", the point is to give less spatial decoding pressure to the patch tokens so that they can almost completely focus on feature learning instead. There is no reason to have the model learn spatial decoding when the structure prior of images is comically strong (especially compared to text), its a waste of training time and parameters

[followup 2)] I haven't measured but it was passable is all I can say (my experiments setup are horrendous right now lol)

(Also don't mind the phrasing, I just wanted to be 100% clear)

I appreciate the clarifications!

If we have extra compute lying around in the coming weeks, we’ll try this out and report back. It’s a good idea :)

I loved to hear! Just ping me on twitter (same account I sent the link with)