imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.

Same person that was mocking the hands in image generation in 2023, is the same person that was saying 'hands are fixed but it can't generate "the red dog jumps over the jump rope held by the blue pelican while juggling 5 balls"' in 2024, is the same person that posted this.

"One-shot learning" is still part of the discipline, actively studied (definitely in the past and surely in the present).