That's not true and its exactly the point - real life data is noisy, difficult to parse, and ultimately often ambiguous. Text is the distilled information. It's significantly easier to distilled text into a likely response than it is to distill into lessons a continuous data stream whose contents are partially controlled by your own actions.

Yeah but we're talking about something that was trained on petabytes of noisy, difficult to parse and ambiguous information.