Yes. Pretty much all the models that don't suck are trained on user data, either directly or via derived synthetic data.

Many upstart Chinese labs got around the user data issue by just buying copious amounts of Claude and ChatGPT session logs from model routers.