> I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline
No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.
They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.
But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.
Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.
Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.