Yes, and also possible they've gone too heavy on verifiable rewards (RLVR) for agentic coding work, and too light on human feedback (RLHF)
Yes, and also possible they've gone too heavy on verifiable rewards (RLVR) for agentic coding work, and too light on human feedback (RLHF)