We know! This is an eval to evaluate which model is best at running a radio station. The purpose is not to build the best AI radio stations. Grok n' Roll is broken because Grok 4.3 is not doing so well.
We know! This is an eval to evaluate which model is best at running a radio station. The purpose is not to build the best AI radio stations. Grok n' Roll is broken because Grok 4.3 is not doing so well.