I couldn't find any details about size or training data for the model.

It looks like they're taking applications for training data (due August 14th), so I think it's safe to say this is just an announcement of intent and a call for involvement vs. something that is readily available. Seems almost quaint in comparison to the strategy of sucking up every piece of data you can find anywhere on the Internet and feeding it to your LLM but I suspect their intent is to be more careful in what they train their model on.

I have no doubt companies like Microsoft, Amazon, and Google will rush to give them all the data they want in order to keep those government contracts flowing.