I thought it was more because of fundamental limitations in the architecture. As in, no matter the training data, it could not be consistently and generally represented