No secrets—all published. Very efficient attention. Excellent kernels. Great caching subsystem. Small and well trained model.