This is really neat! I'd be tempted to try this again targetting ~1B params and the entire cookbook of "current" small model ideas: gated delta nets (or is ~2K context too short to benefit?), per layer or n-gram embeddings, gated residuals etc.
This is really neat! I'd be tempted to try this again targetting ~1B params and the entire cookbook of "current" small model ideas: gated delta nets (or is ~2K context too short to benefit?), per layer or n-gram embeddings, gated residuals etc.