If you follow the interpretability argument, then Mamba and LSTMs would be the scariest thing ever and yet in practice they don't perform as well as transformers that have basically infinite recall within their context window.
If you follow the interpretability argument, then Mamba and LSTMs would be the scariest thing ever and yet in practice they don't perform as well as transformers that have basically infinite recall within their context window.