You can probably even share context between questions by cleverly manipulating the attention mask.

Nice idea! Didn't think about that; a single linear memory allocation could do the trick