Just don't try concurrent requests. It's not fun. It does not seem to process in parallel at all. Instead, I'd see it swap contexts out frequently. Inefficient AF. Makes sense if you consider that each turn between two trying-to-run-in-parallel requests would activate different experts, requiring different parts of the model to be loaded to service different requests.
Maybe there's some whackadoodle way to only use a portion of the VRAM for experts for one request and another portion for experts for the other such that requests could actually run parallel instead of concurrently?
EDIT: I'm wrong! It already can do this with the parallel argument.
However, a few days ago, the dev(s now?) added swapping contexts to and from RAM.
Also, from my above, I don't see wh