That sounds about right for Q4.

it's an extremely sparse MOE, so there is some odds of acceptable performance using a smaller in-memory cache and the rest on flash. ... I don't have a setup to test that right now.

(Of course, if translation is all you want much smaller models will work. Ling-tiny can do summarization, dom manipulation, scripting, etc. too).