I think it depends on usage pattern. You trade speed for lower memory usage. Maybe engine specialisation and faster SSDs is the future for local inference, who knows
I think it depends on usage pattern. You trade speed for lower memory usage. Maybe engine specialisation and faster SSDs is the future for local inference, who knows