You're limited by the manufacturer (CUDA is king, thus NVIDIA is the king right now) and your lack of VRAM will make using a useful model difficult.

I'm not surprised at all.

Context: I have a farm of DGX Sparks and several RTX 6000's, and can run very close to foundational models with ~2 sparks