There are dozens of providers, the model is just really cheap to run thanks to its clever architecture that optimize the compute and memory usage even in long contexts.