It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.
It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.
Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?
I guess it might be relative, but switching from VertexAI endpoint to OpenRouter was like 2-3x faster for us.