Curious about Vulkan overhead on Intel vs AMD/Nvidia for long context. Any benchmarks vs vllm/sglang?