Where to run an LLM gateway
A request through an LLM gateway hangs on for 20 seconds while the upstream model streams tokens, yet the gateway does only about 10 ms of real work. Cloudflare bills those milliseconds and charges nothing for egress, but bills every read and write of state; Alibaba Cloud, Tencent Cloud, AWS, Google Cloud and Azure make state almost free but bill every byte of the prompts the gateway forwards; Oracle, with its first 10 TB of egress free each month, is the one exception. This article queries prices as of 2026-10-09 with the aliyun CLI, tccli, the clouds' public pricing APIs and official pricing pages, works out the bill line by line for seven providers at 10 million, 100 million and 1 billion requests a month, adds latency, reachability and real-world cases, and ends with when to choose which.