I’m sorry but this looks like vibe coded marketing slop that’s highly inaccurate.
For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.
Secondly, it seems to base all benchmarks off a batch size of 1.
Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.
And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.
I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.
For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.
Secondly, it seems to base all benchmarks off a batch size of 1.
Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.
And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.
I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.
Those units…
Also “is” is load-bearing