Guided path · mid-size company
Build for headroom, not just today's traffic.
You have real, growing usage — more simultaneous users, and a product that can't afford to be slow or go down. The cheapest build that technically fits the model is the wrong answer here: you need spare memory for traffic spikes, room to run more than one model, and a machine that keeps serving customers if part of it fails.
Step 1 — which model are you building around?
Step 2 — the recommendation
2x NVIDIA HGX H200 8-GPU Server
- GLM-5.2 needs 893 GB just to load. A startup's minimum build gets you exactly that much room — fine for one model, one modest audience, zero margin for error.
- One server already runs GLM-5.2 comfortably, so a mid-size company doesn't need more memory — it needs more capacity. A second server means you can serve twice the simultaneous users (or run a second model), and it means a hardware fault doesn't take your whole product offline.
- Headroom, quantified: this build carries 2,256 GB against a 893 GB requirement — 2.5x the model's footprint, versus 1.3x for the bare-minimum startup build.
- This is a cluster. Multiple machines wired together by NVLink/InfiniBand act as one larger, faster computer with a shared memory pool. That's a datacenter-grade capability — it is not the same thing as plugging several desktop cards into one PC, which has no way to pool memory across cards at all.
- Total price: $740,000. More upfront than the startup build — the trade you're buying is capacity you won't outgrow in a quarter.
- Total power: 16,000 W — about 13 homes running around the clock. Running continuously, that's a full EV battery's worth of energy (90 kWh) every about 5.6 hours. At this scale, talk to your facility or colocation provider about power capacity before you talk about the invoice.