Guided path · small startup
Run one model, as cheaply as it can honestly be run.
You need to serve exactly one model in production. You don't need extra headroom, redundancy, or room to run a second model — you need the smallest, cheapest hardware that actually holds this model's weights in memory. Overspend here and you're burning runway for nothing.
Step 1 — which model are you running?
Step 2 — the recommendation
2x NVIDIA HGX H200 8-GPU Server
- Kimi K2.7 (Code) needs 1,200 GB of GPU memory. That's 1000 billion parameters × 1 GB, plus 20% working room for the model's short-term scratch space while it's generating a response.
- One HGX H200 server (1,128 GB) is just short of Kimi K2.7's 1,200 GB requirement. Two servers clustered together give 2,256 GB, comfortably enough, at the lowest price that qualifies. A GB300 rack also fits, with far more headroom, at a much higher price.
- This is a cluster: 1 machine type(s) working as one. A cluster just means multiple physical machines wired together and coordinated so they act like a single, bigger computer. The servers here are joined by NVLink/InfiniBand — a tight, purpose-built, high-speed connection inside and between datacenter machines. That is a fundamentally different (and far faster) thing than plugging several desktop RTX cards into one PC, which can't pool memory this way at all.
- Total price: $740,000. This is the lowest-cost combination on our shelf that actually has enough memory — nothing here is oversized for the job.
- Total power: 16,000 W — about 13 homes running around the clock. Running flat-out, this build would burn through a typical EV battery's charge (90 kWh) every about 5.6 hours. Budget for that in your hosting/power costs, not just the sticker price.