Guided path · small startup

Run one model, as cheaply as it can honestly be run.

You need to serve exactly one model in production. You don't need extra headroom, redundancy, or room to run a second model — you need the smallest, cheapest hardware that actually holds this model's weights in memory. Overspend here and you're burning runway for nothing.

Step 2 — the recommendation

1x NVIDIA HGX H200 8-GPU Server

  1. GLM-5.2 needs 893 GB of GPU memory. That's 744 billion parameters × 1 GB, plus 20% working room for the model's short-term scratch space while it's generating a response.
  2. One HGX H200 server already holds 1,128 GB — more than the 893 GB GLM-5.2 needs — so a single server is the smallest thing that actually fits it.
  3. Total price: $370,000. This is the lowest-cost combination on our shelf that actually has enough memory — nothing here is oversized for the job.
  4. Total power: 8,000 W — about 7 homes running around the clock. Running flat-out, this build would burn through a typical EV battery's charge (90 kWh) every about 11.2 hours. Budget for that in your hosting/power costs, not just the sticker price.
Request this quote →