Guided path · small startup

Run one model, as cheaply as it can honestly be run.

You need to serve exactly one model in production. You don't need extra headroom, redundancy, or room to run a second model — you need the smallest, cheapest hardware that actually holds this model's weights in memory. Overspend here and you're burning runway for nothing.

Step 2 — the recommendation

2x NVIDIA HGX H200 8-GPU Server

  1. DeepSeek V4 (Pro) needs 1,920 GB of GPU memory. That's 1600 billion parameters × 1 GB, plus 20% working room for the model's short-term scratch space while it's generating a response.
  2. A single HGX H200 server (1,128 GB) isn't enough for DeepSeek V4 Pro's 1,920 GB requirement. Two servers clustered together give 2,256 GB — enough room, at the lowest price that qualifies. A GB300 rack (20 TB) also fits, with far more headroom, at a much higher price.
  3. This is a cluster: 1 machine type(s) working as one. A cluster just means multiple physical machines wired together and coordinated so they act like a single, bigger computer. The servers here are joined by NVLink/InfiniBand — a tight, purpose-built, high-speed connection inside and between datacenter machines. That is a fundamentally different (and far faster) thing than plugging several desktop RTX cards into one PC, which can't pool memory this way at all.
  4. Total price: $740,000. This is the lowest-cost combination on our shelf that actually has enough memory — nothing here is oversized for the job.
  5. Total power: 16,000 W — about 13 homes running around the clock. Running flat-out, this build would burn through a typical EV battery's charge (90 kWh) every about 5.6 hours. Budget for that in your hosting/power costs, not just the sticker price.
Request this quote →