Zen Flagship Inference Rig — Runs 100B+ Open Models On-Prem

$28,000.00

Starting at $28,000 — a single-GPU flagship system built to bring 100B+ parameter open-weight models fully in-house.

What it runs

  • gpt-oss-120b on a single card — 96GB gives full headroom over its ~80GB planning target
  • Llama 4 Scout at Q4 (~55GB) or Q8 with offload
  • DeepSeek V4-Flash at 4-bit (~45GB) and DeepSeek V4-Pro at 4-bit (~85GB)
  • FLUX.2 dev at full FP16 (~64GB)
  • Every model in the Zen Local LLM Starter, Creator AI Studio, and Mid-Size Model Workstation lines simultaneously

Core specs

  • NVIDIA RTX PRO 6000 Blackwell 96GB
  • AMD Threadripper PRO 9975WX (8-channel memory, 128 PCIe lanes)
  • 256GB DDR5 RDIMM memory
  • 8TB NVMe model vault + 2TB NVMe OS drive
  • WRX90 platform, 1600W+ power supply, liquid cooling

Best for: government, healthcare networks, and enterprise pilots. Pair with a Zen Care managed-IT plan for ongoing model updates, quantization tuning, and inference operations.

Hardware requirement research: YingTu, InsiderLLM, NurAzhar, Newegg pricing. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.

Starting at $28,000 — a single-GPU flagship system built to bring 100B+ parameter open-weight models fully in-house.

What it runs

  • gpt-oss-120b on a single card — 96GB gives full headroom over its ~80GB planning target
  • Llama 4 Scout at Q4 (~55GB) or Q8 with offload
  • DeepSeek V4-Flash at 4-bit (~45GB) and DeepSeek V4-Pro at 4-bit (~85GB)
  • FLUX.2 dev at full FP16 (~64GB)
  • Every model in the Zen Local LLM Starter, Creator AI Studio, and Mid-Size Model Workstation lines simultaneously

Core specs

  • NVIDIA RTX PRO 6000 Blackwell 96GB
  • AMD Threadripper PRO 9975WX (8-channel memory, 128 PCIe lanes)
  • 256GB DDR5 RDIMM memory
  • 8TB NVMe model vault + 2TB NVMe OS drive
  • WRX90 platform, 1600W+ power supply, liquid cooling

Best for: government, healthcare networks, and enterprise pilots. Pair with a Zen Care managed-IT plan for ongoing model updates, quantization tuning, and inference operations.

Hardware requirement research: YingTu, InsiderLLM, NurAzhar, Newegg pricing. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.