Zen Mid-Size Model Workstation — Runs 70B Models On-Prem

$17,500.00

Starting at $17,500 — a dual-GPU workstation built specifically to run dense 70B-class open models entirely on-premise, at usable quality.

What it runs

  • Llama 3.3 70B at Q4_K_M and Q5_K_M — validated at roughly 27 tokens/sec on dual RTX 5090
  • Gemma 4 31B at full Q8 precision
  • Llama 4 Scout at INT4 (~63GB)
  • Wan 2.2 and HunyuanVideo 1.5 video generation at their higher-quality recommended tiers
  • FLUX.2 dev with headroom to spare

Core specs

  • 2x NVIDIA GeForce RTX 5090 32GB (64GB combined VRAM)
  • AMD Threadripper 9970X (32-core / 64-thread)
  • 128GB DDR5 memory, expandable
  • 4TB + 2TB NVMe Gen4 SSD
  • 1600W power supply, dual-GPU cooling, full tower workstation chassis

A single-GPU alternative using one RTX PRO 5000 Blackwell 48GB is available for office/rack deployments needing lower power draw.

Best for: the \"our data never leaves the building\" customer — Nevada law firms, medical groups, and engineering firms running serious local models.

Hardware requirement research: Hardwarepedia, dev.to, HaiMaker, AvenChat. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.

Starting at $17,500 — a dual-GPU workstation built specifically to run dense 70B-class open models entirely on-premise, at usable quality.

What it runs

  • Llama 3.3 70B at Q4_K_M and Q5_K_M — validated at roughly 27 tokens/sec on dual RTX 5090
  • Gemma 4 31B at full Q8 precision
  • Llama 4 Scout at INT4 (~63GB)
  • Wan 2.2 and HunyuanVideo 1.5 video generation at their higher-quality recommended tiers
  • FLUX.2 dev with headroom to spare

Core specs

  • 2x NVIDIA GeForce RTX 5090 32GB (64GB combined VRAM)
  • AMD Threadripper 9970X (32-core / 64-thread)
  • 128GB DDR5 memory, expandable
  • 4TB + 2TB NVMe Gen4 SSD
  • 1600W power supply, dual-GPU cooling, full tower workstation chassis

A single-GPU alternative using one RTX PRO 5000 Blackwell 48GB is available for office/rack deployments needing lower power draw.

Best for: the \"our data never leaves the building\" customer — Nevada law firms, medical groups, and engineering firms running serious local models.

Hardware requirement research: Hardwarepedia, dev.to, HaiMaker, AvenChat. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.