Image 1 of 1
Zen Mid-Size Model Workstation — Runs 70B Models On-Prem
Starting at $17,500 — a dual-GPU workstation built specifically to run dense 70B-class open models entirely on-premise, at usable quality.
What it runs
- Llama 3.3 70B at Q4_K_M and Q5_K_M — validated at roughly 27 tokens/sec on dual RTX 5090
- Gemma 4 31B at full Q8 precision
- Llama 4 Scout at INT4 (~63GB)
- Wan 2.2 and HunyuanVideo 1.5 video generation at their higher-quality recommended tiers
- FLUX.2 dev with headroom to spare
Core specs
- 2x NVIDIA GeForce RTX 5090 32GB (64GB combined VRAM)
- AMD Threadripper 9970X (32-core / 64-thread)
- 128GB DDR5 memory, expandable
- 4TB + 2TB NVMe Gen4 SSD
- 1600W power supply, dual-GPU cooling, full tower workstation chassis
A single-GPU alternative using one RTX PRO 5000 Blackwell 48GB is available for office/rack deployments needing lower power draw.
Best for: the \"our data never leaves the building\" customer — Nevada law firms, medical groups, and engineering firms running serious local models.
Hardware requirement research: Hardwarepedia, dev.to, HaiMaker, AvenChat. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.
Starting at $17,500 — a dual-GPU workstation built specifically to run dense 70B-class open models entirely on-premise, at usable quality.
What it runs
- Llama 3.3 70B at Q4_K_M and Q5_K_M — validated at roughly 27 tokens/sec on dual RTX 5090
- Gemma 4 31B at full Q8 precision
- Llama 4 Scout at INT4 (~63GB)
- Wan 2.2 and HunyuanVideo 1.5 video generation at their higher-quality recommended tiers
- FLUX.2 dev with headroom to spare
Core specs
- 2x NVIDIA GeForce RTX 5090 32GB (64GB combined VRAM)
- AMD Threadripper 9970X (32-core / 64-thread)
- 128GB DDR5 memory, expandable
- 4TB + 2TB NVMe Gen4 SSD
- 1600W power supply, dual-GPU cooling, full tower workstation chassis
A single-GPU alternative using one RTX PRO 5000 Blackwell 48GB is available for office/rack deployments needing lower power draw.
Best for: the \"our data never leaves the building\" customer — Nevada law firms, medical groups, and engineering firms running serious local models.
Hardware requirement research: Hardwarepedia, dev.to, HaiMaker, AvenChat. Pricing subject to current GPU/memory market conditions — final quote confirmed at order.