Low-cost AI homelab computer options
This post is based on a conversation with Gemini about building a low-cost AI homelab using older workstation hardware and surplus GPUs.
The prior context is that I spent a couple of months(!) learning how to host local LLMs using an old datacenter GPU (AMD Instinct MI25). This lead me to wonder if there were better similar alternatives.
As with prior exercises, I find that an LLM can get 50% to 90% of where I want to go, but needs help. I this case, if you count the work needed to get to the point of asking the question, the percentage is lower. Still, using various LLMs did save much time.
Count this as an experiment, about an exercise, containing an experiment, to support further exercises. :)
To extract the conversation had to:
- Save the outerHTML of the subtree containing just the conversation.
- Feed the HTML extract to Gemini, and ask for Markdown.
- Used chat (Copilot) in vscode to draft this article with summary.
- Added the preamble, trimmed out some of the generated text, and tweaked format.
This worked. (Feeding the entire saved HTML page did not.) There might be a better way. :)
Summary
- Older workstation and enterprise surplus hardware can often offer better performance-per-dollar than new consumer DDR4/DDR5 platforms.
- Dual-socket platforms like the HP Z840 or Lenovo ThinkStation P910/P920 provide the PCIe lanes, power, and memory capacity needed for multiple MI25-style GPUs.
- Larger CPU RAM makes sense for model overflow and larger LLM experimentation, even though DDR3/DDR4 system RAM is much slower than GPU HBM2.
- One- or two-GPU setups are the most practical value; three or four GPUs are possible but require strong power delivery and careful slot clearance.
Recommended platforms
HP Z840 workstation
- Dual Intel Xeon E5-2600v3 or v4 (Haswell/Broadwell)
- DDR4 ECC Registered RAM is still cheaper than new consumer DDR4/DDR5 platforms
- Wide PCIe slot spacing and robust power supplies (850W to 1285W)
- Can host two or three double-slot cards like fan-modded MI25s if you manage auxiliary power and clearance
Lenovo ThinkStation P910 / P920
- P910 uses dual Xeon E5-v4 (DDR4), P920 moves to Xeon Scalable
- Strong modular interior design and 1400W redundant power options
- Multiple true x16 slots tied directly to CPU lanes
- Good choice if you want to push to three or four GPUs without lane starvation
What to avoid
- Noisy rackmount server solutions were excluded by application requirements.
- Common desktop platforms are poorly suited as fewer PCIe lanes and present cost for DDR4/DDR5.
What makes sense for system memory?
Larger CPU RAM is useful to run oversized models and allow spillover to system memory.
- Older DDR3-based multi-socket Xeon systems can support 256GB or 512GB of ECC RAM cheaply
- Attractive for experimenting with larger models, even though system RAM bandwidth is far slower than GPU HBM2
- Expect much slower performance when the model overflows from VRAM to system memory
Budget guidance
Estimates assuming a tower workstation chassis, cheap high-capacity ECC memory, and fan-modded AMD Instinct MI25 GPUs.
- 1 GPU – 16 GB HBM2: $250 to $350
- Base workstation + one MI25 + ducted fan + 64GB/128GB RAM
- Good for up to 14B/32B quantized models
- 2 GPUs – 32 GB HBM2: $350 to $500
- Sweet spot for value and useful for 32B models in fast HBM2
- Needs enough PSU capacity and PCIe power headers
- 3 GPUs – 48 GB HBM2: $500 to $700
- Enables heavier model work such as 34B/70B at lower quantization
- Requires 1000W+ PSU and enough slot spacing
- 4 GPUs – 64 GB HBM2: $700 to $1,000+
- Maximum density for a standard tower build
- Watch for PCIe x8 lane limitations, clearance issues, and high power draw
Alternatives to the MI25
NVIDIA Tesla P100 (16GB HBM2)
- Secondary market price around $130 to $150
- Higher memory bandwidth than MI25 and native CUDA support
- Still a passive enterprise card that needs a blower/fan mod
NVIDIA Tesla P40 (24GB GDDR5)
- Often priced between $200 and $400 on the surplus market
- 24GB VRAM per card can ease larger-model fitment
- Slower GDDR5 bandwidth compared to HBM2 cards
Practical notes
- The MI25 board has fan mounting holes and can be cooled with the same style fan as the WX9100.
- With the right duct and fan, the MI25 works well in a desktop chassis and only requires two slots.
- Recent Vulkan support is approaching ROCm performance for LLM workloads, and newer ROCm versions have restored MI25 support.
Original conversation
User:
Given current inflated hardware prices, what is a low cost cost AI homelab computer with the best performance per dollar? For example, an AMD Instinct MI25 can be had for under $100, and with a GPU fan and 3D printed duct, can serve in a desktop. Could get more than one, though need enough power and PCIe lanes. In present, DDR3 is half the price of DDR4, and DDR4 is half the price of DDR5.As example, have an old HP Z820 that I got used several years ago. Has enough power, and with dual (old) Xeons it has enough PCIe lanes to host at least two MI25. The cost to add 256GB RAM was modest. Looking for a similar sort of rig, and the best value at current at current prices.
Gemini: Given the current inflated pricing landscape—where sky-high consumer DRAM prices have turned building new DDR4/DDR5 platforms into a costly endeavor—the surplus enterprise/workstation market remains the absolute king of performance-per-dollar for AI homelabs.
Just like your HP Z820, the goal is to leverage older server platforms that feature abundant PCIe lanes, robust power delivery, and dirt-cheap older generation memory (like DDR3 or early DDR4) to support high-density, low-cost accelerator setups like the AMD Instinct MI25 ($100 for 16GB of HBM2 VRAM!).
The following options represent the best-value multi-GPU AI homelab rigs currently available on the secondary market.
1. The Direct Successor: HP Z840 Workstation
- The Formula: Dual Intel Xeon E5-2600v3 or v4 (Haswell/Broadwell, up to 22 cores per socket), DDR4 ECC Registered RAM (significantly cheaper than consumer DDR4/DDR5 right now), and massive PCIe slot spacing.
- Why it fits:
- It directly mirrors your Z820 experience but steps up to support full PCIe 3.0 natively across all slots at full bandwidth (
x16), maximizing data throughput to cards like the MI25. - It features massive proprietary multi-rail power supplies (often 850W to 1285W) with plenty of PCIe power connectors.
- You can find barebones or lightly configured units very cheaply on eBay, and adding high-density ECC RDIMMs remains relatively affordable compared to consumer retail markup.
- It directly mirrors your Z820 experience but steps up to support full PCIe 3.0 natively across all slots at full bandwidth (
- GPU Capacity: Easily handles two to three double-slot GPUs (like fan-modded MI25s) provided you manage the auxiliary power pin-outs and physical clearance.
2. The Enterprise Route: Dell PowerEdge R730 or Lenovo ThinkServer RD650
- The Formula: 2U Rackmount dual Xeon E5-v3/v4 servers.
- Why it fits: If you don't mind rack noise (or can home-lab it in a garage/closet), enterprise servers offer incredible PCIe lane distribution. A dual-CPU R730 provides up to 7 PCIe slots driven directly by the CPUs.
- The Catch: Enterprise servers are designed for active server-class blower cards (like passive MI25s that rely on high-cfm chassis fans). Dropping desktop-modded blower fans or 3D-printed shrouds onto MI25s inside a 2U or tower server requires creative power tap solutions (since server boards rarely supply standard desktop GPU power from internal headers, requiring breakout boards or PCIe power adapters).
3. The High-Lane AMD Alternative: ThinkStation P910 / ThinkStation P920
- The Formula: Lenovo’s flagship workstation line. The P910 uses dual Xeon E5-v4 (DDR4), while the P920 steps up to Xeon Scalable (Skylake/Cascade Lake).
- Why it fits: Lenovo’s P-series workstations are legendary for their modular, cable-free interior design and massive power supplies (up to 1400W with redundant options).
- PCIe Layout: The P910/P920 layout features multiple true
x16slots tied directly to CPU lanes. You can comfortably slot three or four MI25 cards or Tesla P40s without running into lane starvation or bottlenecking your PCIe switch tree.
Key Bottlenecks & Practical Tips for Multi-MI25 Rigs
- The ROCm Software Tax: While an AMD Instinct MI25 gives you 16GB of ultra-fast HBM2 bandwidth (~484 GB/s) for roughly $100, remember that AMD’s ROCm software stack can occasionally require more configuration patience than NVIDIA's CUDA. Ensure your target LLM frameworks (like
llama.cpporvLLM) are compiled with explicit support for GFX900 (MI25 architecture) targets, as older Vega architectures require specific environment variables (HSA_OVERRIDE_GFX_VERSION=9.0.0) to function smoothly on modern libraries. - Power & Thermal Management: MI25s are power-hungry (~300W peak) and entirely passive out of the box. When 3D-printing your blower ducts, prioritize high-static-pressure fans (such as 40mm or 50mm server grade fans, or blower style fan setups) because HBM stacks thermal-throttle quickly if the heatsink isn't actively blasted with directed airflow.
User:
The MI25 board has mounting holes for a fan (presumably the same board as a WX9100, so can use the same sort of fan), from experience is well-cooled with a fan and 3D printed duct, and needs only 2-slots room.Also recent Vulkan seems to be roughly on par with ROCm for LLM performance. Note recent ROCm version adds back support for the MI25, so no special configuration should be needed.
Exclude the (noisy) rackmount server solutions.
How much memory makes sense? The notion is larger CPU RAM would allow larger LLM to run (if slowly), so perhaps worthwhile. Large CPU memory might argue more for DDR3.
What sort of budget are we looking at for one, two, three, and (if possible) four GPU setups?
Are there alternatives to the MI25 that might offer better price / performance?
Gemini:
System Memory: How Much Makes Sense?
For an AI homelab, leaning into larger CPU RAM (and thus choosing an older DDR3/early DDR4 platform) is a viable strategy for running oversized Large Language Models—with a major caveat regarding performance.
- The CPU Offloading Trade-off: If you want to run a 70B parameter model (which requires roughly 35GB–40GB of space quantized at 4-bit) across a dual-GPU 32GB VRAM setup, overflow tokens can spill into system RAM via frameworks like
llama.cpp. However, system RAM bandwidth (even quad-channel DDR3 or DDR4) is a fraction of HBM2 speed (~50–80 GB/s vs. the MI25's ~484 GB/s). - The Verdict on DDR3 vs. DDR4:
- Going with an older DDR3-based multi-socket Xeon system (like an HP Z820 or v2 workstation) lets you pack 256GB or 512GB of ECC RAM for pennies (often under $50–$100 total for massive capacities). This gives you a playground to experiment with humongous models that would otherwise require multi-thousand-dollar hardware.
- Just expect token generation speeds to crawl significantly whenever the context window or model weights exceed total VRAM capacity and heavily rely on system memory.
Budget Breakdown for 1, 2, 3, and 4 GPU Setups
Assuming you are sourcing a tower workstation chassis (like a Z820/Z840 or ThinkStation equivalent), packing it with affordable high-capacity ECC memory, and utilizing fan-modded AMD Instinct MI25 ($75–$100 each):
| GPU Count | Total VRAM |
Build Cost | Notes |
|---|---|---|---|
| 1 GPU | 16 GB HBM2 | $250–$350 | Base workstation + 1x MI25 + 3D printed duct/fan + 64GB/128GB RAM. Great entry point for up to 14B/32B quantized models. |
| 2 GPUs | 32 GB HBM2 | $350–$500 | Sweet spot for value. Comfortably runs 32B models fully in fast HBM2, or larger models with mild offloading. Requires ensuring the PSU has enough 8-pin PCIe headers. |
| 3 GPUs | 48 GB HBM2 | $500–$700 | Enables full-speed execution of heavy models like 34B/70B at lower quants. Requires a workstation with a 1000W+ PSU and physical slot spacing to breathe. |
| 4 GPUs | 64 GB HBM2 | $700–$1000+ | Maximum density for a standard desktop tower. Bottlenecks: * Motherboards may force the bottom slots down to PCIe 3.0 x8 lanes or hit physical bottom-of-case clearance issues with standard 2-slot cards. Power draw will also approach ~1200W+ under full load, testing the limits of standard household circuits and proprietary workstation power distribution boards. |
Alternatives to the MI25 for Price/Performance
While the MI25 is tough to beat at ~$80–$100 for raw memory bandwidth (484 GB/s HBM2), a few alternative surplus cards are worth cross-shopping depending on your software ecosystem preferences:
- NVIDIA Tesla P100 (16GB HBM2)
- The Price: ~$130 – $150 on the secondary market.
- The Pros: Also uses 16GB of HBM2 (~732 GB/s bandwidth, even faster than the MI25), but operates on native CUDA. If you ever run into quirks with Vulkan or ROCm on niche repository updates, standard CUDA just works out of the box with zero compilation overhead.
- The Cons: Like the MI25, it's a passive enterprise card requiring a fan/shroud mod.
- NVIDIA Tesla P40 (24GB GDDR5)
- The Price: Historically a darling of the homelab community, though pricing fluctuates wildly (often $200–$400 depending on surplus waves).
- The Pros: 24GB of VRAM per card means a two-card setup gives you 48GB, easily fitting larger models completely in VRAM without relying on slow system memory or aggressive quantization.
- The Cons: Uses older GDDR5 memory instead of HBM2, meaning token generation bandwidth is considerably slower (~346 GB/s) compared to the blazing-fast memory fabric of the MI25 or P100.
Conclusion
For a low-cost AI homelab, the best value is usually older workstation hardware with cheap ECC RAM and surplus HBM2 cards.
A dual-socket workstation like the Z840 or a ThinkStation P910/P920 gives you the PCIe lanes and power delivery needed for two to four MI25-style GPUs, while a DDR3/DDR4 platform can give you a lot of system memory for model offloading.