If your team pays a cloud provider $1,500–4,000 a month to keep a model serving around the clock, you already own the hardware; you just have not taken delivery of it. At 80% utilisation and above, a box of your own pays for itself in 12–24 months, and often faster. The question is where to buy it. Most of what goes into a GPU server (chassis, power, risers, cooling, memory-modded cards) is made in Guangdong and sold there at a fraction of US retail. The part that is not made cheaper anywhere, the NVIDIA H200 and B300 systems, is rationed by allocation, and the hard part is getting a slot. This guide covers both: what is actually cheaper from China in 2026, what to buy in the US instead, two builds with landed budgets (a 4-GPU inference box for a small team and an 8-GPU 4U server), the DGX H200 and B300 systems for teams that need HBM and NVLink, and how we deliver all of it tested, duty-paid, with a warranty you can call.
Who this is for
- A 2–10 person AI team running inference 24/7. You serve a model behind an API, a RAG pipeline or an agent, and the cloud bill is $1,500–4,000 a month and climbing. Below about 70% utilisation the cloud still wins; above 80%, your own hardware wins on cost, on latency and on where the data lives. Your build is in the "Inference Box" section below.
- A builder or a small lab that wants a rig at home or in the office. One 240 V circuit, four cards, a model that answers in under a second and never leaves the room. The same build, one honest note: renting the cards out on Vast.ai or RunPod when idle covers electricity and a little more, not the rig.
- A company that needs H200 or B300 class systems now. Training or fine-tuning labs, inference providers, universities and HPC groups, enterprises building private AI, regional and sovereign AI projects, and the decentralised networks that only accept data-centre hardware. For you, allocation and lead time matter more than a 10% discount. See the DGX section.

What is cheaper from China, and what is not
The rule is simple: buy the metal, power and cooling where they are made, buy silicon where the price is the same everywhere, and buy memory-modded cards only from a source that tests them. The table is what we see in September 2026, ex-works Guangdong against US retail.

| Component | China, ex-works | US retail | Verdict |
|---|---|---|---|
| 4U chassis, 8 GPU, full PCIe x16 slots, rails | $120–180 | $600–900 (Chenbro RM41300); Supermicro barebone from $18,000 | China. On Amazon the 8-card 4U class barely exists; what is listed is mining leftovers with USB x1 risers. |
| 4U chassis, 4 GPU | $100–130 | $300–400 (Rosewill RSV-AI01) | China |
| Open frame, 6–8 GPU, flat-packed | $40–60 | $250 (mining-era frames) | China |
| Power kit: 2,400–3,000 W server PSU + breakout board + 12V-2x6 cables | $110–140 | $250–400 (Parallel Miner, still built around 6-pin mining connectors) | China |
| 12V-2x6 600 W cables, 16 AWG, 60–80 cm | $3 each | $25–40 each (CableMod) | China |
| PCIe 4.0 x16 full-speed shielded riser, 30 cm | $14 | $50–90 (LINKUP, effectively the only one) | China |
| 120×38 mm high-static PWM fans | $6 | $15–25 (Delta AFB1212SHE) | China, Delta or Sunon original or clone by choice |
| SP3 4U CPU cooler, thermal pads, rack PDU 30 A, 12U flat-pack rack | $25 / $2 / $60 / $120 | $60–90 / $10–20 / $150–350 / $150–400 | China |
| RTX 4090 48 GB turbo (memory-modded, 2-slot blower) | $3,000–3,800 | $4,500–5,500 from US resellers; $3,838–4,500 from Shenzhen sellers on eBay, often without US shipping and with return-to-China warranty | China, with the test protocol below |
| RTX 4090 24 GB turbo, RTX 5090 32 GB 2-slot | Market | Market ($5,000–5,800 new 5090) | Small margin; worth it inside a full build |
| EPYC CPUs (used), ECC DDR4, NVMe, 10 GbE NICs, UPS | Same or riskier (NVMe fakes) | eBay / Newegg | Buy in the US |
| Used 8-GPU platforms (ASUS ESC8000A-E11 with 2× EPYC) | Clone + duty costs more | ~$4,800 used | Buy in the US |
| RTX 6000 Ada, L40S, A100, RTX PRO 6000 Blackwell | World price | World price | No China margin; buy through normal channels |
The RTX 4090 48 GB: what it is and how to buy one that works
NVIDIA does not make a 48 GB RTX 4090. The card exists because US export controls cut China off from the A100 and H100, and Shenzhen workshops answered with a memory mod: the AD102 chip is moved to a server-style board, the twelve 2 GB GDDR6X modules are replaced with 4 GB modules (the same memory the RTX 6000 Ada uses), the BIOS is reflashed, and a two-slot blower cooler is fitted so eight cards fit in a 4U chassis. The result is 48 GB of VRAM at about $65 per gigabyte, against roughly $150 per gigabyte for an RTX 6000 Ada with the same chip and the same memory, and it is the best VRAM-per-dollar on the market for inference.

| Card | VRAM | Price | $/GB | Notes |
|---|---|---|---|---|
| RTX 4090 48 GB turbo (China mod) | 48 GB | $3,000–3,800 | ~$65 | No NVIDIA warranty; quality varies by workshop; blower, 2-slot |
| RTX 5090 | 32 GB | $2,400+ | ~$75 | Native 12V-2x6, 575 W, 3.5-slot coolers on most models |
| RTX A6000 (used) | 48 GB | ~$4,100 | ~$85 | Ampere; slower, mature drivers |
| RTX PRO 6000 Blackwell | 96 GB | ~$8,600 | ~$90 | Best single card for a serious node; world price |
| RTX PRO 5000 Blackwell | 48 GB | ~$4,800 | ~$100 | World price |
| A100 80 GB (used) | 80 GB | ~$10,000 | ~$125 | HBM2e, NVLink on SXM versions |
| RTX 6000 Ada / L40S | 48 GB | ~$7,200 | ~$150 | Same AD102 as the modded 4090 |
The risks are real and specific. There is no NVIDIA warranty. Solder quality varies between workshops. Some cards are built on the cut-down AD102-250 "D" chip, 10–15% slower than the AD102-300 in the retail 4090. Drivers occasionally need a specific version. Independent teardowns have shown both excellent and sloppy work, which is why buying one from a marketplace listing is a coin toss and buying twenty through a contract is not. What we require from the workshop, in the purchase contract, for every card:
- AD102-300 chip, Samsung or Hynix GDDR6X, PCB revision and memory batch documented.
- A burn-in video per card, nvidia-smi and memtest_vulkan screenshots, serial numbers listed on the packing list.
- A 30-minute four-card vLLM load test before shipment, filmed.
- 12-month factory warranty with RMA at the seller's cost, and 5% spare memory chips shipped with each batch for our own service bench.
- Our own inspection before export, and a 90-day Union Delta warranty on top, handled in the USA, not by returning the card to China.
Build 1: the Inference Box for a small team (4× 4090 48 GB, 192 GB VRAM)
This is the configuration we recommend to most teams that come to us: four 48 GB cards on a single-socket EPYC platform, 192 GB of VRAM, vLLM or Ollama installed, running on one 240 V circuit. It serves a 70B-class model in FP16 with room for context, or a 200B-class model quantised, with concurrency a small product needs. It fine-tunes with LoRA. And the data never leaves your office.
| Item | Source | Budget |
|---|---|---|
| 4× RTX 4090 48 GB turbo, tested, 12-month factory warranty | China, our contract | ~$12,000 |
| 4U 4-GPU chassis with rails, 2× 2,400 W CRPS power, breakout, cables, 8 high-static fans, SP3 cooler | China, consolidated with the cards | ~$600 |
| EPYC 7402 (used), 7× x16 board (ASRock Rack ROMED8-2T class), 128 GB ECC DDR4, 2 TB NVMe, 10 GbE NIC | USA | ~$1,800 |
| Sea freight, duties (~30% on China-origin items), assembly and burn-in | Union Delta | ~$4,500 |
| Landed, assembled, tested | ~$19,000–24,000 | |
| Same box built entirely from US retail parts | $27,000–32,000 |
What it needs at home or in the office: a dedicated 240 V line (a dryer or EV circuit works), about 2.2 kW under full load, which is roughly $220 a month in electricity at $0.13/kWh, a UPS, and a room where a blower rig can be loud. If none of that is possible, a single 4U in a colocation cabinet runs $150–300 a month.
The maths against the cloud: a team paying $2,000–4,000 a month for 24/7 inference pays the box off in 6–12 months at high utilisation, and keeps the hardware. A team that only bursts a few hours a day should stay on the cloud; we say so when the numbers say so.
If you rent it out when idle: a 4090-class card earns $0.30–0.50 an hour on Vast.ai or RunPod's community cloud at 40–80% occupancy, which is $200–450 a month per card net of power. That covers the electricity and a little more. It does not pay back the rig in under two years, so build it for your own workload and treat rental as a subsidy.
Build 2: the 8-GPU 4U server (8× 4090 48 GB, 384 GB VRAM)
The industrial version: a 4U rack chassis with eight two-slot blower cards, dual 2,700 W CRPS supplies, a dual-socket or high-lane EPYC platform, 10 or 25 GbE, on 208–240 V at about 5 kW. This is production inference for a product, a multi-tenant node, or a host on the decentralised networks that pay in real dollars. It belongs in a colocation cabinet, not a spare room: budget $1,200–1,800 a month for power and space.

| Budget | |
|---|---|
| 8× RTX 4090 48 GB turbo, tested | ~$24,000 |
| 4U 8-GPU chassis, 2× CRPS 2,700 W, breakout, cables, fans, rails | ~$900 |
| Platform: EPYC, 7× x16 board or used ESC8000-class barebone, 256 GB ECC, NVMe, 25 GbE | ~$4,000–5,500 |
| Freight, duties, assembly, 8-card burn-in | ~$7,500 |
| Landed, assembled, tested | ~$37,000 |
| Comparable 8× 48 GB systems from US integrators | $55,000–65,000 |
A note on chassis: the 4U 8-GPU class that AI builders need, with eight full-speed PCIe x16 slots, 650–700 mm depth and proper airflow across blower cards, is a standard product for Guangdong server-case factories and almost absent from US retail. We buy it directly from the factories that make it, in batches, which is also why we can sell the chassis, the power kit and the riser set on their own to teams that already have cards.
DGX H200 and DGX B300: when you need HBM and NVLink
Consumer cards top out where model training, long-context serving at scale and multi-node work begin. Above that line there is one product family that matters in 2026: NVIDIA's 8-GPU systems, the DGX H200 on Hopper and the DGX B300 on Blackwell Ultra, and the HGX-based 8-GPU servers built on the same boards by the major OEMs. The constraint on these is not price. It is allocation: production is spoken for quarters ahead, large buyers take priority, and a small lab or a mid-size company placing its first order is quoted lead times in months.
| DGX H200 | DGX B300 | |
|---|---|---|
| GPUs | 8× H200 SXM (Hopper) | 8× B300 SXM (Blackwell Ultra) |
| GPU memory | 141 GB HBM3e per GPU, 1.1 TB total | 288 GB HBM3e per GPU, 2.3 TB total |
| GPU interconnect | 4th-gen NVLink, 900 GB/s per GPU, 7.2 TB/s aggregate | 5th-gen NVLink, 14.4 TB/s aggregate |
| AI performance | FP8 32 PFLOPS | FP4 144 PFLOPS, FP8 72 PFLOPS |
| CPU and system memory | 2× Xeon Platinum 8480C, 2 TB DDR5 (up to 4 TB) | 2× Xeon 6776P, 2 TB DDR5 (up to 4 TB) |
| Networking | 8× ConnectX-7, up to 400 Gb/s InfiniBand or Ethernet, 2× BlueField-3 | 8× ConnectX-8, up to 800 Gb/s, 2× BlueField-3 |
| Storage | 8× 3.84 TB NVMe + 2× 1.92 TB OS | 8× 3.84 TB NVMe E1.S + 2× 1.92 TB OS |
| Power | ~10.2 kW max, 10× 3.2 kW redundant supplies | ~14.5 kW max, 12× 3.3 kW redundant supplies |
| Form factor, weight | 10U, ~226 kg | 10U, ~239 kg |
| Software, warranty | DGX OS, NVIDIA AI Enterprise, Base Command / Mission Control; 3-year limited hardware warranty | |

Who needs them: teams training or fine-tuning models above the 70B class; inference providers serving long contexts to many tenants, where HBM bandwidth and NVLink decide throughput; universities and HPC centres; enterprises building a private AI platform that must stay on-premises; regional and sovereign AI initiatives; and the decentralised inference networks that only admit H100-class and better hardware. A typical first project is 32 GPUs: four 8-GPU H200 servers, 4.5 TB of HBM3e, a 400 Gb/s InfiniBand or Ethernet fabric, rack integration, power distribution, cooling and commissioning. At 10–15 kW per box, this is data-centre hardware; a colocation partner with the right power density is part of the plan from day one.
Contact us about DGX allocation →
Spare parts and consumables for a running fleet
The cheapest place to save money on a GPU fleet is the parts you replace every quarter. Fans, thermal pads, power supplies, cables and risers are commodity items in Guangdong, and a batch of 50–500 units consolidated in our Guangzhou warehouse ships as one box.
| Part | China, ex-works | US retail | Typical batch |
|---|---|---|---|
| 120×38 mm high-static PWM fans (Delta / Sunon or clone) | $6 | $15–25 | 100–500 |
| 12V-2x6 600 W cables, 16 AWG, extended | $3 | $25–40 | 100–1,000 |
| Thermal pads 12.8 W/mK, 1.0–2.0 mm, for VRAM and VRM | $2 | $10–20 | 50–500 |
| 2,400–3,000 W CRPS server PSUs (Delta, Lite-On, HP-class), new | $70–110 | $200–400 | 10–100 |
| PCIe 4.0 x16 shielded risers; x16 to 2×x8 / 4×x4 bifurcation cards; SlimSAS / OCuLink adapters | $14 / $25 / $20 | $50–90 / $40–120 / $50–100 | 50–500 |
| SP3 / SP5 4U coolers, GPU air shrouds for blower cards | $25 / $8 | $60–90 / none available | 20–100 |
| Rack PDUs 30 A C19/C13, 1U shelves, cable management, 12–18U racks flat-packed | $60 / $8 / $120 | $150–350 / $20–40 / $150–400 | 10–50 |
What the landed cost really looks like
Ex-works prices are only half the number. A first mixed parts order we costed in September 2026, twenty 8-GPU chassis, ten 4-GPU chassis, twenty open frames, twenty power kits, cables, risers, fans, coolers and two EPYC platforms, came to about $12,100 ex-works and 845 kg. Sea freight at roughly $2.50 per kg added $2,100, and duty at 30% added $3,600, for a landed total of about $17,800: 1.47× the factory price. The same list at US retail is over $40,000. Cards are lighter but carry the same duty percentage on a much higher value, which is why they only make sense from China in batches of eight or more at wholesale prices. In every quote we send, freight and duty are already inside the number, DDP to your door.
How we deliver hardware
- We buy on China's domestic market as a Chinese company, in RMB, from the factories that make the chassis and power systems and from the workshops that build the 48 GB cards, not from export resellers. How that works and why the price is lower: How a China Sourcing Agent Works.
- Everything is tested before it leaves China. Cards get the burn-in protocol above. Chassis, PSUs and risers are checked against the spec and assembled as a dry run where the order is a full build.
- Consolidation and repacking in our Guangzhou warehouse. Chassis, power, cards and accessories from several factories leave as one shipment, packed to cut volumetric weight.
- Export, freight and duties handled. Sea for chassis and bulk parts, air or express for cards and urgent orders, DDP to your address. Duty classification is confirmed with a licensed broker before we quote.
- Assembly and burn-in in the USA for full builds, with vLLM or Ollama configured, and a 90-day Union Delta warranty on top of the 12-month factory warranty, serviced in the USA.
- Options that remove the risk of a first purchase: rent-to-own on the Inference Box from about $1,200 a month, a remote test drive over SSH before you commit, and trade-in from 24 GB to 48 GB cards.
Fees: for component orders and builds the service is priced as a percentage of the order, agreed before we buy, typically 5–8%, negotiated on larger orders. DGX-class systems are quoted as a project. The sourcing request itself is free: send the workload, the budget and the address, and you get a configuration, a landed price and a lead time within 24 hours.