If your team pays a cloud provider $1,500–4,000 a month to keep a model serving around the clock, you already own the hardware; you just have not taken delivery of it. At 80% utilisation and above, a box of your own pays for itself in 12–24 months, and often faster. The question is where to buy it. Most of what goes into a GPU server (chassis, power, risers, cooling, memory-modded cards) is made in Guangdong and sold there at a fraction of US retail. The part that is not made cheaper anywhere, the NVIDIA H200 and B300 systems, is rationed by allocation, and the hard part is getting a slot. This guide covers both: what is actually cheaper from China in 2026, what to buy in the US instead, two builds with landed budgets (a 4-GPU inference box for a small team and an 8-GPU 4U server), the DGX H200 and B300 systems for teams that need HBM and NVLink, and how we deliver all of it tested, duty-paid, with a warranty you can call.

$65/GBVRAM cost of an RTX 4090 48 GB, vs ~$150/GB for RTX 6000 Ada
3–4×cheaper 4U 8-GPU chassis at the factory than on Amazon
12–24 mopayback on your own hardware at 80%+ utilisation
2.3 TBHBM3e in one DGX B300, 8 GPUs, ~14.5 kW

Who this is for

  • A 2–10 person AI team running inference 24/7. You serve a model behind an API, a RAG pipeline or an agent, and the cloud bill is $1,500–4,000 a month and climbing. Below about 70% utilisation the cloud still wins; above 80%, your own hardware wins on cost, on latency and on where the data lives. Your build is in the "Inference Box" section below.
  • A builder or a small lab that wants a rig at home or in the office. One 240 V circuit, four cards, a model that answers in under a second and never leaves the room. The same build, one honest note: renting the cards out on Vast.ai or RunPod when idle covers electricity and a little more, not the rig.
  • A company that needs H200 or B300 class systems now. Training or fine-tuning labs, inference providers, universities and HPC groups, enterprises building private AI, regional and sovereign AI projects, and the decentralised networks that only accept data-centre hardware. For you, allocation and lead time matter more than a 10% discount. See the DGX section.
NVIDIA DGX Station desk-side AI workstation with a GB300 Grace Blackwell board
For a lab that wants data-centre silicon under a desk: NVIDIA DGX Station with a GB300 Grace Blackwell board. Image: NVIDIA.

What is cheaper from China, and what is not

The rule is simple: buy the metal, power and cooling where they are made, buy silicon where the price is the same everywhere, and buy memory-modded cards only from a source that tests them. The table is what we see in September 2026, ex-works Guangdong against US retail.

8-GPU AI server chassis from a Chinese factory, front view with drive bays and mesh panel
An 8-GPU AI server chassis from a Guangdong factory: full-length PCIe x16 slots, redundant CRPS power bays, 650–700 mm depth. Ex-works $120–180; the same class sells for $600–900 in the USA.
GPU server components: factory price in China vs US retail, September 2026
ComponentChina, ex-worksUS retailVerdict
4U chassis, 8 GPU, full PCIe x16 slots, rails$120–180$600–900 (Chenbro RM41300); Supermicro barebone from $18,000China. On Amazon the 8-card 4U class barely exists; what is listed is mining leftovers with USB x1 risers.
4U chassis, 4 GPU$100–130$300–400 (Rosewill RSV-AI01)China
Open frame, 6–8 GPU, flat-packed$40–60$250 (mining-era frames)China
Power kit: 2,400–3,000 W server PSU + breakout board + 12V-2x6 cables$110–140$250–400 (Parallel Miner, still built around 6-pin mining connectors)China
12V-2x6 600 W cables, 16 AWG, 60–80 cm$3 each$25–40 each (CableMod)China
PCIe 4.0 x16 full-speed shielded riser, 30 cm$14$50–90 (LINKUP, effectively the only one)China
120×38 mm high-static PWM fans$6$15–25 (Delta AFB1212SHE)China, Delta or Sunon original or clone by choice
SP3 4U CPU cooler, thermal pads, rack PDU 30 A, 12U flat-pack rack$25 / $2 / $60 / $120$60–90 / $10–20 / $150–350 / $150–400China
RTX 4090 48 GB turbo (memory-modded, 2-slot blower)$3,000–3,800$4,500–5,500 from US resellers; $3,838–4,500 from Shenzhen sellers on eBay, often without US shipping and with return-to-China warrantyChina, with the test protocol below
RTX 4090 24 GB turbo, RTX 5090 32 GB 2-slotMarketMarket ($5,000–5,800 new 5090)Small margin; worth it inside a full build
EPYC CPUs (used), ECC DDR4, NVMe, 10 GbE NICs, UPSSame or riskier (NVMe fakes)eBay / NeweggBuy in the US
Used 8-GPU platforms (ASUS ESC8000A-E11 with 2× EPYC)Clone + duty costs more~$4,800 usedBuy in the US
RTX 6000 Ada, L40S, A100, RTX PRO 6000 BlackwellWorld priceWorld priceNo China margin; buy through normal channels
Duties change the picture for heavy items less than you would expect. A 4U chassis at $150 ex-works lands at roughly $220–260 after sea freight and about 30% duty, still a third of the US price. Cards are where duty bites: at 4 cards the saving over a US reseller is thin; at 8–16 cards on a wholesale price of $2,900–3,100 it is real. Always confirm the duty rate for your HS code with a licensed broker before you commit; our numbers assume about 30%.

The RTX 4090 48 GB: what it is and how to buy one that works

NVIDIA does not make a 48 GB RTX 4090. The card exists because US export controls cut China off from the A100 and H100, and Shenzhen workshops answered with a memory mod: the AD102 chip is moved to a server-style board, the twelve 2 GB GDDR6X modules are replaced with 4 GB modules (the same memory the RTX 6000 Ada uses), the BIOS is reflashed, and a two-slot blower cooler is fitted so eight cards fit in a 4U chassis. The result is 48 GB of VRAM at about $65 per gigabyte, against roughly $150 per gigabyte for an RTX 6000 Ada with the same chip and the same memory, and it is the best VRAM-per-dollar on the market for inference.

Two-slot blower RTX 4090 turbo graphics card in foam packaging from a Shenzhen workshop
A two-slot blower RTX 4090 as shipped from a Shenzhen workshop. The 48 GB versions use the same cooler; the memory mod is on the board underneath.
VRAM cost per gigabyte, September 2026 street prices
CardVRAMPrice$/GBNotes
RTX 4090 48 GB turbo (China mod)48 GB$3,000–3,800~$65No NVIDIA warranty; quality varies by workshop; blower, 2-slot
RTX 509032 GB$2,400+~$75Native 12V-2x6, 575 W, 3.5-slot coolers on most models
RTX A6000 (used)48 GB~$4,100~$85Ampere; slower, mature drivers
RTX PRO 6000 Blackwell96 GB~$8,600~$90Best single card for a serious node; world price
RTX PRO 5000 Blackwell48 GB~$4,800~$100World price
A100 80 GB (used)80 GB~$10,000~$125HBM2e, NVLink on SXM versions
RTX 6000 Ada / L40S48 GB~$7,200~$150Same AD102 as the modded 4090

The risks are real and specific. There is no NVIDIA warranty. Solder quality varies between workshops. Some cards are built on the cut-down AD102-250 "D" chip, 10–15% slower than the AD102-300 in the retail 4090. Drivers occasionally need a specific version. Independent teardowns have shown both excellent and sloppy work, which is why buying one from a marketplace listing is a coin toss and buying twenty through a contract is not. What we require from the workshop, in the purchase contract, for every card:

  • AD102-300 chip, Samsung or Hynix GDDR6X, PCB revision and memory batch documented.
  • A burn-in video per card, nvidia-smi and memtest_vulkan screenshots, serial numbers listed on the packing list.
  • A 30-minute four-card vLLM load test before shipment, filmed.
  • 12-month factory warranty with RMA at the seller's cost, and 5% spare memory chips shipped with each batch for our own service bench.
  • Our own inspection before export, and a 90-day Union Delta warranty on top, handled in the USA, not by returning the card to China.

Build 1: the Inference Box for a small team (4× 4090 48 GB, 192 GB VRAM)

This is the configuration we recommend to most teams that come to us: four 48 GB cards on a single-socket EPYC platform, 192 GB of VRAM, vLLM or Ollama installed, running on one 240 V circuit. It serves a 70B-class model in FP16 with room for context, or a 200B-class model quantised, with concurrency a small product needs. It fine-tunes with LoRA. And the data never leaves your office.

Inference Box: 4× RTX 4090 48 GB, landed in the USA
ItemSourceBudget
4× RTX 4090 48 GB turbo, tested, 12-month factory warrantyChina, our contract~$12,000
4U 4-GPU chassis with rails, 2× 2,400 W CRPS power, breakout, cables, 8 high-static fans, SP3 coolerChina, consolidated with the cards~$600
EPYC 7402 (used), 7× x16 board (ASRock Rack ROMED8-2T class), 128 GB ECC DDR4, 2 TB NVMe, 10 GbE NICUSA~$1,800
Sea freight, duties (~30% on China-origin items), assembly and burn-inUnion Delta~$4,500
Landed, assembled, tested~$19,000–24,000
Same box built entirely from US retail parts$27,000–32,000

What it needs at home or in the office: a dedicated 240 V line (a dryer or EV circuit works), about 2.2 kW under full load, which is roughly $220 a month in electricity at $0.13/kWh, a UPS, and a room where a blower rig can be loud. If none of that is possible, a single 4U in a colocation cabinet runs $150–300 a month.

The maths against the cloud: a team paying $2,000–4,000 a month for 24/7 inference pays the box off in 6–12 months at high utilisation, and keeps the hardware. A team that only bursts a few hours a day should stay on the cloud; we say so when the numbers say so.

If you rent it out when idle: a 4090-class card earns $0.30–0.50 an hour on Vast.ai or RunPod's community cloud at 40–80% occupancy, which is $200–450 a month per card net of power. That covers the electricity and a little more. It does not pay back the rig in under two years, so build it for your own workload and treat rental as a subsidy.

Build 2: the 8-GPU 4U server (8× 4090 48 GB, 384 GB VRAM)

The industrial version: a 4U rack chassis with eight two-slot blower cards, dual 2,700 W CRPS supplies, a dual-socket or high-lane EPYC platform, 10 or 25 GbE, on 208–240 V at about 5 kW. This is production inference for a product, a multi-tenant node, or a host on the decentralised networks that pay in real dollars. It belongs in a colocation cabinet, not a spare room: budget $1,200–1,800 a month for power and space.

Server rack in a colocation cabinet with network switches and patch cables
Eight cards at 5 kW belong in a cabinet with 208–240 V power and real airflow, not in a spare room. Photo: Yuriy Vertikov / Unsplash.
8-GPU 4U server: landed cost vs US equivalents
Budget
8× RTX 4090 48 GB turbo, tested~$24,000
4U 8-GPU chassis, 2× CRPS 2,700 W, breakout, cables, fans, rails~$900
Platform: EPYC, 7× x16 board or used ESC8000-class barebone, 256 GB ECC, NVMe, 25 GbE~$4,000–5,500
Freight, duties, assembly, 8-card burn-in~$7,500
Landed, assembled, tested~$37,000
Comparable 8× 48 GB systems from US integrators$55,000–65,000

A note on chassis: the 4U 8-GPU class that AI builders need, with eight full-speed PCIe x16 slots, 650–700 mm depth and proper airflow across blower cards, is a standard product for Guangdong server-case factories and almost absent from US retail. We buy it directly from the factories that make it, in batches, which is also why we can sell the chassis, the power kit and the riser set on their own to teams that already have cards.

DGX H200 and DGX B300: when you need HBM and NVLink

Consumer cards top out where model training, long-context serving at scale and multi-node work begin. Above that line there is one product family that matters in 2026: NVIDIA's 8-GPU systems, the DGX H200 on Hopper and the DGX B300 on Blackwell Ultra, and the HGX-based 8-GPU servers built on the same boards by the major OEMs. The constraint on these is not price. It is allocation: production is spoken for quarters ahead, large buyers take priority, and a small lab or a mid-size company placing its first order is quoted lead times in months.

NVIDIA DGX H200 vs DGX B300, key specifications
DGX H200DGX B300
GPUs8× H200 SXM (Hopper)8× B300 SXM (Blackwell Ultra)
GPU memory141 GB HBM3e per GPU, 1.1 TB total288 GB HBM3e per GPU, 2.3 TB total
GPU interconnect4th-gen NVLink, 900 GB/s per GPU, 7.2 TB/s aggregate5th-gen NVLink, 14.4 TB/s aggregate
AI performanceFP8 32 PFLOPSFP4 144 PFLOPS, FP8 72 PFLOPS
CPU and system memory2× Xeon Platinum 8480C, 2 TB DDR5 (up to 4 TB)2× Xeon 6776P, 2 TB DDR5 (up to 4 TB)
Networking8× ConnectX-7, up to 400 Gb/s InfiniBand or Ethernet, 2× BlueField-38× ConnectX-8, up to 800 Gb/s, 2× BlueField-3
Storage8× 3.84 TB NVMe + 2× 1.92 TB OS8× 3.84 TB NVMe E1.S + 2× 1.92 TB OS
Power~10.2 kW max, 10× 3.2 kW redundant supplies~14.5 kW max, 12× 3.3 kW redundant supplies
Form factor, weight10U, ~226 kg10U, ~239 kg
Software, warrantyDGX OS, NVIDIA AI Enterprise, Base Command / Mission Control; 3-year limited hardware warranty
NVIDIA rack-scale AI system with a compute tray pulled out
Rack-scale NVIDIA systems: 10–15 kW per 10U box, high-airflow or liquid cooling, data-centre only. Image: NVIDIA.

Who needs them: teams training or fine-tuning models above the 70B class; inference providers serving long contexts to many tenants, where HBM bandwidth and NVLink decide throughput; universities and HPC centres; enterprises building a private AI platform that must stay on-premises; regional and sovereign AI initiatives; and the decentralised inference networks that only admit H100-class and better hardware. A typical first project is 32 GPUs: four 8-GPU H200 servers, 4.5 TB of HBM3e, a 400 Gb/s InfiniBand or Ethernet fabric, rack integration, power distribution, cooling and commissioning. At 10–15 kW per box, this is data-centre hardware; a colocation partner with the right power density is part of the plan from day one.

Union Delta can supply DGX H200 and DGX B300 systems and 8-GPU HGX servers through our supply partners, with rack integration, networking, delivery and commissioning, in full compliance with US export regulations. If you have a workload and a date, send them: we come back with current allocation, lead time and a landed price rather than a waiting-list email.
Contact us about DGX allocation →

Spare parts and consumables for a running fleet

The cheapest place to save money on a GPU fleet is the parts you replace every quarter. Fans, thermal pads, power supplies, cables and risers are commodity items in Guangdong, and a batch of 50–500 units consolidated in our Guangzhou warehouse ships as one box.

Consumables and spares: what a fleet actually goes through
PartChina, ex-worksUS retailTypical batch
120×38 mm high-static PWM fans (Delta / Sunon or clone)$6$15–25100–500
12V-2x6 600 W cables, 16 AWG, extended$3$25–40100–1,000
Thermal pads 12.8 W/mK, 1.0–2.0 mm, for VRAM and VRM$2$10–2050–500
2,400–3,000 W CRPS server PSUs (Delta, Lite-On, HP-class), new$70–110$200–40010–100
PCIe 4.0 x16 shielded risers; x16 to 2×x8 / 4×x4 bifurcation cards; SlimSAS / OCuLink adapters$14 / $25 / $20$50–90 / $40–120 / $50–10050–500
SP3 / SP5 4U coolers, GPU air shrouds for blower cards$25 / $8$60–90 / none available20–100
Rack PDUs 30 A C19/C13, 1U shelves, cable management, 12–18U racks flat-packed$60 / $8 / $120$150–350 / $20–40 / $150–40010–50

What the landed cost really looks like

Ex-works prices are only half the number. A first mixed parts order we costed in September 2026, twenty 8-GPU chassis, ten 4-GPU chassis, twenty open frames, twenty power kits, cables, risers, fans, coolers and two EPYC platforms, came to about $12,100 ex-works and 845 kg. Sea freight at roughly $2.50 per kg added $2,100, and duty at 30% added $3,600, for a landed total of about $17,800: 1.47× the factory price. The same list at US retail is over $40,000. Cards are lighter but carry the same duty percentage on a much higher value, which is why they only make sense from China in batches of eight or more at wholesale prices. In every quote we send, freight and duty are already inside the number, DDP to your door.

How we deliver hardware

  1. We buy on China's domestic market as a Chinese company, in RMB, from the factories that make the chassis and power systems and from the workshops that build the 48 GB cards, not from export resellers. How that works and why the price is lower: How a China Sourcing Agent Works.
  2. Everything is tested before it leaves China. Cards get the burn-in protocol above. Chassis, PSUs and risers are checked against the spec and assembled as a dry run where the order is a full build.
  3. Consolidation and repacking in our Guangzhou warehouse. Chassis, power, cards and accessories from several factories leave as one shipment, packed to cut volumetric weight.
  4. Export, freight and duties handled. Sea for chassis and bulk parts, air or express for cards and urgent orders, DDP to your address. Duty classification is confirmed with a licensed broker before we quote.
  5. Assembly and burn-in in the USA for full builds, with vLLM or Ollama configured, and a 90-day Union Delta warranty on top of the 12-month factory warranty, serviced in the USA.
  6. Options that remove the risk of a first purchase: rent-to-own on the Inference Box from about $1,200 a month, a remote test drive over SSH before you commit, and trade-in from 24 GB to 48 GB cards.

Fees: for component orders and builds the service is priced as a percentage of the order, agreed before we buy, typically 5–8%, negotiated on larger orders. DGX-class systems are quoted as a project. The sourcing request itself is free: send the workload, the budget and the address, and you get a configuration, a landed price and a lead time within 24 hours.

Frequently asked questions

Is it legal to import memory-modded RTX 4090 48 GB cards into the USA?
Yes. US export controls restrict shipping high-end GPUs to China, not importing consumer-class cards from China into the USA. What applies is import duty on China-origin goods (budget about 30% and confirm the HS code with a licensed broker) and the fact that NVIDIA's warranty does not cover modified boards, which is why the factory warranty and our own testing matter.
How much cheaper is a GPU server built with parts from China?
Chassis, power kits, risers, fans and cables cost three to eight times less at the factory than at US retail, and still a third to a half less after freight and duty. A 4-GPU inference box with 48 GB cards lands at roughly $19,000–24,000 assembled, against $27,000–32,000 built from US retail parts; an 8-GPU 4U server lands at about $37,000 against $55,000–65,000 from US integrators. CPUs, RAM, NVMe and original NVIDIA professional cards cost the same everywhere and are better bought in the USA.
Can Union Delta supply NVIDIA DGX H200 or DGX B300 systems?
Yes, through our supply partners, in compliance with US export regulations, with rack integration, networking, delivery and commissioning. Allocation and lead time change month to month; send the number of GPUs, the workload and your target date and we come back with what is currently available and when.
Can I run a 4-GPU AI rig at home?
Yes, with a dedicated 240 V circuit, about 2.2 kW under load (roughly $220 a month in electricity at $0.13/kWh), a UPS and tolerance for blower noise. Eight cards at about 5 kW belong in a colocation cabinet at $1,200–1,800 a month.
What can a 192 GB VRAM box actually run?
A 70B-class model in FP16 with room for context, a 200B-class model quantised to 4-bit, or several smaller models served side by side with vLLM. It handles LoRA fine-tuning of 7B–70B models. Full training runs and long-context serving at scale are where H200 and B300 systems take over.
Should I buy hardware or stay on the cloud?
Stay on the cloud below about 70% utilisation or if your workload bursts a few hours a day. Above 80% utilisation, 24/7, your own box pays back in 12–24 months and often in 6–12 for a team already paying $2,000–4,000 a month. Data residency, latency and predictable cost are the other reasons teams move on-premises.
What warranty do I get?
A 12-month factory warranty on the cards written into our purchase contract, with RMA at the seller's cost, plus a 90-day Union Delta warranty on full builds serviced in the USA. Every card is burn-in tested before export and again after assembly. DGX systems carry NVIDIA's 3-year limited hardware warranty.