SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

What does an AI server cost to build?

The accelerator sets most of the bill. Learn how VRAM, memory bandwidth, power headroom, cooling and your own tariff decide what the rest costs.

What an AI server costs to build

What an AI server costs to build is decided by one line on the invoice: the accelerator. On a single-card machine the GPU is usually more than half the total, and often more than everything else added together. The rest is ordinary desktop hardware at ordinary desktop prices.

That leaves one part with two questions attached. The VRAM (video memory, the memory soldered onto the card) decides which models will load at all. The memory bandwidth decides how fast tokens come out when one person is using it. Headline compute figures decide much less than the marketing implies, for reasons this guide works through below.

You will not find prices here. Card prices and electricity tariffs move faster than any guide, so the numbers have to come from you: the vendor's spec page for VRAM and bandwidth, a retailer for the price, your last electricity bill for the tariff. The structure is what stays true.

The shape of the bill

Order the parts list by how much each item can hurt you, not by what it costs.

  • The accelerator. The only part where the choice is genuinely hard, and the part that sets the bill.
  • The power supply. The cheapest place to save money and the most expensive place to be wrong, because a supply that fails under load can take the card with it.
  • Board and CPU. For one card, a current mid-range desktop platform is enough. The board matters for slot spacing and PCIe lanes, not for speed.
  • System RAM. Buy at least as much system RAM as the card has VRAM. Loaders read weights through the page cache, and any CPU offload needs somewhere to live.
  • NVMe storage. Model files run to tens of gigabytes each and you will keep several. Disk speed shows up as load time, not as tokens per second.
  • Case, fans and cables. Cheap, and the difference between a machine you can sit next to and one you cannot.

Electricity never appears on a parts list, because nobody buys it once. It is the only line that keeps charging you after the build is finished, and across a few years it can rival a component.

VRAM decides which models will load at all

Weights, the KV cache (key/value cache, the running state for the tokens already in the context), the CUDA context itself and a working margin all live in VRAM. If the total does not fit, one of two things happens. A server like vLLM stops with CUDA out of memory and never answers a request. A loader built on llama.cpp keeps going and puts the leftover layers in system RAM, where the CPU reads them over a small fraction of the bandwidth, so the model replies at a speed that feels broken. In Ollama, ollama ps names the split in its PROCESSOR column: 100% GPU is what you want, and any figure that mentions CPU means you are paying for a card you are not fully using.

So size the VRAM against the model rather than the other way round. Turning a parameter count into gigabytes is arithmetic, and the arithmetic that turns parameters and quantisation into gigabytes is a whole post of its own. Two things get forgotten when people do that sum. The KV cache grows with context length, so a long context can cost more than the weights do. And every concurrent request carries its own cache. VRAM that fits the weights exactly leaves no room for either.

Quantisation is the other lever, and the cheapest one: a lower precision copy of the same model is smaller, so it both fits and reads faster, at some cost in output quality. Where that quality cost becomes visible between q4, q8 and fp16 decides how many gigabytes you can avoid buying. Once you have a VRAM budget, which models that budget can actually serve is the list to check before you order anything.

Memory bandwidth, not compute, decides tokens per second

Generating text for a single user is a memory-bound job. To produce one token the GPU reads every weight it needs, once. Then it does the same thing again for the next token. The ceiling on speed is therefore bandwidth divided by the bytes it must read, and arithmetic throughput does not move that ceiling.

Here is the sum, with round numbers chosen to be easy rather than to describe a real product. Take a model occupying 20 GB of VRAM on a card the vendor rates at 800 GB/s. 800 divided by 20 is 40, so 40 tokens per second is the ceiling. Real output lands below it, because attention reads the KV cache as well as the weights. The ceiling still tells you the shape of the problem: halve the model size and the speed roughly doubles, double the bandwidth and the speed roughly doubles, and doubling the compute changes nothing here.

Compute does matter, in two places. Prompt processing, also called prefill, reads the whole prompt at once and is a dense matrix job, which is why a long document on a weak card is slow before the first token appears. And serving several requests at once reads each weight once for the entire batch, which turns a memory-bound job into a compute-bound one. If the box will serve more than one person, the concurrency settings change the picture, and how OLLAMA_NUM_PARALLEL and the request queue behave under load is where that trade lives.

One trap while measuring. nvidia-smi reporting 100% utilisation during generation does not mean the card is compute-bound. That counter reports whether any kernel was running, not whether the arithmetic units were busy. A memory-starved card reads 100% while doing very little work.

Two cards do not double single-user speed when the layers are split across them. A token passes through the first card's layers, then the second card's, in sequence. Two cards buy capacity, so a larger model fits, and they buy throughput for concurrent users. They do not make one reply appear twice as fast.

Power supply headroom, and the connector nobody seats fully

A card's rated board power is an average, not a limit. Modern cards draw short spikes well above the rating, in the millisecond range, and a supply that looks adequate on paper can trip its own protection on those spikes and switch the machine off. The symptom is a reboot with nothing in the logs, always under heavy load, never at idle. Size the supply above the sum of sustained draws with real room left over, and prefer a unit whose specification says something about transient behaviour.

The connector is its own failure mode. The 12V-2x6 plug, and the 12VHPWR plug before it, carries the card's whole power budget through a small set of contacts. Push it in until the latch clicks. A plug sitting slightly out means the same current flows through fewer contacts, and the contacts that are left get hot enough to discolour the housing. Use the cable that came with the supply instead of a chain of adapters, and leave enough space that the cable is not bent hard right at the plug.

You can also lower the ceiling on purpose. nvidia-smi -q -d POWER prints the card's minimum and maximum enforced limits, and sudo nvidia-smi -pl 300 sets the limit to 300 W. Do not guess at the effect: set the limit, run the same prompt, compare tokens per second. Generation is memory-bound, so the clock speed you give up usually costs less than the watts you save, but measure it on your own model before you trust that. The limit resets when the driver reloads or the machine reboots, so wrap the command in a small systemd unit if it is part of the plan.

nvidia-smi -q -d POWER
sudo nvidia-smi -pl 300
nvidia-smi --query-gpu=name,power.draw,temperature.gpu,clocks.current.sm --format=csv

The last command is the one to keep in a second terminal while a model runs. Power near the limit with the temperature flat is healthy. Clocks falling while the temperature climbs is not.

Cooling and noise, when the machine sits near where somebody sleeps

Heat is how a cheap case becomes expensive. A card pulling 300 W puts 300 W of heat into the room it sits in, and in a small closed room you will feel that within an hour.

The cooler design decides how much freedom you have. A consumer card with open fans dumps heat inside the case, so the case has to move it out, and two intake fans with one exhaust is usually enough as long as the front panel is not solid. A datacentre card is often passive, with no fan at all, because it expects a rack chassis forcing high pressure air straight through its fin stack. Put a passive card in a quiet tower and it heat-soaks and throttles. The fix is loud fans ducted at the card, which defeats the reason you bought a quiet tower.

Check instead of assuming. nvidia-smi -q -d PERFORMANCE lists the reasons the card is holding its clocks down, and a thermal reason is a different problem from a power cap. A card that throttles after ten minutes of steady load has a cooling problem, not a silicon problem.

Noise is a real cost, because the fix is money: a larger case and bigger, slower fans. If the machine has to live in a bedroom, take the power limit instead and accept the lost speed.

PCIe lanes, slot spacing, and the second card you might buy later

For one card doing inference, PCIe link width barely matters. The weights cross the bus once, at load time. After that the card works out of its own memory, so an x8 link instead of x16 shows up as a slower model load and almost nothing else.

It starts to matter with two cards. A mainstream desktop CPU offers a limited number of lanes, so a second card usually drops both slots to x8, or hangs off chipset lanes shared with your NVMe drives. That is fine for splitting layers across cards, where the traffic between them is small. It is not fine for tensor parallelism, which moves data between the cards on every layer and wants the widest link you can give it.

The physical problem is harder than the electrical one. Two triple-slot cards need a board whose x16 slots sit far enough apart, a case tall enough to hold them, and a supply with two native cable runs. Decide this before buying the board, because changing a motherboard later means rebuilding the whole machine.

Check what you actually negotiated, and check it under load rather than at idle, because the link downshifts to save power when the card is quiet:

nvidia-smi --query-gpu=pcie.link.gen.max,pcie.link.gen.current,pcie.link.width.current --format=csv

Electricity is watts times hours times your own tariff

This is the line a build spreadsheet always leaves out, and the only one that never stops.

Do the multiplication with your own figures. A card drawing 350 W, plus 100 W for the rest of the box, is 450 W, so 0.45 kW. Four hours of real work a day is 1.8 kWh. Multiply that by the per-kWh rate printed on your last bill, then by 365, and you have the yearly running cost at that duty cycle.

Idle is the part that surprises people. The same machine left powered on around the clock, idling at 70 W, burns 1.68 kWh before it produces a single token. Across a year that often costs more than the busy hours do, because there are so many more of them. Measure your own idle draw at the wall with a plug meter, since the driver's reading covers the card alone and not the board, the drives or the supply's own losses.

Two honest ways to cut it. Suspend or power the box off when nobody is using it, and accept the wait on first use. Or give the machine a second job so the idle watts are doing something, for example letting the same card handle Jellyfin hardware transcoding between prompts.

Consumer card or datacentre card

Treat this as a trade with named consequences. Each side gives up something specific.

A datacentre card gives you ECC (error correcting code) memory, and nvidia-smi -q -d ECC reports corrected and uncorrected error counts, so you learn that memory is going bad before it ruins work. Most consumer cards return N/A there. Without ECC, a flipped bit in VRAM becomes wrong output with nothing in any log, or a crash you cannot reproduce. For a chat assistant that is a shrug, because you notice and ask again. For a long unattended job you would have to rerun from the start, it is expensive.

Datacentre cards are also built for rack airflow, come with vendor support terms, and often carry more VRAM on one card than anything on the consumer shelf, which is the real reason most people buy them.

Consumer cards give you a retail warranty you can use without a support contract, a deep second-hand market when you upgrade, and cooling designed for the kind of case you already own. The catch is contractual rather than technical: the licence on the consumer driver has long restricted data centre deployment, and the exact terms change over time. Read the licence on the driver download page before you build something other people pay to use. It has nothing to say about one card in your own home.

Budget a day for the driver whichever you pick. On a headless server the driver, the kernel module and Secure Boot do not always agree on the first attempt. Two failures are common enough to have their own guides: the driver that installs but never loads on Ubuntu Server, and the module whose key is rejected because Secure Boot will not trust it.

Used or new

A used card is the largest single saving available on this build, and it moves the risk onto you.

What you give up is the warranty, any claim if it dies next month, and every bit of knowledge about how it was run. A card out of a render farm or a mining rig has spent its life at full load. Fans and thermal pads are consumables, so a card with dried pads throttles under sustained work even though it looks clean in a photograph.

What to do on arrival, inside whatever return window the seller gave you. Load it hard for an hour, not for two minutes, and watch temperature, clocks and power with nvidia-smi dmon. Clocks that drop in the first ten minutes and stay down mean the cooler is not doing its job. Then run a model you already know and read the output: coherent text from a model that works elsewhere is a better memory check than any quick benchmark, and garbage tokens point at bad VRAM. A card that survives an hour of steady load usually has years left in it, because the parts that wear out are the fans and the pads, and both can be replaced.

Resale belongs in this sum too. Consumer cards sell on easily, so the real cost of a card you keep for two years is what you paid minus what you get back. Datacentre cards have a much thinner second-hand market, so plan on keeping one for its whole working life.

A self-built AI box is not a general purpose VPS

It is worth being blunt about what you end up owning. A machine built around an accelerator is a specialised tool. It sits in your home, on your electricity and your internet connection. It has no out-of-band console, no snapshot, no second power feed, and no hands in a data centre at 03:00. It is excellent at the job you built it for, and it is a poor place to put anything that other people need to reach.

The money question is utilisation. Hardware you own costs the same whether it works or idles, so the cost per token falls the more you use it and is at its worst while the machine is powered on doing nothing. Rented capacity behaves the opposite way. If your real usage is a few hours a week, renting a GPU VPS by the month or paying per token wins, and the point where that flips is arithmetic with your own numbers in it. Where a GPU server breaks even against API tokens does that calculation properly, so there is no reason to guess at it here.

Build when you want the model on hardware you control, when the data must not leave the building, when the work is long and steady enough to keep the card busy, or when learning the stack is itself the point. Those are good reasons. Saving money at low utilisation is not one of them.

FAQ

How much does it cost to build an AI server?

There is no single figure, and any guide that prints one is wrong within a few months. The shape of the bill is stable instead: the accelerator is usually more than half the build, the power supply and case are small next to it, and electricity is a recurring cost that the spreadsheet forgets. Price your own build by choosing the VRAM you need first, reading that card's price and memory bandwidth from the vendor's own pages, then adding a mid-range desktop platform around it plus a year of electricity at the rate on your bill.

Is more VRAM better than a faster GPU?

For deciding what you can run, yes. VRAM is a hard gate: a model that does not fit either refuses to load with CUDA out of memory, or spills into system RAM and answers at a speed that feels broken. Speed is a separate axis, and for a single user it is set by memory bandwidth rather than by compute. So read a spec sheet in that order: VRAM as a yes or no test, then bandwidth as the speed number, then compute last.

Do I need ECC memory in an AI server?

Not for personal use. Without ECC, a flipped bit in VRAM becomes wrong output with nothing written to any log, and for an interactive assistant you notice and ask again. It matters when a job runs unattended for hours and silent corruption means rerunning all of it, or when you cannot tell a corrupted answer from a good one by reading it. Cards that support ECC report error counts through nvidia-smi -q -d ECC, and most consumer cards return N/A there.

Is it cheaper to build an AI server or to rent a GPU?

It depends on how many hours the machine actually works. Hardware you own costs the same whether it is busy or idle, so owning wins at high steady utilisation and loses badly at a few hours a week. Renting and per-token pricing charge only for what you use, so they win below that line. Put electricity and your own maintenance time on the owned side of the comparison, because those are the two costs that never show up in a purchase price.

#ai-server#gpu#vram#self-hosted-llm#hardware-cost