Local & private AI · 13 min read

What it actually costs to run AI on your own hardware.

By James Durkin, JDCS Updated 27 July 2026

Most articles about running AI on your own hardware are written by people selling hardware. This one has real Australian prices in it, and it also has the arithmetic that says most small businesses shouldn't buy any. If you're under about 30 seats, the honest answer is usually that the cloud is cheaper, and there are still good reasons to own the box anyway. Both of those things are in here.

The short version: hardware for a serious local setup runs from about $3,499 to $18,499 depending on how much memory you need. Electricity is a rounding error. Maintenance is not. On-premise pays for itself somewhere around 17 seats for a $5,000 build, 32 for a $15,000 build and 65 for a $40,000 build, which means a 10-person business buying hardware to save money will be disappointed. Buy it for control, for a compliance obligation, or because unmetered use changes how your team works.

What the hardware actually costs

Every price below was checked at Australian retailers in August 2026. They move fast at the moment, so treat them as a starting point rather than a quote, and confirm the current figure before you commit to anything.

  • Entry, around $3,500. A Mac Studio with an M4 Max and 36GB of unified memory sits at $3,499. That comfortably runs a good mid-sized open model for one or two people, and it's quiet enough to live on a desk.
  • Middle, around $6,200 to $7,000. A Mac Studio M3 Ultra with 96GB is $6,999. An ASUS Ascent GX10, which is the DGX Spark design with 128GB, was $6,249 at Centre Com. An RTX 5090 with 32GB was about $6,499 for the card on its own, before the rest of the machine.
  • Serious, $18,499. An RTX PRO 6000 Blackwell with 96GB was $18,499 at Umart. This is the tier where a small team genuinely gets concurrent, responsive access rather than taking turns.

A note on the shape of these machines, because it catches people out. The unified-memory boxes (the Macs and the GB10 machines) hold big models cheaply but run dense large models slowly. The discrete GPU cards are far faster per user but hold less. Which one is right depends entirely on the model you plan to run and how many people are running it at once, and getting that wrong is the most expensive mistake available in this category.

The 2026 memory shortage is the elephant in the room

DRAM prices rose by up to 98% in the first quarter of 2026, and further rises have been expected since. That has worked its way through to retail in ways that are hard to miss. Apple raised its Australian prices citing the RAM shortage. The RTX 5090 launched here at a $4,039 recommended price and has been selling around $6,499, roughly 61% above where it started. High-memory Mac Studios, the 256GB configurations, are effectively unavailable at any sensible price.

The practical advice follows from that: if you can defer six to twelve months, you probably should. Buying at the top of a shortage means paying a premium for an asset that depreciates anyway. Where the need is real and immediate, renting capacity or using an Australian-hosted option in the meantime is usually the better call, and it keeps your capital free while prices are unstable.

The power bill is the least interesting number

This is where a lot of the conversation goes, and it deserves about a paragraph. Using roughly 30 cents per kilowatt hour, which is a reasonable planning figure for Australian business electricity, annual running costs look like this:

  • Mac Studio M4 Max: about $273 a year.
  • DGX Spark class machine: about $342 a year.
  • RTX PRO 6000 workstation: about $1,025 a year.

Set that against $4,400 to $54,000 a year of API or per-seat SaaS spend, which is the range these builds are competing with, and the electricity stops mattering. Even the thirstiest option on that list costs less per year than two staff seats of a per-seat AI subscription. Stop over-indexing on the power bill; it isn't the number that decides this.

The break-even, stated honestly

Here's the part that other people leave out. These are worked examples using published prices and typical usage, not audited accounts from a particular client, and you should run yours with your own numbers. The on-premise figures include the hardware spread over three years, electricity, and maintenance time costed at a normal Australian rate. The cloud API figures assume a fairly heavy user, around five million input tokens and 600,000 output tokens a month, on mid-tier list pricing.

  1. Ten people. On-premise loses. Per-seat SaaS at $45 a month comes to about $5,400 a year. Cloud API on a mid-tier model comes to about $4,435. A modest on-premise build, including four hours a month of admin, lands near $7,700. Buying hardware here costs you money and buys you control, which may still be the right trade, but it is not a saving.
  2. Thirty people. Roughly a tie. SaaS around $16,200, cloud API around $13,306, on-premise around $14,050. At this size the answer turns entirely on how much admin actually happens. If maintenance is genuinely four hours a month, you're ahead. If it's ten, you're behind, and ten is common in the first year.
  3. One hundred people. On-premise wins clearly. Around $28,757 on-premise against roughly $44,352 on cloud API and $54,000 on per-seat SaaS. The build pays back in somewhere between 1.6 and 2.6 years, which is a real and defensible business case.

Reduced to seat counts, the crossover sits at roughly 17 seats for a $5,000 build, 32 seats for a $15,000 build, and 65 seats for a $40,000 build, measured against cloud API spend. The comparison against per-seat SaaS is a little kinder, because per-seat pricing is the more expensive way to buy. If you're a 12-person firm and someone has shown you a spreadsheet where a server saves you money, ask them what maintenance figure they used.

The costs that don't appear on the quote

The hardware invoice is the easy part, and it's rarely where projects come unstuck.

  • Maintenance: five to ten hours a month. Model updates, driver and runtime patching, the occasional CVE, someone's request that broke something. At consulting rates this is frequently larger than the hardware amortisation, which is why it's the number most proposals quietly leave out. Either budget for a retainer or train someone internally, and say which one out loud before you buy.
  • A UPS, roughly $400 to $2,500. Depends on what it has to hold up and for how long. A hard power cut mid-write is an unpleasant way to learn this lesson.
  • Circuit capacity. A 1,400 watt workstation is enough to trip a standard 10 amp office circuit once it shares that circuit with anything else, a kettle being the classic culprit. Worth an electrician's opinion before the machine arrives.
  • Backups of the right things. Model weights are re-downloadable and don't need backing up. Your fine-tunes, embeddings, vector databases and prompt libraries are the actual asset, they represent weeks of work, and they're what a sensible backup regime protects.
  • Noise and heat. Anything sustaining more than about 400 watts doesn't belong in a room where people take phone calls, and 1,400 watts will beat your air conditioning on an Australian summer afternoon. If a card comes in a lower-power variant with the same memory, and the RTX PRO 6000 does at roughly 300 watts instead of 600, spec that one for an office.
  • Who fixes it at 2am. The question everyone skips. A single box in a cupboard is a single point of failure, so never let on-premise be the only path. Keeping a cloud API fallback configured costs almost nothing when it's idle and turns an outage into a slow morning instead of a stopped business.

One thing to avoid outright: second-hand data-centre accelerators. Ex-fleet A100 and H100 cards look like a bargain and are a poor fit for a small business, between the power draw, the airflow they assume, the absence of any warranty, and carrier boards that don't go in a normal workstation. That's a false economy with a long tail.

The honest counterpoints, and the one that runs the other way

Three arguments against on-premise deserve to be taken seriously, because they're right.

  • Capability is not equivalent. A cost model that assumes the outputs are interchangeable is a dishonest cost model. Open models you can host are roughly one release cycle behind the best cloud models, which is invisible on some jobs and decisive on others. We've written up where that gap does and doesn't matter.
  • Idle hardware costs the same as busy hardware. A box running at 8% utilisation depreciates exactly as fast as one running flat out, while a cloud bill goes to zero when nobody types. Duty cycle is the variable that decides this, and most small businesses have a much lower one than they imagine.
  • Cloud prices keep falling on some tiers. A three-year model is competing against a future price, not today's. Some things have got dramatically cheaper. Others have gone up sharply, which is the subject of the AI price rises that have already happened, and honest planning has to hold both.

Then there's one counterpoint that runs the other way, and in practice it's the one that changes decisions. Unmetered usage changes behaviour. Under per-token or per-seat billing, teams self-censor. They don't run the whole archive through a summariser because someone will ask about the bill. They check twice before a big job. Once the meter is off, usage tends to climb steeply, and the work people do with it is often worth more than the arithmetic above. That's difficult to model and easy to observe, and it's a legitimate reason to own the hardware even when the spreadsheet says otherwise.

Bottom line: under about 30 seats, buy the hardware for control, for a compliance obligation, or for unmetered use. Don't buy it to save money, because at that size it generally won't. Above roughly 65 seats the financial case stands on its own. In between, the deciding number isn't the price of the machine, it's how many hours a month somebody spends looking after it.

Want the numbers run on your actual usage?

The first conversation is free, and it's just as likely to end with me saying stay in the cloud. You'll get a straight read on your break-even, what the build would involve, and what it would cost to look after. See pricing or the local AI page for how the work is structured.

Start a conversation

Cost questions, answered.

How much does it cost to run AI on your own hardware in Australia?
The hardware itself runs from about $3,499 for a Mac Studio M4 Max with 36GB, through roughly $6,249 to $6,999 for a 96GB to 128GB machine, up to $18,499 for an RTX PRO 6000 Blackwell with 96GB. Prices were checked in August 2026 and are moving quickly because of the memory shortage. Power adds a few hundred dollars a year, and admin time adds considerably more than that.
Is running AI locally cheaper than paying for ChatGPT or Claude?
Below about 30 seats, usually not. A 10-person business paying per-seat SaaS at $45 a month spends around $5,400 a year, or around $4,435 a year on cloud API at mid-tier list pricing, against roughly $7,700 a year for a modest on-premise build once admin time is counted. On-premise wins clearly at 100 seats. Between those two points it depends almost entirely on how much maintenance actually happens.
How much power does a local AI server use?
Less than people expect. At about 30 cents a kilowatt hour, annual electricity runs to roughly $273 for a Mac Studio M4 Max, about $342 for a DGX Spark class machine, and about $1,025 for an RTX PRO 6000 workstation. Against $4,400 to $54,000 a year of API or per-seat spend, the power bill is noise. It is the wrong number to agonise over.
How many staff do you need before on-premise AI pays for itself?
Roughly 17 seats for a $5,000 build, about 32 seats for a $15,000 build, and about 65 seats for a $40,000 build, measured against cloud API spend and including maintenance time. Those numbers move a long way if admin hours are higher or lower than assumed, which is why the maintenance estimate deserves more scrutiny than the hardware quote.
What are the hidden costs of on-premise AI?
Five to ten hours a month of maintenance is the big one, and at consulting rates that is often larger than the hardware amortisation. Then a UPS, enough circuit capacity (1,400 watts will trip a standard 10 amp office circuit), backups for the fine-tunes and vector databases that are the real asset, somewhere to put a noisy warm box, and an answer to who fixes it when it stops at 2am.