Is Running a Local LLM Worth It? A Real Cost Account

Reading Time: 6 minutes

The short answer

Most local LLM cost comparisons argue about the wrong number. A fully loaded Mac Studio with 256GB of unified memory costs about $10,899. Net of what you can resell it for, that works out to roughly $1,388 a year over six years. My own API bill is $776 a year.

So at my current usage, the machine costs about 1.8 times what I am already paying. The discount does not change that conclusion, and neither does shopping around.

That is not the interesting part. The interesting part is which numbers actually drive the decision, because almost every local-versus-API comparison I have read argues about the wrong one. This is the second article in a series. The first one covered memory: how to work out whether a model will run on your machine at all. This one covers money: whether it is worth running.

The electricity question is where most people start, and where most people stop. I will deal with it first so we can put it aside.

Three inputs, all from primary sources

Local LLM cost comparisons go wrong when the inputs are estimates. So all three of mine come from official sources, with the date I pulled them.

The machine. Mac Studio, M5 Ultra, 36-core CPU and 80-core GPU, 256GB of unified memory, 1TB of storage. Apple’s US configuration works out to roughly $10,899. I derived that from two published data points: a 256GB build with a 2TB SSD is $11,299, and the 36-core/80-core unit Apple sent reviewers, with 256GB and 4TB, is $12,299. Backing out the storage differences lands just under $10,900. The base M5 Ultra starts at $5,499.

Electricity. Two numbers, because where you live changes this section more than anything else in the article. In Dubai, where I am, the residential rate is 0.230 dirhams per kWh plus a 0.060 dirham fuel surcharge, plus 5% VAT. That comes to about 8.3 cents per kWh. In the United States the residential average was 18.34 cents per kWh in September 2026. That is 2.2 times my rate. California sits at 34.7 cents. Hawaii at 52.7.

Power draw. Apple publishes wall-measured figures. A Mac Studio with the same 80-core GPU draws 9 watts idle and 270 watts at maximum. Maximum, in Apple’s definition, means a compute test that saturates the processor, which is not what model inference looks like. I use 200 watts below and flag it in the limitations.

Electricity: real, but not the decision

Here is the annual electricity cost at a 200-watt inference load, priced at the US residential average.

  • Two hours a day: 218 kWh, about $40 a year, $240 over six years.
  • Eight hours a day: 637 kWh, about $117 a year, $700 over six years.
  • Around the clock: 1,752 kWh, about $321 a year, $1,928 over six years.

At my Dubai rate those become $18, $53 and $145 a year.

Local LLM cost and electricity: US power costs $40, $117 and $321 a year against Dubai's $18, $53 and $145
Annual electricity at three usage levels, US residential average versus Dubai. 200W inference load; the machine line is the base-case annual cost.

Read that against the machine’s annual cost of $1,388 and you see the problem with starting here. Even in Hawaii, even running flat out, electricity is a minority of the bill. At a normal eight-hour load in the US it is 8% of it.

This matters because it means the electricity argument, in either direction, cannot settle the question. People who say local inference is nearly free because power is cheap are describing 8% of the cost. People who say local inference is expensive because of the power bill are describing the same 8%.

I live in one of the cheaper electricity markets for this kind of work. If the machine does not pay for itself at 8.3 cents, it pays for itself less in California. That direction matters, and it points against local, not for it.

The number that actually decides local LLM cost: resale

Total cost is what you pay minus what you get back. So the resale assumption does almost all the work, and it is the number people guess at.

I would rather not guess. Swappa publishes completed sales. In September 2026, a 2022 Mac Studio with the M1 Ultra chip sold for $2,292 in the 1TB configuration, $3,088 at 2TB and $3,757 at 4TB. That chip launched at $3,999 in 2022. Four and a half years later it is holding about 57% of its price.

That looks encouraging, and it is also misleading, for two reasons.

First, memory and storage upgrades do not come back to you. Used-equipment buyers price the base machine and treat the upgrades as nice to have. A 256GB configuration has a much smaller pool of buyers than a standard one, because the only people who want 256GB of unified memory are people doing exactly the kind of work you are doing.

Second, and this cuts the other way: 2026 brought a global memory shortage on top of a buying wave from people running models locally. Apple discontinued the 128GB and 512GB tiers and raised the price of the 256GB upgrade. Shortages prop up used prices while they last and drop them when they end. High-memory machines with narrow buyer pools fall hardest.

So I use three scenarios instead of one number.

Annual cost of owning the machine

Six years, including six years of electricity at the US rate. Base price $10,899.

  • Optimistic, 45% residual: net $6,695, or $1,116 a year.
  • Base case, 30% residual: net $8,330, or $1,388 a year.
  • Pessimistic, 15% residual: net $9,965, or $1,661 a year.

Now put those next to what you spend on API calls today.

The step most comparisons skip

Buying the machine does not zero out the API bill.

An open-weight model running locally does not match a frontier lab’s model on the hardest tasks. What you actually get is a mix: the routine work goes to the local machine, and the genuinely difficult work still goes out to an API. So the honest comparison is not the machine against your whole API bill. It is the machine against the portion of your API bill it can actually replace.

Call that the substitution rate. Mine is not measured yet, so I run three values through it. Take a $776 annual API bill, six years, base-case resale of $8,330.

  • Spending flat, 70% replaced: you avoid $3,258. The API wins.
  • Spending grows 25% a year, 70% replaced: you avoid $6,113. The API still wins.
  • Spending grows 50% a year, 70% replaced: you avoid $11,283. Now the machine wins.
  • Spending grows 25% a year, 90% replaced: you avoid $7,859. Close to even.

Which gives a clean threshold. At a 70% substitution rate, your API spending needs to grow about 37% a year for the machine to break even. At 90%, the threshold drops to about 27%.

That is the question to ask yourself, and it is a question about your usage, not about hardware.

Local LLM cost break-even: at 70% substitution the machine breaks even at 37% annual spending growth
Substitution rate against annual growth in API spending. Where a curve crosses the grey line, the machine pays for itself.

The part that undercuts the whole exercise

Everything above assumes API prices hold still. They are not holding still.

There is a piece on the front page of Hacker News titled “tokens too cheap to meter.” Its argument is that the price of machine intelligence is falling by orders of magnitude per year with no sign of slowing, that models will become infrastructure rather than products, and that quality and access, not token count, will soon be the binding constraint. It goes further than I can verify, predicting frontier-quality models running locally on commodity hardware within three to six years.

I cannot tell you whether that timeline is right. I can tell you what it implies for this article: if the price of tokens keeps collapsing, the comparison I just did has a shelf life. Falling API prices push the break-even later. Cheaper local hardware pulls it earlier. Both are moving.

There is a concrete example going the wrong way for local. In independent testing by Artificial Analysis, Mercury 2.5 outputs 770 tokens per second at $0.25 per million input tokens and $0.75 per million output, with a 260K context window. On speed and on price, the hosted side is not standing still. Whatever case local deployment has, it will not be won on either of those two axes.

So the order should be reversed

If saving money is the goal, then the order that makes sense is the opposite of the order people usually use.

Measure your substitution rate first. Any machine you already own that can run a model will do for this. Take a month of real tasks, sort them by type, and find out which ones local can absorb and how much worse the output gets. That gives you a percentage rather than a feeling.

Look at your spending trend second. If your API bill has been flat for six months, no hardware price makes the machine rational. If it has been doubling, the machine looks very different.

Look at the price last. And when you do, look at net cost, meaning purchase price minus resale, against the portion of your spending it replaces. Not against the sticker price, and not against a per-token comparison that will be stale in a quarter.

If you are buying the machine for reasons that are not money, say so

Plenty of people run models locally when the arithmetic says not to. The reasons are usually real: privacy, working offline, not having a provider change its pricing or shut off a model you built on, controlling the whole pipeline, or simply not wanting your prompts to leave the building.

Those are legitimate. They are also not savings, and packaging them as savings makes the decision worse, because you end up arguing about a number that does not settle it. Decide which of the two things you are actually buying, and the hardware question gets much easier.

A cheaper version of the same experiment

If you want to run the experiment before committing, the used market is a much smaller bet. A used M2 Ultra machine with 192GB, which is plenty for very large open-weight models including long-context KV cache, sold for around $4,409 from a refurbisher. Net of resale over six years that is about $631 a year, below my current API bill. An M1 Ultra with 128GB lands around $522 a year.

For a US reader at typical electricity prices, that changes the local LLM cost conclusion entirely. The used machine can pay for itself at a flat spending rate, while the new one needs your spending to grow by more than a third a year.

There is a further consideration I want to state plainly because it is easy to skip. Any money spent on hardware is money not spent elsewhere. If you carry debt, compare the machine’s annual cost against what that money would save you by paying the debt down. In my case the two numbers are close enough that the hardware is not obviously the better use of the cash.

Limits

The inference-load figure of 200 watts is my estimate, not a measurement. Apple publishes idle and maximum and nothing in between.

The three resale scenarios are my judgment based on completed sales and used-market conditions, not a forecast. The sample of high-memory machines is small and the error bars are wide.

The substitution rates of 50%, 70% and 90% are not measurements. I picked them to show how much they matter. Yours will be different, and you should measure it rather than assume.

Every API price here is from September 2026. That assumption has a shelf life and it might be short.

This article counts money and not time. Running your own inference means maintaining it, handling version upgrades, and debugging models that will not load. That time is priced at zero here, which flatters the local option.

Sources

Buy me a coffee
Everything here is researched, fact-checked and written by one person. If it has been useful to you, you can buy me a coffee.
Scan with WeChat PayScan with WeChat Pay
Scan with AlipayScan with Alipay
Personal receipt codes · WeChat Pay & Alipay

本文采用 CC BY 4.0 许可。欢迎转载与引用,请注明作者并附上原文链接。
Licensed under CC BY 4.0. Quoting and republishing are welcome with attribution and a link back to this article.

发表评论