Pricing & cost justification

Cloud AI is a bill that never stops. Owning it is a cost that ends.

Every month you rent AI, the meter resets to zero. This page shows exactly what owning your AI costs - four costed build tiers, the honest local-vs-cloud trade-off, and interactive calculators that find your break-even month to the day.

Find the month your build has already paid for itself.

Move the sliders to your team. Watch the cumulative cost diverge and find your break-even month.

Three year cost comparison showing a cloud subscription line climbing continuously while an owned build flattens after the initial spend
A subscription never flattens. An owned build does.
Team size 20
Cloud cost / person / month RM 180
Years to compare 5
One-time local build (RM) RM 27,000

Build cost auto-suggests from the tier you'd need; drag to match a real quote.

You save over 5 years
RM 138,000
Cloud subscriptionLocal (VYROX)Break-even
Break-even
7 months
Cloud over period
RM 216,000

Illustrative estimate. Cloud ≈ ChatGPT Team / Claude Team seats (~RM130-280/user/mo) at ≈ RM4.6/USD; local figure includes hardware + commissioning; electricity ~RM1-6/day. Your exact crossover is calculated in the free audit.

Local, cloud, or hybrid? The straight trade-off.

No single answer is right for everyone. Here's the comparison without the spin - including when cloud or hybrid is genuinely the better call.

Swipe to see all columns

DimensionLocal / On-premCloud APIHybrid
Data privacyHighest - never leaves youLowest - sent to a third partyHigh - sensitive stays local
Recurring costLow & fixed (power + support)Variable - can balloonMixed - base + overflow
Upfront costHigher - hardwareNear zeroModerate
LatencyLow, predictable (your LAN)Internet + provider loadDepends on path
Offline operationYesNoPartial
Frontier reasoning ceilingVery good (hardware-capped)HighestBest of both
Scaling to spikesHardware-limitedElasticBurst to cloud
MaintenanceManaged by VYROXVendor-managedMost complex
When cloud is actually better
  • You need the absolute top frontier model and data isn't sensitive
  • Usage is very low or unpredictable - hardware would sit idle
  • You're prototyping before committing to hardware
  • Sudden massive spikes on-prem can't economically cover
When hybrid wins
  • Routine work runs local; a few hard tasks need a frontier model
  • Sensitive data strictly on-prem, non-sensitive bursts to cloud
  • You're migrating cloud → local and want a gradual cutover
  • You need an overflow valve for seasonal peaks
Our honest take

For most SMEs handling private or regulated data with steady daily volume, local pays for itself in months and removes per-token billing risk. If cloud or hybrid fits you better, we'll tell you - and build that instead.

Laptop and printed cost comparison report showing cloud versus owned build figures
Every recommendation starts with your numbers, not a sales script
Designed builds · costed

Four ready-to-deploy local-AI builds.

Complete, VYROX-commissioned systems - hardware sized, models loaded, runtime and agents wired in, staff trained.

Tier 01 · Solo desk

Desk AI

RM 9k-19k once
  • Mac Mini M4 Pro 48GB or RTX 4090 24GB + 128GB RAM
  • Qwen3-Coder 30B, DeepSeek R1 Distill 32B, Gemma 3 27B
  • ~30-45 tok/s single-stream
  • ~1-3 concurrent · team of 3-8
  • Ollama + Roo Code / Continue.dev
Replaces ~RM12k/yr cloud + Copilot. Payback ~9-19 mo incl. setup, depending on hardware.
Tier 02 · Team - popular

Studio AI

RM 22k-32k once
  • Mac Studio M4 Max 128GB (runs 70B) or RTX 5090 32GB (runs 32B-class)
  • Qwen3 32B, DeepSeek R1 70B (Mac), GLM 4.5 Air
  • ~60-100 tok/s on 32B single-stream
  • ~5-10 concurrent · team of 20-50
  • Ollama + Open WebUI + Cline / Aider
Replaces ~RM33k/yr cloud seats. Payback ~10-14 mo - then no per-seat licence.
Tier 03 · Heavy / agents

Engine AI

RM 55k-75k once
  • 2× RTX 5090 (64GB) or RTX PRO 6000 96GB, Threadripper, 256GB
  • Qwen3.6-35B, GLM-5.1, Kimi K2.6 / DeepSeek V4 (server)
  • 100+ tok/s aggregate (vLLM batching)
  • ~15-30 concurrent · 60-150 staff + agents
  • vLLM + Open WebUI + agent fleet
Agent token-burn alone runs RM100k+/yr in cloud. Payback ~6-9 mo.
Tier 04 · Org / frontier

Rack AI

RM 180k-500k+ once
  • 4-8× RTX 5090 / RTX PRO 6000 / H100 / H200 / DGX
  • DeepSeek V3.1 671B, Qwen3-235B, frontier open models
  • Hundreds tok/s aggregate
  • ~100+ concurrent · whole organisation
  • vLLM cluster + gateway + SSO + monitoring
Replaces RM300k+/yr enterprise AI - with full data sovereignty.

How "users supported" is calculated: concurrent users = free VRAM after model weights ÷ KV-cache per session (≈ 2 × layers × kv-dim × context × precision), capped by GPU throughput ÷ a 15 tok/s per-user floor and by the runtime's parallel slots (Ollama ~8, vLLM many). "Team size" assumes typical intermittent office use (~1 active generation per 5 staff). Numbers shown are at ~8K context with FP16 KV-cache; Q8 KV-cache roughly doubles concurrency and shorter context increases it further. Use the cost configurator to model your exact model, context and concurrency.

Four VYROX local AI hardware tiers lined up side by side, from a small desk mini PC to a full datacentre rack server, showing the physical scale difference between Desk AI, Studio AI, Engine AI and Rack AI builds
Same private AI stack, four sizes: DeskStudioEngineRack

The price on this page is a finished, working system, not a box of parts.

Every tier above is a fully commissioned deployment. Nothing here is an add-on you discover later - it is priced in from the first quote.

Open workstation chassis showing the GPU, memory and storage components included in a build
Every build is assembled & burn-in tested before delivery
Handover documentation, runbook and dual monitors showing the delivered AI system interface
Staff trained to run it day to day, no IT team required

Tell us your purpose. Get your build, cost and payback.

Pick what you'll use AI for and your team - the recommendation, hardware, model and 5-year savings update instantly.

What will you use AI for? pick all
Coding Research Writing Vision / docs 24/7 agents
People using it 10
Cloud cost / person / mo RM 180
Hardware preference
No preference Apple Mac NVIDIA GPU
Recommended build
Studio AI

A shared server for your team - big-model quality for everyone, fully private.

HardwareRTX 5090 / Mac Studio
ModelQwen3 32B / R1 70B
Memory32-128 GB
Users~5-10 concurrent
ToolsOllama + Open WebUI
RM 27k
One-time
11 mo
Break-even
RM 138k
5-yr saved
Get this build quoted free
Cost configurator · build it your way

Adjust the model, the hardware and the usage. See the full cost.

Two ways to use it: set how many concurrent users you need and leave hardware on Auto - we pick the cheapest build that serves them - or drive every variable yourself (LLM size, quantization, GPU, context, hours/day). The itemised build cost, capacity, payback and cloud savings update live. Indicative estimates; your exact quote comes from the audit.

1 · Primary purpose
Coding Research Writing Vision / docs RAG chatbot 24/7 agents
2 · LLM size 24-32B
3 · Quantization Q4_K_M
4 · Context length 8K
5 · Concurrent users needed 4

On Auto, we size the cheapest build that serves this many at your chosen model & context.

7 · Usage 8 h/day
8 · Support plan
Self-managed Standard Premium
Cloud you'd otherwise pay · seats 10
Cloud cost / seat / mo RM 180
Compare over 5 yrs
Fits - Qwen3 32B class on this build
Users this build supports
~6 concurrent · team of ~30
One-time build
RM 27,000
Payback vs cloud
11 mo
Running / yr
RM 2,400
Performance
~60 t/s
Saved / 5 yrs
RM 138k
Total one-timeRM 27,000
Cumulative cost - cloud vs your build
CloudYour buildBreak-even
Get this exact build quoted free

Indicative estimate. Hardware ≈ Malaysia street pricing during the 2026 GPU/DRAM shortage; commissioning, electricity (TNB commercial ~RM0.50/kWh) and optional support included as shown. Cloud baseline ≈ Team-tier seats (+ agent API where selected). VYROX confirms exact specs, prices and a measured savings projection in the free audit.

What each tier costs against cloud seats over 3 years.

Using the same RM 180/person/month cloud-seat assumption as the calculators above, matched to a typical team size per tier. These are worked examples to show the shape of the maths - use the calculators above for your own numbers.

Swipe to see all columns

TierTeam size (example)Cloud, 3 yearsLocal build + running, 3 years3-year saving
Desk AI5 peopleRM 32,400RM 18,200RM 14,200
Studio AI35 peopleRM 226,800RM 36,000RM 190,800
Engine AI100 peopleRM 648,000RM 87,500RM 560,500
Rack AI300 peopleRM 1,944,000RM 460,000RM 1,484,000

Worked example only, not a quote. Cloud = seats × RM 180/month × 36 months (no price rises modelled, which is optimistic for cloud - most vendors raise per-seat pricing over time). Local = mid-point of each tier's build-cost range, plus 3 years of electricity and standard-plan support at the rates used in the cost configurator above. Your exact 3-year number depends on your model, context length and concurrency - run it in the cost configurator or ask for the free audit.

Close-up flat illustration of a rising stack of monthly subscription invoices next to a single flat block representing a one-time local AI purchase, in white, green and teal
Stacking bills vs one flat cost

Cloud pricing is per seat. Local pricing is per machine.

A 5-person clinic and a 300-person operations floor pay the same cloud rate per head, so the bill scales linearly with headcount forever. A local build's hardware cost barely changes once you're serving the same team on the tier above - which is why the saving in the table compounds so much faster for bigger teams, not just proportionally more.

What it costs to run after you have bought it.

Once the build is commissioned there are only four things left that cost money: electricity, optional support, the occasional part, and eventually a hardware refresh. Model upgrades are not on that list. Here is each one, with the arithmetic shown.

1. Electricity, worked through

Illustrative worked example at RM 0.50 per kWh, the same commercial rate used in the cost configurator above. Malaysian commercial tariffs vary by band, by state and by demand charge, so treat this as a shape-of-the-number exercise and put your own last bill's rate in place of RM 0.50. The arithmetic is simply kWh per day x RM 0.50 x days.

Swipe to see all columns

If the machine usesPer dayPer month (30 days)Per year (365 days)Over 3 years
2 kWh / dayRM 1.00RM 30RM 365RM 1,095
4 kWh / dayRM 2.00RM 60RM 730RM 2,190
8 kWh / dayRM 4.00RM 120RM 1,460RM 4,380
12 kWh / dayRM 6.00RM 180RM 2,190RM 6,570

Illustrative example, not a tariff quotation. The RM 1 to RM 6 per day figure quoted elsewhere on this page is the same range seen from the other direction: at RM 0.50 per kWh it corresponds to roughly 2 to 12 kWh a day, which is the span from a small desk machine idling most of the day up to a multi-GPU build under sustained load. A machine that mostly idles between prompts sits near the bottom of this table, not the top, because a GPU only draws its full rating while it is actually generating.

An office electricity meter and a utility bill next to a running local AI server, illustrating the ongoing power cost of an owned AI build
The only meter that keeps running is the electricity one

2. What the tiers on this page imply per year

This is not new pricing. It is the 3-year table above with the build cost subtracted, so you can see what is left over as running cost. Build mid-point is the middle of each tier's quoted range. Running cost here includes both electricity and a standard support plan.

Swipe to see all columns

TierBuild mid-point3-year total (from table above)Implied 3-year runningImplied per year
Desk AIRM 14,000RM 18,200RM 4,200~RM 1,400
Studio AIRM 27,000RM 36,000RM 9,000~RM 3,000
Engine AIRM 65,000RM 87,500RM 22,500~RM 7,500
Rack AIRM 340,000RM 460,000RM 120,000~RM 40,000

Derived arithmetic from the 3-year worked example above, shown for transparency, not a separate price list. Rack AI uses the mid-point of the quoted RM 180k to RM 500k range; a build above that range carries a proportionally larger support line, because the support plan is a percentage of hardware cost. Self-managed teams remove the support percentage entirely and are left with only the electricity column from the first table.

3. The other two lines, and the one that is free

What is not included in the price, said plainly.

The tiers above cover the machine, the models, the runtime, the scoped integrations, training and a year of warranty. They do not cover everything a deployment can touch. Here is the honest boundary, so nothing on this list arrives as a surprise line on a later invoice.

Not in the quoted price
  • Electrical work: a dedicated circuit, socket upgrade or distribution board change if your room needs one
  • Room work: air conditioning, ventilation or a lockable cabinet if you do not already have them
  • A UPS (uninterruptible power supply) and any generator or transfer switching
  • Network cabling, switching and any structural LAN work
  • Your internet line, which the system does not need to answer prompts but does need for the initial model download
  • Third-party software licences for systems we connect to, such as your accounting or ERP platform
  • Digitising paper records or cleaning up a document store before it can be used as a knowledge base
  • Custom software built beyond the integrations agreed in the scope
Also worth knowing
  • Cloud frontier-model subscriptions, if you choose the hybrid route described in the comparison above
  • Ongoing content work: someone in your team still has to decide which documents are authoritative
  • Process change: the system answers questions, it does not rewrite how your team works
  • Prices in the tiers reflect hardware street pricing at quotation time, which moves with GPU and memory supply
  • Taxes and duties, which depend on your entity and are shown separately on the quotation

Optional add-ons, and when you actually need one

Each of these is quoted per project rather than carrying a list price, because the cost depends entirely on your site, your systems and how much of it already exists. Ask for any of them to be priced in the free audit rather than discovering them later.

Swipe to see all columns

Add-onWhen you need itWhen you can skip it
UPS and power protectionContinuous or overnight agent workloads, or a site with unstable powerOffice-hours use where a restart is a minor inconvenience
Second unit for redundancyThe system sits in a customer-facing or clinical path where downtime is not acceptableInternal productivity use with a manual fallback
Extra integrations over MCPYou want the model reading a system that was not in the original scopeThe first two use cases already cover most of the daily work
Document preparation and clean-upYour source material is scanned paper, mixed versions, or scattered across drivesYou already have a tidy, current, single-source document store
Extended support beyond year oneNo internal IT capability, or the workload is business-criticalYou have IT staff and are comfortable self-managing with the documentation
Additional training cohortsRolling out to new departments or high staff turnoverA single team that was trained at handover
Rack, cabinet and cooling workRack AI class builds, or any build going into a shared server roomA Desk or Studio class machine in a normal air-conditioned office

If it is not written in the quotation, assume it is not in the price.

That rule protects you from us as much as from anyone else. Ask any vendor, including VYROX, to put the exclusion list in writing next to the inclusion list. A quotation that only lists what you get is telling you half the story.

What actually moves your number up or down.

Two businesses of the same headcount can land in different tiers. These are the variables that decide it, all of which you can test yourself in the cost configurator above before you speak to anyone.

Swipe to see all columns

VariablePushes the price upPulls the price down
Concurrent usersMany people generating at the same moment, which needs more memory for session cache and more throughputIntermittent office use, where roughly one person in five is actively generating
Model sizeA 70B or larger model for the hardest reasoning workA 27B to 32B class model, which handles most business tasks
Context lengthLong documents held in memory per session, which multiplies cache use by every concurrent userShorter prompts, or retrieving the right passages instead of loading whole files
QuantizationFP16 or Q8 weights when maximum fidelity is requiredQ4_K_M, which is the common default and cuts the memory requirement substantially
Agent workloadAutomations running around the clock rather than only during office hoursHuman-initiated use during working hours only
Support planPremium at 18% of hardware cost per yearSelf-managed at RM 0, with your own IT staff and our documentation
RedundancyA second unit so there is no single point of failureA single machine with an accepted manual fallback
Integration scopeSeveral systems connected, each needing its own mapping and testingStarting with one document store and adding systems later
Platform choiceMulti-GPU NVIDIA builds for batching throughputApple Silicon unified memory, which reaches large models on a smaller budget at lower throughput

Each row corresponds to a control in the cost configurator above. Change one at a time and watch the itemised build cost, the capacity line and the payback figure move together, so you can see which trade-off is buying you what.

How to get this approved internally.

The hardest part of a local AI project is usually not the technology. It is turning a page of specifications into a request a finance lead or a board will sign. Here is the framing that works, the questions you will be asked, and where the answers already are.

The five questions a finance lead will ask

Swipe to see all columns

The questionHow to answer itWhere the number comes from
What are we spending today?Total your current AI and assistant subscriptions, plus any per-token API bills, per monthYour own invoices, then enter that as the cloud figure in the calculators above
When does it pay back?Give a month, not a feeling. The calculators produce a break-even month from your own inputsThe savings calculator and the cost configurator on this page
What is the recurring commitment?Electricity plus an optional support plan, and nothing else. State whether you are choosing self-managed, Standard or PremiumThe running cost section above, including the per-year table
What if it does not get used?Name the two use cases that justify it on their own, and who owns them. Do not promise organisation-wide adoption in month oneYour own pilot scope, agreed before purchase
What is the downside if we stop?You keep the hardware, which has resale or second-workload value. A cancelled subscription returns nothingThe resale question in the FAQ below

The budget line, written the way finance reads it

A one-page business case, in five lines

Copy this shape and fill it with your own figures from the calculators above. It fits on one page, which is the point.

  • Today we spend RM ___ per month on AI subscriptions across ___ people.
  • A one-time build of RM ___ replaces that, with RM ___ per year to run.
  • Break-even falls in month ___, and the ___ year saving is RM ___.
  • Our documents stop leaving the organisation, which addresses ___.
  • ___ owns it, starting with ___ and ___, reviewed at 90 days.

Straight answers on cost, contracts and what can change.

The questions we actually get asked before someone signs off on a build.

Why is the upfront cost so much higher than a monthly subscription?
Because you're buying the machine, not renting access to someone else's. A subscription's low entry price is the cost spread thin over every month forever - a local build's cost is concentrated into one payment, then drops to near zero. Look at the 3-year total, not the first invoice.
Are there hidden recurring costs after the build?
Two, both shown in every calculator on this page: electricity (roughly RM 1-6 per day depending on tier and usage hours) and optional support (Standard 10% or Premium 18% of hardware cost per year, or RM 0 if you self-manage). No per-token, per-seat or per-message billing.
Can we start small and upgrade later?
Yes. Most teams start at Desk or Studio AI and add a second unit or move up a tier as usage grows, rather than committing to Engine or Rack AI up front. Because the model and runtime are the same across tiers, moving up is a hardware swap, not a re-platforming project.
What happens if a newer, cheaper GPU comes out next year?
You benefit, we don't lock you in. The hardware is yours, so if prices drop or a better card launches, your next unit or upgrade simply costs less - there's no contract tying you to today's pricing or today's vendor.
Is financing or leasing available for the larger tiers?
For Engine AI and Rack AI builds we can structure the hardware portion through equipment financing or leasing partners, so the cash-flow profile looks closer to a subscription while you still own the outcome. Ask about this in the free audit.
Does the price include support if something breaks?
Every build carries a 1-year hardware warranty regardless of support plan. Beyond that, Standard and Premium support plans cover monitoring and priority repair; self-managed teams handle warranty claims and swaps themselves with our documentation.
How does the price compare if our team grows past the tier we bought?
You add capacity rather than re-buying everything: a second GPU in the same chassis, a second unit on the network, or a move to the next tier's platform. The 3-year table above shows why this is still cheaper than adding cloud seats at RM 180/person/month indefinitely.
What's the resale or trade-in value if we decommission later?
Server-grade GPUs and Apple Silicon workstations hold meaningful resale value, unlike a cancelled subscription which returns nothing. If you retire a build, the hardware can be resold, repurposed for another workload, or traded toward your next tier.
How much electricity does it actually use?
Work it out from your own tariff rather than a vendor claim. At an illustrative RM 0.50 per kWh, a machine using 2 kWh a day costs about RM 30 a month and one using 12 kWh a day costs about RM 180 a month. That is the same RM 1 to RM 6 per day range quoted in the calculators, seen from the energy side. A GPU only draws its full rating while it is generating, so a machine that idles between prompts sits near the bottom of that range.
Do we have to pay again when a better model comes out?
No. The models used are open-weight, so when a stronger one is released that fits your hardware you download it and swap it in at no licence cost. Support plans include the upgrade work; self-managed teams do it themselves with the documentation.
What is deliberately not included in the price?
Electrical and room work, air conditioning, a UPS, network cabling, your internet line, third-party licences for systems we connect to, digitising paper records, and custom software beyond the agreed integration scope. All of these can be quoted per project, and the exclusion list goes in writing next to the inclusion list on every quotation.
How do I present this to our finance lead or board?
Split it into two budget lines, one for the one-time build and one for the annual running cost, and put your current subscription spend directly beside them. For most teams this is a capital purchase that retires a recurring expense rather than new money in full. Bring a break-even month from the calculators on this page, name the two use cases and the person who owns them, and have your accountant confirm how the hardware is treated in your books.
Why do two companies of the same size get different quotes?
Headcount is not what sizes the machine. Concurrent users, model size, context length, quantization, whether agents run around the clock, redundancy and integration scope all move the number, and every one of them is a control in the cost configurator above. Change one at a time and you can see exactly which trade-off is buying you what.
Your move

Stop renting your AI. Own it by next quarter.

Book a free 45-minute Local-AI Audit. We measure your current cloud spend, spec the exact build, and give you the costed break-even date - in writing, no obligation.

  • Free, 45 minutes
  • Costed break-even date
  • No obligation

No deck pitch. Just engineers sizing your build.

Chat with VYROX AI on WhatsApp Free Local-AI audit