All Webbed Labs

Open-weight models vs API models: which should Australian enterprises use?

Last updated Published by All Webbed Labs How we write

The short answer

In the open-weight vs API LLM decision, use a hosted API model by default: the strongest models are API-only, you pay only for what you use, and several can now process data inside Australia through cloud platforms. Choose an open-weight model when you need prompts to never reach a model vendor, when you must run offline or in an isolated environment, when steady high volume makes owning capacity cheaper, or when you need to fine-tune and control the exact model version. Many enterprises end up using both, routed by data sensitivity.

Key takeaways

  • Open-weight means the trained weights are published so you can run the model on your own infrastructure. It doesn't always mean open source: licences such as Llama's carry conditions.
  • API models still lead on the hardest reasoning and coding tasks, but open-weight models such as gpt-oss, Qwen and Mistral handle most classification, extraction and summarisation work well.
  • Self-hosting gives the strongest sovereignty position: no model vendor ever sees your prompts, and you hold the logs.
  • Self-hosting only beats per-token pricing at steady, high volume. GPUs cost money while idle, and you take on operations, patching and evaluation.
  • A middle path exists: open-weight models served by a cloud platform such as Amazon Bedrock in Sydney, where the model maker never sees your data and you don't run GPUs.

Open-weight vs API LLMs: which should you use?

Start with a hosted API model unless you have a specific reason not to; move to open-weight when sovereignty, isolation, volume or control demands it. The decision is less about ideology and more about four practical questions: how capable does the model need to be, where may the data go, how much volume will you run, and who will operate it.

The terms first. An API model (Claude, GPT, Gemini) runs only on its vendor’s or a cloud partner’s infrastructure; you send a request and pay per token. An open-weight model (Llama, Mistral, Qwen, DeepSeek, OpenAI’s gpt-oss) has its trained weights published, so you can download it and run it wherever you like. “Open-weight” isn’t the same as open source: the training data and code usually aren’t released, and some licences restrict use. Buyers searching for an “open source LLM”, a “self-hosted LLM”, a “local LLM” or a “private LLM” usually mean an open-weight model run on infrastructure they control; API models are sometimes called closed or proprietary models.

How do they compare side by side?

API models win on capability and simplicity; open-weight models win on control and, at scale, on cost. Details reflect vendor pages as at 28 September 2026.

Hosted API modelOpen-weight, served by a cloud platformOpen-weight, self-hosted
ExamplesClaude, GPT, Gemini via vendor API or Bedrock, Azure, Vertex AIgpt-oss, Qwen, Mistral on Amazon BedrockAny open-weight model on your own GPUs
Capability ceilingHighest: frontier models are API-onlyGood to very goodGood to very good; you pick the exact model
Who sees your promptsModel vendor or cloud providerCloud provider, not the model makerOnly you (and your cloud provider’s infrastructure)
Australian processingSome models, on some platforms, at the time of writingSome models listed in SydneyYes, in any Australian region with GPU capacity
PricingPer tokenPer tokenPer GPU hour, whether used or not
OperationsNoneNoneYou run serving, scaling, patching, monitoring
Model version controlVendor retires versions on its schedulePlatform retires versions on its scheduleYou keep a version as long as you like
Fine-tuningLimited, on selected modelsVaries by platformFull control
Offline or air-gappedNoNoYes

Where do API models win?

On raw capability, speed to launch, and low or unpredictable volume. The best models for complex reasoning, long documents, coding and agentic work are, at the time of writing, available only through APIs. You pay per token, so a pilot costs almost nothing, and you inherit the vendor’s scaling, safety work and continuous improvements.

Residency, once the main objection, has narrowed. Several API models can now process data inside Australia through Amazon Bedrock, Azure or Google Vertex AI, though availability varies by model and changes often. Our guide to which models can run onshore tracks the detail, and the Bedrock vs Azure vs Vertex comparison covers the platforms.

Choose an API model if:

  • The task needs top-tier reasoning, writing or coding quality.
  • Volume is low, spiky or unknown.
  • You don’t have, and don’t want, people to operate GPU infrastructure.
  • An Australian-region API option exists for the model, or offshore processing has been assessed and accepted.

When should you use a self-hosted LLM?

When control matters more than the last increment of capability. Open-weight models win on control: over where data goes, which exact model runs, how it’s tuned, and whether it runs at all without an internet connection. When you self-host, prompts and outputs never reach a model vendor, and your own logs prove where every request went.

Situations where that matters:

  1. Data that can’t go to a third-party AI service under contract, client instructions, legal privilege or government classification.
  2. Isolated or offline environments, such as secure facilities, remote sites with poor connectivity, or on-premises deployments.
  3. Version stability. API vendors retire models; a self-hosted model behaves identically until you choose to change it, which matters for validated workflows.
  4. Deep customisation. Full fine-tuning, custom decoding, or distilling a small specialised model.
  5. Steady, high volume of simpler tasks, where owned capacity undercuts per-token prices.

Open-weight quality is good enough for much enterprise work. Many tasks, such as extracting fields from invoices, classifying emails or answering questions over retrieved documents, don’t need a frontier model. OpenAI’s gpt-oss-120b, for example, has 117 billion parameters and fits on a single 80 GB GPU; gpt-oss-20b runs within 16 GB of memory.

What about licences?

Most popular open-weight models allow commercial use, but the terms differ, so check before you build a product on one.

Model familyLicence at the time of writingThings to note
OpenAI gpt-ossApache 2.0Permissive, commercial use allowed
Qwen3 (for example Qwen3-32B)Apache 2.0Check each model card; licences vary across releases
Mistral (Small 4, Large 3, Ministral 3)Apache 2.0Mistral Medium 3.5 uses a modified MIT licence
Meta Llama 4Llama 4 Community LicenseAttribution (“Built with Llama”), acceptable use policy, separate licence above 700 million monthly active users

Also consider provenance. Some organisations, particularly in government and critical infrastructure, have policies on the country of origin of AI models. That’s an organisational risk decision, but it needs making explicitly before you pick a model family.

When does self-hosting actually save money?

Only when a GPU is kept busy. Per-token pricing charges nothing when idle; a GPU instance charges every hour it runs. The break-even is simple arithmetic.

Here is an illustration using round, clearly hypothetical numbers. Use the live Bedrock, model vendor and EC2 pricing pages in the sources for real figures in your region.

StepIllustrative assumptionResult
GPU instance able to serve your chosen model$10 an hour (AUD)
Running 24/7$10 × 730 hours$7,300 a month
Add operations effort0.2 of an engineer at $200,000 a year loadedabout $3,300 a month
Total monthly cost of self-hostingabout $10,600
API price for a comparable model$5 per million tokens, blended input and output
Break-even volume$10,600 ÷ $5 per millionabout 2.1 billion tokens a month

Below the break-even, the API is cheaper; above it, owning capacity wins, provided the instance can actually serve that throughput. Real deployments shift the numbers: you may need redundancy (two instances), savings plans reduce GPU cost, prompt caching reduces API cost, and a smaller open-weight model may do the job on a cheaper GPU. Our LLM running costs guide works through these levers in detail.

The practical conclusion: for most organisations starting out, the API is cheaper. Self-hosting is justified first by control, and only later, if ever, by cost.

Is there a middle path?

Yes: run an open-weight model through a managed cloud platform in an Australian region. At the time of writing, AWS’s Bedrock regional table lists Sydney for several open-weight models, including gpt-oss-120b and models from Qwen and Mistral, while Llama 4 models are listed only in US regions. You pay per token, AWS operates the infrastructure, and the model maker never sees your data. You’re still on a foreign-owned cloud, which matters for some sovereignty definitions but not most commercial risk assessments.

The other common pattern is routing by sensitivity: a single internal gateway sends restricted data to a self-hosted or onshore model and everything else to the most capable API model. That gives most users frontier quality while keeping the sensitive lane under your control.

A decision guide

If this is trueLean towards
You need the best available reasoning or coding qualityAPI model
Volume is low or unpredictableAPI model
An Australian-region option exists for the model you needAPI model via that platform
Prompts must never reach a model vendorSelf-hosted open-weight
You must run offline or in an isolated networkSelf-hosted open-weight
Steady, very high volume of simpler tasksOpen-weight, self-hosted or managed
You need a frozen model version for a validated processSelf-hosted open-weight
You want onshore processing without running GPUsOpen-weight on a managed platform in Sydney
Mixed data sensitivityBoth, behind one routing gateway

If fine-tuning is the reason you’re considering open-weight models, note that some API providers also offer fine-tuning for selected models, though you don’t get the weights and the options are narrower. Read RAG vs fine-tuning first: retrieval often solves the problem more cheaply than either.

How All Webbed Labs approaches model choice

We test candidate models, API and open-weight, against your own examples before recommending one, and we write down the model, platform, region and date in the architecture decision. Where sovereignty requires it, we deploy open-weight models in your own cloud account in an Australian region, with logging you control. Where it doesn’t, we’ll usually recommend an API model and say why. See our private LLM deployment service and LLM integration work.

Frequently asked questions

Are open-weight models as good as GPT, Claude or Gemini?

For the hardest tasks, generally not at the time of writing: the frontier models are API-only. For well-defined tasks such as extraction, classification, routing, summarisation and retrieval-augmented answers over your own documents, a good open-weight model is often good enough. Test on your own data rather than relying on public benchmarks.

Is it legal to use open-weight models commercially in Australia?

Usually yes, but read the licence. gpt-oss, the Qwen3 models and many Mistral models are released under Apache 2.0, which permits commercial use. Meta's Llama 4 Community License allows commercial use but requires attribution, compliance with Meta's acceptable use policy, and a separate licence above 700 million monthly active users.

What is the best open source LLM for business?

There isn't one best model; it depends on the task, the hardware you can run and the licence you can accept. At the time of writing, OpenAI's gpt-oss, the Qwen3 family and Mistral's Apache 2.0 models are sensible starting points because their licences permit commercial use. Shortlist two or three and test them on your own examples.

Is an open-weight model the same as an open source LLM?

Not quite, though people use the terms interchangeably. Open-weight means the trained weights are published so you can run the model yourself. Open source, strictly, would also cover the training code and data, which usually aren't released, and some open-weight licences, such as Llama's, add conditions an open source licence wouldn't.

Can we run an LLM on-premises or locally instead of in the cloud?

Yes, with an open-weight model. You can run a private LLM on your own servers, in a secure facility or fully offline, which API models don't allow. You take on the hardware, serving, patching and evaluation, so it makes sense when isolation is a requirement rather than a preference.

Does self-hosting take care of our Privacy Act obligations?

No single technical choice does that. Self-hosting in an Australian region removes the overseas disclosure question for the model itself, which simplifies your APP 8 analysis, but you still need to meet the other Australian Privacy Principles for collection, use, security and access. The OAIC's guidance on AI products is a good starting point.

What hardware does an open-weight model need?

It depends on size. OpenAI's gpt-oss-20b runs within 16 GB of memory, and gpt-oss-120b fits on a single 80 GB GPU thanks to quantisation. Larger models need several GPUs. Cloud GPU instances are available in Australian regions, subject to quota.

Sources

  1. Regional availability by models (Amazon Bedrock User Guide) , Amazon Web Services
  2. gpt-oss-120b model card , OpenAI on Hugging Face
  3. Llama 4 Community License Agreement , Meta
  4. Models overview , Mistral AI
  5. Qwen3-32B model card , Qwen on Hugging Face
  6. Amazon Bedrock pricing , Amazon Web Services
  7. Amazon EC2 On-Demand pricing , Amazon Web Services
  8. Guidance on privacy and the use of commercially available AI products , Office of the Australian Information Commissioner
Let's Build Something Extraordinary

Ready to Transform Your
Technology Operations?

Join the Australian businesses trusting All Webbed Labs to deliver their most critical software projects. Let's talk about what we can build together.

Free 30-minute strategy call
No commitment required
Response within 1 business day
NDA available on request