[SEC.INSIGHTS-REF.2026]
Back to Insights
Capability

You Don't Pick a Local AI Stack. You Grow Into One.

September 20, 2026
You Don't Pick a Local AI Stack. You Grow Into One.

You don't choose one of these tools. You grow through them.

Bottom Line: The local AI conversation is dominated by tool comparisons that miss the point. You don't pick an inference engine once and declare a winner; you grow through them in stages as your work shifts from rapid experimentation to parameter control and sovereign application pipelines.

There's a genre of local AI article built on one assumption: pick a tool, declare a winner. I published an Ollama vs LM Studio piece earlier this year, and I still stand by its conclusion: the choice is workflow-fit, not quality.

But writing that piece, and everything since, taught me the framing was still wrong. You don't choose one of these tools. You grow through them. Each one answered a question the previous one surfaced, and where you stop depends on how deep your use has gone.

I learned that by going through all three, in order.

Stage one: Ollama got me running

Ollama is the fastest path to a model running on a home PC. Install it, pull a model, you're chatting in minutes. It gives you a simple chat UI, or the terminal if that's your thing, and it hosts models so your apps and tools can connect to them, the same way the bigger engines do.

It also reaches beyond your hardware. Ollama's cloud side hosts larger open models: free rate-limited trials when I started, formal subscription tiers now ($0 to $100 a month for individuals at the time of writing). When a new model drops and you want ten minutes with it before committing, that click-and-go path is hard to beat.

The ceiling showed up when my use did. A model pulled from Ollama comes live with the parameters Ollama set. That's fine for most use cases.

It's less fine when you're building a specific application and need different settings, or want visibility into what's actually happening under the hood. As an entry point it's perfect. For me it was stage one.

The trigger: sovereign applications on small models

The work that pushed me past Ollama wasn't a chatbot. It was building locally run applications that needed real AI capability, the kind of requirements most people would default to a frontier model for.

Here's the problem with that default: attach a frontier API to a locally built or containerised application and it isn't sovereign. The app is yours; the intelligence isn't. And depending on the task, you often don't need a frontier model at all.

So I took the reverse approach. Start from the smallest model that could effectively do the task, then iterate until the outputs were consistent enough to trust in a closed network.

That process lives or dies on how easily you can test, adjust parameters and re-run. Consuming a frontier API, or Ollama's presets, didn't get me there. I needed deeper access through a UI.

That work is what became extract.solved, a sovereign document-intelligence pipeline that runs entirely behind the client's firewall.

Stage two: LM Studio gave me the controls

LM Studio is where I stopped fighting configuration. Pull any quantisation of a model straight from Hugging Face, set temperatures and parameters in profiles, all in a front-end GUI, then serve it so your tools call the API.

For months, it met every operational need. I still use it today for various needs.

If you're coming from an Ollama-type setup and want more control without any interest in training or fine-tuning, LM Studio is more than fine. A lot of people should stop here, and there's nothing wrong with that.

I wanted to go deeper. Two things pulled me on: how the models themselves get prepared, and the fact that fine-tuning had moved from research project to something within reach of a normal workflow.

Stage three: Unsloth became the centre

The pull started with the models. Unsloth takes open-source models and re-bakes them with their own approach: Dynamic GGUF quantisations that allocate bits per layer instead of applying one setting across the board. The variants they ship on Hugging Face work well for me on this rig. And you don't have to use them: any model you've downloaded runs in the engine the same way.

Then Unsloth Desktop landed, a recent release still carrying a beta label in their docs. One app now gives me chat, an OpenAI-compatible server, model discovery, on-device management, parameter control and monitoring. It also does fine-tuning inside the same interface, which no other tool in this trio does.

Because both LM Studio and Unsloth speak the OpenAI-compatible API, the switch was a re-point. My existing applications pointed at the new server, and nothing else changed.

Today, Unsloth is the central point where I serve any and all models. One inference engine instead of three half-used ones.

Ollama still sits on the PC for quick cloud-model trials and a couple of older simple apps that don't need much tweaking. LM Studio gets occasional use. The centre of gravity moved to what the app made possible, not a feature list.

What fine-tuning means for me right now: the capability is there, I can do it, and it's a playground I'll keep working with, especially around small models and specific tasks. But at the moment I get genuinely operational results from smaller models through better instructions, correct parameters and supporting knowledge. The right model for the task, with the right processes around it.

Fine-tuning is a door I've brought within reach, not a path I walk daily.

The real point: depth, not rivalry

Zoom out and the trio stops looking like competitors. They're stages of how deep your engagement with local inference goes:

  • Running: a model on your hardware, serving you or your apps. Ollama gets you there fastest.
  • Controlling: choosing quantisations, dialling parameters, seeing what happens. LM Studio made this a UI instead of a config file.
  • Owning the pipeline: prepared weights, fine-tuning in reach, one engine that also covers media generation and agent connections. Unsloth is where I am now.

This was never about picking a winner. It is about matching the tool to the problem in front of you, and knowing where to go when you outgrow it.

Why this is a sovereignty argument

I run this stack because my clients' data can't leave the building. The same logic applies to yours.

This is a Hard Hat era problem. The experiment phase is over; the work now is building the stack properly.

The sovereign small-model work above is the proof. You get closed-network capability by testing and iterating until a small model is consistent enough to trust.

"Every layer you outsource, including the intelligence itself, is a dependency you'll price-check later."

What to install

All three, in the order your depth needs them. Start where your pain is.

  • If no model is running yet, Ollama gets you there this afternoon. Some people stay there forever, and if it meets their needs that's the right answer.
  • If you need control over quantisations and parameters without any interest in training, LM Studio is the right stop.
  • If you want the deepest domain of control, prepared weights that work on your hardware class, and fine-tuning as an open door rather than a research project, that's the stage Unsloth covers for me now.

The Bottom Line

Your local AI stack isn't something you pick once. It grows as your work deepens, one stage at a time. The question was never which tool is best. It's how deep your work goes, and whether you own the path to the next level when it does.

Next Steps: If you are exploring local models, edge AI, or privacy-first use of AI and want to map out the right stack for your business, reach out to Intent Solved to build your foundation.

Book a Baseline.Solved Session | Explore Our Approach

Related Resources

Steven Muir-McCarey

Steven Muir-McCarey

Director

I'm a seasoned business development executive with impact across digital, cyber, technology and infrastructure sectors; anchors customer and partnership pipelines to boost revenue for key growth.

Expert at navigating diverse business operations across enterprise and government organisations, solving complex challenges using domain experience with innovative technologies to deliver effective solutions, adept at landing cost efficiencies with improved resource utilisations into programs of importance.

I'm known for developing trusted stakeholder relationships, working with teams and partners to foster better joint collaborations that strengthen and elevate the opportunity aligned to business strategy.

With two decades of experience, I bring customers to brand by understanding, engaging and aligning needs that marries the solution from the right technologies so as to arrive at the desired destination in the most cost-effective way.

I bring an open mindset and authentic leadership to everything I do, and I specialise in anchoring good business fundamentals with acumen that orchestrates longevity for market success.

Whether in public or private enterprises, my track record in achieving repeated impact remains visible in industry solutions available today; I thrive in helping customers to leverage and sequence advancements in technologies to achieve better business operations.