[SEC.INSIGHTS-REF.2026]
Back to Insights
Productivity

Using Someone Else's AI Isn't the Same as Running Your Own

September 4, 2026
Using Someone Else's AI Isn't the Same as Running Your Own

Running your own AI doesn't start with a purchase.

Bottom Line: My use of open and local AI models didn't start with a GPU. It started with an API key and someone else's hardware. Four stages got me from there to running and fine-tuning models myself, and none of them required the spend I assumed this needed before I started. Wherever you are right now, start with the stage that matches your problem, not with a purchase.

The Context

There's a difference between using AI and running your own. A subscription to a hosted chatbot is the first. It's useful, but the model still lives on someone else's servers. Running your own means the model sits somewhere you control. People assume that second option starts with buying serious hardware. It doesn't. I built this up in stages, and the first two didn't cost me anything beyond time. The tools changed as my needs got more specific, not because the earlier ones stopped working.

Why this matters: If you've been putting this off because you think it starts with a five-figure hardware purchase, it doesn't. It starts with understanding what you actually need a model to do for you. The hardware question came later for me than I expected it to when I started, and it can for you too.

The Insight

Stage One: An API Key, Not a GPU

I started by accessing a mix of models through OpenRouter, some open source, some frontier, all through the one API key. No hardware, no installation, no commitment. The point wasn't ranking one model against another in a chat window. It was learning how to connect a model to something I was actually building, and getting free of the standard chat interface every frontier provider defaults to.

This stage matters more than people give it credit for. It's where you find out whether a lower-cost or fit-for-purpose model actually meets the need, instead of defaulting to the top-shelf frontier option out of habit. It's also where you can trial smaller, open models to see if they'd deliver on the need before you commit to setting up the edge compute to run them yourself. And it's where you get a real read on the ongoing token cost of serving a given model through an API in your application, before you commit to it in production. You're still going through someone else's servers at this point, so it's not sovereign yet. But it's the fastest, cheapest way to find any of this out before you decide anything else.

Stage Two: Something Running Locally, Finally

Ollama is where local actually started for me. Small models, running on a GPU I already had, answering without a single request leaving my machine. This is the point where "local AI" stops being a concept and becomes something you're actually doing.

It doesn't take exotic hardware to get here. Small, capable open models run comfortably on consumer GPUs that plenty of people already own for gaming or other work. The barrier isn't the machine. It's knowing this step exists.

Stage Three: Knowing What You Actually Need It To Do

Running a model in a chat window teaches you less than you'd think, because you're stuck with whatever behaviour the defaults give you. The real shift was understanding my own needs well enough to know the standard settings weren't enough, and going looking for the controls to actually change how a model behaved, not just what it answered.

That's what pulled LM Studio into the stack. Granular parameter control, well beyond the defaults, the ability to audition a model at different quantisations, and a server I could point other applications at. Once I understood what I actually needed, I could adjust the model to deliver it, not just accept whatever came out of the box.

Stage Four: Shaping the Model, Not Just Running It

Unsloth Desktop is where I am now. Chat, local serving, model discovery, and fine-tuning, in one application. It's free and open-source, and it runs on Windows, macOS, and Linux.

It wasn't only the fine-tuning capability that pulled me in. Unsloth's community also ships its own tuned model variants, optimised for their platform. I wasn't downloading raw weights and hoping they ran well, I was starting from something already tuned for the way I was running it.

I didn't move here because the earlier tools failed. I moved here because my needs grew past what running a model as-is could give me. Fine-tuning changes how a model behaves, not what it knows. It shapes output format, structure, and consistency. That's a different problem to the one Stage One or Stage Two solves, and it's only worth reaching for once you've actually hit that specific limit.

What This Actually Buys You

Here's the part that matters more than any tool in this stack. Small, open models are genuinely capable enough now to handle real daily work. Drafting, summarising documents, extracting information, and a first pass on code. Not every task, but more of my daily load than I expected when I started.

Run those tasks locally instead of through a frontier API, and the documents, client details, and internal information involved in them never have to leave your business's perimeter. That's the actual destination. Your own capability and your own data, staying inside your own walls.

What This Means for You

If You Are...Then This Means...
Curious, with no hardware to speak ofStart with API access through a provider like OpenRouter. You'll learn what these models can do for your work for a few dollars in API calls, not a hardware budget.
Sitting on a GPU you already ownInstall Ollama and run a small model against some real tasks this week. That's Stage Two, and it's closer than you think.
Building tools or workflows that need to call a modelYou've outgrown a chat window. Look at LM Studio or a similar serving layer, and get parameter control before you get frustrated.
Hitting limits that prompting cannot resolveTraining or fine-tuning an open-source model yourself is a real, accessible option now, not something you need an ML team for. That's when it's worth reaching for. Not before.

Each stage solved a specific problem I actually had. None of them were about owning better hardware for its own sake.

How to Act on This

Immediate Actions (This Week)

  1. Work out which stage you're actually at. Not which tool sounds most advanced. Match the stage to the problem you have right now.
  2. If you have zero dedicated hardware, start with API access. Run some real work tasks through various open and commercial models. You can stay measured on cost with pay-as-you-go access, no subscription to a model provider required, and it tells you more than any article will.
  3. If you already have a GPU sitting idle, install Ollama today. That was the whole barrier for me: not knowing this step existed.

Strategic Actions (This Quarter)

  • Build the serving layer once you actually need one. Not before. A chat window is fine until you're plugging a model into something else you've built.
  • Treat fine-tuning as a last resort, not a starting point. It solves a specific problem: inconsistent output, the wrong format, a structure your process can't use. If that's not your problem, you don't need it yet.
  • Measure what this actually saves you. Time, cost per task, and what stays inside your business instead of going through a third party. That's your Return on Intent, not the tool you're running.

The Bottom Line

None of this required a hardware decision to get started, and most of it still didn't for me. I moved through four stages because my needs got more specific, not because a bigger machine was the goal. Start at whichever stage matches the problem you actually have. If that's an API key and a free afternoon, that's a legitimate place to start, it just isn't the destination on its own. Keep moving through the stages and you get there: models capable enough for your daily work, running close enough to you that your data stays where it belongs.

Next Steps: Read The Right AI Model for the Job for how to match models to actual business functions once you're running locally, or reach out to Intent Solved if you want help working out which stage your team is actually at.

Related Resources

Steven Muir-McCarey

Steven Muir-McCarey

Director

I'm a seasoned business development executive with impact across digital, cyber, technology and infrastructure sectors; anchors customer and partnership pipelines to boost revenue for key growth.

Expert at navigating diverse business operations across enterprise and government organisations, solving complex challenges using domain experience with innovative technologies to deliver effective solutions, adept at landing cost efficiencies with improved resource utilisations into programs of importance.

I'm known for developing trusted stakeholder relationships, working with teams and partners to foster better joint collaborations that strengthen and elevate the opportunity aligned to business strategy.

With two decades of experience, I bring customers to brand by understanding, engaging and aligning needs that marries the solution from the right technologies so as to arrive at the desired destination in the most cost-effective way.

I bring an open mindset and authentic leadership to everything I do, and I specialise in anchoring good business fundamentals with acumen that orchestrates longevity for market success.

Whether in public or private enterprises, my track record in achieving repeated impact remains visible in industry solutions available today; I thrive in helping customers to leverage and sequence advancements in technologies to achieve better business operations.