Local LLM

A Local LLM for Your Business, Without Running a Server

You can run a language model on a computer in your office. Plenty of people do, on a spare PC, for the afternoon it takes to break. This page is about doing it for a team that needs it to work on Monday.

£500 per month for the software and the hardware rental. You rent the machine — you don’t buy it.

A local LLM is a large language model that runs on your own hardware instead of a provider’s cloud, so prompts and documents never leave your premises. We install open-weight local LLM models on a dedicated machine in your office, add document search and assistants, and manage the updates and support.

Find out which model suits the job

A short call to match the model and machine to the work. If a DIY setup would genuinely do, we will tell you.

What Runs On It
A local LLM running on a dedicated business machine with chat, document spaces and assistants

One machine, an open-weight model sized for your team, and the chat, document and assistant tools around it. Nothing goes to a cloud provider.

Local LLMs installed for UK professional firms, clinics and back offices

The model choice, the machine, the install and the support, from one team.

Also From Us · Cloud AI Agents

The Phone and the Inbox Can Stay in the Cloud

Local AI is for the confidential, document-heavy work. For answering calls, managing the inbox and chasing leads, our hosted agents are up and running in days with no hardware at all — and plenty of clients run both.

See the Local AI product page →
The Honest Comparison

Do It Yourself, or Have It Done

Running a local LLM is not hard for an afternoon. Running one for a team, every working day, with backups and access control and someone to ring, is a different job.

 DIY on a spare machineCloud LLM subscriptionManaged local LLM (ours)
Where prompts and files goYour machineThe provider’s serversYour machine
Choosing local LLM modelsTrial and error, eveningsNo choiceTested against your documents by us
HardwareWhatever is spareNoneSpecified and supplied for the team
Reads your document libraryIf you build itOnly what is pasted inDocument spaces with citations
Access controlIf you build itPer-seat accountsYour Microsoft or Google sign-in
Model upgrades and updatesYou, when you rememberAutomaticUs, as part of support
When it breaksYouTheir status pageRestored from backup by us
Model Choice

Which Local LLM Models, and Why It Depends

The question we hear most is which is the best local LLM. There is no single answer, and anyone giving you one has not asked what you will use it for. The open-weight families — Llama, Mistral, Gemma, Qwen and others — each come in several sizes, and a model that is excellent at drafting English correspondence may be an indifferent choice for pulling structured facts out of four hundred contracts.

Three things decide it. What the work is: drafting and summarising, question-answering across a document library, or a specific structured task like extracting dates and parties. What hardware the model will run on: larger models give better answers and need a more capable machine, so the model and the machine are chosen together, not separately. And what your documents actually look like: scanned PDFs, long contracts, spreadsheets and clinical letters all behave differently, so we test candidate models against a sample of yours during setup rather than against a public leaderboard.

An on premise LLM also has to be kept current. New model versions arrive regularly, and a good upgrade is worth having; a careless one changes the answers your team has learned to expect. Model upgrades are part of our support arrangement precisely so that someone is making that call on purpose.

If your interest is a self hosted AI LLM you run yourself, we are not going to talk you out of it — the open-weight ecosystem is genuinely good, and for one technical person it can be the right call. Our offer is for the team that wants the same privacy without owning the maintenance.

What You Get

The Model Is the Engine. This Is the Car.

A bare model is a demo. A team needs the machine, the tools around the model, the sign-in and someone to maintain it.

🧠

A model matched to the work

An open-weight local LLM chosen and tested against your own documents, on a machine specified to run it well for your team size.

📂

Document spaces around it

The model on its own answers from what you paste in. Spaces let it search, summarise and compare across hundreds of your files, citing the page each answer came from.

💬

Chat and assistants on top

A private chat assistant for the team and custom assistants for the recurring jobs, all running on the same machine as the model.

IncludedNot included
A dedicated machine, sized for your team, rented to you (not sold) and installed by usA cloud subscription — nothing runs on our servers
£500 per month for the software and the hardware rentalBuying the hardware — the machine is rented, not sold
Private chat, document spaces and custom assistants on that machineClinical, legal or financial advice — it drafts, you decide
Single sign-on with your Microsoft 365 or Google WorkspaceA replacement for your practice, case or CRM system
Import of your existing documents, and training for the people using itPhone answering — see AI Receptionist
Software updates, model upgrades and support from the people who installed itInbox automation — see Email Manager
Model selection tested on a sample of your own documentsA guarantee that a local model matches the largest cloud services on every task
How It Works

Three Steps to a Working Local LLM

Weeks, not months. Most clients start with one team and one document set, then widen it.

1

Match model to job

A short call on what the model will read and for whom. We shortlist candidate models and specify the machine to run them.

2

Install and test

The machine is installed on your network, your sign-in connected, your documents imported. We test the shortlisted models against your real files and settle on one.

3

Run and maintain

We train your team, then handle updates, model upgrades and support. Adding people later is configuration, not a new project.

Straight Answers

Where Local AI Is the Wrong Answer

We would rather lose the enquiry than the trust. Three things we tell every prospect before they sign anything.

It is not a frontier model

The largest cloud models are still ahead of any local LLM on hard, novel reasoning. If your work needs that, a local model is the wrong tool and we will say so. For summarising, drafting and answering questions about your own documents, most offices are happy with it.

It is slower than the big cloud models

Open-weight models on a single machine are capable for summarising, drafting and answering questions about your own documents. For frontier reasoning on hard, novel problems, the largest cloud models are still ahead. That is the trade.

It is not for the phone or the inbox

A local LLM is built for confidential, document-heavy work. Call answering and inbox triage are better served by our hosted agents, which need no hardware and are live in days.

Why Business AI Agents

We Run Our Own

Chosen against your files

Model selection is tested on a sample of your own documents, not a public leaderboard.

Sized with the model

The machine and the model are specified together. A model the hardware cannot run well is the commonest DIY mistake.

Upgrades made on purpose

Model upgrades are part of support, so someone is deciding when the answers your team relies on are allowed to change.

UK data protection built in

No third-party AI processor and no restricted transfer for the AI step. UK GDPR, the Data Protection Act 2018 and the ICO as regulator, with a data processing agreement for the support access you grant us.

Which Page Answers Your Question

If your question is…The short answerRead more
What is the full system around the model?One machine: chat, spaces, assistants, supportOn-premises AI solution
Is this the same as private AI?Yes — the model is what makes it privatePrivate AI
What does it change for UK GDPR?No processor, no transfer, for the AI stepGDPR compliant AI
We want to own the hardwareYou do; we maintain itSelf-hosted AI
We are in LondonInstalled on site across the capitalLocal AI London
We are a dental practicePatient documents stay in the practiceLocal AI for dental practices
Why local at all?The case for Local AI, in fullLocal AI product page

Local LLM by City and Sector

Local LLM London Local LLM Manchester Local LLM for dental practices Local LLM for solicitors Local LLM for estate agents Local LLM for gp practices

Local LLMs for Business: Common Questions

What is a local LLM?

A large language model — the kind of model behind chat assistants — running on hardware you control rather than accessed through a provider’s service. Your prompts, uploaded documents and chat history stay on that machine.

Which is the best local LLM?

It depends on the job. Drafting and summarising, answering questions across a document library, and working in a language other than English each favour different open-weight models, and the machine has to fit the model. We pick and test the model against your actual documents during setup.

Can we run one on an ordinary office PC?

Small models will run on a well-specified desktop, and that is fine for one person experimenting. For a team using document search all day you need a machine specified for it. We size and supply that machine after a short call.

Does it need the internet?

No. The model runs on the machine, so chat and document search keep working if your connection drops. The internet is used only for signing in and, if you enable it, linking your own Microsoft 365 or Google Workspace account.

How does a local LLM compare with ChatGPT?

Capable for everyday business tasks — summarising, drafting, extracting, answering questions about your own documents — and somewhat slower. The largest cloud models remain ahead on hard, novel reasoning. You trade a little capability and speed for your data never leaving the building.

Who keeps the model up to date?

We do. Model upgrades and software updates are part of the ongoing support arrangement, from the people who installed the machine. You are not left maintaining a server.

How much does Local AI cost?

£500 per month for the software and the hardware rental. The dedicated machine is rented to your business, not sold: you never buy the hardware. It is installed in your building and runs the private chat, document spaces and assistants. Local AI is for business customers only.

Tell Us What the Model Would Read

One short call to shortlist the model and size the machine. If a spare PC and a free afternoon would genuinely do the job, we will tell you that instead.

Book a Consultation

Prefer email? sghaith@businessaiagents.co.uk