A Local LLM for Your Business, Without Running a Server
You can run a language model on a computer in your office. Plenty of people do, on a spare PC, for the afternoon it takes to break. This page is about doing it for a team that needs it to work on Monday.
£500 per month for the software and the hardware rental. You rent the machine — you don’t buy it.
A local LLM is a large language model that runs on your own hardware instead of a provider’s cloud, so prompts and documents never leave your premises. We install open-weight local LLM models on a dedicated machine in your office, add document search and assistants, and manage the updates and support.
Find out which model suits the job
A short call to match the model and machine to the work. If a DIY setup would genuinely do, we will tell you.
One machine, an open-weight model sized for your team, and the chat, document and assistant tools around it. Nothing goes to a cloud provider.
The Phone and the Inbox Can Stay in the Cloud
Local AI is for the confidential, document-heavy work. For answering calls, managing the inbox and chasing leads, our hosted agents are up and running in days with no hardware at all — and plenty of clients run both.
- ☁ Cloud-based AI Receptionist
- ☁ Cloud-based Email Manager
- ☁ Cloud-based Leads Outreach
- 🏢 On-premise Local AI
Do It Yourself, or Have It Done
Running a local LLM is not hard for an afternoon. Running one for a team, every working day, with backups and access control and someone to ring, is a different job.
| DIY on a spare machine | Cloud LLM subscription | Managed local LLM (ours) | |
|---|---|---|---|
| Where prompts and files go | Your machine | The provider’s servers | Your machine |
| Choosing local LLM models | Trial and error, evenings | No choice | Tested against your documents by us |
| Hardware | Whatever is spare | None | Specified and supplied for the team |
| Reads your document library | If you build it | Only what is pasted in | Document spaces with citations |
| Access control | If you build it | Per-seat accounts | Your Microsoft or Google sign-in |
| Model upgrades and updates | You, when you remember | Automatic | Us, as part of support |
| When it breaks | You | Their status page | Restored from backup by us |
Which Local LLM Models, and Why It Depends
The question we hear most is which is the best local LLM. There is no single answer, and anyone giving you one has not asked what you will use it for. The open-weight families — Llama, Mistral, Gemma, Qwen and others — each come in several sizes, and a model that is excellent at drafting English correspondence may be an indifferent choice for pulling structured facts out of four hundred contracts.
Three things decide it. What the work is: drafting and summarising, question-answering across a document library, or a specific structured task like extracting dates and parties. What hardware the model will run on: larger models give better answers and need a more capable machine, so the model and the machine are chosen together, not separately. And what your documents actually look like: scanned PDFs, long contracts, spreadsheets and clinical letters all behave differently, so we test candidate models against a sample of yours during setup rather than against a public leaderboard.
An on premise LLM also has to be kept current. New model versions arrive regularly, and a good upgrade is worth having; a careless one changes the answers your team has learned to expect. Model upgrades are part of our support arrangement precisely so that someone is making that call on purpose.
If your interest is a self hosted AI LLM you run yourself, we are not going to talk you out of it — the open-weight ecosystem is genuinely good, and for one technical person it can be the right call. Our offer is for the team that wants the same privacy without owning the maintenance.
The Model Is the Engine. This Is the Car.
A bare model is a demo. A team needs the machine, the tools around the model, the sign-in and someone to maintain it.
A model matched to the work
An open-weight local LLM chosen and tested against your own documents, on a machine specified to run it well for your team size.
Document spaces around it
The model on its own answers from what you paste in. Spaces let it search, summarise and compare across hundreds of your files, citing the page each answer came from.
Chat and assistants on top
A private chat assistant for the team and custom assistants for the recurring jobs, all running on the same machine as the model.
| Included | Not included |
|---|---|
| A dedicated machine, sized for your team, rented to you (not sold) and installed by us | A cloud subscription — nothing runs on our servers |
| £500 per month for the software and the hardware rental | Buying the hardware — the machine is rented, not sold |
| Private chat, document spaces and custom assistants on that machine | Clinical, legal or financial advice — it drafts, you decide |
| Single sign-on with your Microsoft 365 or Google Workspace | A replacement for your practice, case or CRM system |
| Import of your existing documents, and training for the people using it | Phone answering — see AI Receptionist |
| Software updates, model upgrades and support from the people who installed it | Inbox automation — see Email Manager |
| Model selection tested on a sample of your own documents | A guarantee that a local model matches the largest cloud services on every task |
Three Steps to a Working Local LLM
Weeks, not months. Most clients start with one team and one document set, then widen it.
Match model to job
A short call on what the model will read and for whom. We shortlist candidate models and specify the machine to run them.
Install and test
The machine is installed on your network, your sign-in connected, your documents imported. We test the shortlisted models against your real files and settle on one.
Run and maintain
We train your team, then handle updates, model upgrades and support. Adding people later is configuration, not a new project.
Where Local AI Is the Wrong Answer
We would rather lose the enquiry than the trust. Three things we tell every prospect before they sign anything.
It is not a frontier model
The largest cloud models are still ahead of any local LLM on hard, novel reasoning. If your work needs that, a local model is the wrong tool and we will say so. For summarising, drafting and answering questions about your own documents, most offices are happy with it.
It is slower than the big cloud models
Open-weight models on a single machine are capable for summarising, drafting and answering questions about your own documents. For frontier reasoning on hard, novel problems, the largest cloud models are still ahead. That is the trade.
It is not for the phone or the inbox
A local LLM is built for confidential, document-heavy work. Call answering and inbox triage are better served by our hosted agents, which need no hardware and are live in days.
We Run Our Own
Chosen against your files
Model selection is tested on a sample of your own documents, not a public leaderboard.
Sized with the model
The machine and the model are specified together. A model the hardware cannot run well is the commonest DIY mistake.
Upgrades made on purpose
Model upgrades are part of support, so someone is deciding when the answers your team relies on are allowed to change.
UK data protection built in
No third-party AI processor and no restricted transfer for the AI step. UK GDPR, the Data Protection Act 2018 and the ICO as regulator, with a data processing agreement for the support access you grant us.
Which Page Answers Your Question
| If your question is… | The short answer | Read more |
|---|---|---|
| What is the full system around the model? | One machine: chat, spaces, assistants, support | On-premises AI solution |
| Is this the same as private AI? | Yes — the model is what makes it private | Private AI |
| What does it change for UK GDPR? | No processor, no transfer, for the AI step | GDPR compliant AI |
| We want to own the hardware | You do; we maintain it | Self-hosted AI |
| We are in London | Installed on site across the capital | Local AI London |
| We are a dental practice | Patient documents stay in the practice | Local AI for dental practices |
| Why local at all? | The case for Local AI, in full | Local AI product page |
Local LLM by City and Sector
Local LLMs for Business: Common Questions
What is a local LLM?
A large language model — the kind of model behind chat assistants — running on hardware you control rather than accessed through a provider’s service. Your prompts, uploaded documents and chat history stay on that machine.
Which is the best local LLM?
It depends on the job. Drafting and summarising, answering questions across a document library, and working in a language other than English each favour different open-weight models, and the machine has to fit the model. We pick and test the model against your actual documents during setup.
Can we run one on an ordinary office PC?
Small models will run on a well-specified desktop, and that is fine for one person experimenting. For a team using document search all day you need a machine specified for it. We size and supply that machine after a short call.
Does it need the internet?
No. The model runs on the machine, so chat and document search keep working if your connection drops. The internet is used only for signing in and, if you enable it, linking your own Microsoft 365 or Google Workspace account.
How does a local LLM compare with ChatGPT?
Capable for everyday business tasks — summarising, drafting, extracting, answering questions about your own documents — and somewhat slower. The largest cloud models remain ahead on hard, novel reasoning. You trade a little capability and speed for your data never leaving the building.
Who keeps the model up to date?
We do. Model upgrades and software updates are part of the ongoing support arrangement, from the people who installed the machine. You are not left maintaining a server.
How much does Local AI cost?
£500 per month for the software and the hardware rental. The dedicated machine is rented to your business, not sold: you never buy the hardware. It is installed in your building and runs the private chat, document spaces and assistants. Local AI is for business customers only.
Tell Us What the Model Would Read
One short call to shortlist the model and size the machine. If a spare PC and a free afternoon would genuinely do the job, we will tell you that instead.
Prefer email? sghaith@businessaiagents.co.uk
