The assistant that never leaves your machine.

Thornhaver is free and runs entirely on hardware you already own. Your questions, your documents, your work — none of it is sent anywhere, because there is nowhere to send it.

Free forever. No account, no API key, no network call.

9versions to choose from
0.3sto the first word
£0.00per message, forever
0 bytessent off your machine

Advertisement

Ad slot — 728×90

How it works

No cloud in the loop.

Most assistants send everything you type to a company's servers, meter it by the token, and change the model underneath you whenever they like. Thornhaver does none of that, because the model is a file on your disk and the software runs on your own processor.

01

Private by construction

Not a privacy policy — an architecture. The app listens on your own machine only. There is no server to send your data to, so there is nothing to leak, subpoena, or train on.

02

Nothing is metered

Ask a thousand questions or a hundred thousand. Retry as often as you like. Nothing is counted, because nobody is paying for it but your electricity bill.

03

Pinned, not moved

A hosted model can change overnight and quietly break what you built on it. Your version stays exactly as it is until you decide otherwise.

04

It learns your work

Rate answers as you go. The good ones become training data for the next version, so it grows toward what you actually do rather than the average of the internet.

Benchmarks

Nine local models, measured on one 8 GB card.

Every version below was run on the same machine — a Radeon RX 6600 with 8 GB — with the same questions, on the same day, through the same software. Generation speed, the delay before the first visible word, and factual accuracy on deliberately obscure questions.

VersionBest forSpeedFirst wordAccuracyStatus
Thornhaver 1.0Programming only 42 tok/s 0.1sdeclines other topics Available
Thornhaver 1.1Long-form writing 8 tok/s 17.9s4/4 Available
Thornhaver 1.2Careful reasoning 42 tok/s 6.9s4/4 Available
Thornhaver 1.3Reasoning, larger 11 tok/s 26.1s5/5 Available
Thornhaver 1.4Raw speed 59 tok/s 0.3sinvents answers Available
Thornhaver 1.5Code and review 17 tok/s 0.2s5/5 Available
Thornhaver 1.6General conversation 15 tok/s 0.2s3/4 Available
Thornhaver 1.7General work 15 tok/s 0.2s5/5 Available
Thornhaver 1.8 recommended General work 8 tok/s 0.3s4/4 Available
Thornhaver 1.9Fine-tuned on our own corpus. In training. Coming soon
Thornhaver 2.0Long context, tool use and retrieval. In development. Coming soon

Speed is tokens generated per second once the model is warm in graphics memory. First word is the delay before anything appears on screen — a reasoning model can generate quickly yet feel slow, because it thinks silently before speaking. Accuracy is the share of obscure-term questions answered correctly on inspection, not by keyword matching.

One finding worth the whole table. The fastest model here was demoted from the recommended slot for confidently inventing answers — it defined a ship's keelson as a piece of fishbone, four separate times, in fluent and entirely convincing prose. Speed is easy to measure. Truthfulness is not. Check anything that matters, whichever version you use.

Advertisement

Ad slot — 336×280

Questions

The things worth asking first.

Does anything I type get sent anywhere?

No. Thornhaver listens only on your own machine, and the model runs on your own processor and graphics card. There is no account, no API key, and no outbound request carrying your text. Disconnect from the internet entirely and it keeps working.

Is it really free?

Yes. The model runs on your hardware, so there is no per-message cost for anyone to pass on to you. This site carries advertising; the software costs nothing and there is nothing to subscribe to.

What hardware do I need?

A reasonably modern computer with a dedicated graphics card of 8 GB or more. It runs on less, just more slowly, because a model that does not fit in graphics memory falls back to the processor. Every measurement on this page comes from a Radeon RX 6600 with 8 GB — deliberately ordinary hardware.

Is it as capable as the big hosted assistants?

Not at everything, and anyone telling you otherwise is selling something. A model that fits on your desk is smaller than one running on a datacentre rack. What you get instead is privacy, no metering, no rate limits, and a model that never changes without your say-so.

Can it be wrong?

Yes, and it will be. Every language model states wrong things in a confident voice — see the fishbone note above. That is exactly why each version is published with its measured accuracy, and why the rating buttons exist in the app.

Which version should I use?

The one marked recommended, unless you have a reason not to. It is chosen by measurement rather than preference: the most accurate version that still starts answering immediately. If you are writing code, the code-focused versions will serve you better.

What happens to my ratings?

They are written to a file on your own disk and stay there. You can read them, export them as a training set, or delete them. They are raw material for improving your own copy — not a data collection programme.

Free, private, and running in about five minutes.

No account to make, nothing to pay for, nothing leaving your computer.

Get Thornhaver