Thornhaver is free and runs entirely on hardware you already own. Your questions, your documents, your work — none of it is sent anywhere, because there is nowhere to send it.
Free forever. No account, no API key, no network call.
Advertisement
How it works
Most assistants send everything you type to a company's servers, meter it by the token, and change the model underneath you whenever they like. Thornhaver does none of that, because the model is a file on your disk and the software runs on your own processor.
Not a privacy policy — an architecture. The app listens on your own machine only. There is no server to send your data to, so there is nothing to leak, subpoena, or train on.
Ask a thousand questions or a hundred thousand. Retry as often as you like. Nothing is counted, because nobody is paying for it but your electricity bill.
A hosted model can change overnight and quietly break what you built on it. Your version stays exactly as it is until you decide otherwise.
Rate answers as you go. The good ones become training data for the next version, so it grows toward what you actually do rather than the average of the internet.
Benchmarks
Every version below was run on the same machine — a Radeon RX 6600 with 8 GB — with the same questions, on the same day, through the same software. Generation speed, the delay before the first visible word, and factual accuracy on deliberately obscure questions.
| Version | Best for | Speed | First word | Accuracy | Status |
|---|---|---|---|---|---|
| Thornhaver 1.0 | Programming only | 42 tok/s | 0.1s | declines other topics | Available |
| Thornhaver 1.1 | Long-form writing | 8 tok/s | 17.9s | 4/4 | Available |
| Thornhaver 1.2 | Careful reasoning | 42 tok/s | 6.9s | 4/4 | Available |
| Thornhaver 1.3 | Reasoning, larger | 11 tok/s | 26.1s | 5/5 | Available |
| Thornhaver 1.4 | Raw speed | 59 tok/s | 0.3s | invents answers | Available |
| Thornhaver 1.5 | Code and review | 17 tok/s | 0.2s | 5/5 | Available |
| Thornhaver 1.6 | General conversation | 15 tok/s | 0.2s | 3/4 | Available |
| Thornhaver 1.7 | General work | 15 tok/s | 0.2s | 5/5 | Available |
| Thornhaver 1.8 recommended | General work | 8 tok/s | 0.3s | 4/4 | Available |
| Thornhaver 1.9 | Fine-tuned on our own corpus. In training. | — | — | — | Coming soon |
| Thornhaver 2.0 | Long context, tool use and retrieval. In development. | — | — | — | Coming soon |
Speed is tokens generated per second once the model is warm in graphics memory. First word is the delay before anything appears on screen — a reasoning model can generate quickly yet feel slow, because it thinks silently before speaking. Accuracy is the share of obscure-term questions answered correctly on inspection, not by keyword matching.
Advertisement
Questions
No. Thornhaver listens only on your own machine, and the model runs on your own processor and graphics card. There is no account, no API key, and no outbound request carrying your text. Disconnect from the internet entirely and it keeps working.
Yes. The model runs on your hardware, so there is no per-message cost for anyone to pass on to you. This site carries advertising; the software costs nothing and there is nothing to subscribe to.
A reasonably modern computer with a dedicated graphics card of 8 GB or more. It runs on less, just more slowly, because a model that does not fit in graphics memory falls back to the processor. Every measurement on this page comes from a Radeon RX 6600 with 8 GB — deliberately ordinary hardware.
Not at everything, and anyone telling you otherwise is selling something. A model that fits on your desk is smaller than one running on a datacentre rack. What you get instead is privacy, no metering, no rate limits, and a model that never changes without your say-so.
Yes, and it will be. Every language model states wrong things in a confident voice — see the fishbone note above. That is exactly why each version is published with its measured accuracy, and why the rating buttons exist in the app.
The one marked recommended, unless you have a reason not to. It is chosen by measurement rather than preference: the most accurate version that still starts answering immediately. If you are writing code, the code-focused versions will serve you better.
They are written to a file on your own disk and stay there. You can read them, export them as a training set, or delete them. They are raw material for improving your own copy — not a data collection programme.
No account to make, nothing to pay for, nothing leaving your computer.
Get Thornhaver