An Assistant That Speaks for Me
A narrow brief, one source of truth, and a fixed set of real questions to test every change against.
Finch is the assistant on this site. Ask it what I do, what I’ve built or how to reach me, and it answers, typed or out loud.
The first article covered keeping it from being broken: prompt injection, cost attacks, poisoned history. This one covers keeping it accurate. An assistant can be secure and still get facts about me wrong, and it states wrong facts as confidently as right ones.
Why a chat window
I spend hours every day inside Hotline, the desktop app I build where a team of AI assistants work alongside me like colleagues in a shared room. Much of Hotline is now built and improved from inside Hotline. Because I work in a conversation all day, it’s the interface I reach for first, so when I rebuilt this site I made the homepage one too. Instead of a page of links, you ask what you want to know.
A narrow brief
Finch talks about me in the third person: “George” the first time, “he” after that. It covers what a visitor to a personal site usually wants: what I do, what I’ve built, what I’ve written and how to get in touch. For anything else, it says it can’t help with that here and suggests something it can.
I kept the scope small on purpose. Every topic it can discuss is another place it can be wrong about me.
Only what it’s been told
The main rule: use only the facts you’ve been given. No invented numbers, no stretched durations, no guessing whether I’d take a particular job.
The second rule covers what to do without an answer. Models tend to fill gaps with plausible guesses, so Finch has a set response: say I haven’t shared that here, then give the closest fact it does have.
Visitor: does he have a CISSP?
Finch: Not one he’s listed. His certifications are AWS Certified Cloud Practitioner, Cisco Meraki CMNA and Cisco CEN, on top of twenty years of hands-on work making sure only the right people can get into a company’s systems.
Exact details come from the page
A model can write a good sentence and still get one character of an email address wrong, so Finch never types them.
The idea comes from how an agent loop calls tools. In Hotline, an assistant doesn’t carry out an action itself. It asks for a tool by name, and the app runs it. Finch does a small version of the same thing. When someone asks how to reach me, it writes one short sentence and ends its reply with a tag naming a tool. The page reads the tag and shows a contact card with the real address, a copy button and my links. My resume, open-source projects and recent articles work the same way.
It was the practical option. The page already has all of that data, so a tag is enough to display it, and I didn’t need to build a backend that runs tools for the model. The model only chooses when to show a card; everything on it comes from the site’s own data.
One source of truth
Finch’s facts, the resume page and the old About page were written at different times. When I added my full resume to the site, the three disagreed. The About page had the wrong year for my first job, the wrong founding year for Leveldesk and the wrong device count, and Finch used a different job title from the resume.
I corrected Finch’s facts from the resume, confirmed which title is official, and removed the About page. Everything on it was already covered by the resume or by Finch.
What it knows about my articles
Each article has a one-line main point and a short list of key takeaways in its metadata. Finch gets those with the title, not the full text. Asked about an article, it answers in two or three sentences, starting from the main point.
That keeps the prompt small, so answers stay fast and cheap, and Finch summarises the point I wrote down rather than one it chose itself.
Choosing the model
Typed and spoken questions both go to Google’s Gemini 3.5 Flash Lite, through the same Cloudflare gateway as everything else. Finch started on Qwen 3.8, but Gemini was faster and more consistent on the same questions, and one model means one set of behaviour to test. Spoken answers have extra rules: one or two short sentences, no formatting, and numbers written the way people say them, like “more than two hundred clients” instead of “200+”.
Before settling on it, I tested Gemini against fourteen other setups. There were the cheap models that kept coming up (DeepSeek V4 and V4.1 Flash, GLM 5.3 Flash), fast hosts like Groq and Relace, Anthropic’s Claude Haiku 4.5 and Sonnet 5.5, Meta’s Muse Spark, Google’s older Gemini 2.5 Flash Lite, and a free stealth model on OpenRouter. DeepSeek V4.1 Flash and GLM 5.3 Flash ran on more than one host, since the same model can behave very differently depending on who serves it. Each one got Finch’s real instructions and the same 24 questions, typed and spoken. The questions covered facts from my resume, things I haven’t shared (a CISSP, my salary, whether I’m looking for a job), the tags that bring up cards, and a couple of attempts to break it. The whole comparison cost under five dollars.
Accuracy barely separated them. Every setup answered 96 to 100 percent of the questions correctly, and the mistakes were small but telling. GLM said my customers spend more than ten million dollars “a year” with Cisco, which my resume doesn’t say. Gemini 2.5 typed out my email address instead of bringing up the contact card. Letting the models think before answering made them slower and no more accurate.
The slowest answers separated them more than the typical ones. GLM 5.3 on Relace usually answered in 0.58 seconds, faster than Gemini, but one answer in ten took more than four. MiMo’s took almost nine. On a call, that sounds like the line went dead. Gemini’s slowest one in ten came back in 0.83 seconds.
The host mattered as much as the model. The same DeepSeek V4.1 Flash answered in 0.57 seconds on Ollama, 1.1 on Fireworks and 1.5 through OpenRouter, and the first two had the extra distance from my machine to cover. So did the way I measured. My first test called Gemini through Cloudflare’s general-purpose API and timed it at 2.1 seconds. Sent the way the site actually sends it, it answers in under 0.7.
I also checked every spoken answer against the voice rules: no formatting, no digits, no more than about forty words. This is where the models differed most. The Claude models knew every answer but bolded names and read years out as digits in about a third of them; Muse Spark did the same. Gemini kept to the rules in 94 percent of answers. To be fair to the others, the rules were written and tuned while Finch ran on Gemini.
The costs surprised me. Every Finch request sends about 4,700 tokens of instructions and gets back 30 to 70, so the input price, and whether the host reuses those repeated instructions from a cache, decides the bill far more than the output price. Claude Haiku lists at more than three times Gemini’s input price, but it reads cached instructions at a tenth of the price, so 1,000 answers came to $0.91 against Gemini’s $1.49. The DeepSeek V4.1 hosts cache too, and came in at $0.16 to $0.38.
Here is everything, with Gemini first and the rest in order of typical answer time:
| Model | Right | Natural | Typical | Slowest 1 in 10 | Per 1,000 |
|---|---|---|---|---|---|
| Gemini 3.5 Flash LiteGoogle | 100% | 94% | 0.66s | 0.83s | $1.49 |
| Gemini 2.5 Flash LiteGoogle | 96% | 85% | 0.50s | 0.61s | $0.30 |
| DeepSeek V4.1 FlashOllama | 100% | 92% | 0.57s | 0.73s | $0.21 |
| GLM 5.3Relace | 100% | 81% | 0.58s | 4.12s | $1.01 |
| Qwen 3.8 27BGroq | 100% | 79% | 0.61s | 0.74s | $3.87 |
| Muse Spark 1.3Meta | 98% | 67% | 0.72s | 1.19s | $3.68 |
| Claude Haiku 4.5Anthropic | 100% | 65% | 0.77s | 0.94s | $0.91 |
| MiMo v2.6 FlashRelace | 100% | 73% | 1.06s | 8.72s | $0.45 |
| Claude Sonnet 5.5Anthropic | 100% | 67% | 1.08s | 1.76s | $2.44 |
| DeepSeek V4.1 FlashFireworks | 96% | 94% | 1.13s | 1.91s | $0.38 |
| DeepSeek V4 FlashWorkers AI | 100% | 75% | 1.22s | 2.16s | $1.79 |
| Space Bunny Alphastealth | 100% | 87% | 1.39s | 2.19s | free |
| GLM 5.3 FlashFireworks | 98% | 81% | 1.42s | 4.59s | $0.45 |
| DeepSeek V4.1 FlashOpenRouter | 100% | 92% | 1.52s | 2.55s | $0.16 |
| GLM 5.3 FlashWorkers AI | 98% | 83% | 3.36s | 13.7s | $0.21 |
The tests ran over a few days while Finch’s instructions were still changing, so the earlier runs sent slightly shorter prompts. Qwen on Groq answered 43 of its 48 questions before it hit Groq’s free daily limit.
I chose Gemini. A few setups beat it on one measure: Gemini 2.5 Flash Lite and DeepSeek V4.1 on Ollama were about a tenth of a second faster, and several were cheaper. None of them matched it on all three things a spoken assistant needs: right every time, fast every time, and sounding like speech rather than text. Switching would also mean a second provider for the words while Google still does the voice. At this site’s traffic, the cheapest option would save a few dollars a year.
Testing with real questions
Before any change to Finch’s instructions ships, I run a fixed set of questions: where did he go to school, what certifications does he have, how big did Leveldesk get, what’s cymbal, how do I reach him. Then I read every answer.
An automated check can confirm that an answer mentions Leveldesk. It can’t tell whether the number next to it is right, or whether the answer quietly adds something I never said.
What I’d tell someone building this
- Keep the scope small. Each extra topic is another chance to be wrong.
- Give it a set response for gaps. Say so, then offer a related fact.
- Don’t let the model type anything that has to be exact. Tags in the reply, cards on the page.
- Keep one copy of the facts.
- Test prompt changes against real questions, and read the answers.