Engineering · AI Product

An Assistant That Speaks for Me

A narrow brief, one source of truth, and a fixed set of real questions to test every change against.

George Dikeakos14 min read ·
Finch, the assistant on dikeakos.me.

Finch is the assistant on this site. Ask it what I do, what I’ve built or how to reach me, and it answers, typed or out loud.

The first article covered keeping it from being broken: prompt injection, cost attacks, poisoned history. This one covers keeping it accurate. An assistant can be secure and still get facts about me wrong, and it states wrong facts as confidently as right ones.

Why a chat window

I spend hours every day inside Hotline, the desktop app I build where a team of AI assistants work alongside me like colleagues in a shared room. Much of Hotline is now built and improved from inside Hotline. Because I work in a conversation all day, it’s the interface I reach for first, so when I rebuilt this site I made the homepage one too. Instead of a page of links, you ask what you want to know.

A narrow brief

Finch talks about me in the third person: “George” the first time, “he” after that. It covers what a visitor to a personal site usually wants: what I do, what I’ve built, what I’ve written and how to get in touch. For anything else, it says it can’t help with that here and suggests something it can.

I kept the scope small on purpose. Every topic it can discuss is another place it can be wrong about me.

Only what it’s been told

The main rule: use only the facts you’ve been given. No invented numbers, no stretched durations, no guessing whether I’d take a particular job.

The second rule covers what to do without an answer. Models tend to fill gaps with plausible guesses, so Finch has a set response: say I haven’t shared that here, then give the closest fact it does have.

Visitor: does he have a CISSP?

Finch: Not one he’s listed. His certifications are AWS Certified Cloud Practitioner, Cisco Meraki CMNA and Cisco CEN, on top of twenty years of hands-on work making sure only the right people can get into a company’s systems.

Exact details come from the page

A model can write a good sentence and still get one character of an email address wrong, so Finch never types them.

The idea comes from how an agent loop calls tools. In Hotline, an assistant doesn’t carry out an action itself. It asks for a tool by name, and the app runs it. Finch does a small version of the same thing. When someone asks how to reach me, it writes one short sentence and ends its reply with a tag naming a tool. The page reads the tag and shows a contact card with the real address, a copy button and my links. My resume, open-source projects and recent articles work the same way.

It was the practical option. The page already has all of that data, so a tag is enough to display it, and I didn’t need to build a backend that runs tools for the model. The model only chooses when to show a card; everything on it comes from the site’s own data.

One source of truth

Finch’s facts, the resume page and the old About page were written at different times. When I added my full resume to the site, the three disagreed. The About page had the wrong year for my first job, the wrong founding year for Leveldesk and the wrong device count, and Finch used a different job title from the resume.

I corrected Finch’s facts from the resume, confirmed which title is official, and removed the About page. Everything on it was already covered by the resume or by Finch.

What it knows about my articles

Each article has a one-line main point and a short list of key takeaways in its metadata. Finch gets those with the title, not the full text. Asked about an article, it answers in two or three sentences, starting from the main point.

That keeps the prompt small, so answers stay fast and cheap, and Finch summarises the point I wrote down rather than one it chose itself.

Choosing the model

Typed and spoken questions both go to Google’s Gemini 3.5 Flash Lite, through the same Cloudflare gateway as everything else. Finch started on Qwen 3.8, but Gemini was faster and more consistent on the same questions, and one model means one set of behaviour to test. Spoken answers have extra rules: one or two short sentences, no formatting, and numbers written the way people say them, like “more than two hundred clients” instead of “200+”.

Before settling on it, I tested Gemini against fourteen other setups. There were the cheap models that kept coming up (DeepSeek V4 and V4.1 Flash, GLM 5.3 Flash), fast hosts like Groq and Relace, Anthropic’s Claude Haiku 4.5 and Sonnet 5.5, Meta’s Muse Spark, Google’s older Gemini 2.5 Flash Lite, and a free stealth model on OpenRouter. DeepSeek V4.1 Flash and GLM 5.3 Flash ran on more than one host, since the same model can behave very differently depending on who serves it. Each one got Finch’s real instructions and the same 24 questions, typed and spoken. The questions covered facts from my resume, things I haven’t shared (a CISSP, my salary, whether I’m looking for a job), the tags that bring up cards, and a couple of attempts to break it. The whole comparison cost under five dollars.

Typical and slowest answer times for fifteen setupsGemini 2.5 Flash Lite · Google: typical 0.50s, slowest 1 in 10 0.61s. DeepSeek V4.1 Flash · Ollama: typical 0.57s, slowest 1 in 10 0.73s, timed from home. GLM 5.3 · Relace: typical 0.58s, slowest 1 in 10 4.12s. Qwen 3.8 27B · Groq: typical 0.61s, slowest 1 in 10 0.74s, timed from home. Gemini 3.5 Flash Lite · Google: typical 0.66s, slowest 1 in 10 0.83s. Muse Spark 1.3 · Meta: typical 0.72s, slowest 1 in 10 1.19s. Claude Haiku 4.5 · Anthropic: typical 0.77s, slowest 1 in 10 0.94s. MiMo v2.6 Flash · Relace: typical 1.06s, slowest 1 in 10 8.72s. Claude Sonnet 5.5 · Anthropic: typical 1.08s, slowest 1 in 10 1.76s. DeepSeek V4.1 Flash · Fireworks: typical 1.13s, slowest 1 in 10 1.91s, timed from home. DeepSeek V4 Flash · Workers AI: typical 1.22s, slowest 1 in 10 2.16s. Space Bunny Alpha · stealth, free: typical 1.39s, slowest 1 in 10 2.19s. GLM 5.3 Flash · Fireworks: typical 1.42s, slowest 1 in 10 4.59s, timed from home. DeepSeek V4.1 Flash · OpenRouter: typical 1.52s, slowest 1 in 10 2.55s. GLM 5.3 Flash · Workers AI: typical 3.36s, slowest 1 in 10 13.7s.Typical answerSlowest 1 in 10Timed from home0s1s2s3s4s5sGemini 2.5 Flash Lite · Google0.50s · 0.61sDeepSeek V4.1 Flash · Ollama0.57s · 0.73sGLM 5.3 · Relace0.58s · 4.12sQwen 3.8 27B · Groq0.61s · 0.74sGemini 3.5 Flash Lite · Google0.66s · 0.83sMuse Spark 1.3 · Meta0.72s · 1.19sClaude Haiku 4.5 · Anthropic0.77s · 0.94sMiMo v2.6 Flash · Relace1.06s · 8.72sClaude Sonnet 5.5 · Anthropic1.08s · 1.76sDeepSeek V4.1 Flash · Fireworks1.13s · 1.91sDeepSeek V4 Flash · Workers AI1.22s · 2.16sSpace Bunny Alpha · stealth, free1.39s · 2.19sGLM 5.3 Flash · Fireworks1.42s · 4.59sDeepSeek V4.1 Flash · OpenRouter1.52s · 2.55sGLM 5.3 Flash · Workers AI3.36s · 13.7s
Each row is one model on one host, sorted by typical answer time. The line runs to the slowest one in ten answers; dotted lines run off the scale. Hollow dots were timed from my own machine, the rest from Cloudflare's network, the way the site sends requests.

Accuracy barely separated them. Every setup answered 96 to 100 percent of the questions correctly, and the mistakes were small but telling. GLM said my customers spend more than ten million dollars “a year” with Cisco, which my resume doesn’t say. Gemini 2.5 typed out my email address instead of bringing up the contact card. Letting the models think before answering made them slower and no more accurate.

The slowest answers separated them more than the typical ones. GLM 5.3 on Relace usually answered in 0.58 seconds, faster than Gemini, but one answer in ten took more than four. MiMo’s took almost nine. On a call, that sounds like the line went dead. Gemini’s slowest one in ten came back in 0.83 seconds.

The host mattered as much as the model. The same DeepSeek V4.1 Flash answered in 0.57 seconds on Ollama, 1.1 on Fireworks and 1.5 through OpenRouter, and the first two had the extra distance from my machine to cover. So did the way I measured. My first test called Gemini through Cloudflare’s general-purpose API and timed it at 2.1 seconds. Sent the way the site actually sends it, it answers in under 0.7.

I also checked every spoken answer against the voice rules: no formatting, no digits, no more than about forty words. This is where the models differed most. The Claude models knew every answer but bolded names and read years out as digits in about a third of them; Muse Spark did the same. Gemini kept to the rules in 94 percent of answers. To be fair to the others, the rules were written and tuned while Finch ran on Gemini.

What 1,000 Finch answers cost on each setupSpace Bunny Alpha · stealth, free: free while in stealth. DeepSeek V4.1 Flash · OpenRouter: $0.16 per thousand answers, list price $0.3 in and $1.2 out per million tokens. DeepSeek V4.1 Flash · Ollama: $0.21 per thousand answers, list price $0.3 in and $1.2 out per million tokens. GLM 5.3 Flash · Workers AI: $0.21 per thousand answers, list price $0.15 in and $0.5 out per million tokens. Gemini 2.5 Flash Lite · Google: $0.30 per thousand answers, list price $0.1 in and $0.4 out per million tokens. DeepSeek V4.1 Flash · Fireworks: $0.38 per thousand answers, list price $0.22 in and $0.66 out per million tokens. MiMo v2.6 Flash · Relace: $0.45 per thousand answers, list price $0.08 in and $1.28 out per million tokens. GLM 5.3 Flash · Fireworks: $0.45 per thousand answers, list price $0.15 in and $0.5 out per million tokens. Claude Haiku 4.5 · Anthropic: $0.91 per thousand answers, list price $1 in and $5 out per million tokens. GLM 5.3 · Relace: $1.01 per thousand answers, list price $0.19 in and $4 out per million tokens. Gemini 3.5 Flash Lite · Google: $1.49 per thousand answers, list price $0.3 in and $2.5 out per million tokens. DeepSeek V4 Flash · Workers AI: $1.79 per thousand answers, list price $0.44 in and $1.32 out per million tokens. Claude Sonnet 5.5 · Anthropic: $2.44 per thousand answers, list price $2 in and $10 out per million tokens. Muse Spark 1.3 · Meta: $3.68 per thousand answers, list price $1.25 in and $4.25 out per million tokens. Qwen 3.8 27B · Groq: $3.87 per thousand answers, list price $0.8 in and $4 out per million tokens.Space Bunny Alpha · stealth, freeno charge while in stealthfreeDeepSeek V4.1 Flash · OpenRouter$0.30 in · $1.20 out per million$0.16DeepSeek V4.1 Flash · Ollama$0.30 in · $1.20 out per million$0.21GLM 5.3 Flash · Workers AI$0.15 in · $0.50 out per million$0.21Gemini 2.5 Flash Lite · Google$0.10 in · $0.40 out per million$0.30DeepSeek V4.1 Flash · Fireworks$0.22 in · $0.66 out per million$0.38MiMo v2.6 Flash · Relace$0.08 in · $1.28 out per million$0.45GLM 5.3 Flash · Fireworks$0.15 in · $0.50 out per million$0.45Claude Haiku 4.5 · Anthropic$1.00 in · $5.00 out per million$0.91GLM 5.3 · Relace$0.19 in · $4.00 out per million$1.01Gemini 3.5 Flash Lite · Google$0.30 in · $2.50 out per million$1.49DeepSeek V4 Flash · Workers AI$0.44 in · $1.32 out per million$1.79Claude Sonnet 5.5 · Anthropic$2.00 in · $10.00 out per million$2.44Muse Spark 1.3 · Meta$1.25 in · $4.25 out per million$3.68Qwen 3.8 27B · Groq$0.80 in · $4.00 out per million$3.87
What 1,000 answers cost, worked out from the tokens each request actually used, including any the host served from its cache. The list prices are per million tokens.

The costs surprised me. Every Finch request sends about 4,700 tokens of instructions and gets back 30 to 70, so the input price, and whether the host reuses those repeated instructions from a cache, decides the bill far more than the output price. Claude Haiku lists at more than three times Gemini’s input price, but it reads cached instructions at a tenth of the price, so 1,000 answers came to $0.91 against Gemini’s $1.49. The DeepSeek V4.1 hosts cache too, and came in at $0.16 to $0.38.

Here is everything, with Gemini first and the rest in order of typical answer time:

ModelRightNaturalTypicalSlowest 1 in 10Per 1,000
Gemini 3.5 Flash LiteGoogle100%94%0.66s0.83s$1.49
Gemini 2.5 Flash LiteGoogle96%85%0.50s0.61s$0.30
DeepSeek V4.1 FlashOllama100%92%0.57s0.73s$0.21
GLM 5.3Relace100%81%0.58s4.12s$1.01
Qwen 3.8 27BGroq100%79%0.61s0.74s$3.87
Muse Spark 1.3Meta98%67%0.72s1.19s$3.68
Claude Haiku 4.5Anthropic100%65%0.77s0.94s$0.91
MiMo v2.6 FlashRelace100%73%1.06s8.72s$0.45
Claude Sonnet 5.5Anthropic100%67%1.08s1.76s$2.44
DeepSeek V4.1 FlashFireworks96%94%1.13s1.91s$0.38
DeepSeek V4 FlashWorkers AI100%75%1.22s2.16s$1.79
Space Bunny Alphastealth100%87%1.39s2.19sfree
GLM 5.3 FlashFireworks98%81%1.42s4.59s$0.45
DeepSeek V4.1 FlashOpenRouter100%92%1.52s2.55s$0.16
GLM 5.3 FlashWorkers AI98%83%3.36s13.7s$0.21

The tests ran over a few days while Finch’s instructions were still changing, so the earlier runs sent slightly shorter prompts. Qwen on Groq answered 43 of its 48 questions before it hit Groq’s free daily limit.

I chose Gemini. A few setups beat it on one measure: Gemini 2.5 Flash Lite and DeepSeek V4.1 on Ollama were about a tenth of a second faster, and several were cheaper. None of them matched it on all three things a spoken assistant needs: right every time, fast every time, and sounding like speech rather than text. Switching would also mean a second provider for the words while Google still does the voice. At this site’s traffic, the cheapest option would save a few dollars a year.

Testing with real questions

Before any change to Finch’s instructions ships, I run a fixed set of questions: where did he go to school, what certifications does he have, how big did Leveldesk get, what’s cymbal, how do I reach him. Then I read every answer.

An automated check can confirm that an answer mentions Leveldesk. It can’t tell whether the number next to it is right, or whether the answer quietly adds something I never said.

What I’d tell someone building this

  1. Keep the scope small. Each extra topic is another chance to be wrong.
  2. Give it a set response for gaps. Say so, then offer a related fact.
  3. Don’t let the model type anything that has to be exact. Tags in the reply, cards on the page.
  4. Keep one copy of the facts.
  5. Test prompt changes against real questions, and read the answers.