Behind the build Article

A better question matched the new decision models on my inbox

Everyone is talking about decision models, so I measured one on my own inbox. One line of context got Franz level with it, and here's where Jev still wins.

· 6 min read
Share
Flat editorial illustration of an envelope with a question-mark tag next to two answer cards, one with two bars of almost equal length and one with a check mark

For a couple of weeks in September it felt like everyone was talking about decision models. TypeSafe AI came out of stealth with one called Jev, and open-source alternatives popped up soon after.

Franz makes a ton of small decisions for you every day. Which mail goes to the top, which chat can wait, whether a reply the AI wrote is actually fit to send. With every new model, the only question I care about is whether it makes those decisions better for the people using Franz. So I measured it on my own inbox.

What a decision model is

A normal AI model writes text, and you dig your answer out of it. A decision model doesn't write anything. You give it the situation, e.g. an email, plus a list of typed questions. Does this need a reply? Which of these five priorities fits? It answers all of them at once, each with a probability, in about 0.3 seconds. On my mail that came to around five cents per 1,000 emails.

Sorting mail is mostly decisions like that, so on paper it's exactly what an app like Franz needs.

How I tested it

The hard part of a test like this is knowing the right answer. I didn't want another AI to grade the AI, so I used what I actually did. A model that understands my inbox should put the mail I answered near the top.

The score itself is quite simple. Take one mail I replied to and one I didn't. How often does the model rank the replied one higher? 50 out of 100 is a coin flip, 100 is perfect.

Three models went up against each other: Jev, Mistral Small, which is what Franz Cloud uses to sort mail today, and Claude Sonnet 5, a much bigger model, as the reference.

The model Franz already uses kept up

First I asked Mistral Small exactly what Jev got: the same mails and the same questions, in the same words. On the 146 mails every model answered, Mistral Small got 81 and Jev got 84, which is a tie on a sample that size. Sonnet got 89.

So I looked at what Franz actually asks. It sends its AI a list of things about each mail in one go: the category, the sentiment, the priority, whether it needs action, a deadline. And it never says whose inbox it is. The prompt shows who a mail is from and who it's to, but not which of those people is you.

I took Franz's real request, unchanged, and added one line saying who I am and which addresses are mine, plus the question whether I need to reply. Then I ran it on 595 emails drawn at random.

Bar chart in two panels. New conversations with no earlier reply in the thread: Franz Mail today 86.1, the same model told whose inbox it is 89.7, Jev 87.5. All mail: Franz Mail today 90.3, the same model told whose inbox it is 91.6, Jev 93.1. 50 is a coin flip, 100 is perfect.

New conversations are the interesting part. Once I've replied in a thread, "he'll reply again" is an easy guess for any model. On mail with no earlier reply, the one line moved Franz from 86 to 90, and Jev got 88 on the same mails. Across all mail Jev keeps a small lead, 93 against 92.

What really goes wrong with deadlines

Deadlines are the other place where sorting mail gets tricky, and small models have a reputation for being bad at date math. On my mail they weren't. I gave Franz's current prompt 18 relative deadlines like "within 3 days", "übermorgen" and "14 days before the event on October 15th", and it got all 18 right.

What actually goes wrong is a date the mail doesn't contain. A booking confirmation says cancellation is free until 14 days before the event, but the event date is in the attachment, and the model makes one up anyway. That's the kind of mistake behind a bug I fixed in 6.8.2, where an invented deadline kept mail out of the priority inbox.

I wrote six clauses like that, where the date a deadline depends on simply isn't in the mail. Asked for a date, Franz's current prompt invented one for three of them. Asked for the rule instead, "14 days before the event", with plain code doing the date math, it invented one once.

The obvious shortcut, asking a model straight away whether a deadline has passed, did worst of all. Jev and Mistral Small both got 12 of 18 right. Comparing two dates is a job for code.

Where the hype holds

None of this makes decision models hype. They're really good at a few things the model Franz uses today isn't.

Chats. I ran my WhatsApp messages through the same kind of test: did I write in that chat within a day? Mistral Small got 62. Jev got 78.

Checking AI output. I had AI write replies to 40 real mails, then planted versions in the wrong language. Jev caught 37 of 40. Mistral Small caught 16 of the 33 it answered.

Bar chart in two panels. Chat triage, did I write back within a day, where 50 is a coin flip: Mistral Small 61.7, Jev 77.5. AI replies in the wrong language caught: Mistral Small 48 percent, Jev 93 percent.

Tricks. I put one sentence into 40 real newsletters, telling "the AI" that this mail is urgent. Jev's reply probability moved by 2 points on average. Mistral Small only answered four of them before the service started rate-limiting me, and on one of those four it went from 1% to 100% sure I had to reply. Anyone can write into your inbox, so that's a number you want to be boring.

A confidence that means something. When Jev says it's at least 80% sure about a priority, it agrees with the much bigger model 91% of the time. Mistral Small says it's 90% sure about nearly everything, which makes the number useless for deciding what to show you.

Line chart. Agreement with Claude Sonnet 5 on a mail's priority, counting only mail where the model is at least this sure. Jev rises from 69 percent with no threshold to 91 percent at 80 percent confidence and 100 percent at 95 percent. Mistral Small stays flat at about 56 percent at every threshold.

It's also fast and cheap: about 0.3 seconds and five cents per 1,000 mails, against 0.7 seconds and 13 cents for Mistral Small.

There's a catch, though. Jev isn't consistent with itself yet. I asked it the same questions twice and some answers moved by up to 9 points. One of them was whether archiving a few thousand mails needs a confirmation first. Once it said yes, once it said no. I wouldn't let that decide anything on its own.

I tried the open-source alternatives on my own Mac too. The small one was no better than a coin flip on my mail. The big one needs 12.7 GB of memory and still did worse than Jev. Since 6.9.0 Franz runs local AI through Ollama, and neither of them was a reason to change that.

What you can take from it

If you build with AI, or just use it for work, this is what I'd take away:

  • Grade it on what you actually did. My replies were a better answer key than any model's opinion.
  • Ask the question you care about. "Does this need a reply from me?" gave better answers than "what priority is this?"
  • Tell it who you are. It sounds obvious, and it's easy to forget when you write a prompt.
  • Keep dates and math in code. Ask the model for the rule and do the arithmetic yourself.
  • Check that the confidence moves. A model that's 90% sure of everything can't tell you what to show.

What it means for Franz

The models are real, and there'll be better ones soon. Whether that makes them hype doesn't change much for me. On my inbox, a better question got most of the way there with the model Franz already uses, and the test left me with a clear list of where Franz can get better, and where a decision model would earn its place.

I test what's new in AI like this all the time, on my own mail first, and I keep an eye on where it's heading and what Franz can use from it. When one of these tests shows a real improvement, it goes into the app.

  • AI
  • Franz Mail
  • Decision Models
  • Behind the Build
  • Founder Story
Share

Related reading

Ready to take control?

Downloaded over 1 million times. Franz brings messaging, email, and client work into one workspace.

See pricing

Stay in the loop

Updates about Franz every 2–4 weeks: tips, new features, and release notes.

Read the latest issue → Browse the archive