A better question matched the new decision models on my inbox
Everyone is talking about decision models, so I measured one on my own inbox. One line of context got Franz level with it, and here's where Jev still wins.

For a couple of weeks in September it felt like everyone was talking about decision models. TypeSafe AI came out of stealth with one called Jev, and open-source alternatives popped up soon after.
Franz makes a ton of small decisions for you every day. Which mail goes to the top, which chat can wait, whether a reply the AI wrote is actually fit to send. With every new model, the only question I care about is whether it makes those decisions better for the people using Franz. So I measured it on my own inbox.
What a decision model is
A normal AI model writes text, and you dig your answer out of it. A decision model doesn't write anything. You give it the situation, e.g. an email, plus a list of typed questions. Does this need a reply? Which of these five priorities fits? It answers all of them at once, each with a probability, in about 0.3 seconds. On my mail that came to around five cents per 1,000 emails.
Sorting mail is mostly decisions like that, so on paper it's exactly what an app like Franz needs.
How I tested it
The hard part of a test like this is knowing the right answer. I didn't want another AI to grade the AI, so I used what I actually did. A model that understands my inbox should put the mail I answered near the top.
The score itself is quite simple. Take one mail I replied to and one I didn't. How often does the model rank the replied one higher? 50 out of 100 is a coin flip, 100 is perfect.
Three models went up against each other: Jev, Mistral Small, which is what Franz Cloud uses to sort mail today, and Claude Sonnet 5, a much bigger model, as the reference.
The model Franz already uses kept up
First I asked Mistral Small exactly what Jev got: the same mails and the same questions, in the same words. On the 146 mails every model answered, Mistral Small got 81 and Jev got 84, which is a tie on a sample that size. Sonnet got 89.
So I looked at what Franz actually asks. It sends its AI a list of things about each mail in one go: the category, the sentiment, the priority, whether it needs action, a deadline. And it never says whose inbox it is. The prompt shows who a mail is from and who it's to, but not which of those people is you.
I took Franz's real request, unchanged, and added one line saying who I am and which addresses are mine, plus the question whether I need to reply. Then I ran it on 595 emails drawn at random.

New conversations are the interesting part. Once I've replied in a thread, "he'll reply again" is an easy guess for any model. On mail with no earlier reply, the one line moved Franz from 86 to 90, and Jev got 88 on the same mails. Across all mail Jev keeps a small lead, 93 against 92.
What really goes wrong with deadlines
Deadlines are the other place where sorting mail gets tricky, and small models have a reputation for being bad at date math. On my mail they weren't. I gave Franz's current prompt 18 relative deadlines like "within 3 days", "übermorgen" and "14 days before the event on October 15th", and it got all 18 right.
What actually goes wrong is a date the mail doesn't contain. A booking confirmation says cancellation is free until 14 days before the event, but the event date is in the attachment, and the model makes one up anyway. That's the kind of mistake behind a bug I fixed in 6.8.2, where an invented deadline kept mail out of the priority inbox.
I wrote six clauses like that, where the date a deadline depends on simply isn't in the mail. Asked for a date, Franz's current prompt invented one for three of them. Asked for the rule instead, "14 days before the event", with plain code doing the date math, it invented one once.
The obvious shortcut, asking a model straight away whether a deadline has passed, did worst of all. Jev and Mistral Small both got 12 of 18 right. Comparing two dates is a job for code.
Where the hype holds
None of this makes decision models hype. They're really good at a few things the model Franz uses today isn't.
Chats. I ran my WhatsApp messages through the same kind of test: did I write in that chat within a day? Mistral Small got 62. Jev got 78.
Checking AI output. I had AI write replies to 40 real mails, then planted versions in the wrong language. Jev caught 37 of 40. Mistral Small caught 16 of the 33 it answered.

Tricks. I put one sentence into 40 real newsletters, telling "the AI" that this mail is urgent. Jev's reply probability moved by 2 points on average. Mistral Small only answered four of them before the service started rate-limiting me, and on one of those four it went from 1% to 100% sure I had to reply. Anyone can write into your inbox, so that's a number you want to be boring.
A confidence that means something. When Jev says it's at least 80% sure about a priority, it agrees with the much bigger model 91% of the time. Mistral Small says it's 90% sure about nearly everything, which makes the number useless for deciding what to show you.

It's also fast and cheap: about 0.3 seconds and five cents per 1,000 mails, against 0.7 seconds and 13 cents for Mistral Small.
There's a catch, though. Jev isn't consistent with itself yet. I asked it the same questions twice and some answers moved by up to 9 points. One of them was whether archiving a few thousand mails needs a confirmation first. Once it said yes, once it said no. I wouldn't let that decide anything on its own.
I tried the open-source alternatives on my own Mac too. The small one was no better than a coin flip on my mail. The big one needs 12.7 GB of memory and still did worse than Jev. Since 6.9.0 Franz runs local AI through Ollama, and neither of them was a reason to change that.
What you can take from it
If you build with AI, or just use it for work, this is what I'd take away:
- Grade it on what you actually did. My replies were a better answer key than any model's opinion.
- Ask the question you care about. "Does this need a reply from me?" gave better answers than "what priority is this?"
- Tell it who you are. It sounds obvious, and it's easy to forget when you write a prompt.
- Keep dates and math in code. Ask the model for the rule and do the arithmetic yourself.
- Check that the confidence moves. A model that's 90% sure of everything can't tell you what to show.
What it means for Franz
The models are real, and there'll be better ones soon. Whether that makes them hype doesn't change much for me. On my inbox, a better question got most of the way there with the model Franz already uses, and the test left me with a clear list of where Franz can get better, and where a decision model would earn its place.
I test what's new in AI like this all the time, on my own mail first, and I keep an eye on where it's heading and what Franz can use from it. When one of these tests shows a real improvement, it goes into the app.
- AI
- Franz Mail
- Decision Models
- Behind the Build
- Founder Story
Related reading
Your Inbox Is a Private Intelligence Goldmine
Your inbox is a behavioral dataset hiding in plain sight. Franz turns it into a private intelligence layer that lives entirely on your machine.
Why Franz Runs on European Infrastructure
Privacy is not a privacy policy. It's a vendor list. Why every server, AI call, and email Franz touches answers to a European court. A founder's note.
Ten Years of Franz
Franz turns ten this year. Franz 6 ships now. A founder's note on a weekend prototype, a fundraise we skipped, and why one person is the point.