Start with the thing nobody in this category actually owns. Origin, an AI financial advisor, publishes its own stack in detail: Claude 4.1 Opus as the "primary reasoning agent", OpenAI GPT as "the system’s technical backbone", Gemini 2.5 for "real-time grounding", and Perplexity Sonar Pro as the "specialized retrieval backbone". Cleo names GPT-4o and Gemini Flash inside its own engineering posts. Breadmaxxer runs on frontier models too.
Which tells you the model is the commodity here. If it were the product, every app in the category would give you the same answer, because they are all calling the same handful of APIs. The question worth asking is what each one builds on top — and the honest news is that the serious ones have arrived at very similar answers.
The clearest illustration comes from a competitor. In a benchmark Cleo published, general-purpose models were given a real month of transactions and asked ordinary questions about them. Asked what the user spent on bills, Gemini Flash 2.5 answered $28,813. The correct figure was $3,047 — an error of more than $25,000, delivered fluently and with no hedging.
Across 21 scored data points, Cleo’s own system got 17 right, GPT-4o 13, Claude Sonnet 4 seven, and Gemini Flash 2.5 six. Treat the scores with the caution any vendor-run benchmark deserves — Cleo is measuring Cleo. But the failure mode it demonstrates is real, well documented elsewhere, and the whole reason this category exists: a language model does not calculate, it predicts text that looks like a calculation. Cleo’s own line is the best summary anyone has written of it: "If your AI can’t add, you probably shouldn’t ask it to manage your budget."
Their fix is the fix everyone lands on. Cleo describes an architecture that "routes all calculations through deterministic tools built for financial data", where "the LLM’s role is strictly interpretive — parsing the query, formatting the output — not doing the math". Origin describes a deterministic computational engine for the same reason. Breadmaxxer does not let a figure reach you unless it traces back to a real record. Three teams, one conclusion.
Read enough of these engineering posts and the same four decisions show up at companies that have never spoken to each other. In practice, this is what "vertical AI" means: a specific set of structures built around a rented model.
This is what building AI into any data-heavy profession looks like once the demo is over.
The pattern underneath all three is position. The product that wins is the one standing where the record gets made, because it ends up holding a copy of something nobody else can buy.
Here is the part most "vertical AI" writing skips. Once everyone makes the same four moves, those moves are the price of entry. What separates two well-built products is simpler and much harder to copy: what each one can see.
Look at what the money apps say their systems read, in their own words:
Read those three lists again. Every signal in them comes out of a bank feed or a chat window. That is the entire supply of data available to a general consumer money app, and all three are using it well. It is also the ceiling. A transaction can tell you $3,739 left your account across twelve days. Nothing in it says whether you worked them.
If you are building vertical AI into your own product, the four moves above are the buildable part — hard work, but known work. The strategic decision is upstream of all of it, and these are the questions worth answering before you write a prompt:
That last one is worth sitting with, because connectors are the fashionable answer right now and they are frequently the wrong one. Origin, arguing against a rival that plugs financial data into a general chatbot, makes the point better than we could: "a data layer, not a reasoning system. What it gives Claude or ChatGPT is access to your financial information. It doesn’t change how those models reason about it." Connecting more sources to a model that still cannot add gets you wrong answers about more things.
Breadmaxxer is built for people whose pay changes every week — servers, drivers, shift workers. It makes the same four moves as everyone else in this piece. The difference is the second input: alongside the bank feed, it has the shifts you log — hours, tips, tip-out, miles — which means it can join what you spent to what you earned and hand back one event instead of two unrelated charts.
That is the whole argument of this article, made concrete: not a better model, and not a better prompt, but a signal that was not available to anyone reading only a bank feed. If you want to see exactly what that buys — including a real card from a real account, and the guarantee we will be judged on — the next piece is the worked example.
Vertical AI means an AI product built for one industry or job rather than for everything. In practice it is not a custom model — nearly everyone rents the same foundation models — but a set of structures around one: calculations done in code, memory that persists between sessions, an agent loop that can take actions, and a layer that turns raw records into meaning. What ultimately separates two vertical products is which data each one owns.
Because a language model predicts text, it does not calculate. In a benchmark Cleo published, Gemini Flash 2.5 reported $28,813 of bill spending when the real figure was $3,047, and across 21 scored data points the general models got between six and 13 right. The fix every serious money app uses is the same: run the arithmetic in ordinary code and let the model only explain the result.
Anthropic draws the line clearly: agents are "systems where LLMs dynamically direct their own processes and tool usage", typically "using tools based on environmental feedback in a loop". A chatbot waits for you to ask and answers from what it was handed. An agent decides which tools to call, checks what came back, and can surface something you never asked about.
Yes, and not conversation history. Origin describes it as "structured financial context — if you reclassify a transaction, that correction persists and gets applied automatically going forward". Cleo built a retrieval system after finding that pasting whole transcripts into the prompt buried the important details. Without persistent memory you re-teach the product the same facts every week.
The architecture is not — the four common moves are all buildable by a competent team, and the leading apps have independently built them. What is defensible is a proprietary input: a record that exists only because your users create it inside your product. That is why the same pattern shows up in medicine and law, where the winning products sit where the record is made.
A second input. Other apps read a bank feed, and read it well; Breadmaxxer also reads the shifts you log — hours, tips, tip-out, miles — so it can join spending to earnings and tell you that a spending run and an earnings dip were the same twelve days. It also refuses to state a number it cannot trace back to your own data.