If you already use ChatGPT or Claude every day, it is reasonable to wonder why you would need a separate AI product for bookkeeping. These models are remarkably capable. They can write software, analyze documents, explain difficult concepts, help plan a business, generate marketing ideas, and reason through problems that would have seemed impossible for software just a few years ago. Gemini, Grok, and other leading AI models are becoming more capable for many of the same reasons.
We think these tools are extraordinary, and Bonnie itself would not be possible without the advances happening in modern artificial intelligence. But there is an important distinction between an AI model that can discuss bookkeeping and a system that can actually keep your books.
Much of that distinction comes down to consistency.
Bookkeeping has very little room for an AI that changes its mind
Large language models such as ChatGPT and Claude are designed to generate answers. They evaluate the information you give them and produce a response based on what they determine is likely to be useful or appropriate.
That flexibility is one of their greatest strengths. It is why ChatGPT can give you ten different marketing ideas, rewrite an email five different ways, help you think through a difficult decision, or explain the same concept differently depending on who is asking.
The same flexibility becomes a problem when you apply a general-purpose AI model directly to bookkeeping.
In technical terms, systems like these are not inherently deterministic. In plain English, that means you should not assume that asking the same question multiple times will always produce exactly the same answer. That may not matter much when you are brainstorming a company name. It matters quite a bit when you are deciding how a financial transaction should appear in your books.
This is well documented, and the companies building these models say so themselves. Anthropic’s documentation for Claude states that even at the setting designed to make output as repeatable as possible, “the results will not be fully deterministic,” and that identical inputs may produce different outputs. (Anthropic) OpenAI describes its chat completions as non-deterministic by default, and notes that even when developers use the tools built for reproducible output, the system makes a best effort and determinism is not guaranteed. (OpenAI)
Why it changes its mind
Here is the shortest honest explanation of what a large language model does.
Your phone’s keyboard suggests the next word as you type. It has no idea what you mean. It has simply seen enormous amounts of text and learned which words tend to follow which other words. A large language model works on that same principle, scaled up almost unimaginably far. At that scale, predicting the next word well begins to require something that looks a great deal like understanding, which is why the comparison to your keyboard stops being fair very quickly.
The underlying shape of the thing stays the same, though. The model produces what is most likely to come next, given everything it has seen. There is no rulebook it consults, and no record of the decisions your business has already made.
So when you ask it to categorize a transaction, the answer gets generated fresh each time. Ask again tomorrow, in a new conversation, with the words arranged slightly differently, and the answer can come out differently. Generating is the job it was built to do.
There is a second cause that surprises most people. Researchers at Thinking Machines Lab set out to explain why these models vary even when everything appears identical, and found that the largest factor has to do with how these services run behind the scenes. Your request gets bundled together with requests from other people, and the size of that bundle changes depending on how busy the servers are at that moment. Very small differences in the math ripple forward and occasionally change the final answer. The response you get can depend, in part, on how many other people happened to be using the service when you pressed send. (Thinking Machines Lab)
Everyone has watched their phone confidently suggest exactly the wrong word. The same basic mechanism explains why a general AI model is an uncertain place to store an accounting decision.
What that looks like in your books
Imagine hiring a bookkeeper and showing them a $247 purchase from Home Depot. The first time they review it, they classify the transaction as Supplies & Materials. A week later, you show them the same transaction and they decide it belongs in Cost of Goods Sold. The following month, another identical purchase gets classified as Repairs & Maintenance.
None of those categories is necessarily absurd. Depending on the business and the purpose of the purchase, each might even be defensible. That is exactly what makes it a problem. A good bookkeeping system should not repeatedly reconsider decisions that have already been made.
If your business has determined that a particular type of Home Depot purchase belongs in Supplies & Materials, you should be able to rely on that treatment continuing consistently unless something about the transaction changes or you decide to change the rule. You would expect that from a human bookkeeper, and you should expect it from an AI bookkeeper too.
The effect reaches well past any single line item. Bookkeeping earns its keep by letting you compare things. When the same purchase lands in three different categories over three months:
- Your profit and loss statement shows spending patterns that never happened.
- Comparing this quarter to last quarter tells you very little, because the categories shifted underneath you.
- A category that suddenly spikes sends you looking for a problem that does not exist.
- A category that quietly shrank hides one that does.
The hardest part is that the resulting numbers look completely reasonable. An inconsistent system rarely announces itself. It simply leads you, quietly, toward the wrong conclusion.
Consistency is one of several gaps
Even setting consistency aside, a chat assistant and a bookkeeping system are built to do structurally different jobs.
| General AI assistant | Purpose-built bookkeeping system | |
|---|---|---|
| Remembers your past decisions | Rarely, and usually not across conversations | Yes, by design |
| Sees your bank activity | Only what you paste in | Connected directly to your accounts |
| Records what changed and why | No | Yes, with a full audit trail |
| Maintains a double-entry ledger | No | Yes, underneath everything |
| Same question, same answer | No guarantee | Guaranteed where it matters |
| Built to hold financial data | It is a chat box | Yes, with controls around it |
The memory row deserves some attention. Every new conversation starts over. Where these tools do offer memory features, those features were designed to remember that you prefer short emails. Serving as the authoritative record of how your business has treated eleven months of vendor payments is a much larger responsibility.
There is also a practical question about what you are pasting into a general chat product to begin with. Your bank statements are among the most sensitive documents your business has, and they belong somewhere built to hold them.
What a general AI model is genuinely good for
We are not arguing that you should stop using ChatGPT or Claude. We use them constantly.
They are excellent at the work that surrounds your books:
- Explaining what a term on your financial statements actually means.
- Drafting the uncomfortable email to a customer who is sixty days late.
- Helping you think through whether to raise your rates.
- Summarizing a contract before you sign it.
- Reasoning through a business decision with you at eleven at night.
That is real value, and it is available to you today. The place these tools run into trouble is serving as the system of record for your business, where a decision gets made once and then holds.
What purpose-built actually means
The useful move is to stop asking one thing to do two very different jobs.
Some parts of bookkeeping call for judgment. What kind of business is this? What was this purchase actually for? Does this vendor mean something different for a general contractor than for a consultant? Those are exactly the sort of questions AI is good at.
Other parts call for certainty. Once a judgment has been made, the same transaction should be treated the same way every time, until you decide otherwise. That is the work of an ordinary rule, and ordinary rules have been dependable for a very long time.
A system built for bookkeeping does both. It uses AI to reach a decision, and it uses deterministic logic to make that decision stick. Your corrections become rules. Your rules apply consistently. The AI handles what is genuinely new and leaves settled questions settled.
A general-purpose chat model has none of that scaffolding around it. Nobody wired a rulebook to it, gave it your chart of accounts, connected it to your bank, or arranged for it to remember what you told it back in March.
How Bonnie incorporates deterministic rules and learns the more you use it
Bonnie splits the work along the line described above. The AI does the thinking on transactions Bonnie has not seen before, drawing on your chart of accounts and what it knows about the kind of business you run. Once a decision has been made, an ordinary rule takes over and holds it. When you recategorize something, Bonnie writes a rule tied to that vendor and to your business, and from that point on the rule gets checked before any model is asked anything. Matching transactions land in the same category every time. If you later correct that same vendor differently, the older rule is retired so there is only ever one answer in force. The effect compounds. Every correction you make is a question Bonnie will not bring back to you, so the amount of work waiting for your attention shrinks the longer you use it.
Rules you create outrank the ones Bonnie learns on its own, and yours are applied first. Every rule is visible to you on the Rules screen, including how many transactions each one has matched, and you can edit or switch off any of them. If you would rather Bonnie stop deciding a particular vendor altogether, you can tell it to always ask you, and transactions from that vendor will wait for you in Review with no model involved at all.
We made the rules visible on purpose. A lot of software advertises automatic categorization while keeping the logic behind it out of sight, and you can usually tell when you are dealing with one of those. The same wrong category keeps coming back month after month, and there is nothing on the screen you can reach to stop it. You find yourself correcting the same transaction in March that you already corrected in January, with no way to learn what the software believes about that vendor and no way to tell it otherwise. Bookkeeping is your record of your own business. You should be able to look at the rule behind any decision Bonnie makes, and change it.
The bottom line
A general-purpose AI model can talk about bookkeeping with real fluency. Keeping books asks for a few other things:
- The same transaction treated the same way every time.
- Decisions that stay made once you have made them.
- A connection to where your money actually moves.
- A record of what changed and why.
- An accounting structure underneath that holds it all together.
Every one of those is a property of a system rather than a property of a conversation.
Next up in this guide: is it worth building your own bookkeeper with AI?