Short version: AI is good at reading messy documents and drafting text. It's unreliable at arithmetic that has consequences, and it's bad at knowing when it's wrong. So in the software we build, for paperwork and anything touching money, the AI reads, a person approves, and the software does the math. Here's that rule, the few places we deliberately let AI act on its own, the real failures that taught me, and five questions to ask anyone selling you an AI feature.

The house rule

The rule fits in one sentence: the model transcribes the numbers; the server does the arithmetic. In practice, that means three things.

  • AI reads and drafts. It pulls the vendor, date, and amounts off a crumpled receipt. It reads the terms out of a grant award letter. It drafts a summary, a listing, a reply.
  • Code does the math. Totals, tax, allocations, balances — anything that adds up or moves money — is handled by ordinary software that gives the same answer every time and can be tested.
  • A person approves the paperwork and anything touching money. An invoice, a quote, a report to a board, a change to live software. AI can prepare it. A person signs off.

None of this is anti-AI. It's how you'd treat a sharp new hire: let them do the reading and the first draft, check their work before it goes out, and don't hand them the checkbook in week one.

What AI is good at, and what it isn't

AI is good at the work small businesses drown in: reading documents that don't follow a format. A receipt photographed on a truck dashboard. A scanned PDF with a coffee ring on it. A certificate of insurance where every agency puts the expiration date somewhere different. A text from a crew member that says "can't do Tuesday but Wed works." Until recently, each of those needed a person, or a template that broke whenever the format changed. Now a model can read them in seconds, for a fraction of a cent each — the math is in what AI actually costs a small business.

It's weaker in two places that matter. It can misread a number or get a sum wrong. And it sounds exactly as sure of itself when it's wrong as when it's right — a wrong total doesn't come with a warning label. That combination is why we don't let it do the math, and why, outside the few logged steps below, it doesn't send paperwork or anything touching money on its own.

Where it shows up in software we've built

The pattern shows up across very different businesses:

  • Receipts and vendor invoices are read into expense line items that a person checks before they count. The Umbrella HQ case study shows this for a nonprofit bookkeeping firm, along with grant terms pulled from award letters.
  • Estimates. In Mr. Camera HQ, the system I built for the production company I run, you can type "add two camera ops for three days" and the AI turns it into estimate line items. It also reads uploaded vendor bids into the estimate. The totals come from the system, not the AI.
  • Certificates of insurance and W-9s. In Mr. Camera HQ, Claude reads uploaded certificates and W-9s for expiration dates, coverage, carrier, and tax details, and flags anything that needs attention for the office to check.
  • Crew replies to booking texts — "can't do Tuesday but Wed works" — are read as yes, no, or needs time. That's one of the few steps that runs on its own; more on it below.
  • A monthly financial summary for a nonprofit's board. The board financials app we built for a performing-arts nonprofit is built to draft a 150–250-word monthly summary for the accountant to edit and approve, from that month's figures only, before the board sees it.
  • Bug reports from client sites are sorted by AI in our client portal: severity, category, a short summary. An AI agent drafts the fix; nothing deploys until a person approves it.

Sometimes the right amount of AI is none. A payroll allocation tool we built for a grant-funded social-services nonprofit splits each employee's wages, taxes, workers' comp, and health insurance across about two dozen funding sources, and it uses no AI at all. It won't proceed until the imported payroll totals tie out, and the journal entry it produces has to balance to the penny with no hand-typed adjustments. When the whole job is arithmetic and rules, ordinary code is the right tool. That one gets its own case study.

Where we let AI act on its own, and why

The rule isn't "a person clicks every button." A few steps run on their own, every one of them is logged, and we chose each one on purpose:

  • Sorting bug reports. When a client reports a bug, Claude assigns a severity and category and writes a summary in our client portal without waiting for anyone. A mislabeled report is quick for a person to re-sort; making every report wait for someone to sort it would slow down every fix. Fixing is a different matter: an AI agent drafts the fix, and nothing deploys until a person approves it.
  • Off-script replies to crew booking texts. In Mr. Camera HQ, the software that runs my production company, and in Stickman HQ, crew answer booking texts. When someone replies off-script — "can't do Tuesday but Wed works" — the AI reads it as yes, no, or needs time, records it as accepted or declined (which can set off the next step, such as the deal memo), and texts a short reply on its own. Anything unclear goes to a person. Crew are answering a question we asked, there are only a few possible answers, and a prompt reply keeps a booking moving.
  • Clearance heights on WillIFit. Readings that pass the exact-match checks described below are published without a person; readings that conflict wait for one.

What those have in common: a small job, a handful of possible answers, a log a person can check, and a person on anything unclear. Invoices, quotes, payroll, board reports, and changes to live software don't fit that description, so a person approves them. If you're adding AI to your own business, start stricter than we do: don't let it auto-send anything to your customers at first.

What happened when I trusted the low-cost model too much

The clearest evidence I have comes from my own product, WillIFit, which checks parking-garage clearance heights for RV and truck drivers. It pulls street-level photos of garage entrances and has Claude find and read the clearance sign. Three lessons from building it shaped how we build everything else.

1. It read a speed-limit sign as a clearance height

In a $5 pilot, I let the lowest-cost model do the whole job. It got 4 of 11 readings wrong. One of them was a speed-limit sign it reported as the garage's clearance. On a clearance site, that's the kind of mistake that ends with an RV wedged under a concrete beam.

The fix wasn't to give up on the low-cost model. It was to give it a smaller job. Now it only answers one question — is there a clearance sign in this photo? — and a stronger model does the reading. It's never trusted to read a number.

2. The answer has to match its own transcription

The stronger model still has to show its work. It writes out the sign's text as it sees it, and the height it reports has to match that transcription exactly. If the number and the transcription disagree, the reading is thrown out. It's a simple check, written in ordinary code, aimed squarely at the kind of misread that got through in the pilot.

3. "High confidence" isn't the same as right

Later, I looked at cases where a new reading the model called high-confidence disagreed with an earlier reading of the same sign. About half of those high-confidence readings were wrong. So conflicting readings now go into quarantine for a person to settle, instead of letting the newest answer win.

That's the whole approach in miniature. The AI does the tedious looking. Code checks whatever can be checked. A person handles the cases where the machine disagrees with itself.

The patterns that make it safe

These are the building blocks we use. Not every feature needs all of them, but anything that touches money or customers should have most.

  • A review queue. AI output lands as a draft in a queue, not in your books. Nothing counts until someone approves it.
  • The source next to the answer. The reviewer sees the receipt photo or the page of the letter right beside each value the AI pulled from it, so checking is a glance, not a trip to the filing cabinet.
  • Flags on low confidence and on conflicts. Anything unreadable, unusual, or in disagreement with an earlier value gets marked for a closer look — keeping in mind the WillIFit lesson that the model's own confidence is a hint, not a guarantee.
  • Drafts, not sends, for customer paperwork. AI can draft the email, the quote, or the invoice. A person reviews it and clicks send.
  • Spending limits. Our client portal meters and caps the AI it runs for each client; AI inside your software runs under a spending limit stated in your quote, so a stuck job can't quietly run up a bill.
  • An audit log. What the AI proposed, what the person changed, who approved it, and when. When your accountant asks where a number came from, there's an answer.
  • Nothing deploys without a person's approval. In our own client portal, AI sorts bug reports and an AI agent drafts the fix; nothing deploys until a person approves it.

When not to use AI at all

If you handle five receipts a month, type them in. If the job is pure arithmetic or rules — payroll splits, tax calculations, inventory counts — use ordinary code. If a wrong answer could hurt someone and there's no practical way for a person to check it first, don't automate it. Our list of 12 jobs AI can do for a small business, and 4 it shouldn't walks through more of these.

Five questions to ask anyone selling you an AI feature

  1. What does the AI do, and what does regular code do? If the AI is calculating totals or anything that moves money, ask why.
  2. What happens when it's wrong? Who sees the output before it reaches your books or a customer, and how do they fix it?
  3. Can I see the source next to the answer? If checking means digging up the original document, nobody will check.
  4. What does it cost per use, and is there a cap? Ask who pays when usage spikes, and whether you'll be told first.
  5. Where does my data go? Which AI company processes it, and is it used to train their models? We cover this in is it safe to put customer data into ChatGPT or Claude.

A vendor who answers all five plainly has probably thought about what happens when the AI gets it wrong. One who answers with a demo probably hasn't. For the bigger picture on where AI fits in a small business, see our practical guide to AI for small business.

Frequently asked questions

Will AI make mistakes in my books?

It can misread a document, and it doesn't always know when it has. That's why AI output headed for your books lands as a draft for a person to check, with the source document beside it, and why totals are calculated by ordinary code rather than by the AI.

Why not let AI do the math?

Ordinary code gives the same answer every time and can be tested; AI guarantees neither. So the AI reads the numbers off the page and the software adds them up. Our rule is that the model transcribes the numbers and the server does the arithmetic.

Does a person have to approve everything the AI does?

Not everything, and where it doesn't is deliberate. In paperwork and anything touching money, AI drafts and a person approves: invoices, quotes, board reports, and changes to live software. A few steps run on their own and are logged — sorting bug reports, and in our production-company software, reading a crew member's off-script text as yes, no, or needs time and sending a short reply — with anything unclear going to a person.

What is a review queue?

A list of AI-prepared items waiting for a person. Each item shows what the AI extracted or drafted next to the original document, with anything unusual flagged. The person approves, edits, or rejects it, and only approved items count.

What should I ask a vendor selling an AI feature?

Ask what the AI does versus what regular code does, what happens when it's wrong, whether you can see the source next to each answer, what it costs per use and whether there's a cap, and where your data goes and whether it's used to train anyone's models.