← All posts

How we categorize with a confidence score, and what to do about the red ones.

The Marginmoth agent scores its own work on every transaction. Here is what the score means, what lands in your review queue, and how the feed back improves over months.

Most "AI bookkeeping" products advertise their accuracy in aggregate and quietly re-classify the wrong ones in the background. That is exactly the failure mode a CPA will (rightly) refuse to sign off on. We do it differently: every transaction carries a per-line confidence score in the close, and anything below your threshold lands in a review queue instead of being silently overridden.

The score is calibrated against the prior twelve months of transactions in your books. Stripe payouts against your standard revenue categories run above 98 percent. AWS sees the same pattern month after month and sits comfortably above 95. Uber Eats, Lyft, and WeWork show up in slightly different shapes depending on the merchant ID and the period, so they sit in the 84-92 range. Anything below 90 lands in the queue with the agents reasoning spelled out so you can tag it in seconds.

This is what makes the close auditable: the close itself contains a confidence column per line, and your bookkeeper signs off knowing which lines need a humans eye. As you tag the queue, the corrections feed the category model for next month, so your scores trend up over months, not down. After three or four months of corrections most clients are clearing the queue in a few minutes a week.

We refuse to publish a single accuracy percent across all clients and all categories because that number is meaningless. The honest number is per line, per client, per category, and the close you receive shows both the score and the tag.