Teaching an AI to Sound Like Your Business
Guide

Reading the AI Audit Log: Cost, Tools and Quality

4 أغسطس 2026 · 4 دقائق قراءة
A spreadsheet of figures being reviewed on a desk
Photo: Unsplash

Teams spend a week configuring an AI and then never look at what it does, on the reasonable-sounding grounds that no customer has complained. That is a poor signal: most people who get a bad automated answer do not complain, they go quiet. Our guide to teaching an AI to sound like your business ends here because the log is the only place the truth is recorded.

What the log actually contains

Every AI reply leaves a record: the customer's message, the reply, which tools were called and with what arguments, which documents were retrieved, how many tokens were used and roughly what that cost. Together these answer the question you cannot answer from the transcript alone — not just what the AI said, but why.

That distinction matters. A wrong answer that came from a correctly retrieved document is a content problem. The same wrong answer with nothing retrieved is a coverage problem. They look identical to the customer and need completely different fixes.

A weekly half hour

You do not need to read everything. Sample deliberately.

  1. Every conversation that ended in a handoff — these are the AI's own admissions of failure and the fastest route to missing documents.
  2. Five random conversations it completed alone, read end to end. Completed does not mean correct, and this is where you find the confidently wrong ones.
  3. Anything unusually long. Fifteen exchanges to answer one question means the customer kept rephrasing because the answer was not landing.
  4. The most expensive conversations of the week, which usually reveal a loop or a tool being called repeatedly for no reason.

Thirty minutes covers all four in a small business, and the second item is the one worth protecting when you are busy.

What the patterns mean

Individual conversations are anecdotes; the shape across a week is information. A rising handoff rate on one topic means a gap in your documents. Tools that never get called mean either the AI does not know when to use them or they are not needed — either way, a decision to make. Retrieval returning nothing on common questions means your knowledge base and your customers' vocabulary have drifted apart, which is a rewrite job of the kind described in building a knowledge base your AI can actually use.

Watch the cost line as a diagnostic rather than a bill. A conversation that costs five times the average is usually not a complicated customer; it is a loop, a retry, or an enormous amount of retrieved context being sent on every turn. The economics of that are covered in what an AI reply actually costs you.

Nobody complains about a mediocre automated answer. They just stop replying — which is why the log matters more than the inbox.

be digital ai team

Score a sample properly

Reading with a gut feeling produces a gut conclusion. Score each sampled conversation on three things instead: was it factually correct, did it sound like your business, and did it move the customer forward. Three binary judgements, twenty conversations, ten minutes.

Track the totals month to month. It is a crude measure and it is enough — the value is in noticing that correctness dropped from nineteen out of twenty to fifteen after you changed a prompt, which no amount of impressionistic reviewing would have caught.

Close the loop

A review that does not produce changes is a hobby. Each session should end with a short list: documents to write, rules to adjust, capabilities to turn off. Make one change at a time and check the effect in the next review, for the same reason described in writing system rules and a style guide for your AI — batch changes teach you nothing about cause.

Who should read it

In most small businesses this lands on the owner, and it should not stay there. The person who answers customers all day recognises a wrong answer faster than anyone, and they are the ones who know what customers actually meant. Give an experienced agent the sampling job, let them flag the bad ones, and keep the configuration decisions with whoever owns the setup. It takes half an hour a week and it is the difference between an AI that improves and one that quietly gets worse.

See it live

A 20-minute walkthrough of the be digital ai conversation log — tools, cost and quality in one place.

Book a Demo

المزيد من هذه السلسلة