Why a Chatbot Cannot Be Tested Like a Credit Score

In early 2024 the parcel firm DPD switched off part of its online chat assistant after a frustrated customer coaxed it into swearing and writing a poem about how useless it was. A few weeks later a Canadian tribunal ordered Air Canada to compensate a passenger who had relied on its website chatbot’s assurance that a bereavement fare could be claimed retrospectively, which the airline’s actual policy did not allow. Neither system had been hacked. Both did what conversational software occasionally does, only in front of a customer.

Businesses that have long used predictive models often assume the same testing habits will carry over. A credit model takes a fixed set of inputs and returns a number, and the same inputs always produce the same number, so it can be checked against thousands of past cases where the right answer is known. A language model breaks each of those assumptions. Ask it the same question twice and the wording will differ, and now and then so will the substance. There is no single correct reply to compare against, and a response can be perfectly fluent and still wrong. Testing therefore means running the same prompts many times, scoring the replies against written criteria, and reporting how often things go wrong.

Invented answers are the best-known failure. Most customer-facing assistants are connected to the company’s own policies and help pages so that replies are grounded in real documents, which reduces fabrication without removing it. A good test checks whether each claim in an answer can be traced back to a source, what the assistant does when the documents are silent or contradict one another, and whether it hands over to a person at the right moment. Often the culprit is an out-of-date PDF in the knowledge base, not the model.

Then there is a weakness that predictive models do not have. A credit score cannot be argued with, but a chat assistant receives its instructions and its conversation through the same channel, so a user can tell it to ignore its rules, and so can text buried in an email or web page it has been asked to summarise. No complete cure for this is known. An independent AI audit of a generative system therefore borrows as much from security testing as from statistics, with reviewers playing both the awkward customer and the deliberate attacker. An assistant that only talks can embarrass its owner, while one that can issue refunds or open account records can cause real loss.

A conventional model drifts because the world around it changes. A chat assistant usually runs on a model rented from an outside provider, which updates and retires versions on its own timetable, so behaviour can shift even though nobody in the business has touched a line of code. DPD attributed its incident to an error following a system update. The practical defence is a fixed library of test conversations that is rerun whenever the underlying model, the instructions or the knowledge base changes. The same library is worth running on a schedule even when nothing appears to have changed, because providers do not always announce their adjustments.

An outside review can say how often the assistant failed across the scenarios tried, which safeguards held, and whether the firm would notice a problem itself. It cannot certify that the assistant will never say something untrue, because the number of possible conversations has no limit and the findings apply only to the version examined. What management receives is a set of failure rates and known weak spots, to be judged against the job in hand. An assistant that gives out opening hours and one that explains mortgage terms may share the same error rate and still deserve very different verdicts.

Similar Articles

Comments

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular