AIARTICLE

The chatbot started making mistakes. The problem was in the company.

The chatbot started making mistakes. The problem was in the company.
Image: Cezar Taurion

I closely followed a project that showed me clearly, in practice, why a good AI demo is very far from being a system ready to serve real customers.

It was at a furniture and appliance retail company. The goal seemed simple: create a chatbot capable of answering questions about products and helping the consumer choose a refrigerator, TV, sofa or washing machine.

In controlled tests, it worked very well. The problem started when real customers began talking to it. And it was precisely real usage that revealed problems that traditional tests had not captured.

The first shock came from the data. The company began noticing customer complaints about product characteristics. When reconstructing the conversations and comparing the responses with corporate sources, curious situations were discovered. The bot recommended a refrigerator saying it had 480 liters, while another database recorded 462. In some cases it confused packaging dimensions with product dimensions. In others, it found different attributes in the ERP, in the digital catalog and in the description provided by the manufacturer.

The LLM had not created these inconsistencies. It simply made public and conversational a data disorganization that already existed.

The first fix, therefore, did not happen in the model. They had to define authoritative sources for each attribute, review taxonomies, identify missing and conflicting fields and establish rules about which information could be used in responses. When there was not sufficiently reliable information, the bot should assume the uncertainty, not fill it in with a plausible answer.

Then integration problems appeared. The customer would ask for a washing machine suited to their needs and available for delivery in their city. The bot would find an excellent product in the catalog, but stock, logistics and delivery time were in different systems. Some recommendations were technically good and commercially useless: the product was not available in that region or the delivery time shown had already changed.

This was discovered by comparing recommendations logged in the records with actual orders, stock and delivery simulations.

It was necessary to separate two things that initially seemed to be one: recommending a product and stating its availability. The first could use the catalog and characteristics. The second required a real-time query to the transactional systems. When that query failed, the agent could not simply continue as if it had obtained the information.

Then came the questions that no test script had imagined.

"Can this sofa handle my three kids jumping on it every day?"

"Does this refrigerator fit in my building's elevator?"

"Can I put this TV on the balcony where the afternoon sun hits?"

That's when another problem in the tests was noticed. What had mainly been tested was what they expected customers to ask. Real customers had not taken part in drafting their test script.

They then began analyzing real conversations, grouping questions that produced problematic responses and creating evaluation sets with ambiguous, incomplete, adversarial and out-of-scope cases. Failures found in production began to continuously feed the evals.

Response classes were also defined. For certain objective attributes, the bot could only answer when it found evidence in an authorized source. For subjective situations or ones with insufficient information, it should state the limitation. For topics with higher risk or ambiguity, it would transfer the conversation to a person.

This seems trivial. It is not. LLMs are very good at producing plausible answers even when the available information is insufficient. In customer service, a polite, convincing and wrong answer can be more dangerous than a simple "I don't know."

Unexpected behaviors also emerged when customers insisted, rephrased questions, contradicted the bot or tried to take it to subjects completely out of context. They discovered that guardrails based predominantly on prompts were insufficient.

Additional scope controls, validation of responses in specific situations, limits on certain statements and rules external to the model were necessary. The principle became simple: not everything the model can answer means the system should allow it to answer.

And another problem emerged that had not been considered: latency. When instrumenting the end-to-end flow, they noticed that an apparently simple question could trigger several operations: intent classification, context retrieval, product search, API calls, validation and response generation.

The customer saw none of this. They just waited. Observability showed where the time was being consumed. Some queries could run in parallel. Others were redundant. Certain questions did not need the more sophisticated model. Some information could use caching with appropriate validity criteria. Retries needed to be limited.

At certain points, the best optimization was simply to eliminate a call to the LLM. The combination of incorrect answers, inconsistencies, integration problems and excessive latencies led to the decision to temporarily suspend automated service and go back to engineering.

What changed. Testing stopped just evaluating whether "the chatbot answers well." It began measuring factual quality, source coverage, integration availability, latency, tool failures, out-of-scope behavior, transfer to humans and the ability to recover when some dependency did not work.

And something even more important was created: a failure taxonomy. Model error, retrieval error, inconsistent data, unavailable integration, inadequate business rule, insufficient information and orchestration failure stopped being lumped into the same box generically called "AI error."

That changed the discussion. Putting an LLM in front of the customer is relatively easy. Putting the company in front of the customer through an LLM is much harder. The chatbot ends up translating into natural language everything that exists behind it: data, integrations, rules, processes, legacy systems and governance decisions.

And that was the biggest lesson from that project. When the chatbot started making mistakes, at first it seemed like there was an AI problem. After investigating, something much more interesting was discovered: the AI was making visible, at scale and in front of the customer, problems the company already had for a long time.

Translated from the Brazilian Portuguese original · Read the original

More from Cezar Taurion
View profile →
Read also
AI

Does your site exist for AI? Start with the terminal

A site can work normally in the browser and still be inaccessible to AI crawlers. This article presents terminal tests to identify blocks by firewall or CDN and to check whether the content is in the HTML sent by the server. It also explains the order of Technical GEO checks: access, reading, interpretation, and measurement.

Leandro Vieira··1 min