Livia Sannaro

Back to all sessions

Lecture 3

Spot invented details in plausible studio summaries

ModelsEvidence

Prerequisites: Lectures 1 and 2. You should already know how to record an ordinary query, classify the first shape of the answer, and build a small source trail from public evidence. In this lecture we use those habits to look at a narrower problem: the invented detail that sounds too ordinary to notice.

A shop owner asks an AI assistant for “un commercialista per aprire una piccola società e gestire la contabilità.” The answer names a local studio and says, in clean Italian, that it “specializes in company setup and ongoing accounting for new SRLs.” The sentence has no fireworks. It sounds like the sort of thing any studio might put on a service page. The owner reads it once, accepts it, and sends a message asking for help with incorporation.

In the teaching example I use for this lecture, the studio does handle recurring accounting for small companies. It does not present itself as a company-setup specialist. Its website mentions new clients only in a paragraph about document intake. A directory lists “business consulting” as a broad category, and one review says the studio “helped us after we opened the company.” The AI answer has not invented a purple elephant. It has added one small professional badge. That is why the mistake is dangerous.

The most convincing inventions are boring

When people hear the word hallucination, they often imagine something spectacular: a fake address in another city, a partner who never existed, a service no accounting studio would offer. Those things can happen, but the everyday problem is usually duller. A model adds a service emphasis, a client type, a credential flavor, or a location cue that fits the surrounding language so well that nobody stops to check it.

Before we use the term, we need a smaller unit. A business fact is a stable, checkable claim about name, address, service scope, credentials, or contact details. “The studio is in Verona” is a business fact if it can be checked. “The studio supports payroll coordination” can be a business fact if the public material says that clearly. “The studio is known for helping new SRLs” is trickier. It may be a claim about reputation, service scope, and client type at once. That mixed shape makes it easy to repeat and hard to verify.

A hallucination is a confident model statement that is unsupported, wrong, or filled in from pattern. I am not using the word to scold the machine. I am using it as a practical label for a check: does this sentence have support in the source trail, or did the model make a smooth bridge over a gap?

The boring inventions matter because local professional services are built from boring words. Accounting, payroll, VAT, declarations, advisory, company setup, compliance. These words live near each other. If a studio’s public evidence says three of them clearly and one of them faintly, an answer may write a sentence that treats all four as equally firm. The sentence still smells like an accounting studio. That is exactly the problem.

Look for the extra hinge in the sentence

A useful way to inspect a plausible summary is to find the hinge word: the word or phrase that turns a visible fact into a stronger claim. In the opening example, “ongoing accounting” may be supported. “For new SRLs” may be partly suggested by the user’s query. “Specializes in company setup” is the hinge. It takes a nearby possibility and makes it sound like the studio’s identity.

The hinge is often small. “Also helps with” becomes “specializes in.” “Has experience with” becomes “is known for.” “Located near” becomes “serving the whole province.” “Supports payroll coordination” becomes “payroll office.” If the reader is busy, these shifts pass like a cashier sliding an extra receipt into the bag. Later, when the wrong client calls, the studio has to sort the receipt from the purchase.

In a recurrent pattern from local service descriptions, AI answers tend to prefer complete phrases over careful fragments. A studio page might say, “We assist companies during ordinary accounting management, including VAT deadlines and coordination with payroll consultants where required.” That sentence is cautious and accurate, but it is not easy to compress. A model may produce: “The studio offers bookkeeping, VAT, and payroll services.” The last phrase is not always wrong, yet it may remove the boundary around payroll coordination.

To check the hinge, write the answer sentence in plain pieces. Name. Location. Service. Client type. Credential. Reputation phrase. Contact detail. Then ask which pieces are directly visible in the public evidence and which pieces are only implied. This does not require a complicated audit. It requires a pencil and a little suspicion toward tidy grammar.

There is a small human discomfort here. The sentence may feel close enough. Many professionals do not want to argue with a phrase that is only slightly inflated. But “close enough” has a cost in accounting work. The wrong service frame attracts the wrong expectation, and expectation is already part of client intake.

Pattern filling is not the same as finding

A large language model writes by continuing patterns. That is useful when the pattern is language: it can explain, summarize, translate, and organize. It becomes risky when a pattern starts to stand in for a checked business fact. The model knows that “accounting studio,” “small company,” “SRL,” “tax declarations,” and “company setup” often travel together in text. If the public evidence is thin, the answer may lean on that familiar cluster.

This is different from finding a fact in a source trail. Suppose the studio website has a page titled “Costituzione società e apertura partita IVA” and the body names the service clearly. In that case, an answer mentioning company setup has visible support. Suppose instead the website only says, “We help new clients organize the first documents for accounting management,” while the user’s query asks about opening a company. If the answer calls the studio a setup specialist, the support is weak. The model has filled the space between the query and the studio’s general service language.

The distinction is not always neat. Retrieval can bring in a page that contains broad category language, and the model can still overstate it. Training data can give the model a general pattern, and the answer can still resemble a true service. We do not need to solve the whole hidden process to do the practical check. We need to ask whether the public evidence can carry the weight of the claim.

Here is a simplified teaching example. The answer says, “Studio X is suitable for restaurants that need accounting and payroll.” The source trail shows a review from a restaurant owner, a website page about small companies, and a directory category for payroll. The restaurant part is based on one client memory. The small-company part is supported. The payroll part may need closer reading. The generated sentence has stitched three levels of evidence into one smooth recommendation.

Unsupported does not always mean completely false. A studio may coordinate payroll through a consultant, answer client payroll questions, and keep documents aligned for monthly processing. The phrase “specialized in payroll” may still be too strong if the studio does not present payroll as its main service. The issue is not that the word came from nowhere. The issue is that the generated answer gave it the wrong role.

I prefer marking these cases in three ordinary phrases: supported, partly supported, unsupported. “Supported” means the public evidence clearly carries the claim. “Partly supported” means some words are visible, but the answer strengthens or reframes them. “Unsupported” means the public trail does not give the claim a safe footing. These are not formal scores. They are working labels a studio can actually use.

The query can donate details too

A strange thing happens in many AI checks: the user’s question lends words to the answer. If someone asks, “Which studio helps new companies with accounting in my town?”, the model may answer in a way that makes “new companies” sound like a property of the named studio. The question has carried a detail into the summary.

That does not make the answer useless. It means we must read it as an answer to that exact query. In Lecture 1 we began recording the ordinary query because first discovery depends on real client language. Here the same habit protects us from a false accusation. A model may not have invented “company setup” from public evidence; it may have accepted the user’s frame and attached it too firmly to the studio.

Try a pair of teaching queries. First: “commercialista per piccola SRL con contabilità continuativa.” Second: “studio commercialista specializzato in apertura società per stranieri.” If the same studio appears in both answers with different descriptions, we should not immediately conclude that the public evidence says both things equally. The query may be pulling the answer toward different service frames.

This is why the first check should not be a courtroom. It is more like checking whether a label stuck to the correct folder. You record the query, the answer, and the public evidence. Then you ask: is the disputed phrase visible in the evidence, donated by the query, or filled in from a common pattern? Sometimes you can only say “unclear.” That is an honest result, not a failure.

One small warning: do not clean the query so much that it stops sounding like a client. If you ask only perfect consultant questions, you will learn how AI answers perfect consultant questions. A studio needs to know what happens when a tired owner writes a messy request after dinner.

What to do in the first pass

In the first pass, do not rewrite the website. Do not message every directory. Do not ask the AI assistant ten more questions until it gives a better sentence. Begin with one answer and mark every business fact. Put a small sign beside each one: visible, partly visible, or not visible in the source trail.

Then look at the unsupported or partly supported phrases. Are they service claims? Location claims? Credential claims? Client-type claims? The category matters because some mistakes are more urgent than others. A wrong contact detail can block a client immediately. A wrong service emphasis may waste intake time. A vague reputation phrase may be annoying but less actionable.

A teaching example might end like this. The answer correctly names the studio and town. It says the studio handles recurring accounting, which the website supports. It adds “company setup specialist,” which is only weakly suggested. It says “for new SRLs,” which came mostly from the query. It says “with payroll services,” where the evidence shows coordination but not direct payroll office work. The answer is not garbage. It is a tidy little cabinet with two drawers mislabeled.

That is the level of diagnosis we need at Lecture 3. We are not yet deciding how to correct every source. We are learning to notice where a plausible sentence stops being a careful description and starts becoming a confident completion.

What matters to remember

Invented details in AI summaries are often ordinary professional phrases, not dramatic fantasies. Their danger is that they sound like normal studio language.

A business fact is checkable. When an AI answer adds service scope, client type, credential language, location, or contact detail, mark whether the source trail supports that exact claim.

Hallucination, in this course, is a practical label for a confident statement that lacks safe support. The useful question is not “did the model behave badly?” but “can the public evidence carry this sentence?”

The same course anchor still holds: four ways an AI answer reshapes a small accounting studio — names the practice, narrows the service, borrows nearby evidence, or leaves the firm unmentioned. In this lecture, the main risk sits inside narrowing and borrowing, where a plausible extra detail can change the studio’s role.

A first pass should classify claims as supported, partly supported, or unsupported. That is enough to begin without pretending we know the model’s whole hidden process.

Check yourself

Explain in your own words why a plausible invented detail can be harder to catch than an obviously false one.

A plausible invented detail is hard to catch because it fits the normal language of the profession. If an AI answer says a studio helps with company setup, payroll, or tax declarations, nothing sounds strange at first. Those are ordinary accounting-related phrases. The problem appears only when we compare the sentence with the studio’s actual public evidence. A wildly wrong address would attract attention immediately. A slightly overstated service may pass unnoticed, then shape the client’s expectation before the studio can explain its real work.

Give an example from a small accounting studio where one hinge phrase changes the meaning of an AI summary.

A studio website might say it helps small companies with recurring accounting and coordinates payroll documents with an external consultant. An AI answer could summarize that as “a payroll specialist for small businesses.” The hinge phrase is “specialist.” It turns a supporting activity into the main identity of the studio. The word payroll is not invented from nothing, but the sentence gives it too much importance. A client who needs direct payroll processing may call, while a client who needs broader recurring accounting may think the studio is too narrow.

How would you distinguish a supported business fact from a partly supported claim in an AI answer?

I would look for the exact claim in the source trail and ask whether the public evidence can carry the whole sentence. If the answer says the studio is located at a certain address and the website and listings show that address, the fact is supported. If the answer says the studio specializes in new company setup, but the evidence only mentions organizing documents for new accounting clients, the claim is partly supported at best. Some words are nearby, but the answer has strengthened them into a bigger service identity.

When should you blame the query itself for donating a detail to the answer?

I would suspect the query has donated a detail when the answer repeats the user’s wording as if it were a fact about the studio. For example, if the question asks for a commercialista “for opening a new SRL,” the answer may describe a named studio as suitable for new SRLs even when the public evidence mainly supports ordinary accounting. That does not prove the studio’s sources are wrong. It shows that the answer must be read together with the exact query, because the user’s frame may have shaped the summary.

How would you explain hallucination to a studio partner without making it sound mystical?

I would say that a hallucination is a confident sentence that the evidence does not safely support. It is not magic and it is not necessarily a ridiculous invention. The model may use common accounting patterns and produce a phrase that sounds normal, such as “company setup specialist” or “payroll office.” Our job is to check whether that phrase is visible in public evidence, partly suggested, or unsupported. This keeps the discussion practical: we inspect claims, instead of arguing with the machine’s personality.