Livia Sannaro

Back to all sessions

Lecture 6

Compare engine differences without chasing certainty

ModelsEvidence

Prerequisites: Lectures 4 and 5. You should already be able to sketch a studio view behind one AI answer and record whether the answer came from memory or live search. This lecture adds comparison: what changes when more than one AI system answers the same studio query.

I once printed four short AI answers about the same small accounting studio and put them side by side with a red pencil. One answer named the studio and called it “practical for payroll.” One named a larger firm in the next town. One produced a tidy list of three providers but left the studio out. One gave the right name, then added a line about company formation that the studio never claimed. The strange part was how calm they all sounded.

This is a teaching example, not a formal audit. Still, it captures a recurrent pattern in this work. When a studio owner sees four different answers, the first reaction is often to ask which engine is “right.” I understand the impulse. A practice has limited time; it wants a verdict. But the useful lesson usually sits before the verdict: which parts of the studio become easy to name, which parts shift, and which parts disappear when the tool changes.

Variation is a signal, not a scoreboard

Engine variation means differences between AI systems on the same studio query. The practical meaning is simple. If three AI systems answer the same question differently, you have learned something about the instability of the studio view. You have not learned which system is the judge of reality.

A small accounting studio can be especially exposed to this. Its public evidence is often thin, local, and uneven. The website may be careful, but short. A directory may use a broad category. A review may praise one concrete task, like payslips, because clients remember the thing they touched. A chamber record may hold the legal form but no warm explanation of the work. Different systems can lean on different fragments, or on no visible fragment at all.

The trap is turning comparison into a race. “Engine A found us, engine B did not, so engine A is better.” Maybe for that query, on that day, in that mode. But for our course, the better question is more modest: what changed between the answers? Did the name remain stable? Did the town remain stable? Did the service scope narrow? Did a vague phrase appear in every answer? Did only the live-search answer find the current wording?

That modesty is not timid. It is more useful than a dramatic conclusion. A studio cannot build a monthly habit on outrage. It can build one on repeated observations.

Hold the query still before you compare

If you change the query, the engine, and the mode all at once, the comparison collapses. You may still feel that you have done research, but the result is a drawer full of mixed receipts. For this lecture, hold the query as still as possible.

A simple comparison starts with one client-style question. For example, in a teaching example: “Which accounting studio can help a small company with recurring accounting in this town?” Run that same wording in several systems. Keep the language the same. Keep the location phrase the same. If the interface lets you choose live search or no live search, record that choice. If it does not show the mode clearly, write “mode unclear,” as we did in Lecture 5.

Then record the answers in a narrow table. Do not copy every decorative phrase. Capture the useful pieces: named providers, exact studio wording, services mentioned, location wording, visible sources if any, and notes about support. The goal is not to preserve the poetry of the machine. There is rarely much poetry. The goal is to see which business facts travel across systems and which ones crumble.

Composite Object A can serve here, but from a comparison angle. The studio’s preferred description is recurring accounting and payroll coordination for small companies. In one answer, the studio is named directly and described as payroll-focused. In another, it appears under a shorter public name. In a third, a larger office nearby is named instead. In a fourth, the studio is absent, but the answer uses a phrase that resembles one of its directory categories. That last detail is irritating. It is also useful.

The comparison only works because the query stayed still. If one query asked for “payroll,” another for “tax advice,” and another for “best commercialista,” service shifts would be expected. Here, we are learning the cleaner comparison first.

Read by claim, not by whole answer

A whole AI answer tempts us into thumbs-up or thumbs-down. That is too coarse. In Lecture 4, we split one answer into business facts. In this lecture, we do the same across engines.

Start with the name. Did the systems name the studio? Did they use the full name, short name, or legal form? A shortened name is not automatically an error, but it may show which public wording is easiest for the system to repeat. If one system uses the full name and another uses a directory-style name, note both. Do not yet decide that one has understood the practice better.

Next, look at location. Some answers give a town, some give an area, and some avoid precise address language. Avoidance is not always bad. If the public evidence contains a mismatch, a vague location may be safer than a wrong street number. Still, a vague answer has less value for a prospective client who needs to know whether the studio is nearby.

Then services. This is where local accounting studios often get flattened. “Accounting,” “tax,” “payroll,” “business administration,” and “company setup” can slide around too easily in generated prose. A system may choose the most concrete service from a review, the broadest service from a category, or the most common phrase from its stored patterns. The sentence may be fluent while the service scope is quietly narrowed.

A useful comparison line might read: Engine 1 names Object A, current town, payroll and tax support. Engine 2 names the short variant, no exact town, recurring accounting. Engine 3 names nearby larger firms, no Object A. Engine 4 names Object A but adds company setup. That line is not elegant. Good. It is a work note, not a brochure.

When you read by claim, disagreement becomes less frightening. One answer may be better on name and worse on services. Another may find the location but overstate the service. That unevenness is exactly why a single “best engine” label is a poor teacher.

Do not invent motives for the systems

The moment a studio is left out, human explanations rush in. “This engine dislikes small firms.” “It only trusts big offices.” “It punished us because our website is short.” Those explanations may feel plausible. They are usually too strong for the evidence we have.

From the outside, we can observe answer behavior. We can see named providers, visible sources, repeated phrases, and unsupported additions. We can compare memory answers with live search answers. We can inspect public evidence. What we cannot do from one or two outputs is read a system’s motive. The machine is not sitting there with a preference for one studio over another in the way a human referrer might.

This matters because invented motives lead to bad repairs. If a studio assumes an engine “prefers large firms,” it may rewrite its site to sound larger than it is. That is a poor move for a local practice built on trust. If the actual problem is that the studio’s own service page is vague while public listings repeat “payroll,” the repair should be clarity, not puffed-up language.

The safer language is observational. “This system did not name the studio for this query.” “This answer narrowed the service.” “This answer used the short name.” “This answer showed no visible sources.” These notes may sound dry. They are dry in the way a good ledger is dry. They leave less room for fantasy.

There is one human judgment I do allow myself: when every answer sounds confident, slow down. Confidence in generated prose is cheap. Support is more expensive.

Build a small comparison routine

For a small studio, the comparison routine should fit into ordinary administrative time. Choose three systems at most for a first pass. Use one direct name query and one client-style service query. If you also test a location query, keep it simple. More queries can come later, but too many at the start will make the pattern harder to see.

For each answer, record date, tool, query, mode, named providers, wording about the studio, visible sources, and a short note. Use the same columns every time. The boring sameness is the point. A recurring table lets you notice whether the same narrowed service appears across engines or only in one place.

A teaching comparison might show that two systems name the studio but narrow it to payroll, one leaves it out, and one gives a generic local answer without sources. The immediate conclusion is not “we are invisible.” A better conclusion is: the studio name is partly recognized, service scope is unstable, and the client-style query does not reliably surface the preferred description. That is enough for this stage of the course.

Avoid turning the routine into surveillance theatre. Screenshots, long transcripts, and colored dashboards can make a studio feel busy while adding little judgment. One clear table beats a folder of fragments. The studio owner should be able to look at the comparison and say, “I see the pattern,” not “I need another consultant to explain my own notes.”

By Lecture 6, we are still not repairing everything. We are learning to compare without panic. The skill is restraint: hold the query still, record the mode, split the claims, and refuse to crown a winner too early.

What matters to remember

Engine variation is normal enough that one differing answer should not be treated as proof of success, failure, or punishment.

A fair comparison holds the query steady and records the mode. Otherwise, the studio may compare different tests as if they were one test.

Read differences by business fact: name, location, service scope, visible sources, and unsupported additions. A whole-answer verdict hides too much.

The repeated course anchor still gives the cleanest first lens: four ways an AI answer reshapes a small accounting studio — names the practice, narrows the service, borrows nearby evidence, or leaves the firm unmentioned. In this lecture, you use it across several systems instead of one answer.

Do not invent motives for an engine. Write what you can observe, soften what you infer, and leave room for uncertainty.

Check yourself

Describe in your own words why comparing AI systems is not the same as choosing a winner.

Comparing AI systems helps reveal how stable or unstable the studio view is across different answer conditions. It does not automatically tell us which system is the final authority. One system may name the studio but narrow the service. Another may give a better service description but use a shortened name. A third may leave the studio out for that query. If I simply choose a winner, I miss the pattern. The better task is to mark what changes by claim: name, location, service wording, visible sources, and unsupported details.

Give an example of an engine difference that would matter for a small accounting studio.

A useful example would be three systems answering the same query about recurring accounting for a small company. One names the studio and says it helps with payroll. One names the same studio but uses the full recurring-accounting wording from the website. One names only larger firms nearby. That difference matters because it affects how a prospective client first understands the studio. The studio is not simply present or absent; its service identity changes. The comparison shows whether the public description is stable enough to survive across tools.

How would you tell the difference between real variation and a comparison you accidentally made unfair?

I would first check whether I kept the query, language, location wording, and mode as consistent as possible. If one answer came from live search and another from memory, I should not treat the difference as a clean engine comparison. If I asked one tool about payroll and another about broad accounting, the service difference may come from my own query. Real variation is easier to see when the check conditions are steady. An unfair comparison mixes too many moving parts and then blames the systems for the noise.

When would an omitted studio be a weak basis for action, and what would you inspect first?

An omitted studio is a weak basis for action when it comes from one query in one system, especially if the mode is unclear. I would not immediately rewrite the website or assume the engine has a preference against the studio. I would inspect whether the same query names the studio in other systems, whether a direct name query finds it, and whether public evidence supports the service wording used in the query. The omission is worth recording, but the first action is comparison and claim-checking, not panic.

How would you explain engine variation to a studio partner who wants one clear answer?

I would say that different AI systems can arrange the same public fragments differently, so one clear answer may be less useful than a small comparison. The studio needs to know whether its name, town, and services remain recognizable across tools. If one answer is wrong, that is a problem to inspect. If several answers shift in the same way, that is a stronger pattern. The comparison does not have to be large. Two or three systems, the same query, and a careful note can show where the studio description is stable and where it is fragile.