The founding edition is open — Join the first readers shaping The Agentic Observer.Join The Briefing →

Methodology

Evidence, not star ratings

A five-star review system would make the agent market look easier to understand. It would not necessarily make it more intelligible.

Author
James Lawrence
Reading time
6 minutes

The easiest version of an agent review site is also the least useful.

Collect a list of products. Repeat their feature descriptions. Add screenshots, affiliate links and a five-star score. Publish “best AI agent” pages for every category that search engines can identify.

That model may produce traffic. It does not produce much confidence.

Agents are difficult to evaluate because performance depends on the task, data, tools, model, instructions, permissions, operating environment and cost of failure. The same system can look exceptional in one demonstration and fail repeatedly when placed inside a changing business workflow.

A universal rating conceals those conditions.

An output is not an outcome

An agent may produce an excellent email without selecting the correct buyer. It may complete a refund while violating policy. It may reconcile an account accurately but require so much human preparation that the promised economic benefit disappears.

Evaluating the visible output alone excludes the system around it.

For consequential work, the relevant questions include:

  • Did the agent complete the intended task?
  • Was the result correct and commercially useful?
  • Could it repeat the result?
  • Which data and tools did it require?
  • What authority was it given?
  • When did a person intervene?
  • Could its actions be reconstructed?
  • What happened when information was incomplete?
  • How much did the complete deployment cost?
  • What was the measurable outcome?

Those answers cannot always be compressed into a single number.

Claims need visible labels

The Agentic Observer will use four evidence states.

Verified means we independently observed, tested or received direct primary evidence supporting the claim.

Supported means credible evidence from more than one source supports the claim, although we did not independently reproduce it.

Vendor claim means the company made the assertion and we have represented it accurately, but we have not verified it.

Unknown means material information was unavailable, incomplete or undisclosed.

These labels are deliberately simple. Their purpose is to stop a reader mistaking a vendor statement, customer anecdote, benchmark result and independently observed outcome for equivalent evidence.

Suitability matters more than abstract quality

There may be no single “best” customer-service agent.

One buyer may need a system that answers low-risk questions across a large knowledge base. Another may need an agent authorised to amend an account, apply a policy and issue money. Their security, evaluation and supervision requirements are not comparable.

Our guides will therefore try to answer:

  • Best for whom?
  • Best for which work?
  • Under which operating constraints?
  • Supported by what evidence?
  • With which material limitations?

Where a score genuinely helps, we will show the underlying dimensions and weighting. Where it creates false precision, we will not use one.

Independence must be operational

“Independent” cannot mean that a publication has no commercial model. It must mean that the commercial model cannot purchase the conclusion.

The Agentic Observer may earn revenue through subscriptions, sponsorship, commercial research, events and deployment assessments. Those relationships will be disclosed where relevant.

Companies may submit evidence, brief us on their products and correct factual errors. They may not buy a ranking or approve an editorial finding.

This standard will sometimes make the work slower. It may make certain commercial opportunities less attractive. It is nevertheless the foundation of the business we intend to build.

A methodology that can improve

No launch methodology will answer every question in a market changing this quickly.

We will publish our evaluation criteria, record review dates, explain material changes and correct substantive errors. We will also say when evidence is too limited to support a confident conclusion.

That may be less satisfying than a definitive score.

It is more useful than certainty we have not earned.

Have evidence we should examine?

Submit a company, deployment or correction for research consideration.

Related reading