Buyer's Guide
How we will evaluate AI sales agents
Sales-agent claims should be tested against the commercial workflow: not merely whether the agent can write, but whether it can create and progress credible revenue opportunities.
- Author
- James Lawrence
- Reading time
- 8 minutes
Sales agents are an obvious place to begin observing the agent economy.
The work is valuable, measurable and connected to systems companies already use. The category is attracting investment and customer interest. It is also filled with terms that can make very different products sound interchangeable.
An “AI sales agent” might be a research assistant, an outbound sequencer, a conversational website agent, an autonomous SDR, a meeting-preparation tool or an orchestration layer coordinating an entire revenue workflow.
Before comparing products, the work must be defined.
The outcome is not activity
An agent can generate thousands of contacts and messages without creating credible pipeline.
It can increase reply volume while damaging positioning. It can book meetings that salespeople would never have chosen to attend. It can produce immaculate CRM records for opportunities that should not exist.
Activity matters only when it contributes to a valuable commercial outcome.
Our first research programme will therefore examine the complete path from market selection to qualified commercial progress.
The ten evaluation dimensions
1. Market and account selection
Can the agent identify accounts with a plausible reason to buy, or does it simply reproduce database filters?
We will examine the quality of selection logic, use of relevant signals, exclusions and the agent's ability to explain why an account belongs in the target market.
2. Buyer identification
Can it distinguish the economic buyer, operational owner, technical evaluator and other relevant participants?
Finding a valid email address is not the same as identifying the correct buying role.
3. Research and context
Does the agent collect information that changes the commercial approach, or merely generate a generic company summary?
Useful research should influence the problem hypothesis, message, channel or next action.
4. Commercial judgment and message quality
Can the agent connect a credible business problem to a relevant proposition without inventing facts, using false familiarity or producing empty personalisation?
We will assess clarity, relevance, specificity, accuracy and the strength of the proposed reason to engage.
5. Execution and coordination
Can the agent execute the intended sequence across the required channels and systems? Does it avoid duplication, respect timing and maintain state across interactions?
6. Reply handling and qualification
Can the agent interpret interest, objection, ambiguity and rejection? Does it know when to answer, ask a question, stop or involve a person?
Where qualification is part of the workflow, we will examine whether the agent collects commercially meaningful evidence rather than simply categorising sentiment.
7. CRM and workflow discipline
Does the agent create accurate, usable records? Can another person understand what happened, what was learned and what should happen next?
Automation that corrupts the system of record creates downstream cost.
8. Permissions, compliance and reputation
Which systems, data and communication rights does the agent require? Can its authority be limited? Does it respect consent, exclusion rules, brand standards and relevant policies?
The risk is not only regulatory. Poor agent behaviour can damage sender reputation and buyer trust.
9. Human supervision and exception handling
Where must a person intervene? How are exceptions surfaced? Can the agent recognise uncertainty and escalate before taking an inappropriate action?
The true operating cost includes the people preparing, monitoring and correcting the system.
10. Commercial outcome and economics
What did the deployment produce?
Relevant measures may include accepted opportunities, qualified meetings, conversion progression, pipeline value, revenue contribution, time saved and total cost per commercially valuable outcome.
We will not treat message volume or meetings booked as sufficient evidence on their own.
How the initial research will work
The first edition will combine several forms of evidence:
- structured vendor submissions;
- product documentation and demonstrations;
- buyer and operator interviews;
- available customer evidence;
- hands-on workflow tests where access permits;
- commercial-model and deployment analysis;
- explicit recording of unknowns and limitations.
Products will not be forced into one ranking when they perform materially different jobs.
The output will identify category structure, buyer suitability, deployment requirements, evidence quality, notable strengths, material limitations and questions buyers should ask.
Who should participate
We want to hear from:
- companies that have deployed an AI sales agent;
- revenue leaders currently evaluating one;
- founders building sales-agent products;
- RevOps and security leaders responsible for implementation;
- practitioners with evidence of a deployment that failed or underperformed.
Positive and negative evidence are both useful. Participation does not guarantee coverage or endorsement.
The larger purpose
Sales agents are the first field, not the limit of the publication.
The same underlying questions apply wherever agents are asked to perform consequential work: What can they complete? What do they need? How are they controlled? What does a person still do? What evidence supports the result?
Answering those questions carefully is how a review becomes intelligence.
Contribute to the first AI sales-agent buyer's guide.
Tell us about a product, deployment or buying decision we should examine.
Related reading
Research
The agent economy needs more than agents →Methodology
Evidence, not star ratings →The Briefing
Issue Zero: The economy forming around the agents →Buyer's Guide
The complete buyer's guide to AI sales agents →Research
The agent trust and control stack →Deployment Playbook
How to run a controlled AI-agent pilot →