Buyer's Guide
The complete buyer's guide to AI sales agents
AI sales agents should be bought against a defined commercial workflow—not a label, a demonstration or a promise to replace headcount.
- Author
- James Lawrence
- Reading time
- 24 minutes
- Evidence note
- Buyer framework. The candidate landscape records vendor positioning and does not constitute independent testing or endorsement.
The phrase AI sales agent is being applied to products that perform materially different work.
One researches accounts. Another selects prospects and writes outbound email. Another responds to website visitors. Another updates the CRM. Another coordinates several tools while a salesperson remains responsible for every consequential decision.
These products should not be placed into one league table simply because their marketing uses the same category label.
The correct buying process begins before the vendor list.
It begins with the work.
1. Define the commercial job
Do not begin with:
We need an AI SDR.
Begin with:
We need to identify 200 high-fit accounts each month, find the correct buying roles, create a credible reason to engage, execute approved outreach, recognise meaningful replies and pass accepted opportunities to a salesperson with complete context.
Or:
We need to respond to qualified website visitors within 30 seconds, answer approved product questions, determine whether the account meets our criteria and book the correct meeting without overpromising functionality.
A useful job definition contains six elements:
- 1.Trigger: What starts the work?
- 2.Inputs: Which data, systems and instructions are required?
- 3.Actions: Which steps should the agent perform?
- 4.Boundaries: Which actions must it never take without approval?
- 5.Output: What must be passed to a person or system?
- 6.Outcome: Which commercial result would justify the deployment?
If the job cannot be described, the agent cannot be evaluated properly.
2. Identify which category you are actually buying
Category A: Research and preparation agents
These systems gather account, buyer and market information; summarise context; create hypotheses; and prepare people for outreach or meetings.
Typical value: Time saved and better preparation.
Primary risk: Generating plausible but irrelevant or inaccurate research.
Suitable early use: Human-reviewed account briefs, contact mapping and meeting preparation.
Category B: Prospecting and outbound agents
These systems may source contacts, enrich data, identify signals, create messaging, schedule sequences and handle some replies.
Typical value: Greater prospecting capacity and faster experimentation.
Primary risk: Poor targeting, repetitive messaging, compliance failure, sender damage and meetings without commercial value.
Suitable early use: A narrow segment with clear exclusions, approved claims and human review at defined points.
Category C: Inbound qualification agents
These systems engage website visitors or inbound enquiries, answer questions, collect qualification information and route or book meetings.
Typical value: Faster response and better coverage outside working hours.
Primary risk: Incorrect answers, missed high-value buyers, weak qualification and inappropriate commitments.
Suitable early use: Well-defined enquiry types backed by a maintained knowledge source and immediate escalation.
Category D: Conversation and voice agents
These systems hold live or asynchronous conversations through voice, email, messaging or chat.
Typical value: Coverage, speed and handling of repeatable conversations.
Primary risk: Brand damage, consent issues, failure to understand nuance and escalation that happens too late.
Suitable early use: Narrow, disclosed, low-risk conversations with recorded outcomes and fast human hand-off.
Category E: Revenue workflow agents
These systems coordinate actions across research, communication, calendar, CRM and sales-engagement tools.
Typical value: Less operational friction and more consistent execution across a complete workflow.
Primary risk: The agent gains broad access while responsibility becomes unclear.
Suitable early use: Internal coordination and recommendations before external autonomy.
Category F: Sales-assistance layers inside existing platforms
CRM, sales-engagement, data and marketing platforms increasingly embed agent functions into products a company already uses.
Typical value: Faster adoption, existing data context and fewer new integrations.
Primary risk: Assuming that bundled availability proves superior task performance or lowers total cost.
Suitable early use: Compare the embedded function with the specialist alternative using the same test.
3. Decide whether the workflow is ready
An agent will not repair a sales motion whose human version has never worked.
Before buying, score the workflow against these readiness questions.
Target market
- Can the team describe the accounts it should pursue?
- Are there observable reasons those accounts may need the product now?
- Are exclusions documented?
- Can a skilled salesperson explain why one account is preferable to another?
Proposition
- Is the business problem clear?
- Are approved product claims available?
- Can the company explain why a buyer should change now?
- Are proof points documented and current?
Process
- Is there a defined path from first interaction to accepted opportunity?
- Are qualification and hand-off criteria explicit?
- Does the CRM contain usable fields and stages?
- Are owners accountable for the next action?
Data
- Are source permissions understood?
- Can duplicates, customers, active opportunities and excluded contacts be suppressed?
- Is the data current enough for the action?
- Can the team identify which sources produced each material claim?
Governance
- Is there an accountable owner for the agent?
- Can the organisation define actions requiring approval?
- Is communication consent and legitimate use understood for each market?
- Is there a response plan for inappropriate outreach or data exposure?
If several answers are “no”, buy process improvement before autonomy.
4. Evaluate the complete operating system
The visible message or conversation is only one component. Evaluate twelve connected dimensions.
1. Account selection
Can the system identify accounts with a plausible reason to buy? Ask it to explain inclusions and exclusions. Test edge cases, subsidiaries, existing customers, competitors and poor-fit accounts.
2. Buyer identification
Can it distinguish the economic buyer, operational owner, evaluator, champion and blocker? A valid contact is not necessarily a relevant contact.
3. Signal interpretation
Does the system use a signal that changes the action, or merely mention one? A funding announcement, job posting or technology change matters only when connected to a credible business implication.
4. Research accuracy
Can every material statement be traced to a source? Test names, roles, product use, corporate relationships and recent events. Plausible invention is still invention.
5. Commercial judgement
Does the agent connect an observable situation to a relevant problem and proposition? Look for clarity, restraint and a reason to engage—not superficial personalisation.
6. Execution discipline
Can it maintain timing, channel logic, exclusions, suppression and state? Inspect duplicates, messages sent after a reply, cross-campaign conflicts and unauthorised changes.
7. Reply handling
Test clear interest, polite refusal, unsubscribe, referral, objection, ambiguity, sarcasm, out-of-office, wrong person and legal complaint. The agent must know when to stop.
8. Qualification and hand-off
Does the system collect evidence that a salesperson can use? Define exactly what makes an opportunity accepted and inspect whether the hand-off contains source, context, conversation and next step.
9. System-of-record quality
Does the CRM tell the truth after the agent acts? Inspect field accuracy, stage changes, ownership, activity records, contact creation and rollback.
10. Authority and control
Can administrators limit audiences, claims, channels, daily volume, discounts, calendar availability, data sources and system actions? Can access be revoked immediately?
11. Observation and recovery
Can a reviewer reconstruct the decision and action? Look for traces, evidence, change history, versioning, alerts, exception queues and a process for correcting downstream records.
12. Commercial economics
Measure accepted opportunities, progression, pipeline quality, revenue contribution, time saved and total cost. Do not accept emails sent, replies or meetings booked as sufficient outcomes.
5. Calculate the real cost
The subscription price is not the cost of deployment.
Use this model:
**Total monthly cost = platform + data + sending/telephony + models + integrations + implementation + supervision + remediation + displaced tools**
Include:
- licence or platform fee;
- contact and enrichment data;
- email domains, inboxes and deliverability infrastructure;
- calling, messaging and recording;
- model or usage charges;
- integration and workflow work;
- prompt, policy and knowledge preparation;
- human review and exception handling;
- security, legal and procurement time;
- cost of incorrect records, poor meetings and brand damage;
- savings from tools genuinely removed.
Then calculate cost per accepted commercial outcome, not cost per activity.
Useful denominators include:
- sales-accepted opportunity;
- opportunity reaching a defined stage;
- qualified pipeline created;
- retained or expanded revenue;
- hours of skilled work genuinely removed.
6. Ask the vendor for evidence
Use these questions in every sales process.
Work and suitability
- 1.Which exact steps does the product complete without a person?
- 2.Which steps normally require review or intervention?
- 3.Which customer profiles and sales motions are a poor fit?
- 4.What happens when the agent is uncertain?
- 5.What does the customer still need to design and maintain?
Data and integrations
- 1.Which systems must the product read from and write to?
- 2.How are source permissions and provenance recorded?
- 3.How are customers, open opportunities, unsubscribes and exclusions suppressed?
- 4.Which actions can be rolled back?
- 5.What happens when an integration fails part-way through a workflow?
Models and reliability
- 1.Which model providers are used, and can that change without notice?
- 2.How is a workflow tested after a model or prompt change?
- 3.Which failure modes occur most often in production?
- 4.Can we inspect task-level traces and reasons for escalation?
- 5.How do you detect deterioration over time?
Authority, security and compliance
- 1.Can permissions be defined by agent, action, data source and user?
- 2.How are credentials stored, rotated and revoked?
- 3.Can the agent be prevented from making unapproved claims or commitments?
- 4.Which security reports, subprocessors and data-retention controls are available?
- 5.How do you support regional communication and privacy obligations?
Evidence and economics
- 1.What proportion of customer deployments progress beyond pilot?
- 2.What is the median time to a usable production workflow?
- 3.How many hours per week do customers typically spend supervising it?
- 4.What is the denominator behind the performance claim being presented?
- 5.Can you provide evidence of accepted opportunities or revenue progression, not only messages and meetings?
- 6.What additional tools or data services are normally required?
- 7.Which costs increase with contacts, messages, tasks, tokens or outcomes?
- 8.What can we export if we leave?
- 9.Will you agree the pilot’s acceptance criteria before signature?
- 10.Can the pilot be ended without converting into a long annual commitment?
7. Design a pilot that can fail honestly
A pilot should produce a buying decision, not preserve vendor momentum.
Define the test cohort
Use a bounded account segment with clear inclusion and exclusion rules. Preserve a comparison group where practical.
Freeze the success definition
Agree the primary outcome, guardrail measures and disqualifying incidents before the test.
Example primary outcome: Sales-accepted opportunities created from the test cohort.
Example guardrails: Complaint rate, unsubscribe handling, factual accuracy, duplicate-contact rate, CRM accuracy and human-review time.
Build a gold set
Create historical examples covering good targets, poor targets, strong messages, weak messages, meaningful replies, objections, disqualifications and escalation cases.
Start without external authority
Test replay and shadow recommendations before allowing live communication or system changes.
Release authority in stages
Permit only the actions supported by evidence. Keep higher-risk segments, claims, discounts, contract language and strategic accounts behind human approval.
Inspect the misses
Do not only sample successful outputs. Review false positives, ignored signals, inappropriate replies, records changed incorrectly and cases the agent failed to escalate.
Set the decision rule
End with one of three decisions:
- Deploy: Evidence supports bounded production use.
- Remediate: The workflow has value, but defined failures must be corrected and retested.
- Reject: The system does not produce sufficient value or control for this job.
8. Candidate landscape
This is a research starting point, not a ranking. Descriptions reflect how the companies position their products on their own websites. Verify current capability, pricing, integrations and ownership directly. Inclusion is not endorsement.
Autonomous and assisted outbound
- 11x — positions digital workers for sales work, including prospecting and voice. 11x
- Artisan — positions Ava as an AI business-development agent within a broader outbound platform. Artisan
- AiSDR — positions an AI SDR around buyer signals, research, outreach and booked meetings. AiSDR
- Amplemarket — offers an AI sales platform and Duo as an AI sales assistant/agent within its prospecting system. Amplemarket
- Regie.ai — positions an AI sales-engagement platform combining AI agents and human-led prospecting. Regie.ai
- Unify — positions a go-to-market platform around intent signals, workflows and outbound execution. Unify
- Salesmotion — positions an AI-native outbound and account-intelligence workflow around market signals. Salesmotion
- Alta — positions AI revenue-workforce products across prospecting and related commercial tasks. Alta
- Persana AI — positions a sales agent and data platform around research, signals and outreach. Persana AI
- Relevance AI — provides a platform for creating specialist agents, including sales workflows. Relevance AI
Data, research and orchestration
- Clay — provides data enrichment and workflow tooling widely used to construct AI-assisted go-to-market systems. Clay
- Common Room — positions customer intelligence and signal-based go-to-market workflows. Common Room
- Apollo — combines prospect data, engagement workflows and AI assistance inside a sales platform. Apollo
- ZoomInfo Copilot — applies account and buying-signal data within ZoomInfo’s go-to-market platform. ZoomInfo
- 6sense — positions revenue intelligence and AI around account identification, prioritisation and buying stages. 6sense
- Demandbase — positions account intelligence and orchestration for account-based go-to-market work. Demandbase
Inbound qualification and conversation
- Qualified — positions Piper as an AI SDR for website pipeline generation and qualification. Qualified
- Salesforce Agentforce — includes agents for sales and service workflows inside the Salesforce platform. Salesforce Agentforce
- HubSpot Breeze — includes AI agents and assistants within HubSpot’s customer platform. HubSpot Breeze
- Intercom Fin — primarily a customer-service agent, but relevant wherever inbound product conversations and hand-offs cross service and sales. Intercom Fin
Voice infrastructure used in sales workflows
- Vapi — provides infrastructure for building and operating voice agents. Vapi
- Retell AI — provides a platform for building, deploying and monitoring voice agents. Retell AI
- Bland AI — provides infrastructure and products for automated phone interactions. Bland AI
- Synthflow AI — provides a no-code platform for voice-agent workflows. Synthflow AI
The presence of a company in more than one functional category is expected. The market is converging faster than its labels.
9. Red flags
Pause the buying process when:
- the demonstration uses a different workflow from yours;
- performance claims omit the denominator or time period;
- meetings booked are presented without acceptance or progression;
- the vendor cannot describe frequent failure modes;
- the implementation burden is described only as “connect your CRM”;
- the agent requires broad credentials before a limited test;
- there is no clear answer on suppression, consent or unsubscribe handling;
- output review is possible but action reconstruction is not;
- pricing excludes required data, sending, telephony or service work;
- a pilot automatically becomes an annual contract before agreed acceptance;
- the vendor refuses to test known edge cases;
- the proposed ROI depends on eliminating people while the workflow still requires extensive supervision.
10. The buying decision
The best AI sales agent is not the product that appears most autonomous.
It is the system that produces the most valuable accepted outcome for your defined workflow—at a controllable cost, with an authority model and evidence trail appropriate to the risk.
For many teams, the correct first deployment will be less autonomous than the demonstration.
That is not failure.
It is how an organisation earns the right to delegate more.
Contribute to the first AI sales-agent category study.
Submit a product, deployment or buying decision we should examine, or ask us to assess a shortlist against your workflow.
Related reading
Research
The agent economy needs more than agents →Methodology
Evidence, not star ratings →Buyer's Guide
How we will evaluate AI sales agents →The Briefing
Issue Zero: The economy forming around the agents →Research
The agent trust and control stack →Deployment Playbook
How to run a controlled AI-agent pilot →