AI agents8 min read

What an AI sales agent actually does (and the two jobs it hands back to you)

An AI sales agent automates three of the five jobs in outbound: finding, qualifying and drafting. Here is what it does with the other two, and where it fails.

Outbound is five jobs; an AI agent reliably does the first threeFive jobs, sold as oneFindQualifyWriteSendReplyruns unattendedstill yoursThe reply is where an unattended agent loses the deal.

An AI sales agent is software that does the top of an outbound sales job. It finds accounts that match your profile, decides which of them are worth contacting, and writes the message. That is three of the five jobs a human SDR does. The other two, sending without wrecking your domain and handling the reply, are the ones vendors are quietest about, and they are where most deployments fall apart.

Short answer

An AI sales agent automates finding, qualifying and drafting. It does not reliably handle deliverability or a real human reply. Buy it for the first three jobs, keep a person on the last two, and the maths works. Buy it as a replacement for a rep and you are generating apologies at scale.

The five jobs vendors sell as one

Outbound is not one activity. It is five, and they have almost nothing in common with each other. A tool that is excellent at the first is not thereby any good at the fourth, but the pricing page sells all five as a single seat.

Which stages of outbound an AI SDR can actually do unattended Five stages of an outbound sequence shown left to right: find accounts, qualify them, write the message, send it, and handle the reply. A confidence bar under each stage shows that finding and writing are largely solved, qualifying is partial, and handling a real reply is not. Find pull accounts thatmatch the profile runs unattended Qualify decide this one isworth a message needs your input Write draft the messagefrom the evidence runs unattended Send deliver it withoutlanding in spam needs your input Reply read the answer andsay the right thing needs you

HOW MUCH OF THE STAGE SURVIVES WITHOUT A HUMAN Buy the first three. Keep your hands on the last two, or the tool is generating apologies at scale.

The pipeline is five jobs, and vendors price all five as one. Finding and drafting are close to solved. Qualifying is only as good as the evidence it is given. The reply is where an unattended agent still loses you the deal.

Every evaluation gets easier once you stop asking "is this AI SDR good" and start asking "which of these five is it good at, and what happens to the other ones."

Finding: solved, and cheaper than you expect

Pulling a list of companies that match a set of attributes is a solved problem and has been for years. What changed is that the filters stopped being fields in a database.

An older tool could find you SaaS companies in the US with 50 to 200 staff. A current one can find you companies whose engineering blog mentioned migrating off a specific tool in the last quarter. That is a real difference in kind, not degree, because the second thing is a reason to write and the first thing is a demographic.

This stage is worth buying. It is repetitive, it is high volume, and being wrong is cheap.

Qualifying: only as good as the evidence you hand it

Here is where the category gets uneven.

Qualification means deciding that this particular company, this week, is worth a message. Doing that well needs evidence. Doing it badly needs only a confidence score, and confidence scores are easy to generate and impossible to check.

Four levels of evidence behind a sales lead A four-rung ladder. The bottom rung is a guess with no source. Above it, a plausible inference. Above that, a fact from a database that cannot be checked. The top rung is a direct quote with a working link to where it was said. Quote + link “We are still doing this in spreadsheets.” plus the URL send it Sourced fact a database row you cannot open or verify today do not send Inference they grew headcount, so they probably need this do not send Guess they are in our target industry do not send Ask of any lead: can I paste the sentence, and does the link still open?
Only the top rung survives being challenged. If your CRM cannot show the sentence and the URL it came from, the lead is somebody's hunch wearing a confidence score.

The failure is specific and it is worth naming, because it is the single most common reason AI outbound produces mail nobody answers. Give a language model a thin record with empty fields and ask it to explain why this account is a good fit, and it will not tell you the record is thin. It will write something plausible. "They are scaling their engineering team" is a sentence a model can produce from a company name and a category alone, with no evidence whatsoever behind it.

Then that invented reason becomes the first line of the email. The prospect reads a confident claim about their own company that is not true, and now you have done worse than sending nothing.

So the test for this stage is one question: can the tool show you the source for a qualified lead, as a quote and a URL that still opens? This is the same evidence test you should already be applying to your own leads. A tool doing this properly can. A tool that returns a score and a paragraph of reasoning cannot, and you should assume the reasoning was generated rather than found.

Writing: solved, with one catch

Language models write a competent cold email. That has been true for a while and it is not the differentiator anyone still pretends it is.

The catch is that a competent cold email is not what gets replies. What gets a reply is a first sentence that could not have been written about any other company, and that sentence is a function of the evidence from the previous stage, not the prose from this one. Give the writer good evidence and a mediocre model produces something that works. Give it nothing and the best model available produces fluent, well-structured, completely generic mail.

Which means the writing stage is rarely the thing to evaluate. If the output reads generic, the problem is almost always two stages upstream.

Sending: where your domain reputation gets spent

This is the stage that gets undersold, and the one that does damage you cannot undo in a quarter.

Volume is the whole appeal of automating outbound, and volume is exactly what mailbox providers watch. A tool that makes it trivial to send two thousand messages a week makes it equally trivial to burn a domain that took two years to build. Once your main domain is filtered, your invoices and your password resets go to spam too.

The things that actually matter here are unglamorous:

  • Send from a separate domain, never the one your company runs on.
  • Authenticate it properly. SPF, DKIM and DMARC, all three.
  • Keep per-mailbox volume low and add mailboxes instead of raising it.
  • Watch reply rate as a health metric, not just a success metric. A collapsing reply rate is usually the first visible sign of a filtering problem, and it shows up before the bounce rate does.

None of this is AI-specific. It is just that automation removes the natural friction that used to cap the damage.

The reply: still yours

A reply is where the deal actually starts and it is the stage least suited to an unattended agent.

Most replies to cold outreach are not "yes" or "no". They are "what is this", "we already use something else", "not now, ask me in March", "who gave you my email", and "I am not the right person but Sam is". Each one needs a different response, several need a judgement call about whether to push, and one of them needs an apology.

An agent handling that queue unattended does the same thing it does with a thin record: it produces something plausible. Worth reading next to how a follow-up sequence should actually behave. At the reply stage, plausible and wrong costs you the account rather than one email.

Seven questions to ask a vendor

Run these in the demo. The answers separate the category quickly.

Question What a good answer sounds like
Show me the source for one qualified lead. A quote and a working URL, produced immediately.
What happens when the evidence is thin? The lead is skipped or flagged, not written up anyway.
Which of the five stages do you actually cover? A direct answer, without redefining the stages.
What is the per-mailbox daily send limit? A number, and a reason for it.
Who handles replies? A person, or a clearly bounded set of reply types.
What does this cost including data and mailboxes? A total, not a seat price.
What does your worst customer outcome look like? An actual story. Anyone who has shipped this has one.

The last question is the useful one. A vendor who has never seen this go wrong has not had enough customers to be worth buying from.

What it costs when it goes wrong

The advertised failure is wasted subscription spend. The real ones are worse and slower:

A burned sending domain. Recoverable, over months, if you were sending from a separate domain. If you were not, your company email is now unreliable.

A market you have already annoyed. In a narrow vertical, two thousand generic emails can reach a meaningful share of everyone who could ever buy from you. There is no second first impression, and the people you burned talk to each other.

Numbers that lie in the right direction. Open rates go up when you send more. Meetings do not. A team can run this for two quarters, watch activity metrics rise, and only notice at the pipeline review that none of it converted.

Who should not buy one

If you have not yet written fifty cold emails by hand and got replies to some of them, do that first. Not out of principle. It is that an AI sales agent scales a process, and you cannot tell whether the process it is scaling works until you have run it slowly enough to watch.

If your total addressable market is under a thousand companies, the maths does not favour automation either. At that size the constraint is not how many messages you can produce, it is how many of those thousand you can afford to get wrong.

And if the reason you want one is that outbound is not working, the tool will not fix it. It will do the same thing faster.

What to do this week

Pick ten accounts you would genuinely like as customers. If you cannot say quickly which ten, that is an ideal customer profile problem rather than a tooling one. For each one, try to write a single sentence naming something specific that happened there recently. Not "they are growing". Something with a date on it.

Count how many of the ten you managed. That number is what an AI sales agent is going to automate. If it is two, the tool will find you two good messages and eight generic ones, and you now know exactly what to test in a demo.

Doing this by hand is the slow way

Revtive runs a team of agents that find the venues, read their rules, qualify leads on real evidence and draft the outreach for you to approve. Free to start, no card.

Start free

Keep reading