How to evaluate AI SDR tools (five questions a demo will not answer)
Every AI SDR tools listicle is written by a vendor that ranks itself first. Here are the five questions that decide whether one works, and how to test them.
Evaluating AI SDR tools comes down to five questions, and a scripted demo answers only the first two. Can it find accounts that fit, and does it write a clean sentence? Both are close to solved across every vendor. The three that decide whether you keep the thing are about sources, replies and whose domain absorbs the damage.
Judge an AI SDR tool on three things a demo cannot show you: whether every claim it makes carries a source link you can open, what it does when a prospect replies with a real question, and whether the sending reputation it risks is yours or its own. Finding accounts and writing sentences are table stakes now.
We make one of these, so read this the way you would read any vendor writing about its own category. The difference is that this page does not end with a ranked list where we are number one. Every listicle currently ranking for AI SDR tools is published by a company selling one, and in most of them the publisher holds first place.
The five questions
If you are still working out what the category even is, start with the definition and come back.
1. Can it find accounts that fit? Every tool does this. It is a filtered query against a data provider, and most vendors resell the same two or three providers. Low signal for a buying decision.
2. Does it write a clean sentence? Also solved, and also low signal. Modern models write fluently. Fluency was never the reason cold email fails.
3. Can it show a source for every claim? Here is where tools separate. Ask for one drafted message and then ask, for each factual claim in it, where that came from. You want a URL that opens and a date. What you often get is a confidence score, which is a guess with a decimal point. The evidence standard is the same one you would apply to a human researcher.
4. What happens when someone replies? Demos end at send. Real outbound begins at the reply, and replies are rarely yes or no. They are "we looked at something like this in March, what is different", or a question about a compliance requirement. Ask to see the reply handling on a real thread, not a happy path.
5. Whose domain burns when it is wrong? If it sends from your domain, your reputation is the collateral for the tool's error rate. That is a fair trade if the error rate is low. It is worth knowing it is the trade.
Questions one and two get answered in twenty minutes on a call. Three, four and five need your own data and a few weeks. Which is the actual argument for running a trial instead of taking four more demos.
What it costs to run, as opposed to what it costs to buy
The quote covers the licence. Three other lines show up later.
Enrichment and data. Usually billed per record resolved, and you pay for attempts that fail to resolve. A list of 5,000 does not cost 5,000 credits, it costs whatever the attempt count is.
Sending infrastructure. Mailboxes, domains and warmup, billed per mailbox. Anyone running meaningful volume runs several domains, because putting it all on your primary is how you lose the primary.
Human review. The line nobody quotes, and often the biggest. Someone reads the drafts before they go. If nobody reads them, you have not bought an AI SDR, you have bought a way to be wrong faster.
Ask for all four numbers before comparing two vendors on the first one.
A two week test that beats four demos
| Day | What you do | What you are looking for |
|---|---|---|
| 1 | Point it at 20 accounts you already know well | Does it tell you anything you did not know, or restate the firmographics |
| 2 to 3 | Open every source link in every draft | How many 404, how many are a year old, how many were inferred rather than found |
| 4 to 7 | Send only what you would have sent yourself | The unedited-send rate. Under 30 percent means you bought a drafting tool |
| 8 to 14 | Reply to it from a second address, with a question | Whether it answers, escalates, or sends the next sequence step regardless |
Run it on a separate domain. Keep volume flat, because raising volume during a quality test tells you nothing about quality and something you will not like about deliverability.
The unedited-send rate is the number that predicts whether you keep the tool. It is also the one no vendor publishes.
Where these tools genuinely earn their money
Not in sending. In the reading.
A person researching an account properly spends most of the time verifying that what a list claimed is still true. That is where the hour actually goes in prospecting, and it is genuinely compressible, because reading forty pages to find the two facts that matter is exactly what a model is good at. A tool that hands you fifteen well-sourced accounts and a reason for each has done the expensive part.
The stages after that are where the claims get loose. Which stages are actually solved and which are still yours is worth being precise about.
Who should not buy one
Anyone without a message that already works. These multiply an existing process. If ten manual messages a week produce nothing, a hundred automated ones produce nothing on four domains.
Anyone under about fifty target accounts. At that size the research is a morning a week and the tool is overhead. Do it yourself and keep the knowledge.
Anyone who cannot spare the review time. The review is not optional, and a tool bought to save time that instead creates a daily reading task is a tool that gets cancelled in month three.
Anyone whose real problem is that nobody has heard of them. Outbound into total obscurity has a low ceiling regardless of how good the research is. Being present where your buyers already ask questions changes the reply rate of the cold message more than any tool does.
Revtive sits in the first stage of that pipeline on purpose. It researches and drafts, it has to cite a source for every signal it reports, and it does not send anything on its own. You can see what it does and does not do without a call.