0effort

← All articles

How AI SDRs Handle Replies (and When They Hand Off to You)

· 9 min read · by the 0effort team

TL;DR: An AI SDR sorts replies into a few buckets: interested, not interested, out-of-office, or unclear — and only drafts or sends where a wrong call is cheap to undo. Pricing, legal, and opt-out replies always escalate to a human, regardless of classifier confidence, because the real risk is a silent misroute, not a bad sentence.

Everyone selling an AI SDR wants to talk about the copy — the subject lines, the tone that doesn't read like a template. Copy is the easy part. The hard part is what happens the instant a reply lands: what kind of message is this, who gets to answer it, and does a wrong call here quietly cost you a deal six weeks from now. Routing is the whole game — it's the stage of the AI SDR workflow most pitches gloss over fastest.

What actually happens in the 30 seconds after a prospect hits reply?

The reply doesn't wait on an AI to think about anything. The moment it lands — by webhook, inbox sync, or IMAP poll — a rule-based classifier reads it and assigns one of a small set of intents: fast, mechanical pattern-matching, not a language model mulling the sentence. Only then do consequences fire — updating the contact's status, pausing or ending the sequence, deciding whether anyone gets alerted. Drafting an actual reply, if that layer is even on, is queued separately and later, kept out of the critical path so a slow AI call can never delay what matters: making sure nobody keeps emailing someone who just said no.

Which reply types does an AI SDR sort into, and which does it confuse most often?

Four buckets cover most of it: interested, not interested, out-of-office, and everything too ambiguous to call — plus a quieter fifth underneath: automated noise like a ticket system's acknowledgment, filtered out before it reaches the other four because it isn't a reply from a person at all.

Bucket What it looks like The defensible default
Interested "Let's talk," a pricing question, a forward to a colleague Sequence stops, lead status advances, someone gets alerted
Not interested "Please stop," "not a fit," hostility Sequence ends everywhere — email, calling, LinkedIn — the same day
Out-of-office Auto-reply headers, "back on the 14th" Sequence pauses and resumes itself around the stated date
Unclear One-word replies, ambiguous phrasing Sequence stops; queued for a human to read rather than guessed at

The misses cluster in predictable places. Negation is the first: "not interested" contains the word "interested," and a classifier that scores keywords without checking what's next to them will read that as enthusiasm — skipping that check is a common way a rejection ends up filed as a lead. Automation is the second: a ticket system's "your request has been received, ref #4471" isn't a reply from a human, and treating it like one can silently end a sequence for someone who never read the email.

Ambiguous text is the third, and it's less a bug than a forced choice. The defensible default leans positive on a genuine tie between interested and not: one extra follow-up to someone lukewarm is a mildly annoying email; filing a real prospect as a rejection is a dead lead that never gets a second look. That's a deliberate choice, and any vendor should be able to tell you which way theirs leans.

Which replies should an AI never answer alone — and why confidence scores are the wrong gate?

The standard pitch is that low-confidence replies escalate to a human. True, and not the important part. Confidence measures how sure the model is about the words it just generated — it says nothing about how expensive being wrong would be. A model can sound extremely sure while inventing a meeting time nobody offered, or promising to send an attachment it has no ability to attach to anything. Gating only on confidence lets both through — the model isn't unsure, it's wrong in a way a probability score doesn't capture.

The more defensible design gates on the category of the reply, independent of how sure the model sounds. Anything touching price, contract terms, legal threats, or a request to stop being contacted routes to a human before a draft is even attempted — not because the model can't write a plausible answer, but because a wrong one there is expensive and hard to walk back. Whatever it does draft gets checked against what's actually true before it goes near a send button — real meeting times only, no promises the system can't keep, no padded non-answers. Call it asymmetric autonomy: full independence where a mistake is cheap and reversible, none at all where it isn't — and the split has nothing to do with how the model feels about its own sentence.

How does the handoff actually trigger?

It's a stack, not a single switch. The first layer runs before any drafting happens: a reply that trips a pattern for pricing, legal language, or an opt-out request never reaches the model — it routes to a human immediately, since a false positive costs one cheap extra read, while a false negative could cost far more. The second layer runs after a draft exists: the system checks the output against what was actually offered this turn, and separately checks whether the model's confidence was actually high enough to justify sending — either check failing pulls the draft back. The third is a human closing the loop directly: rejecting a draft flags the whole conversation as handed off, so the AI stops negotiating a thread a person has taken over. No single one of these is "the" trigger; a system with only one is missing the other two.

What should a good handoff contain by the time it reaches your inbox?

A handoff that says "a prospect replied, log in to see it" barely qualifies — it adds a click between the alert and the information you need. The version worth having tells you who replied, from which company, over which channel, what the thread was about, and what they actually wrote, then links straight into that conversation instead of a dashboard you have to search. Keep two jobs separate, too: an internal heads-up isn't the same as putting a salesperson on a hot thread with context to answer without asking the prospect to repeat themselves. Bundle both into one generic notification and a genuinely hot reply sits unanswered for two days.

What happens to the replies it does answer, and how do you audit them a week later?

A defensible system logs every drafted message, sent or not — including what a human rejected, and why it escalated instead of answering. That matters more a week in than on day one: the useful audit question isn't "how many replies went out," it's "how many drafts did a human rewrite before approving, and what changed." A draft that consistently goes out untouched is doing its job; one that's heavily edited every time says the prompt needs work, not that the feature is broken — but that's only visible if the system keeps discarded drafts next to the sent ones, instead of counting successes alone.

How are out-of-office, wrong-person, and "talk to my colleague" replies handled?

Out-of-office is the easiest bucket to get right, and one of the most common: an auto-reply gets detected from the headers, subject, and body together, and the sequence goes on hold rather than stopping. State a return date, and the sequence should resume on its own. "Wrong person, try my colleague" deserves more attention than it gets, because it isn't a rejection — it's a warm signal wearing a boring costume. Someone who forwards you to a colleague is doing you a favor: they've done internal routing you'd otherwise have to guess at. Treating that like a hard no throws away a better lead than the one you started with.

What breaks reply handling in the real world?

Understanding what a reply says is only half the problem; knowing which contact, campaign, and conversation it belongs to is the other half, and it's less glamorous. A shared inbox (sales@theircompany.com instead of a named person) muddies who you're talking to. A forwarded thread can arrive with mangled headers, forcing a fallback to subject-line and participant matching. A prospect who switches to a personal address mid-conversation can look, to a naive matcher, like a brand-new contact. None of this is exotic — it's the everyday mess of how people use email — but it's where a system that only demos well on clean threads starts quietly losing context. Ask any vendor what happens when a prospect replies from their phone's default mail app, different signature, no quoted history — a shrug is worth knowing before you rely on it. Our vendor due-diligence checklist has more worth asking.

How do you tell whether reply handling is actually working?

"How many replies did the AI handle" is the number every dashboard shows first, and it's close to useless alone — it counts activity, not correctness. Better: what share of replies in each bucket needed a human, and is that share moving the direction you'd expect as targeting settles in. A spike in escalations is worth investigating; so is a flat 100% approval rate with zero edits, since that usually means nobody's reading drafts before approving. The other number that matters is boring but load-bearing: can a workspace owner see, at a glance, whether the automation has fired recently. An automation nobody can observe reads as broken even when it's working.

FAQ

Can an AI SDR send replies on its own, or does it only draft them for approval?

Both modes exist, and it should be your choice, not a default you discover later. The safer start is draft-only: every reply is queued for approval, edit, or rejection before it sends. Even in auto-send mode, pricing, legal, and opt-out replies still route to a human.

How quickly does an AI SDR respond to an inbound reply, and does speed actually matter?

Classification and the actions that carry real risk — stopping a sequence, honoring an opt-out — happen within moments of the reply arriving, before any drafting begins. Drafting runs on a short, separate cycle, and draft mode adds a human approval step on top. Speed matters far more on the suppression side than the writing side.

What happens when a prospect asks about pricing or a technical detail the AI doesn't know?

Pricing and similarly loaded topics are designed to skip the AI and go straight to a human — not because it can't produce a plausible number, but because a wrong one is expensive to walk back. The same logic covers any technical claim the system can't verify: guessing isn't the design, escalating is.

Does an AI SDR handle opt-out and "take me off your list" replies automatically?

It can detect the request and stop sending automatically, across every channel — not just the one the reply arrived on. But detection isn't the same as honoring it: someone who asks to stop, in whatever words they use, expects it to stick whether or not a classifier flagged the sentence. Our GDPR guide covers the compliance side in more detail.

Will it keep sending the rest of the sequence after someone replies?

No reply should let a sequence keep running blind. Any real reply pauses or ends it, and what happens next depends on the bucket: a hard no ends it everywhere, an out-of-office holds it temporarily, and even an unclear reply stops automated sends until a human reads it.

Can it tell an out-of-office apart from a real reply, and what does it do with each?

Yes — out-of-office has distinctive signals in the headers, subject, and body that a real reply rarely shares, making it one of the more reliable buckets to catch. It holds the sequence and, when a return date is stated, resumes automatically around it. A real reply of any sentiment stops the sequence outright.

Put your outbound on autopilot

0effort sources your buyers, writes every touch, answers replies over email and phone, and books the meetings. You just show up.

Start free See how it works

Outbound tips, monthly

One email a month with what's actually working in cold outbound. No spam, unsubscribe anytime.