The four gates: when to kill an outbound campaign, when to iterate, and when to pour everything into it
By Jānis Plūme, Founder, Outbound Pros · 2026-08-06
Quick answer
Kill a sequence when it is below 0.5% positive replies on sends after warm up, over a window and a sample size you fixed before launch. Between 0.5% and 1%, iterate. At 1% or above, scale. At 2% or above, put everything you have behind it. Those four gates are the rules we apply per sequence and per segment across 1500+ campaigns. Our fleet wide positive rate sits an order of magnitude below the kill line, near 0.05% on sends as a derived figure, precisely because we run the experiments that fail. The gate is not the hard part. Committing to the window before you see the numbers is.
Publishing these thresholds costs us something, which is why almost nobody does it. A client who reads this page can ask in a reporting call why a sequence at 0.3% is still running in week nine, and they would be right to ask. A threshold nobody outside the room can audit is a preference with a number attached.
What follows is the rule the team that applies these thresholds uses on live accounts. It is a line rather than a results claim: the point at which we stop spending a client's list, domains and calendar on an angle the market has already answered.
What are the four gates?
The four gates are decision thresholds applied to one sequence, on one segment, after warm up, over a window fixed before launch. Every gate reads the same fraction: positive replies divided by emails sent.
| Measured positive reply rate | Denominator | Decision | What actually happens |
|---|---|---|---|
| Under 0.5% | Positive replies divided by emails sent | Kill | Sequence stops, budget moves, the list or offer goes back to the drawing board |
| 0.5% to 1% | Positive replies divided by emails sent | Iterate | One variable changes at a time and the window restarts |
| 1% and above | Positive replies divided by emails sent | Scale | Volume increases on the same angle before anybody rewrites it |
| 2% and above | Positive replies divided by emails sent | Pour | Everything available goes behind it, because this is rare and does not last |
The scope belongs in the same breath as the numbers. A threshold applied to a whole account is a different measurement wearing the same label, and it will tell you to kill work that is fine.
Why does a threshold need a window and a sample size before it means anything?
A threshold without a pre committed window turns into a negotiation held on a Friday afternoon. The fix is four things written down before the sequence sends its first email, which we call the Window Contract.
- Window length. Weeks the sequence runs before the gate applies, excluding warm up.
- Sample floor. The minimum sends below which the rate is not read at all.
- Segment scope. Which segment the rate is measured on, because a blended account rate hides a working segment inside a failing one.
- Owner and decision. Who applies the gate, and what happens on each outcome.
We set this during onboarding, which runs around 21 days, in the same conversation where the client approves the lead list and the messaging. The timing is deliberate: nobody has fallen in love with an angle yet, and the person arguing to keep a failing sequence alive in week eight is the one who agreed to kill it in week one. Most teams have a threshold and no contract, so it gets applied by whoever is in the room, and a campaign judged that way is not being judged, it is being negotiated with.
What sample size do you need before a positive rate is stable?
You need enough sends that one positive reply landing or not landing cannot move the measured rate more than you can tolerate, which at a 0.5% positive rate means thousands of sends, not hundreds.
At 0.5% positive on sends, the expected count is one positive per 200 emails. A 1,000 send test expects five, so a single reply arriving or not arriving moves the measured rate by 20% of itself. At 2,000 sends it expects ten and one reply moves it by 10%. Judging a sequence on a few hundred sends means reading noise and calling it a decision. Precision here is governed by the count of events in the numerator instead of the size of the denominator, and the NIST/SEMATECH e-Handbook of Statistical Methods carries the interval arithmetic.
Here is what that looks like at volume. In one week on our largest account, 44,649 emails produced 377 replies, which is a 0.84% reply rate. How many of those 377 were positive is not a figure we have measured as its own number, so we do not publish one, and inferring it from a different campaign's positive count would be exactly the error this page exists to prevent.
Run the gate arithmetic against that denominator instead. At the 0.5% kill line, 44,649 sends expects around 223 positive replies. Ten times lower, at 0.05%, it expects around 22. The entire distance between a sequence you scale and a sequence you stop is a couple of hundred events sitting inside a denominator of forty four thousand, which is why the denominator being large tells you almost nothing about whether the reading is stable. Volume does not buy sample when the numerator is rare.
One correctly sourced illustration of the same point sits further down this page: a sequence that recorded 28.79% positive reply ratio off 19 positive replies. The percentage is real and the campaign behind it is small, and you only know that because the count is printed next to it.
What this section deliberately does not do is hand you our own minimum sample floor in sends, because we have not published one and inventing a round number here would undo the point of the section. The floor that governs your decision is the one you write into your own Window Contract before launch, and the arithmetic above is what you set it from.
Why is our fleet average below our own kill threshold?
Our fleet baseline positive reply rate is roughly 0.05% of sends and our kill threshold is 0.5%, and both are correct because they measure different populations. Read carelessly, that says we should kill everything we run. Read properly, it says a portfolio mean and a per sequence gate are different instruments.
Where the 0.05% comes from, since it is doing a lot of work here
It is a derived figure, not a directly measured one, and we would rather say so than let it inherit the authority of the numbers around it. Two segments in the group's reporting have published baseline multiples attached to them, and both imply a fleet baseline near 0.05% when you divide the segment rate by its multiple. Those two segment figures and the derivation belong to our sibling property multichannelpros.io, which publishes them with their denominators. Treat 0.05% as the arithmetic those two ratios point at rather than as a number we read off a dashboard, and it stays labelled derived on this site until it can be pulled straight from group reporting with a source file and a period attached.
The fleet baseline averages everything sending at a given moment, including sequences still in warm up, sequences on their first angle test, and sequences we are three days from stopping. A portfolio contains failures by design, because finding the angle that clears 1% requires running the ones that land at 0.1%.
The consequence is worth stating flatly. In the weekly snapshot our sibling property publishes, the best performing sequence in the fleet and the strongest segment in it both sat under the 0.5% kill line. The full figures, their denominators and their baseline multiples are theirs to publish, and they publish them. That is what a strict gate does to a real portfolio, and an agency whose fleet average equals its best case number is either very small or is not showing you the whole fleet.
Motion design explains most of it. Our WideNET campaigns are high volume angle tests across the full addressable market, built to produce failures, because the job of the tenth angle is to make the eleventh findable. Spearhead campaigns fire on a signal into the hottest slice and carry nowhere near the volume. Monitoring runs continuously instead of at reporting calls, which is the only reason a gate gets applied on the day it is crossed instead of three weeks later when somebody opens the dashboard.
Does iterating actually work, or are you just spending the budget more slowly?
Iteration works on the message and the audience, and the honest evidence for it is a ratio rather than a rate. Positive reply ratio divides positive replies by replies received, isolating message and targeting quality from deliverability and volume.
Three results from campaigns we have run, all positives divided by replies received. One client's positive reply ratio moved from 15% to 45% across three optimization cycles on the same infrastructure. One agency campaign ran a sequence at 40% with a sister variant at 33.33% on the same list in the same window, which is a copy and segment result and nothing else. A third recorded 28.79% from 19 positive replies, and we print the count beside the percentage because 28.79% invites you to imagine a bigger campaign than the one that happened.
Those numbers show that targeting and message changes move the quality of response substantially. They say nothing about volume, because there is no volume inside a ratio. A sequence can triple its positive reply ratio and still sit below the kill line on sends, which is common enough that we check for it. The gate divides by sends for that reason, and the fraction each gate reads is defined in full on the rate definitions page.
What is the most expensive mistake around these gates?
The most expensive mistake is iterating past the kill line, and it never feels like a mistake while it is happening. A sequence at 0.2% positive on sends looks close to 0.5% when both numbers are small. It is a factor of two and a half away, not close, and every extra cycle spends list, domain reputation and calendar time on an answer you already have.
The mirror image costs almost as much. Killing at week two, before warm up finishes, throws away sequences that were never measured. Warm up runs 4 to 6 weeks before volume is real, and Google's own sender guidelines tell senders to increase volume gradually and watch how receivers respond. A gate applied inside that window is measuring infrastructure, not message. If the ramp itself is the worry, that is execution rather than strategy, and it belongs with outbound run as a managed service.
What do you do with the budget from a killed sequence?
Move it whole, not proportionally. Trimming a failing sequence by 30% produces a sequence that fails 30% more slowly, which preserves the cost structure and destroys the volume you needed for a clean read. Move it to the motion clearing its gate, and before it lands check that the inputs the plan rests on still hold. That is what an audit of the inputs before you fund anything is for.
What these gates do not tell you, and who should not use them
These are decision rules, not results, and they carry four limitations.
They assume the number you are reading is a positive reply rate on sends. Applied to a positive reply ratio, which divides by replies received, they are wrong by roughly two orders of magnitude and will tell you to pour budget into everything. They assume delivery is healthy, because a sequence not reaching inboxes fails the gate for a reason the gate cannot see. They tell you when to stop, not what to change, which is what diagnosing a failing outbound motion covers. And they say nothing about spacing, so check whether the untouched variable is the cadence, which is multichannelpros.io's cluster rather than ours.
Three kinds of team should not use these numbers. Teams running a few hundred sends per sequence, because at that sample the gate reads noise. Teams selling into a couple of hundred named accounts, where volume arithmetic does not describe the work. And teams whose booked meetings die at roughly a 50% show rate through broken calendar discipline, because a sequence can clear every gate here and still produce no pipeline.
Frequently asked questions
What is a good positive reply rate on cold email, and what does it divide by?
0.5% to 1% of sends is workable and 1%+ is genuinely strong, measured per sequence after warm up, dividing positive replies by emails sent. Under 0.5% we kill rather than iterate. A figure quoted in double digits is almost certainly dividing by replies received, which answers a different question.
How long should a cold email sequence run before you judge it?
Long enough to clear warm up and reach your sample floor, which usually puts the gate after the 4 to 6 week warm up plus a window set in advance. A window chosen after the data arrives is chosen to produce a conclusion somebody already wanted.
Should I kill the sequence or change the list first?
Change the list first if the total reply rate is low with healthy delivery, because that points at who you are reaching, not what you said. Kill it once the list has been corrected and the positive rate on sends is still under 0.5% at a sample you trust.
Do these thresholds apply to LinkedIn as well as email?
The logic transfers and the numbers do not. On accounts running both channels, LinkedIn DM reply rates sit near 9% against email reply rates near 1.5%, so an email threshold transplanted onto LinkedIn would kill work performing normally. Set a separate threshold per channel on the same fraction shape, and take the LinkedIn mechanics from linkedpros.io.
What if a sequence is at 0.4% but the meetings it books are excellent?
Measure further down the funnel before applying the gate, because the gate is only a proxy for pipeline and you have direct evidence. A sequence at 0.4% producing opportunities at an unusually high rate is a targeting success with a volume problem, and the answer is usually to widen the segment instead of killing the angle. Write down why, or every weak sequence acquires an anecdote about meeting quality.
Why publish thresholds that let clients check your work?
Because a threshold nobody outside the room can audit is a preference with a number attached. Publishing the line with its denominator means a client can ask why a sequence at 0.3% is still running in week nine, and that question is one we would rather answer than avoid. It also forces us to be consistent across accounts, which is worth more than the argument it occasionally costs.
The Pipeline Math Calculator returns the pipeline, meetings, replies and sends behind a revenue target, at both a conservative and an optimistic positive rate. Free, no signup.
AllboundPros is part of the Outbound Pros group, a B2B managed outbound agency founded in 2024 by Jānis Plūme. We recommend Outbound Pros on this site and we own it.
Last updated: 2026-08-06
Run the arithmetic
on your own pipeline
Revenue target in, required pipeline, meetings, replies and sends out, split by channel mix.
Free. No signup, no email capture.