Convin — internal training

From VoiceBot
to Convin Sense

A product with real traction that churned anyway. Where a voice bot genuinely worked, the ceiling it kept hitting, the arithmetic that made clients leave, and what was rebuilt to close the gap.

Session01
SpeakerVaibhav Gupta
TopicThe business journey
Recorded05 Mar 2026
Duration~30 min

Contents

Why a working product still lost its customers.

  • 1The VoiceBot betThe premise, the wedge, and the pitch that sold03
  • 2Where it genuinely workedFour patterns, with the numbers behind each07
  • 3The ceiling60 to 80% of a human, and the two reasons why10
  • 4Why clients leftThe arithmetic that broke the ROI case13
  • 5The reframeThree changes that became Convin Sense14
  • 6Where that landed80 to 85%, and what it unlocked commercially18
  • 7PM lens and prepTakeaways, metrics, likely questions, reference19
Sessions 2 to 5 cover the mechanics. This one covers the argument.02
1 · The VoiceBot bet

A voice AI agent does a job a human agent used to do.

That is the whole premise, and it is deliberately broad. The question was never whether the technology worked; it was which of those jobs it could actually take.

Sales

Dialling lists, pitching a product, trying to close.

Support

Inbound questions, service queries, account issues.

Collections

Chasing an overdue payment, again and again.

…and more

Any repetitive telephone conversation a business pays people to have.

The finding that shaped everything after it

A voice AI agent probably cannot replace everything a human agent does. That was understood early, and it turned the work from “build a better bot” into “find precisely where a bot beats the alternative, and sell only that.”

The constraint was accepted early, which is why the pitch was narrow rather than grand.03
1 · The VoiceBot bet

So the work became finding the wedge.

Which use cases, which industries, which audiences. Where does a voice agent actually replace a human agent, and where does the business see ROI from it?

StartVoiceBot launches. Pilots sign, and convert to contracts.
3–4 monthsThe working use cases are identified.
6–8 monthsDeployed at scale, and the ceiling becomes visible.

The first hypothesis, and it held

Human agents spend most of their time on leads that will never convert. If 25,000 people are spoken to and 500 buy, then 99% of that effort was spent on people who were never going to. Why should a human do that work at all?

Lead qualification was the first use case picked up, and it worked immediately.04
1 · The VoiceBot bet

The funnel every outbound team is quietly wasting.

A worked example from the session: an ed-tech client selling JEE and NEET courses, dialling every month.

50,000 dialledEvery lead the client generates in a month
25,000 connectedRoughly half pick up. The rest never answer
2,000–3,000Conversations that actually go somewhere
500Purchases. About 1% of those dialled
What that means

Human agents held 25,000 conversations to produce 500 purchases.

~1–2%

of connected calls ended in a sale. The other 98% was time spent on people who were never going to buy.

That 98% is the opportunity. It is also the entire pitch.05
1 · The VoiceBot bet

The bot takes the first 25,000 conversations.

Client hands over25,000 leads a month
Bot dials all of themsimilar connect rate to a human
Asks five or six questionsclass, stream, exam, format
Passes ~5,000 throughthe genuinely interested 20%
Humans closestill 500 purchases

What the client is actually buying
Unchanged
Top line. Still 500 purchases a month
80%
Of human agent time freed
30–40%
Cost reduction across the function
Why the bot stops at qualification

On a ₹5,000 or ₹10,000 course there is negotiation and reassurance a bot does not have. So it politely closes: “thank you, one of our senior counsellors will call you back.” It hands over a warm lead rather than attempting the sale.

Same revenue, materially lower cost. That was the pitch, and it worked.06
2 · Where it genuinely worked

Four patterns, found across eight months of pilots.

01 · Lead qualification

The funnel on the previous page. Reliable, and the default pitch.

02 · Surplus leads

Leads the client generates but has no capacity to call. Pure upside.

03 · Non-tech-savvy audiences

People who do not recognise a bot, and so engage with it properly.

04 · Collections

The first few of many reminder calls. Weaker, but real.


The rule underneath all four

A bot wins where the alternative is nobody calling at all, or a human burning time on a lead that was never going to convert.

And where it loses

Wherever it is competing head to head with a good human agent on the same conversation. That is the whole of act 3.

Collections is fourth for a reason: humans still drive most of the recovery.07
2 · Where it genuinely worked

The cleanest case: leads nobody was calling.

A consumer brand generating leads from app downloads, website forms and social ads, with a team of fifteen to twenty agents.

Monthly leads against monthly capacity
25,000–30,000 dialled
~20,000 never touched
the team's actual capacity generated, then left

50,000 leads generated every month


Why this is the easiest sale in the book

Those 20,000 leads produce nothing today. There is no cannibalisation risk and no top-line to protect, because the counterfactual is zero. Spend ₹100 and generate ₹500.

And the qualifying question it gives you

“How many leads do you generate but never dial?” If the answer is a large number, the efficiency debate in act 3 never has to happen.

Wherever a client has surplus leads, the ROI case is arithmetic rather than argument.08
2 · Where it genuinely worked

And the audiences who never realised.

Agri-commerce selling to farmers in small towns. Logistics reaching truck drivers. Ride-hailing reaching taxi drivers.

Average call duration
Typical audience45–50 s The caller works out what it is, and the call ends early.
Tier-3 and driver bases2 min+ They assume they are speaking to a person, so they keep talking.

The one case where the bot won outright

In agri-commerce the voice bot beat the client's own human agents on conversion. A rare result, and worth being precise about: it proves the ceiling is set by the listener, not by the technology.

The uncomfortable corollary

Call a technically literate person and they identify a bot in seconds. The same product performs very differently across two audiences with identical scripts.

Audience, not industry, is the variable that predicts whether a voice bot performs.09
3 · The ceiling

Then the pilots ran properly, and found the ceiling.

Across clients, the same pattern. If a human agent is 100% efficient, the bot ran somewhere between 60 and 80%, depending on the use case.

Per 100 leads
Human only10 buy 50 leads spoken to, 10 purchases. Expensive, and it works.
Bot qualifies, human closes8 buy 20–25 leads passed through, 8 purchases. Cheaper, and it costs revenue.

60–80%
Bot efficiency against a human agent
−20%
Purchases, on the same lead set
Consistent
The pattern held across clients

The first assumption was that this was an implementation problem: automate the lead flow, add real-time transfer, tighten the nuances. It was not.

It took some time to accept that this was structural rather than a bug.10
3 · The ceiling

Detection kills the lead, and then buries it.

Customer answersan unknown number, picked up
First ten secondsthey work out it is a bot
They hang up20–30% of everyone connected
Marked not interestedno qualifying answer was given
Nobody calls backthe human agent skips them too

Why it compounds rather than just costing

The lead is not merely lost, it is mislabelled. The bot records no interest, so the client's human agents never call them either. A silent graveyard forms inside the client's own CRM, and nobody can see it.

The arithmetic on 100 leads

50 connect. Ten of them hang up on detection. Had a human called those ten, perhaps one would have bought. That one purchase is most of the gap between 10 and 8.

This is the primary reason the numbers came out where they did. It is a software problem, and it is the one Sense was built to attack.

“I don't want to speak with a bot” is a ten-second decision with a permanent consequence.11
3 · The ceiling

The second cause, and it is not fixable in software.

Human nuance

Persuading someone to commit to a ₹5,000 or ₹10,000 purchase involves negotiation, reassurance and reading hesitation. A bot does not have that, and no amount of prompting fully supplies it.

Which is why the two causes are treated differently

Detection is a workflow problem: call again, carry the context, try another channel. Nuance is a capability gap. Convin invested in the first and designed around the second.


The exception that proves the rule

In agri-commerce, where the audience never identified the caller as a bot, conversion beat the human agents outright. Remove detection from the equation and the ceiling largely disappears — which is strong evidence that detection, not intelligence, was the dominant loss.

Separating what is fixable from what is structural is the whole analytical move in this session.

Fundamentally, a bot is not a human agent. That was accepted, not argued with.12
4 · Why clients left

Good clients left a product that worked.

Every client faced the same choice: human only, or a hybrid where the bot qualifies and humans close.

Cost falls
40–50%

Fewer agent hours, and a bot is cheaper per conversation.

Top line falls
20–30%

Sometimes 10%, sometimes 40%. The delta was always there.

Verdict
Not worth it

Clients ran the numbers and stopped using the product.

The threshold, and it is worth quoting exactly

At 90% of a human the argument works: take a 10% revenue hit, save 40% of cost, and the client stays. At 60% it does not, and no discount fixes it. Clients with real scale on VoiceBot stopped using it, not because it failed, but because the arithmetic did.

The product worked. The economics did not.13
5 · The reframe

“If the bot is 60% of a human,
what else can we do to make it 85%?”

Note what the question does not ask. It does not ask for a better voice, a smarter model, or more natural speech. It asks what surrounds the call.


Insight 01

Persistence

Stop treating one call as the whole attempt.

Insight 02

Multi-channel

Voice alone was never how a salesperson works.

Insight 03

Intelligence

Decide which, when and what — per lead.

Three changes, conceptualised together. Together they are Convin Sense.14
5 · The reframe

Insight 01 — the call is not the attempt.

What used to happen

The bot called. The customer picked up, spoke for fifteen or twenty seconds, said they were busy, and hung up. Nobody called that person again. The attempt was over.

What happens now

Call again. Carry the context of the previous conversation, open from there, and ask the question that was never answered. Keep trying until the goal is achieved or the configured limits are hit.


Bounded, not infinite
Goal sete.g. three data points
Attemptcall or message
Not achieved?carry context, retry
Limits hitattempts, days, channels

Limits on how many times to call and how many times to message are configured per campaign. Persistence is a setting, not a personality.

This is the direct answer to the detection problem on page 11.15
5 · The reframe

Insight 02 — replicate how a salesperson actually works.

The design brief, almost verbatim from the session: think about what you would do if you had to reach someone.

Callno answer
Call againfour or five hours later
WhatsApp“trying to reach you, can we chat?”
They reply“busy, let's talk tomorrow”
Call at 2pmthe time they named

Why WhatsApp specifically

A lot of people will not answer an unknown number, and some cannot be called at all because of do-not-disturb registration. Almost everyone is on WhatsApp.

The word that matters

Contextual omni-channel. Omni-channel alone is what every competitor claims. Carrying the same context across every surface is the part that is hard.

Voice, WhatsApp, SMS, email, and WhatsApp calls. One conversation, several surfaces.

Nobody sells by calling once and giving up. The product should not either.16
5 · The reframe

Insight 03 — something has to decide.

Persistence and channels are mechanics. A human salesperson is also weighing four questions constantly, and that judgement is what had to be built.

Which channel?

Call, or message? What does this lead respond to?

What content?

What should this specific message or opening say?

What time?

When is this person most likely to respond?

When to follow up?

If there is no answer, how long before trying again?

This layer is the product

A human salesperson weighs all four without thinking about it. Building a system that weighs them per lead, and can explain its answer, is what took the efficiency from 60% to 80–85%. Session 2 shows it working, under the name Next Best Action.

Everything else was already available to buy. This was not.17
6 · Where that landed

60% → 80–85%. And that number sells.

~9 mo
Since Convin Sense was started
2–3 mo
From starting to the first pilot
80–85%
Of a human agent, against 60% before
9–10
Annual deals signed in two to three months

Why that range is enough

At 80–85% the original argument finally works. A 15–20% top-line hit against a 40% cost saving is a trade a client will take, and it is the same arithmetic that failed at 60%.

From churn to renewals

The VoiceBot business signed pilots and lost them. Sense converts them. Nine or ten annual deals in a quarter is the difference the twenty-five points bought.

The rest of this training series is how those twenty-five points were actually engineered.

Sessions 2 to 5 cover the pipeline, the Meta layer, agent configuration and the pilot motion.18
7 · PM lens

Five takeaways worth stealing for any product.

  • 01
    Find the wedge before scaling the pitchThree or four months were spent identifying exactly where a voice bot beats the alternative. The narrow pitch is why the early pilots converted at all.
  • 02
    Retention is the honest metricPilots signed and contracts converted. The product still failed, because none of that survived contact with the client's own arithmetic a few months later.
  • 03
    Separate fixable from structuralDetection was software and most of the loss. Nuance was a capability gap. Attacking the first and designing around the second is the entire strategy.
  • 04
    Benchmark honestly against the incumbentNaming the bot as 60% of a human, out loud, is what made the 85% target meaningful. A vaguer claim would have produced a vaguer product.
  • 05
    Sell the delta, not the featureClients did not buy omni-channel. They bought a cost saving that no longer cost them revenue. The feature list is downstream of that.
PM framing. Analysis layered on the session, not claims made in it.19
7 · PM lens

The dashboard a PM should be reading.

MetricWhat it countsHumanVoiceBotWhy it matters
Connectivity rateCalls actually connected~50%~50%Roughly unchanged. The bot dials the same list.
First-10-second dropHang-ups on detectionn/a20–30%The metric that did not exist before, and mattered most.
Qualification ratePassed through as interestedn/a~20%Of everyone connected. The bot's actual output.
Conversion deltaPurchases against human-onlybaseline−20%The number that decided every renewal.
Agent time savedHours returned to the team80%The saving being sold.
Cost per outcomeAll-in cost per purchasebaseline−30–40%Good, and not good enough on its own.
Metrics named in the session; the framing of them is PM analysis.20
7 · Interview prep

Questions an interviewer will actually ask.

Product sense“You had traction and annual contracts. Why did clients still leave?”The ROI arithmetic: a 40–50% cost saving against a 20–30% top-line hit. The product worked; the economics did not.10 · 13
Diagnosis“Why was the bot only at 60%?”Detection, not intelligence. 20–30% of connected callers hang up in ten seconds, and those leads are then mislabelled and never re-contacted.11
Prioritisation“You are at 60%. Where do you invest first?”Split fixable from structural. Detection is software and most of the loss; nuance is a capability gap. Spend on the first.11 · 12
Go-to-market“Qualify a prospect in one question.”“How many leads do you generate but never dial?” Surplus leads remove the efficiency-gap risk entirely.08
Metrics“One metric to run this product on.”Conversion delta against human-only. Cost saving is easy to move and easy to flatter; the delta decided every renewal.20
Strategy“What is actually defensible here?”Not the voice. The layer deciding channel, timing and content per lead — everything else in the stack can be bought.17
Session 1 rewards judgement more than recall. Have a number ready for each answer anyway.21
7 · Reference card

The numbers, in one place.

Leads dialled50,000per month
Connected25,000roughly 50% pick up
Meaningful conversations2,000–3,000of those connected
Purchases500–600about 1% of dialled
Qualified and passed on~20%of those connected
Agent time saved80%the core of the pitch
Cost reduction30–40%across the function
Bot efficiency60–80%of a human agent
First-10-second drop20–30%of connected callers
Top-line hit20–30%range 10–40%
Cost saving, hybrid40–50%not enough to offset
Surplus leads~20,000generated, never dialled
Call duration, tier-32 min+against 45–50 s
Sense efficiency80–85%the target that was hit
Annual deals9–10in two to three months
The threshold90%where the argument works
All figures from Session 1, 05 Mar 2026. Ranges preserved as stated.22
Where the series goes next

This was the argument. The rest is the machinery.

02How it actually worksThe voice pipeline, the vendor stack, the Next Best Action engine, and every campaign setting.
03The Meta integrationWhatsApp end to end: portfolios, limits, templates, payment, and two accounts that got banned.
04Configuring an agentPrompting from decision trees to the Eight Pillar Framework, and the call-control panel.
05Debugging and the pilotA live teardown, the pricing model, and the 21-day pilot SOP.

Every mechanism in those four sessions traces back to one number on page 13. If you can hold that arithmetic in your head, the rest of the product explains itself.

Five sessions, one argument.23