A product with real traction that churned anyway. Where a voice bot genuinely worked, the ceiling it kept hitting, the arithmetic that made clients leave, and what was rebuilt to close the gap.
That is the whole premise, and it is deliberately broad. The question was never whether the technology worked; it was which of those jobs it could actually take.
Dialling lists, pitching a product, trying to close.
Inbound questions, service queries, account issues.
Chasing an overdue payment, again and again.
Any repetitive telephone conversation a business pays people to have.
A voice AI agent probably cannot replace everything a human agent does. That was understood early, and it turned the work from “build a better bot” into “find precisely where a bot beats the alternative, and sell only that.”
Which use cases, which industries, which audiences. Where does a voice agent actually replace a human agent, and where does the business see ROI from it?
Human agents spend most of their time on leads that will never convert. If 25,000 people are spoken to and 500 buy, then 99% of that effort was spent on people who were never going to. Why should a human do that work at all?
A worked example from the session: an ed-tech client selling JEE and NEET courses, dialling every month.
Human agents held 25,000 conversations to produce 500 purchases.
of connected calls ended in a sale. The other 98% was time spent on people who were never going to buy.
On a ₹5,000 or ₹10,000 course there is negotiation and reassurance a bot does not have. So it politely closes: “thank you, one of our senior counsellors will call you back.” It hands over a warm lead rather than attempting the sale.
The funnel on the previous page. Reliable, and the default pitch.
Leads the client generates but has no capacity to call. Pure upside.
People who do not recognise a bot, and so engage with it properly.
The first few of many reminder calls. Weaker, but real.
A bot wins where the alternative is nobody calling at all, or a human burning time on a lead that was never going to convert.
Wherever it is competing head to head with a good human agent on the same conversation. That is the whole of act 3.
A consumer brand generating leads from app downloads, website forms and social ads, with a team of fifteen to twenty agents.
50,000 leads generated every month
Those 20,000 leads produce nothing today. There is no cannibalisation risk and no top-line to protect, because the counterfactual is zero. Spend ₹100 and generate ₹500.
“How many leads do you generate but never dial?” If the answer is a large number, the efficiency debate in act 3 never has to happen.
Agri-commerce selling to farmers in small towns. Logistics reaching truck drivers. Ride-hailing reaching taxi drivers.
In agri-commerce the voice bot beat the client's own human agents on conversion. A rare result, and worth being precise about: it proves the ceiling is set by the listener, not by the technology.
Call a technically literate person and they identify a bot in seconds. The same product performs very differently across two audiences with identical scripts.
Across clients, the same pattern. If a human agent is 100% efficient, the bot ran somewhere between 60 and 80%, depending on the use case.
The first assumption was that this was an implementation problem: automate the lead flow, add real-time transfer, tighten the nuances. It was not.
The lead is not merely lost, it is mislabelled. The bot records no interest, so the client's human agents never call them either. A silent graveyard forms inside the client's own CRM, and nobody can see it.
50 connect. Ten of them hang up on detection. Had a human called those ten, perhaps one would have bought. That one purchase is most of the gap between 10 and 8.
This is the primary reason the numbers came out where they did. It is a software problem, and it is the one Sense was built to attack.
Persuading someone to commit to a ₹5,000 or ₹10,000 purchase involves negotiation, reassurance and reading hesitation. A bot does not have that, and no amount of prompting fully supplies it.
Detection is a workflow problem: call again, carry the context, try another channel. Nuance is a capability gap. Convin invested in the first and designed around the second.
In agri-commerce, where the audience never identified the caller as a bot, conversion beat the human agents outright. Remove detection from the equation and the ceiling largely disappears — which is strong evidence that detection, not intelligence, was the dominant loss.
Separating what is fixable from what is structural is the whole analytical move in this session.
Every client faced the same choice: human only, or a hybrid where the bot qualifies and humans close.
Fewer agent hours, and a bot is cheaper per conversation.
Sometimes 10%, sometimes 40%. The delta was always there.
Clients ran the numbers and stopped using the product.
At 90% of a human the argument works: take a 10% revenue hit, save 40% of cost, and the client stays. At 60% it does not, and no discount fixes it. Clients with real scale on VoiceBot stopped using it, not because it failed, but because the arithmetic did.
Note what the question does not ask. It does not ask for a better voice, a smarter model, or more natural speech. It asks what surrounds the call.
Persistence
Stop treating one call as the whole attempt.
Multi-channel
Voice alone was never how a salesperson works.
Intelligence
Decide which, when and what — per lead.
The bot called. The customer picked up, spoke for fifteen or twenty seconds, said they were busy, and hung up. Nobody called that person again. The attempt was over.
Call again. Carry the context of the previous conversation, open from there, and ask the question that was never answered. Keep trying until the goal is achieved or the configured limits are hit.
Limits on how many times to call and how many times to message are configured per campaign. Persistence is a setting, not a personality.
The design brief, almost verbatim from the session: think about what you would do if you had to reach someone.
A lot of people will not answer an unknown number, and some cannot be called at all because of do-not-disturb registration. Almost everyone is on WhatsApp.
Contextual omni-channel. Omni-channel alone is what every competitor claims. Carrying the same context across every surface is the part that is hard.
Voice, WhatsApp, SMS, email, and WhatsApp calls. One conversation, several surfaces.
Persistence and channels are mechanics. A human salesperson is also weighing four questions constantly, and that judgement is what had to be built.
Call, or message? What does this lead respond to?
What should this specific message or opening say?
When is this person most likely to respond?
If there is no answer, how long before trying again?
A human salesperson weighs all four without thinking about it. Building a system that weighs them per lead, and can explain its answer, is what took the efficiency from 60% to 80–85%. Session 2 shows it working, under the name Next Best Action.
At 80–85% the original argument finally works. A 15–20% top-line hit against a 40% cost saving is a trade a client will take, and it is the same arithmetic that failed at 60%.
The VoiceBot business signed pilots and lost them. Sense converts them. Nine or ten annual deals in a quarter is the difference the twenty-five points bought.
The rest of this training series is how those twenty-five points were actually engineered.
| Metric | What it counts | Human | VoiceBot | Why it matters |
|---|---|---|---|---|
| Connectivity rate | Calls actually connected | ~50% | ~50% | Roughly unchanged. The bot dials the same list. |
| First-10-second drop | Hang-ups on detection | n/a | 20–30% | The metric that did not exist before, and mattered most. |
| Qualification rate | Passed through as interested | n/a | ~20% | Of everyone connected. The bot's actual output. |
| Conversion delta | Purchases against human-only | baseline | −20% | The number that decided every renewal. |
| Agent time saved | Hours returned to the team | — | 80% | The saving being sold. |
| Cost per outcome | All-in cost per purchase | baseline | −30–40% | Good, and not good enough on its own. |
| Product sense | “You had traction and annual contracts. Why did clients still leave?” | The ROI arithmetic: a 40–50% cost saving against a 20–30% top-line hit. The product worked; the economics did not. | 10 · 13 |
| Diagnosis | “Why was the bot only at 60%?” | Detection, not intelligence. 20–30% of connected callers hang up in ten seconds, and those leads are then mislabelled and never re-contacted. | 11 |
| Prioritisation | “You are at 60%. Where do you invest first?” | Split fixable from structural. Detection is software and most of the loss; nuance is a capability gap. Spend on the first. | 11 · 12 |
| Go-to-market | “Qualify a prospect in one question.” | “How many leads do you generate but never dial?” Surplus leads remove the efficiency-gap risk entirely. | 08 |
| Metrics | “One metric to run this product on.” | Conversion delta against human-only. Cost saving is easy to move and easy to flatter; the delta decided every renewal. | 20 |
| Strategy | “What is actually defensible here?” | Not the voice. The layer deciding channel, timing and content per lead — everything else in the stack can be bought. | 17 |
| Leads dialled | 50,000 | per month |
| Connected | 25,000 | roughly 50% pick up |
| Meaningful conversations | 2,000–3,000 | of those connected |
| Purchases | 500–600 | about 1% of dialled |
| Qualified and passed on | ~20% | of those connected |
| Agent time saved | 80% | the core of the pitch |
| Cost reduction | 30–40% | across the function |
| Bot efficiency | 60–80% | of a human agent |
| First-10-second drop | 20–30% | of connected callers |
| Top-line hit | 20–30% | range 10–40% |
| Cost saving, hybrid | 40–50% | not enough to offset |
| Surplus leads | ~20,000 | generated, never dialled |
| Call duration, tier-3 | 2 min+ | against 45–50 s |
| Sense efficiency | 80–85% | the target that was hit |
| Annual deals | 9–10 | in two to three months |
| The threshold | 90% | where the argument works |
| 02 | How it actually works | The voice pipeline, the vendor stack, the Next Best Action engine, and every campaign setting. |
| 03 | The Meta integration | WhatsApp end to end: portfolios, limits, templates, payment, and two accounts that got banned. |
| 04 | Configuring an agent | Prompting from decision trees to the Eight Pillar Framework, and the call-control panel. |
| 05 | Debugging and the pilot | A live teardown, the pricing model, and the 21-day pilot SOP. |
Every mechanism in those four sessions traces back to one number on page 13. If you can hold that arithmetic in your head, the rest of the product explains itself.