A live teardown of a configured bot that went wrong in four ways, then the commercial story: how the pilot shrank from ₹1.5 lakh and 10,000 leads to ₹10,000 and 1,000, and the 21-day process that replaced it.
The most useful half hour in the series: a configured agent run against a real number, with each fault diagnosed live. Every one of them is a configuration error rather than a bug.
Agent named Arjun, configured male. A female voice called.
The bot re-introduced itself and never advanced, then hung up.
Why this channel, why this message, why the goal is still open.
Human calling enabled on the campaign, and still no way to dial.
Session 4 said a PM who understands the system can diagnose a fault themselves and hand the product team a cause rather than a symptom. This session is what that looks like in practice, and every fault above was found without opening a ticket.
Setting agent gender to male determines how the bot sounds, so the voice should have followed.
The voice ID in the TTS configuration, which is a separate field. Gender is metadata used in the prompt; the voice ID is the actual audio. A female voice ID had been selected, so a female voice called.
Two fields, two purposes, no validation between them. Nothing warns you that the persona and the voice disagree.
Settings → voice cloning. Both routes go through Cartesia; they differ in how much audio you feed them and how good the result is.
Five to ten seconds of recording, a name, a description and a language. The clone is ready almost immediately and picks up intonation surprisingly well. This is what the product uses today.
One to two hours of audio uploaded into Cartesia's environment, which they train on. Materially higher quality, because the dataset is far larger.
Worth trying on your own voice before an interview. It is the fastest way to understand what the TTS layer is doing.
An empty transcript points at the speech-to-text layer, not the model. If STT works everywhere else and fails on one agent, it is a configuration difference on that agent.
Deepgram Nova 3 had been selected directly rather than Convin STT managed. That specific model may not be integrated to this app domain. Switching to the managed option restores transcription.
Which is a useful default to carry: unless you have a reason to pin a specific model, choose the managed option. It resolves to whichever model is actually wired up for that environment.
All three are small information icons, easy to miss, and each one answers a question a PM will otherwise raise as a ticket.
| Icon location | Answers | What it shows |
|---|---|---|
| On a message | Why this message, on this channel | The full decision context: which channel was chosen and why, why this action was taken, and what came from the knowledge base. |
| Above lead metrics | Why the goal is still open | The reasoning behind goal achievement, lead qualification status and interest level for this specific lead. |
| Beside the message timestamp | When the next action fires | The cooldown schedule and the time the next best action is due. Hidden until you hover slightly to the right of the icon. |
The cooldown popup only appears when the pointer is dragged a fraction to the right of the timestamp icon. It took the room several minutes to find it, and the speaker said plainly that this is not the intended experience and will change.
Testing. Rather than waiting hours for a cooldown to expire, execute it and watch the next action fire immediately.
It does not simply resend the last message. Executing the cooldown triggers a fresh NBA, which recalculates and may pick a different channel entirely.
Quoted from the session: “This will be changed eventually, not the ideal UI/UX. Right now we have kept it hidden.” A team that names its own rough edges is easier to work with than one that does not.
Switch to human agent, then use the message icon beside the lead's number. Two options appear, alongside a countdown labelled session active.
Something to do with the Meta limits on the account or the portfolio.
The 24-hour rolling window from Session 3. Because this customer replied at 16:32 on the 13th, the window runs to 16:32 on the 14th and free-form is allowed. Outside that window, only a template can be sent, by a human or by the bot.
Enabling human call as a channel on the campaign is not sufficient. Three separate things must all be true before the click-to-call icon appears.
Human call must be switched on in the agent configuration, not only on the campaign. It was inactive here.
Settings → user management. Calling is off per user by default, and needs admin access to change.
One of the Exotel numbers procured for human calling is assigned to the individual. Their calls dial out from it.
The human call channel also asks for a speech-to-text selection, which seems wrong for a human-to-human call. It is not: Convin transcribes human calls too, and feeds that transcript back into the lead's context. When the AI picks the lead up again it knows what the human already discussed.
Recording begins the moment the call is in progress, which is before the customer picks up. Anything said during the ringing period is captured in the transcript. It surprised the room, and it is worth knowing before you speak over a dialling tone.
Once the human ends the call, the lead returns to Convin AI. The next best action is generated with the human conversation in context, so the follow-up references what was actually said.
Which is the substantive difference between a transfer and a handoff. The lead does not leave the system; it comes back better informed.
The original motion was a full deployment in miniature: integrate the client's CRM in and out, work around their existing dialer, and prove the concept at scale before anyone had proved it at all.
Contextual calling, WhatsApp and an omni-channel loop, with no guardrails at all. The speaker describes even the current product as an MVP. The first two clients were a logistics company recruiting driving partners and a lending business calling tier-2 and tier-3 customers about loans.
| Client | Leads | Ticket | Notes |
|---|---|---|---|
| Snabbit | 10,000 | ₹1,50,000 | The first paid 10,000-lead pilot. Much of the product was built from their feedback. |
| Miles Education | 10,000 | ₹1,50,000 | Ed-tech. Same shape, same ticket. |
| Cashify | — | ₹1,50,000 | Ran alongside the others as volume built up. |
| ABSLI | 25,000 | ₹1,50,000 | Insurance. Two and a half times the leads for the same money. |
₹1.5 lakh did not make money on a month of work. It existed to prove the concept. Which meant every commercial problem with the model was a timing problem, not a pricing one.
Bigger pilots are not more convincing. A different lead set behaves differently for reasons that have nothing to do with the product: where the leads came from, how old they are, what mindset they are in. Running three times as many leads adds duration and variance rather than proof.
Fifteen times less money and ten times fewer leads, deliberately. The token is not revenue; it is a filter.
To maintain sanity. A free pilot attracts companies with no intention of buying, and each one costs a PM several weeks. A small cheque filters for seriousness without pretending to be a revenue line.
Because 1,000 is enough to show ROI, and everything past that adds elapsed time and variance without adding evidence.
₹1 per lead uploaded. It covers the LLM cost of generating next best actions and computing metrics and entities for that lead. A hundred leads is a hundred rupees.
Charged on what actually goes out: voice minutes, and messages by category. During a pilot the client pays only the fixed token; consumption is used to compute the ROI figure.
| Lead management | ₹1 | per lead uploaded |
| Voice | ₹1 | per 15-second pulse, on connect only |
| Utility template | ₹0.12 | per message, paid to Meta |
| Marketing template | ₹0.94–0.95 | per message, paid to Meta |
| Free-form message | see note | billed by Convin, covers the LLM |
The dashboard shows a voice figure of 1.1025; the working number is 1.1.
It matters more than it looks. Free-form messages are the ones sent inside the 24-hour window, which is where most of the conversation happens.
“The cost is point five, charged from our platform.”
Reads as ₹0.50 per message.
“If it's a free-form message I incur five paise.”
Reads as ₹0.05 per message.
On a thousand-lead campaign with a handful of free-form messages per engaged lead, the difference between five paise and fifty paise is the difference between a rounding error and a real line in the ROI calculation you are about to present to a client.
Closure deck, then handback to sales for the annual contract.
Any delay must be raised immediately in the Convin Pilot group. If you own a pilot you own its outcome, and slippage is the default failure mode.
Asked whether 1,000-lead pilots actually finish in fourteen days, the answer was no. The deadline still works, because it creates urgency on both sides.
Note the arithmetic: 100 leads for UAT plus 450 and 450 makes the thousand. The pilot is designed so the first hundred are a rehearsal.
What is this client, what is the use case, and how many leads do they have per month. Low monthly volume caps the revenue, which is the moment to ask whether a second use case exists. Human-agent involvement is agreed here, before anyone promises anything.
Whatever the internal team agreed is worthless if the person signing has a different picture. The DM needs to hear that Sense qualifies rather than closes, and that human agents remain part of the loop.
The minimum for eight-pillar generation. Their best calls, same use case.
Someone with portfolio admin, free for a fifteen-minute call to do the Meta integration.
What their human agents currently achieve. Without it you cannot prove anything.
A pilot that returns 10x sounds like a success. It is not, if their own agents were returning more. The comparison, not the number, is what closes the deal.
Matching the human benchmark is already a win, because Sense does the same job at a fraction of the cost of five or six agents. Beating it is upside. Falling well short means the client should hire instead, and you should know that before they do.
Once the client is testing, the loop does not close. It can run a month or two. There is always another imperfection, and the client is not wrong to see them.
The calls go out under their brand, from a Truecaller-registered number carrying their name. A bad call is their reputational problem, and it can end up on LinkedIn or Instagram. For a large brand that is a real risk, so they will push for perfect.
The honest position to hold: there is always a delta, and it cannot be perfect. The pilot is judged on business outcome, not on whether every call was flawless.
The first hundred leads of the thousand are a rehearsal. They are audited twice, by two teams looking for different things.
You configured the agent, so you know how the conversation was meant to go. You are auditing whether it went that way: whether the flow held, whether objections were handled, whether the bot stayed inside its guardrails.
If the UAT result is strong, the friction disappears for the rest of the pilot. The client hands over the remaining leads without argument. If it is weak, every subsequent campaign is negotiated.
Leads uploaded, connectivity achieved, conversion achieved, cost incurred, and the ROI that falls out of them, set against the benchmarks captured at kickoff.
Sales re-enters and negotiates the annual contract from the pilot numbers. The PM's job ends at a defensible result, not at a signature.
Feedback during campaign one is incorporated while it runs rather than deferred, which is what makes two campaigns better than one long one.
A client sees one demo of contextual omni-channel outreach and concludes it will replace their agents. They are not being unreasonable; the demo is genuinely impressive. But the expectation has to be corrected before the pilot, not after it.
Every AI currently can assist a human being, not replace one. Said at kickoff it sets a bar the pilot can clear. Said at closure it sounds like an excuse.
During internal testing clients will report that a word was misheard, that the pacing was uneven, that the bot did not stop when it should have, that it hallucinated at minute eight. All of that is real, and some of it is unavoidable.
Not on call quality, on business sense. Ten thousand rupees in and seven lakh out is roughly seventy times the investment, and that argument survives a handful of imperfect calls. The quality conversation is real and it belongs to the roadmap, not to the pilot decision.
The 20–25% figure is the speaker's rule of thumb for a single call, not a measured error rate for the product.
| When | What | Why it matters |
|---|---|---|
| Before the client kickoff | Use case, monthly lead volume, human-agent plan | Agreed with sales at internal kickoff |
| At the client kickoff | Decision maker on the call | Non-negotiable. Alignment without them does not hold |
| From the client | 20 call recordings, a WABA contact, benchmark data | All three, or the pilot cannot be judged |
| Before any client testing | Internal testing complete | Client testing loops do not close on their own |
| At 100 leads | Data Labs quality audit, PM conversation audit | Run in parallel, then incorporate |
| At 550 leads | Campaign 1 reviewed, feedback applied | Mid-flight, not deferred to the end |
| At 1,000 leads | Closure deck against the benchmarks | Present to the DM, then hand back to sales |
| Commercial | “Why would you shrink your own pilot?” | ₹1.5 lakh and 10,000 leads took a month and proved nothing extra. ₹10,000 and 1,000 leads proves the same thing in two weeks and closes faster. | 14 · 15 |
| Rigour | “How do you know a pilot succeeded?” | Against benchmarks captured at kickoff: their human connectivity, identification and conversion. ROI alone is not evidence. | 20 |
| Pricing | “What does a client actually pay?” | ₹1 per lead as a management fee, plus consumption: ₹1 per 15-second pulse on connect, and per-message rates by template category. | 16 |
| Debugging | “The bot is not responding. Where do you start?” | Empty transcript means STT. Check whether a specific model was pinned instead of the managed option. Then the three reasoning icons. | 06 · 07 |
| Judgement | “A client wants to keep testing. What do you do?” | Move them to UAT on 100 leads with a real audit. Client testing loops do not close, and the calls carry their brand, so the anxiety is legitimate. | 21 · 22 |
| Honesty | “Will this replace our agents?” | No. It absorbs roughly half the roles in a team and the rest stay human. Say it at kickoff, where it sets a bar, not at closure where it sounds like an excuse. | 24 |
| Pilot token | ₹10,000 | two weeks, 1,000 leads |
| Previous pilot | ₹1,50,000 | one month, 10,000 leads |
| Lead management | ₹1 / lead | covers NBA and metrics |
| Voice | ₹1 / pulse | 15 s, on connect |
| Utility template | ₹0.12 | per message, to Meta |
| Marketing template | ₹0.94–0.95 | per message, to Meta |
| Free-form | unresolved | ₹0.05 or ₹0.50; verify |
| Recordings needed | 20 | minimum, eight pillar |
| UAT size | 100 leads | audited twice |
| Campaign split | 450 + 450 | completes the thousand |
| SOP length | 21 days | behind a two-week promise |
| Instant voice clone | 5–10 s | audio needed |
Loadshare and a lending client for unpaid validation. · Snabbit, the first paid 10,000-lead pilot. · Miles Education, 10,000 leads. · Cashify. · ABSLI, 25,000 leads at the same ₹1.5 lakh ticket.
Carried across all five sessions. Each is a place where the recordings disagree, or where the audio was not clear enough to be sure.
| Free-form rate | Session 3 said point five; Session 5 said five paise. A tenfold gap on the message type sent most often. |
| Marketing rate | ₹0.90 in Session 2, ₹0.95 in Session 3, ₹0.94–0.95 in Session 5. Converging, but ask for the rate card. |
| LLM providers | Session 2 described open-source models; Session 4 described OpenAI, Gemini, Claude or closed-source. |
| Sign-in method | Session 2 said Google or Microsoft; Session 3 said Microsoft only. |
| The new dialer | Named in Session 4 as being trialled to replace Exotel. The name was not audible. |
| Second validation client | Named alongside Loadshare in Session 5. A lending business; the name was not clear. |