The voice pipeline, the vendor stack, the Next Best Action engine, and every lever in campaign settings. Session 1 explained why Sense exists; this is how it runs.
The voice pipeline in this section
Template first, then free-form
Where the segment responds
Live transfer, or full handoff
The call is a component. Sense is the system that decides when to place it, and that decision is the subject of sections 6 through 8.
The four compound. A fast bot that mishears is not human, and a perfectly accurate one that takes four seconds is not either. Get one wrong and the other three cannot rescue the call.
The session named these four as the factors that combine; it did not weight them. Latency is drawn longest because it is the one the speaker singled out.
A real objection, a question, a correction. Stop immediately.
A filler word, a cough, background speech, an agreeing “mm-hm”.
Telling those two apart in under a second, without a transcript yet.
Wiring three APIs together takes days. Interruption, turn-taking and the rest of the nitty-gritty took eighteen months. It is also the part a competitor cannot copy from an architecture diagram.
The lesson that produced Sense: transcribing speech, calling an LLM and playing audio back is not a voice product. It is the easiest quarter of one.
Interruptions, silences, barge-in
What happens after a call that went nowhere
Carrying what was said into the next attempt
Choosing the channel, time and message at all
It removes a telecom migration from the sale, and the client keeps their existing numbers and compliance posture. The integration burden moves to Convin, which is where it is cheapest to carry.
The deciding question in both cases was not which option is better. It was whether the quality delta justifies the price delta at Convin's call volume. In a consumption-priced product, that arithmetic belongs to the PM.
The Physicswallah case. A lead is uploaded, the bot calls to find out whether they want a JEE or NEET course, and must mark them hot or warm back into the client's system.
A disposition. Not an answer to the question it was sent to ask. The lead is neither qualified nor disqualified, and nothing downstream knows the difference.
A retry mechanism existed under VoiceBot. But a retry without context repeats the same cold opening, and the odds that any given lead is free at the exact moment you dial are, as Sudesh put it, very low.
“I'm Sudesh from Physicswallah, about your recent website visit. Are you exploring JEE or NEET courses?”
“Yes, but I'm driving. Can you call me back at 5pm?”
“We spoke this morning. Can we speak now about whether the course is right for you?”
No pickup? Retry tomorrow, and the day after, until the engagement window closes.
“This context will get retained, and the next call will be automatically generated for 5pm with that specific context.”
The lead never repeats themselves, which is precisely what makes the second call read as a follow-up rather than a fresh cold call.
Defined per campaign. For example: identify whether this lead is genuinely interested in a JEE or NEET course.
Persistence is bounded. The engagement window and the per-channel limits end it, not a human deciding to give up.
“Until and unless my goal gets achieved.” That is the condition. Not until the call is made.
the voice pipeline
where it performs
Held at the lead level, not the call level
template, then free-form
longer-form follow-up
+ WhatsApp calls, being added
“Hey Animesh, we actually spoke about this. Are you still interested in the JEE course?” is a WhatsApp message that only works because the call before it is remembered.
of usage cost
of revenue returned
The pilot is not a discount. It is the evidence that makes the annual number defensible: the client commits for the year against an upfront lead volume, then uses Sense as they need it.
Roughly a 3:1 ratio of running pilots to signed contracts. That funnel is the growth engine, and the ~100x figure is a strong case rather than a median.
There is no dormant-seat revenue to hide behind. A client who stops finding value stops spending, and the decline shows up in usage months before it shows up in a contract.
The upside is symmetrical: a client who finds Sense working expands lead volume without renegotiating anything.
Which is the argument for usage dashboards, and the reason the campaign-level cost levers in section 8 are a product concern rather than an ops one.
A subdomain per tenant. Email and password sign-in, so throwaway accounts can be created for testing.
One shared URL for everyone. SSO only, Google or Microsoft. No passwords, no dummy accounts. An organisation ID then selects the client workspace.
Live orgs: PW · TNSDC · Cordelia · Mosaic Wellness · Miles Education
Not “triggering lots of calls at once”. That answer was given in the room and corrected. A campaign is a dataset, plus an agent, plus a time box, run to produce an outcome you can read.
Widths are illustrative, not measured outcome rates.
When leads are uploaded, each receives a default system message: the campaign's opening context for the agent. It is set once, and applied to every lead.
There is no prior interaction with this lead. They came from the website and are looking for a CPA or CMA course. Start engaging through an AI call.
Use when you already know how this segment responds.
There is no prior interaction with this user. You figure out how to engage.
The agent picks voice or WhatsApp itself, using the Next Best Action logic in section 7.
This is where campaign intent meets per-lead context, which makes it the highest-leverage text field in the product.
From a live campaign of roughly 500 leads with two channels active. The lead was exploring CPA and CMA courses.
Any first contact on WhatsApp must be a Meta-approved template in one of three categories. The category decides both what it costs and how many people receive it.
Authentication — OTPs and verification codes. Approved as a third category, not used by Sense.
A marketing template costs seven times a utility one and reaches half as many people, so a message sent in the wrong category is roughly fourteen times less efficient. Choosing the category is an economic decision, not a compliance formality.
It swings with the audience, and it moves in the opposite direction to what most people assume.
More likely to have blocked promotional senders, or to run message filtering. Delivery drops.
Fewer filters in place, promotional messages land. Delivery can reach around 70%.
Reachability, not creative, is the binding constraint on a marketing template.
Send the template again. Currently the same one, though varying it is being explored.
Meta's API in both directions: a GET to retrieve what the customer sent, a POST to send the reply.
Every guardrail in campaign settings exists to buy that first reply economically.
Which reframes the funnel: on WhatsApp the first reply, not the first send, is the real conversion step. Everything before it is paid and unread; everything after it is cheap and human.
if customer_says(X):
send(Y)
Predictable and controllable, and immediately recognisable as automation. It also cannot handle anything the author did not anticipate.
“Got it. You are a student with accounts experience. Are you currently in college as well?”
Casual vocabulary and loose grammar, kept intentionally. A flawless paragraph reads like a machine; this reads like someone typing.
Occasionally an artifact slips through: a stray double dash gives away that a line was generated. In the campaign reviewed live, that happened three times across an entire conversation.
Naturalness here is a product decision, not a model limitation. That distinction is worth making explicitly if the topic comes up.
Before every AI call or WhatsApp message, the system generates a Next Best Action. There is nothing rule-based in this decision.
everything said so far
campaign intent
the agent's prompt
an LLM call, per lead
channel · time · content
an AI call, or a WhatsApp message
why this channel, recorded
why now, recorded
“AI call chosen for direct engagement; initial contact, no channel fatigue.” · “The data shows prior interest; an AI call enables personalised engagement and early trust building.”
The boss is the NBA. The two nodes are the AI call and WhatsApp. The boss decides the mode, the channel, the content and the time, then hands the instruction to whichever node carries it out. The nodes only execute. All judgment sits with the boss.
“The cooldown is determined automatically by an AI, through an LLM call. There is nothing rule-based over here.”
The practical consequence: two leads in the same campaign can be contacted on different channels, at different intervals, with different messages, without anyone configuring that.
| Goal achievement | configurable, not automatic | A “not interested” answer counts as goal-achieved only if you explicitly configure it that way. Otherwise the lead stays open. |
| Interest level | hot · warm · cold | Separate from goal achievement. You can achieve the goal and learn the lead is cold. |
| External ID | immutable unique key | Identifies the lead in Convin's system. Archive and re-add the same person and you must use a different external ID. |
| Lead owner | Convin AI, or a human | Ownership can move to a human, who then works the lead from the same interface with full history. |
Live call transfer hands a conversation to a human mid-call via a tool call, and the call never drops. Lead handoff moves ownership of the whole lead. Different things.
Campaign name and description · agent, with its prompt, attached
Initial language · real-time switching if the customer changes
Default system message applied to every lead on upload
Time zone, start hour, end hour · DND hours the campaign must respect
Engagement days: 3, 5 or 7 · read from user behaviour, not guessed
100 messages per lead, 50 per day, stop after 5 unanswered · 15 calls per lead, 3 to 5 per day
Post-completion inbound behaviour and concurrent call allocation, the slot and queue mechanism, were named as still being built. Both are session 3 material.
Cost is incurred whether or not anyone reads it. An unread marketing template is pure loss, which is why the unanswered-message cap matters so much.
An unanswered call costs nothing. Billing starts when the customer picks up, at ₹1 per pulse, where a pulse is 15 seconds of connected conversation.
Pickups concentrate in the first three days. A lead who has not answered by day three is unlikely to answer on day seven, so days four onward add cost without adding revenue.
One number against 10,000 leads gets marked as spam, and then nobody answers at all. Six or seven numbers are rotated per campaign, each with its own daily call cap.
6–7 numbers · per-number daily cap
A cost lever disguised as a deliverability feature: a spam-flagged number wastes the whole campaign.
| Lever | Metric it moves | How to think about it |
|---|---|---|
| Engagement days | Cost per lead | Shorten when a segment converts early. Lengthen only on evidence from later-day pickups. |
| Max calls per day | Connect rate | More attempts lift reach up to a point, then start burning goodwill and numbers. |
| Number rotation | Connect rate | Protects against spam labelling, the failure mode that silently kills a campaign. |
| Template category | Delivery and cost | Utility reaches ~100% at ~₹0.12; marketing ~50% at ~₹0.90. Use utility wherever the context genuinely allows. |
| Unanswered cap | Wasted spend | Every send before the first reply is paid for and unread. Five is the current default. |
| Initial language | Engagement depth | Switching mid-call is automatic, but the opening still sets the tone. |
| System message | Qualification accuracy | The highest-leverage text in the product. It frames every downstream NBA. |
| Architecture | “Walk me through what happens when a customer says hello.” | Dialer, STT, LLM with the agent prompt, TTS, dialer, inside roughly 2.5 seconds. Name the vendor at each layer. | 04 · 09 |
| Tradeoffs | “Why Cartesia and not ElevenLabs?” | ElevenLabs is more natural but materially more expensive. At Convin's volume the quality delta did not justify the price delta. | 10 |
| Depth | “What is actually hard about building a voice agent?” | Not the three layers. Turn-taking and interruption: meaningful versus non-meaningful interrupts took eighteen months. | 07 |
| Product | “What is defensible? Anyone can wire up these APIs.” | The NBA loop. An LLM choosing channel, timing and content per lead, with its reasoning recorded for audit. | 25 |
| Economics | “A campaign is over budget. What do you change?” | Engagement days first, then template category, then the unanswered cap. Voice bills only on connect, so calls are rarely the leak. | 29 · 31 |
| Detail | “How does a client's workspace get created?” | Not manually. SSO at activate.convin.ai, and the email domain becomes the organisation ID. One org, one client. | 17 |
| Latency budget | ~2.5 s | one full turn, end to end |
| Voice billing | ₹1 / pulse | a pulse is 15 s, only on connect |
| Utility template | ₹0.12 · ~100% | cost, delivery rate |
| Marketing template | ₹0.90 · ~50% | range 30–70% |
| Free-form window | 24 h | from the last customer message |
| Message caps | 100 · 50 · 5 | lead, day, unanswered |
| Call caps | 15 · 3–5 | per lead, per day |
| Number rotation | 6–7 | each daily-capped |
| Engagement window | 3 · 5 · 7 days | from observed behaviour |
| Business | 7–8 · 25+ | contracts, active pilots |
NBA the Next Best Action; the LLM decision on channel, timing and content. · Cooldown the LLM-set wait before the next NBA. · WABA WhatsApp Business Account, attached per campaign. · External ID the immutable lead key. · Pulse 15 seconds of connected call time.