Convin — internal training

Convin Sense
How it works

The voice pipeline, the vendor stack, the Next Best Action engine, and every lever in campaign settings. Session 1 explained why Sense exists; this is how it runs.

Session02
SpeakerSudesh
RoleProgram Manager, Founder's Office
Recorded06 Mar 2026
Duration65 min

Contents

Nine sections, thirty‑three pages.

  • 1How a VoiceBot worksPipeline, latency budget, the four factors, interruption03
  • 2The real stackVendor per layer, and the two cost tradeoffs behind it09
  • 3One call becomes a systemContext retention, persistence, omni-channel11
  • 4The business motionPilots, annual contracts, consumption pricing15
  • 5Inside the dashboardAccess, campaigns, system message, a real lead17
  • 6Meta and WhatsAppTemplate categories, delivery economics, the 24-hour window21
  • 7The Next Best Action engineThe loop, cooldown, lead state25
  • 8Campaign settingsThe control surface and the cost levers28
  • 9PM lens and interview prepTakeaways, levers, likely questions, reference card30
Session 3 continues with slot and queue mechanics, then Meta in depth.02
1.0 · The reframe

A VoiceBot is not the product.
It is one channel inside Sense.

Convin Sense — the orchestration layer
AI call

The voice pipeline in this section

WhatsApp

Template first, then free-form

SMS · Email

Where the segment responds

Human call

Live transfer, or full handoff

The call is a component. Sense is the system that decides when to place it, and that decision is the subject of sections 6 through 8.

Session 1 called this the intelligence layer. This session shows what it is choosing between.03
1.1 · How a VoiceBot works

Three layers, one round trip.

DialerExotel, or the client's own
STTspeech becomes text
LLMtext plus the agent prompt
TTSresponse becomes voice
Dialervoice reaches the customer
Traced concretely — the customer says “hello”

  1. 01
    The dialer receives the speech and hands it to the STT layer, which transcribes it.
  2. 02
    That text goes to the LLM together with the prompt configured for this agent.
  3. 03
    TTS wraps the generated response in a voice, one of several available.
  4. 04
    The dialer plays it back: “Hi, I'm Aaradhya from Convin AI.”
The prompt is configured per agent, so the same three layers behave differently across campaigns.04
1.2 · Latency

The whole round trip has about 2.5 seconds.

One turn, end to end
Dial
STT
LLM
TTS
transporttranscribe generatesynthesiseplay
0.0s2.5s

What the caller perceives
Human reply ~1.0s The benchmark. Anything near this reads as a person.
Sense today ~2.5s Acceptable. The pause exists but does not break the illusion.
Degraded 3–4s Naturalness collapses. The caller knows.
Latency is the largest single lever on whether a bot reads as human.05
1.3 · What makes it feel human

Latency is the biggest factor. It is not the only one.

Latency
Lower is more human. The 2.5 second budget on the previous page.
Transcription accuracy
How faithfully the STT layer captures what was actually said.
Response relevance
Whether the LLM answer addresses the thing the customer asked.
Voice quality
Robotic, casual or professional, and whether that suits the brand.

The four compound. A fast bot that mishears is not human, and a perfectly accurate one that takes four seconds is not either. Get one wrong and the other three cannot rescue the call.

Bar lengths are relative emphasis, not measured values

The session named these four as the factors that combine; it did not weight them. Latency is drawn longest because it is the one the speaker singled out.

Latency is measurable per call. Relevance needs sampling and human review.06
1.4 · The part nobody sees

A voice suite is not just STT, LLM and TTS.

Barge-in, handled correctly
Bot speaking
Bot stops, listens
Responds to what it heard
When to stop

A real objection, a question, a correction. Stop immediately.

When not to stop

A filler word, a cough, background speech, an agreeing “mm-hm”.

The hard part

Telling those two apart in under a second, without a transcript yet.

Why this matters commercially

Wiring three APIs together takes days. Interruption, turn-taking and the rest of the nitty-gritty took eighteen months. It is also the part a competitor cannot copy from an architecture diagram.

Fleetex was the first client, January 2025. The suite has been hardening since.07
1.5 · Eighteen months in

What building the suite actually taught them.

Jan 2025VoiceBot business starts. Fleetex is the first client.
~18 monthsVoice AI suite hardens: turn-taking, retries, context.
~8 months agoConvin Sense exists as its own product.

The lesson that produced Sense: transcribing speech, calling an LLM and playing audio back is not a voice product. It is the easiest quarter of one.

What was missing, and became Sense
Turn-taking

Interruptions, silences, barge-in

Retry logic

What happens after a call that went nowhere

Context

Carrying what was said into the next attempt

Orchestration

Choosing the channel, time and message at all

Each gap on this row became a component of Convin Sense.08
2.0 · The real stack

Who sits in each layer today.

DialerExotelOn annual contracts the client's own dialer is used instead: Ozonetel, Acefone, Genesys, Ubona.
STTDeepgramAn in-house STT exists, but Deepgram is the best in market, so it carries production traffic.
LLMOpenAI4.1-mini through 5.1. The tier is chosen per use case against prompt size and cost.
TTSCartesiaMigrated from ElevenLabs: comparable naturalness at materially lower cost.
Why the client's own dialer matters more than it looks

It removes a telecom migration from the sale, and the client keeps their existing numbers and compliance posture. The integration burden moves to Convin, which is where it is cheapest to carry.

Know this table cold. It is the most likely factual question in a Sense interview.09
2.1 · The two live tradeoffs

Every layer choice is a cost decision in disguise.

Tradeoff 01 — LLM tier buys prompt length
4.1-minilong Cheap per token, so the prompt can be much larger
5.1 / minidefault The working default across campaigns
Top tiershort Same budget buys far less instruction
Tradeoff 02 — TTS quality against price
ElevenLabshigher Genuinely human-like, and priced accordingly
Cartesiachosen Close enough on naturalness, materially cheaper

The deciding question in both cases was not which option is better. It was whether the quality delta justifies the price delta at Convin's call volume. In a consumption-priced product, that arithmetic belongs to the PM.

Both migrations were driven by unit economics, not by benchmark scores.10
3.0 · The limit of a single call

One call gets you a disposition, not an answer.

The Physicswallah case. A lead is uploaded, the bot calls to find out whether they want a JEE or NEET course, and must mark them hot or warm back into the client's system.

Lead uploadedfrom the client CRM
Bot dialsasks the qualifying question
“Call me at 5”the lead is driving
Call endsno answer captured
Pushed backunresolved, to the CRM
What the system actually has

A disposition. Not an answer to the question it was sent to ask. The lead is neither qualified nor disqualified, and nothing downstream knows the difference.

Why retrying alone does not fix it

A retry mechanism existed under VoiceBot. But a retry without context repeats the same cold opening, and the odds that any given lead is free at the exact moment you dial are, as Sudesh put it, very low.

Session 1 described this failure from the business side as the silent lead graveyard.11
3.1 · Contextual calling

The context survives the disconnect.

Call 1 — morning

“I'm Sudesh from Physicswallah, about your recent website visit. Are you exploring JEE or NEET courses?”

“Yes, but I'm driving. Can you call me back at 5pm?”

context
carried →
Call 2 — 5pm, generated automatically

“We spoke this morning. Can we speak now about whether the course is right for you?”

No pickup? Retry tomorrow, and the day after, until the engagement window closes.


“This context will get retained, and the next call will be automatically generated for 5pm with that specific context.”

The lead never repeats themselves, which is precisely what makes the second call read as a follow-up rather than a fresh cold call.

Context retention is the single feature that turns a retry into a relationship.12
3.2 · Goal-driven persistence

The agent is given a goal, not a call list.

Attemptcall or message
Observewhat came back
Carry contextinto the next try
Choose nextchannel and time
repeat until the goal is met,
or the window closes
The goal

Defined per campaign. For example: identify whether this lead is genuinely interested in a JEE or NEET course.


The stop condition

Persistence is bounded. The engagement window and the per-channel limits end it, not a human deciding to give up.


“Until and unless my goal gets achieved.” That is the condition. Not until the call is made.

Goal state replaces call count as the unit of completion.13
3.3 · Contextual omni-channel

Same context. Whichever channel fits.

AI call

the voice pipeline

SMS

where it performs

Shared context

Held at the lead level, not the call level

WhatsApp

template, then free-form

Email

longer-form follow-up

+ WhatsApp calls, being added

“Hey Animesh, we actually spoke about this. Are you still interested in the JEE course?” is a WhatsApp message that only works because the call before it is remembered.

Omni-channel without shared context is just spam on more surfaces. Use the full term.14
4.0 · Where Sense stands

Pilots prove it. Annual contracts bank it.

25+ active pilotsRunning in parallel at any time
~1,000 leads eachA typical pilot size
7–8 annual contractsConverted and live today
The pilot arithmetic
₹10,000

of usage cost

₹4,00,000

of revenue returned

≈ 100x

The pilot is not a discount. It is the evidence that makes the annual number defensible: the client commits for the year against an upfront lead volume, then uses Sense as they need it.

Roughly a 3:1 ratio of running pilots to signed contracts. That funnel is the growth engine, and the ~100x figure is a strong case rather than a median.

Have the pilot math ready, and concede up front that it is a strong case.15
4.1 · The pricing model

Consumption pricing means value has to keep compounding.

Monthly usage
expanding account silently churning
Why this shape matters

There is no dormant-seat revenue to hide behind. A client who stops finding value stops spending, and the decline shows up in usage months before it shows up in a contract.


The upside is symmetrical: a client who finds Sense working expands lead volume without renegotiating anything.


Which is the argument for usage dashboards, and the reason the campaign-level cost levers in section 8 are a product concern rather than an ops one.

Bar heights illustrate two trajectories. They are not measured account data.16
5.0 · Getting in

One URL, many organisations.

Post-call and VoiceBot
cordelia.convin.ai

A subdomain per tenant. Email and password sign-in, so throwaway accounts can be created for testing.

Sense
activate.convin.ai

One shared URL for everyone. SSO only, Google or Microsoft. No passwords, no dummy accounts. An organisation ID then selects the client workspace.


How a tenant is created — automatically, from the email domain
Client gets the URLactivate.convin.ai
Signs in with SSOsomeone@pw.com
Domain is read“pw” becomes the org ID
Org createdone org, one client

Live orgs: PW · TNSDC · Cordelia · Mosaic Wellness · Miles Education

Internal access is requested per organisation, not granted globally.17
5.1 · The unit of work

A campaign is an experiment you run on leads.

Not “triggering lots of calls at once”. That answer was given in the room and corrected. A campaign is a dataset, plus an agent, plus a time box, run to produce an outcome you can read.

Configure agentprompt, goal, knowledge base
Create campaignattach that agent
Upload leadson an ongoing basis
Run N daysthe engagement window
Read outcomesthree buckets, below
What you can read after three days
Showed interest
Explicitly not interested
Never engaged
passed to human agents a real answer, cheaply got no contact inside the window

Widths are illustrative, not measured outcome rates.

Note the vocabulary: you configure an agent, not a bot. Sense uses “agent” throughout.18
5.2 · The system message

The standing instruction every lead starts with.

When leads are uploaded, each receives a default system message: the campaign's opening context for the agent. It is set once, and applied to every lead.

Option A — you name the opening move

There is no prior interaction with this lead. They came from the website and are looking for a CPA or CMA course. Start engaging through an AI call.

Use when you already know how this segment responds.

Option B — the agent chooses

There is no prior interaction with this user. You figure out how to engage.

The agent picks voice or WhatsApp itself, using the Next Best Action logic in section 7.


This is where campaign intent meets per-lead context, which makes it the highest-leverage text field in the product.

Option B is the more interesting one: it hands channel selection to the model.19
5.3 · A real lead journey

One lead, walked end to end.

From a live campaign of roughly 500 leads with two channels active. The lead was exploring CPA and CMA courses.

  • System
    No prior interaction. Start engaging through an AI call.
  • AI call
    The lead talks, mentions experience in accounts payable, then asks to be emailed instead.
  • AI call
    One nudge: “no problem, I'll keep it short.” Then a graceful close.
  • WhatsApp
    Templated message referencing the accounts-payable experience it just heard.
  • WhatsApp
    The lead replies. That reply unlocks free-form LLM messages, and a real conversation follows.
Every step is visible per lead in the dashboard: transcript, messages, and the reasoning behind each action.20
6.0 · Meta and WhatsApp

The first message is never free-form.

Any first contact on WhatsApp must be a Meta-approved template in one of three categories. The category decides both what it costs and how many people receive it.

Delivery rate
Utility~100% Transactional. Reaches essentially everyone.
Marketing~50% Promotional. Half the list never sees it.
Cost per message
Utility₹0.12 Roughly a seventh of the marketing rate.
Marketing₹0.90 Costs 7x more, reaches half as many.

Authentication — OTPs and verification codes. Approved as a third category, not used by Sense.

The asymmetry is the whole point

A marketing template costs seven times a utility one and reaches half as many people, so a message sent in the wrong category is roughly fourteen times less efficient. Choosing the category is an economic decision, not a compliance formality.

Figures approximate. The session gave them as rough per-message rates.21
6.1 · Reachability

Marketing delivery is not a fixed number.

It swings with the audience, and it moves in the opposite direction to what most people assume.

Marketing template delivery rate
50%
30%
70%
0%100%

Tier-1, highly educated
lower

More likely to have blocked promotional senders, or to run message filtering. Delivery drops.

Tier-3 base
~70%

Fewer filters in place, promotional messages land. Delivery can reach around 70%.

Reachability, not creative, is the binding constraint on a marketing template.

Counter-intuitive and therefore memorable: the wealthier audience is the harder one to reach.22
6.2 · The free-form window

One reply, even a single dot, opens 24 hours.

Templated sends
Free-form window — LLM writes like a person
Template again
  24 hours from the customer's last message  

No reply?

Send the template again. Currently the same one, though varying it is being explored.

Under the hood

Meta's API in both directions: a GET to retrieve what the customer sent, a POST to send the reply.

Why it matters

Every guardrail in campaign settings exists to buy that first reply economically.

Which reframes the funnel: on WhatsApp the first reply, not the first send, is the real conversion step. Everything before it is paid and unread; everything after it is cheap and human.

The window resets from the customer's last message, not from your first send.23
6.3 · A deliberate design choice

The messages are not perfectly written. On purpose.

Rejected — rule-based

if customer_says(X):
    send(Y)

Predictable and controllable, and immediately recognisable as automation. It also cannot handle anything the author did not anticipate.

Chosen — LLM generated, every message

“Got it. You are a student with accounts experience. Are you currently in college as well?”

Casual vocabulary and loose grammar, kept intentionally. A flawless paragraph reads like a machine; this reads like someone typing.


The tell

Occasionally an artifact slips through: a stray double dash gives away that a line was generated. In the campaign reviewed live, that happened three times across an entire conversation.

Naturalness here is a product decision, not a model limitation. That distinction is worth making explicitly if the topic comes up.

Three artifacts in a whole conversation is a usable quality bar to quote.24
7.0 · The engine

Next Best Action decides channel, timing and content.

Before every AI call or WhatsApp message, the system generates a Next Best Action. There is nothing rule-based in this decision.

Inputs
Prior interactions

everything said so far

System message

campaign intent

Knowledge base

the agent's prompt

NBA

an LLM call, per lead

channel · time · content

Output
The action

an AI call, or a WhatsApp message

Channel reason

why this channel, recorded

Action reason

why now, recorded


It shows its work

“AI call chosen for direct engagement; initial contact, no channel fatigue.”  ·  “The data shows prior interest; an AI call enables personalised engagement and early trust building.”

The reasoning is a first-class recorded field, not a debug log. That is a deliberate product choice.25
7.1 · Cooldown

How the system decides when to act next.

Action firescall or message
Cooldownno activity at all
NBA regenerateswith what just happened
Next action chosenchannel and time
no rule sets the interval;
an LLM sets it per lead
Sudesh's own analogy
A boss with two nodes.

The boss is the NBA. The two nodes are the AI call and WhatsApp. The boss decides the mode, the channel, the content and the time, then hands the instruction to whichever node carries it out. The nodes only execute. All judgment sits with the boss.

“The cooldown is determined automatically by an AI, through an LLM call. There is nothing rule-based over here.”


The practical consequence: two leads in the same campaign can be contacted on different channels, at different intervals, with different messages, without anyone configuring that.

Per-lead adaptation is the mechanism behind the efficiency gain session 1 measured.26
7.2 · Lead state and ownership

What the system tracks per lead.

Status
Activecommunication in progress
PausedNBA generation stopped
Completedgoal met, no further contact
Archivedmanually deactivated
Goal achievementconfigurable, not automaticA “not interested” answer counts as goal-achieved only if you explicitly configure it that way. Otherwise the lead stays open.
Interest levelhot · warm · coldSeparate from goal achievement. You can achieve the goal and learn the lead is cold.
External IDimmutable unique keyIdentifies the lead in Convin's system. Archive and re-add the same person and you must use a different external ID.
Lead ownerConvin AI, or a humanOwnership can move to a human, who then works the lead from the same interface with full history.

Live call transfer hands a conversation to a human mid-call via a tool call, and the call never drops. Lead handoff moves ownership of the whole lead. Different things.

Goal achievement being configurable is the subtlety most people miss, and the one most likely to be probed.27
8.0 · Campaign settings

Every guardrail in the product lives in one screen.

Identity

Campaign name and description · agent, with its prompt, attached

Language

Initial language · real-time switching if the customer changes

Context

Default system message applied to every lead on upload

Schedule

Time zone, start hour, end hour · DND hours the campaign must respect

Duration

Engagement days: 3, 5 or 7 · read from user behaviour, not guessed

Channel limits

100 messages per lead, 50 per day, stop after 5 unanswered · 15 calls per lead, 3 to 5 per day

Post-completion inbound behaviour and concurrent call allocation, the slot and queue mechanism, were named as still being built. Both are session 3 material.

This is the PM's actual control surface. Name a field and the metric it moves.28
8.1 · Cost control

Two billing models, and the levers that keep them honest.

WhatsApp — charged on send

Cost is incurred whether or not anyone reads it. An unread marketing template is pure loss, which is why the unanswered-message cap matters so much.

Voice — charged on connect

An unanswered call costs nothing. Billing starts when the customer picks up, at ₹1 per pulse, where a pulse is 15 seconds of connected conversation.


Lever — engagement days, read from behaviour
D1D2D3D4D5D6D7

Pickups concentrate in the first three days. A lead who has not answered by day three is unlikely to answer on day seven, so days four onward add cost without adding revenue.

Lever — number rotation

One number against 10,000 leads gets marked as spam, and then nobody answers at all. Six or seven numbers are rotated per campaign, each with its own daily call cap.

6–7 numbers · per-number daily cap

A cost lever disguised as a deliverability feature: a spam-flagged number wastes the whole campaign.

Bar heights show where pickups concentrate. They illustrate the pattern, not measured data.29
9.0 · PM lens

Five things the mechanics tell you about the product.

  • 01
    The moat is orchestration, not voiceSTT, LLM and TTS are all bought in. What cannot be bought is the NBA loop and eighteen months of turn-taking work.
  • 02
    Explainability was built in earlyEvery action records a channel reason and an action reason. That is a deliberate choice, and it is what makes the system tunable rather than merely autonomous.
  • 03
    Constraints are the productMessage caps, engagement days, number rotation, DND hours. The guardrails are what make autonomy safe enough to sell to an enterprise.
  • 04
    Cost per outcome is a PM metricVendor tier, template category and window length all move margin. Under consumption pricing that number belongs to the PM, not to finance.
  • 05
    Configuration is where value leaksThe same product tuned badly burns budget on days four to seven. Onboarding quality is therefore a retention lever.
PM framing. Analysis layered on the session, not claims made in it.30
9.1 · Levers and the metrics they move

What to reach for when a campaign underperforms.

LeverMetric it movesHow to think about it
Engagement daysCost per leadShorten when a segment converts early. Lengthen only on evidence from later-day pickups.
Max calls per dayConnect rateMore attempts lift reach up to a point, then start burning goodwill and numbers.
Number rotationConnect rateProtects against spam labelling, the failure mode that silently kills a campaign.
Template categoryDelivery and costUtility reaches ~100% at ~₹0.12; marketing ~50% at ~₹0.90. Use utility wherever the context genuinely allows.
Unanswered capWasted spendEvery send before the first reply is paid for and unread. Five is the current default.
Initial languageEngagement depthSwitching mid-call is automatic, but the opening still sets the tone.
System messageQualification accuracyThe highest-leverage text in the product. It frames every downstream NBA.
Levers from the session; the mapping to metrics is PM analysis.31
9.2 · Interview prep

Questions an interviewer will actually ask.

Architecture“Walk me through what happens when a customer says hello.”Dialer, STT, LLM with the agent prompt, TTS, dialer, inside roughly 2.5 seconds. Name the vendor at each layer.04 · 09
Tradeoffs“Why Cartesia and not ElevenLabs?”ElevenLabs is more natural but materially more expensive. At Convin's volume the quality delta did not justify the price delta.10
Depth“What is actually hard about building a voice agent?”Not the three layers. Turn-taking and interruption: meaningful versus non-meaningful interrupts took eighteen months.07
Product“What is defensible? Anyone can wire up these APIs.”The NBA loop. An LLM choosing channel, timing and content per lead, with its reasoning recorded for audit.25
Economics“A campaign is over budget. What do you change?”Engagement days first, then template category, then the unanswered cap. Voice bills only on connect, so calls are rarely the leak.29 · 31
Detail“How does a client's workspace get created?”Not manually. SSO at activate.convin.ai, and the email domain becomes the organisation ID. One org, one client.17
Session 2 answers are factual. Precision on vendors, numbers and terms is what separates a good answer.32
9.3 · Reference card

The numbers, in one place.

Latency budget~2.5 sone full turn, end to end
Voice billing₹1 / pulsea pulse is 15 s, only on connect
Utility template₹0.12 · ~100%cost, delivery rate
Marketing template₹0.90 · ~50%range 30–70%
Free-form window24 hfrom the last customer message
Message caps100 · 50 · 5lead, day, unanswered
Call caps15 · 3–5per lead, per day
Number rotation6–7each daily-capped
Engagement window3 · 5 · 7 daysfrom observed behaviour
Business7–8 · 25+contracts, active pilots

Vocabulary

NBA the Next Best Action; the LLM decision on channel, timing and content.  ·  Cooldown the LLM-set wait before the next NBA.  ·  WABA WhatsApp Business Account, attached per campaign.  ·  External ID the immutable lead key.  ·  Pulse 15 seconds of connected call time.

All figures from Session 2, 06 Mar 2026. Approximations preserved as given.33