How prompting evolved from decision trees to the Eight Pillar Framework, how an agent is assembled from a knowledge base, and what every knob in the call-control panel actually does.
An ed-tech client wants to reach people who browsed their CPA and CMA pages and left without enquiring. The message: “you were seen on our website, can we speak now to know more?”
Utility. The customer initiated contact by visiting the website, so the conversation is transactional rather than promotional.
Consent, not who moved first. If there is a written terms and conditions the customer accepted, obliging them to receive WhatsApp or AI calls, utility is defensible. Browsing a page is not consent.
Terms the customer accepted that oblige them to receive WhatsApp or AI calls. Utility is defensible, and you have something to show when Meta asks.
Browsing is not agreement. Send marketing, accept the higher rate and the lower delivery, and keep the account safe.
This sharpens the Session 3 figure that roughly 99% of use-case templates land in marketing. The 1% that legitimately qualify are the ones with consent behind them, which makes consent capture a commercial lever rather than a legal formality.
Leads, engagement days, channel limits, DND hours. The thing you actually run.
Named, given a goal, a personality, channels and tools. Created once, reused across campaigns.
The knowledge base is a brain. It is placed on an agent's head, and that agent is attached to a campaign. Build the brain first.
Meta integration and template configuration are secondary. If you are handed a platform and told to run a pilot, agent configuration is the first real work.
Qualify first, give surface-level information, just enough to get them interested, then pass a qualified prospect to a human who can actually close.
Correct, and it is a property of the category rather than a limitation of Convin. No voice bot currently closes sales. Nobody trusts an AI enough to pay on the same thread.
Which is why every pilot follows the same shape: the bot qualifies, the qualified leads go to human agents, and the humans take the sale from there.
Around January 2025 the bot was built on nodes: a greeting node, then a yes branch and a no branch, and so on down the tree.
Real people do not follow the branch you authored. Every unanticipated question is a dead end, and returning to the flow afterwards is clumsy and obvious.
A lot of voice-bot companies still ship this architecture. It is the clearest technical difference to point at when asked how Convin differs.
Instead of a tree of nodes, the whole conversation design becomes a single prompt: the flows, the branches, and everything a customer might reasonably ask. That prompt goes to the LLM to generate every single response.
Full context on every turn. When a customer asks something unexpected, the answer is already inside the prompt and can be referenced without leaving the flow.
Tokens, on every single turn. The entire prompt is re-sent for each response, and the conversation so far is sent with it.
The speaker's definition, worth using verbatim: hallucination here is when the prompt contains the answer and the model still fails to give it properly. The cause is not missing information. It is too much of it.
Say the thing once, in the fewest words that carry it.
Every character is re-sent on every turn. Length is a recurring cost, not a one-off.
If something is stated in one pillar, it does not need restating in another. Delete the duplicate.
Longer prompts cover more edge cases, and longer prompts degrade both cost and answer quality. There is no setting that resolves this. It is resolved by editing, which is why prompt authoring is a skill the team has to build rather than a task to hand off.
Both answers to “what is the interest rate on this loan?” are correct and contain identical information.
“8% with a processing fee of this, and you get an additional discount of this.”
“That's a great question. You can get 8% and this, with an additional discount of this.”
The difference is a single acknowledging clause. It carries no information and it is the whole reason the second reads as a person rather than a lookup. Fillers of exactly this kind are what the framework mines out of real call recordings.
Naturalness is not a model capability you buy. It is an authoring output, extracted from how the client's own best agents actually speak.
The acknowledging phrases a good agent uses before answering.
The order a real call actually moves in.
What customers push back with, and how the agent responds.
How this specific business talks to this specific audience.
A client's written script says what they intend to say. The recordings show what their best agents actually say, including the objection handling nobody wrote down. The framework captures the second.
Live example shown in the session: a full eight-pillar prompt generated for Aditya Birla Sun Life Insurance.
Take the client's script and their sample human calls, generate a knowledge base, and use it to train the model.
The process is right and the word is wrong. Nothing is trained. A prompt is run against an already-trained model. Training means feeding a dataset repeatedly so the model stops repeating mistakes; this is prompting.
Worth being precise about in an interview. Saying “we train a model per client” overstates what the product does and invites a question you cannot answer.
| Agent name | e.g. Tara |
| Company name | the client |
| Agent gender | male or female |
| Business outcome | the one-line goal |
“Your goal is to qualify those who are interested in booking a medical visit.”
One sentence. It is an instruction to the agent, not a specification.
Overload the goal with detail and it stops being satisfiable. The agent never registers it as achieved, so it keeps calling and messaging the same lead. An over-specified goal does not produce a thorough agent; it produces a spammy one, and spammy behaviour is what gets a WABA reported.
Which connects this field directly to the account-health material in Session 3. A badly written goal is a deliverability risk.
Yes. Personality settings sit next to the voice settings, so it reads as though they shape delivery.
No. These travel with the system prompt and steer the LLM's word choice. A professional setting produces more formal vocabulary and a consultative phrasing. The voice itself is unchanged.
Consultative · Informative · Persuasive · Supportive · Analytical
Persuasive ends most responses with a question. Informative attaches information to each response.
Friendly · Casual · Formal · Enthusiastic
Chosen against the audience rather than the brand's preference about itself.
| Channel | What it is |
|---|---|
| AI call | The outbound voice bot. The default channel for most pilots. |
| Through Meta, as covered in Session 3. | |
| SMS | Used for a US customer, Ducati, over Twilio. |
| WhatsApp AI call | Inbound only. Meta does not permit outbound calls at all. |
| Human call | Transfer to a human agent working inside the same system. |
| Website widget | Runs on a session ID, with no phone number and no name. |
Human in the loop was described as still being built, with a release expected within about a week of the session.
That WhatsApp calling is another outbound channel: upload leads, and the agent rings them through WhatsApp as well as the phone network.
Inbound only. A call icon appears on the business chat page and the customer initiates. Meta forbids outbound calling outright, on the same reasoning as its other disciplinary rules: it will not have customers bombarded with calls.
A website widget conversation is identified by a session ID generated in the browser. There is no phone number, no name, no CRM record. Because the only handle on the visitor is that session ID, widget conversations cannot be handed across to an AI call or to WhatsApp the way a normal lead can.
Convin runs a hybrid widget on its own site that works both as chat and as an AI call. Worth trying before an interview; it is the one part of the product you can experience without access.
The same first line on every call. “Hi, I am Tara from Convin. Are you available for a quick chat?” Predictable, and editable.
Personalised, with context carried from previous calls. The trap: a dynamic opener cannot be edited in place. Switch to static, edit, switch back.
| Setting | Default | What it protects |
|---|---|---|
| Hangup delay | 2 s | After the closing line, how long to wait before disconnecting. If the customer speaks inside the window, the agent answers instead of hanging up. |
| Max call duration | 300 s | A hard cap. The call ends at five minutes even if it is going well. It exists because the bot sometimes goes silent, and a dead call still costs money. |
| Idle timeout | 7 s | How long before the agent asks “are you still there?”. Raise it to 20–30 s for app-onboarding calls where the customer is doing something. |
| Max idle messages | 3 | How many times it may ask that before delivering a closing line and disconnecting. |
All four are use-case dependent. A two-minute qualification call should not carry a five-minute cap, and an onboarding call should not be nudging the customer every seven seconds.
Shown as “Convin managed” and described as a proprietary model. In practice OpenAI, Gemini, Claude, or a closed-source model, chosen per use case.
Not editable from the main dashboard. Changing it requires internal dashboard access.
Currently dynamic, switching between Hindi and English mid-call without configuration.
Verify: Session 2 described the LLM layer as “a lot of open source models”. Session 4 describes it as OpenAI, Gemini, Claude or closed-source. Confirm which is current before quoting either.
Neutral, happy or sad. Choosing happy or sad appends emotion blocks to the conversation, and the bot then literally laughs mid-call. Not a written “haha”, an actual laugh. An agent laughing at a customer is never what you want, so neutral is the standing setting.
Voice ID is worth knowing too: every voice carries an ID, and a voice not listed in the picker can still be used by pasting its ID directly.
Asked in the room: for a collections use case, the script was rewritten to be direct, but the bot keeps delivering it politely. What actually controls this?
Changes the words. The delivery stays polite.
Makes it louder. Loud is not firm.
May help a little. Worth experimenting with.
This is the actual control.
The Next Best Action generation prompt instructs the agent to be professional and to avoid harsh language. That instruction sits above the knowledge base, so no amount of editing the client-specific prompt overrides it. It lives in the internal dashboard under prompt management, per tenant.
Do not edit system prompts directly. They affect a great deal downstream; route changes through the platform owner. An idea raised in the session: maintain use-case-specific system prompts, so collections is not fighting a sales default.
Call-centre chatter, coffee-shop noise or office ambience can be played under the call, on a gain scale of 0 to 10, so the line does not sound suspiciously clean.
Exotel and its peers were built over a decade for human agents: one channel out, one channel back. Nothing in that design anticipated injected audio.
A newer dialer built specifically for voice AI was named in the session as being trialled. If the channel model is right, background audio becomes usable again.
Until then the feature stays configured and unused, because the realism it buys costs more in latency than it returns.
Per million characters of text converted to speech. Which means prompt and response length feed straight into voice cost, not only into LLM cost.
Raising the probability that a specific word is transcribed correctly, for example “eleven” being heard as “seven”. It cannot be done from the agent screen; typing words there has no effect. It is configured in Deepgram by the product team.
The word-boosting detail is a small but real product gap: a field that appears editable, does nothing, and requires a ticket to another team to actually change.
WhatsApp is a text channel, so speech settings look like a stray field left over from the voice configuration.
Voice notes. If a customer replies with one, it is transcribed with Deepgram, and the agent replies with a voice note of its own generated through Cartesia. Same vendors, same configuration.
A file, a plain-language description of when to use it, and the format requirements per type: audio, document, image, video.
The WhatsApp system prompt box on this screen is controlled entirely from the internal dashboard. Anything typed into it is not applied.
Name and description · the agent, attached
Initial language for the first turn
Default system message, applied per lead
Time zone · DND start and end
Communication limits · how many days engagement runs
Name, description, open or closed · eligible values, as in post-call
Single or bulk. Phone number, then an external ID which is mandatory and must be unique. Language can be skipped when set at campaign level. The initial system message can be overridden for one lead, and custom fields referenced in the prompt are supplied here.
Asked in the room, and answered in pieces. Three situations end communication with a lead while goal achievement remains no.
“Do not contact me” is not the same as “I am not interested”. Interest level is marked not interested and the bot will not re-attempt. But goal achievement stays no, because the agent never learned whether they wanted the product. Unless the goal explicitly says to mark a do-not-contact as complete, the system treats the question as still open.
| Setting | Range | Start at | Notes |
|---|---|---|---|
| Temperature | 0 – 1 | 0.7 | Lower to follow the prompt, higher for freer phrasing. 0 unused. |
| Speed | 0 – 2 | 1.2–1.3 | Slower for tier-3 audiences, faster for tier-1 and tier-2. |
| Volume | 0 – 2 | ~1.0 | 0.5 is very quiet, 2.0 effectively shouts. |
| Emotion | 3 options | Neutral | Happy or sad make the bot audibly laugh. Leave it alone. |
| Hangup delay | seconds | 2 s | Grace period after the closing line. |
| Max call duration | seconds | 300 s | Hard cap. Lower it for short use cases. |
| Idle timeout | seconds | 7 s | Raise to 20–30 s for onboarding calls. |
| Max idle messages | count | 3 | Nudges before the agent gives up. |
| Background gain | 0 – 10 | off | Currently not recommended; it raises latency. |
| Competitive | “How is your bot different from everyone else's?” | Most competitors still ship node-based decision trees. Convin sends one whole prompt every turn, so an off-script question is answerable without leaving the flow. | 07 · 08 |
| Method | “How do you write the prompt?” | The Eight Pillar Framework, built from about twenty transcripts of the client's own best calls. Fillers, flow, objections and tone are mined out, not invented. | 11 · 12 |
| Precision | “Do you train a model per client?” | No. Nothing is trained. A prompt runs against an already-trained model. Getting this wrong is the fastest way to lose credibility. | 13 |
| Economics | “Where does prompt length show up?” | Every turn re-sends the whole prompt plus the conversation so far, and TTS is billed per million characters. Length hits LLM cost, voice cost and answer quality. | 08 · 09 |
| Depth | “A collections bot is too polite. What do you change?” | Not the script and not the volume. The NBA system prompt in the internal dashboard overrides the knowledge base. Route it through the platform owner. | 22 |
| Detail | “When does contact stop but the goal stay open?” | Channel limits, an invalid number, or the customer saying do not contact me, which is distinct from not interested. | 27 |
| 1 | Agent persona |
| 2 | Overall goal |
| 3 | State machine flow |
| 4 | Tone, style and language |
| 5 | Behavioural rules and guardrails |
| 6 | Task logic |
| 7 | Objection handling |
| 8 | Output formatting |
| Transcripts needed | ~20 |
| Large prompt size | ~85k chars |
| Turns in a 10-min call | 50+ |
| Temperature default | 0.7 |
| Max call duration | 300 s |
| Idle timeout | 7 s |
| Voicemail accuracy | 80–85% |
| Background gain scale | 0 – 10 |
Deferred to a further session: the pilot SOP itself, and agent tools. Homework set: build a knowledge base, an agent and a campaign for an assigned use case in the app organisation.
| Pillar | What goes in it | Example | |
|---|---|---|---|
| 01 | Agent persona | Who the agent is, and what it is there to do. | “You are Arjun, a counsellor at Physicswallah.” |
| 02 | Overall goal | The single outcome the conversation heads toward. | “Qualify those interested in a JEE course.” |
| 03 | State machine flow | The ordered steps from opening to goal. | Greet, confirm identity, establish need, offer a callback. |
| 04 | Tone, style and language | From speech analytics on the real calls. | Consultative, warm, short sentences, no jargon. |
| 05 | Behavioural rules and guardrails | What to say, what never to say, what to hand off. | “Do not quote fees; the counsellor handles that.” |
| 06 | Task logic | For each customer response, what happens next. | If the parent answers, re-confirm before continuing. |
| 07 | Objection handling | Real objections, with the answers that worked. | “It is too expensive” and the reply a good agent gave. |
| 08 | Output formatting | How the model must shape what it returns. | Structured fields the platform parses, not prose. |
Carried forward from the earlier decks and added to here. Confirm internally rather than picking a version.
Session 2: “a lot of open source LLMs and open source models.”
Session 4: Convin-managed, described as OpenAI, Gemini, Claude or a closed-source model.
These are not the same claim. Ask which is running in production today.
Named in Session 4 as purpose-built for voice AI and currently being trialled.
The name was not clearly audible on the recording.
Worth confirming before mentioning it, since it signals a live infrastructure decision.
Session 3: roughly 99% of use-case templates land in marketing.
Session 4: consent on file is what makes utility defensible.
Compatible, and the second explains the first. Consent capture is the lever that moves the rate.