QVCCS innovation · Diagnostics

What callers actually said: bot journeys, turn-by-turn replay and the no-match backlog

A bot flow is only as good as its understanding of what real callers say. Genesys Cloud records every bot session and every reporting turn. We built Bot Flow Diagnostics, known in the app as Bot Sankey, to harvest them into a lasting warehouse and show three things: where conversations go, what each caller said turn by turn, and which phrases the bot failed to understand.

QVCCS Innovation teamBot Flow Diagnostics user guide →

What callers actually said to the bot A bot flow tile on the left splits into flow bands of different widths, like a Sankey journey map, ending at three outcomes: transfer to ACD, self-served and recognition failure. Above the bands, a caller's speech bubble shows the literal speech-to-text the bot did not understand, and a No match chip marks it. QVCCS INNOVATION · BOT ANALYTICS What callers actually said BOT FLOW Ask · intent Transfer to ACD Self-served Recognition failure “pna please” No match
  • Did you know the Utterance History view and Bot Conversation Library in the Genesys Optimization dashboard retain only the last 10 days of data, whatever date range the dashboard is set to?

  • Did you know the Genesys Platform API documents bot flow sessions and reporting turns as deleted after approximately, but not before, 10 days, so turn-level history must be collected if you want to keep it?

  • Did you know the Optimization dashboard shows metrics for core bot flows only, not for subflows that the core bot flow invokes?

  • Did you know the Bot Performance Detail view in Analytics Workspace accepts custom date ranges of up to 6 weeks for its aggregate bot metrics?

01

09:10: the bot flow looks healthy, the queue says otherwise

A parts and ordering line runs a Genesys Dialog Engine voice bot. It asks what the caller needs, captures a part number and either answers or transfers to the right queue. On Monday at 09:10 the queue supervisor notices more callers arriving at agents without a captured part number, and asks the bot owner a fair question: what are callers actually saying to the bot, and where is it losing them? Is the problem the intent model, the part-number capture, the reprompt wording, or callers who simply ask for a person?

Answering that needs three views at once. The aggregate journey, so you can see which ask is leaking. One caller's session, turn by turn, so you can hear the failure in context. And the full list of things callers said that the bot did not understand, ranked by how often they said them, because that list is the tuning backlog. Bot Flow Diagnostics puts all three on one harvested population.

02

What Genesys Cloud gives you natively, and does well

Genesys provides a strong set of bot insights. The Insights and Optimizations menu in Dialog Engine bot flows and digital bot flows brings together the Optimization dashboard, Intent Miner, Journey Flows, bot performance reports and the Virtual Agent dashboard. The Optimization dashboard shows total interactions, average duration, average turns and end states such as contained, transferred, agent escalation, recognition failure and abandoned, with trend lines against the previous period, a Top Bot Flow Actions with Issues list, an intents breakdown, the Utterance History view and a Bot Conversation Library.

Intent Miner mines intents, utterances, topics and phrases from chat and call recordings, which is the right tool for designing new intents. The Bot Performance Detail view in Analytics Workspace reports entries, durations, turns and intents over custom ranges of up to six weeks. Our app does not replace any of this. It adds what an analyst tuning one bot keeps asking for: a full-population journey diagram, an exact replay of any session, and a ranked corpus of failed utterances that outlives the platform's turn-level retention.

03

The insight: every session is a path, every turn a conversation

Genesys collects Bot Performance data automatically for every Architect bot flow, with no per-flow opt-in, for voice and digital bots alike. Each bot session ends with a bot result, such as TransferToACD or DisconnectRecognitionFailure, grouped into a result category. Inside it, each reporting turn carries the bot's prompts, the caller's utterance, the ask action that ran, the matched intent with its confidence and slot values, and the ask result: SuccessCollection, NoMatchCollection, NoInputCollection, AgentRequestedByUser and so on. On voice, the utterance is the speech-to-text output, exactly the text the language model judged.

That makes every session a path and every turn a small conversation. Project each session into a bounded path of prompt, ask and result steps ending at a terminal, and the population becomes a directed graph. Keep the turns, and any path can be replayed. Collect the no-match utterances, and you have the backlog. Third-party bots do not report through these APIs, so they are out of scope; for Architect bots the coverage is complete.

Bot Performance data converges on one bot-session warehouse On the left, the Genesys Cloud bot data model: an Architect bot or digital bot flow has sessions, each ending with a bot result and result category, and each session holds reporting turns carrying the bot's prompts, the caller's utterance, the ask action, the matched intent with confidence, slot values and the ask result. In the middle, the read-only Genesys Public API calls Bot Sankey makes: the flow list, two cursor walks over bot flow sessions and reporting turns, a best-effort bot aggregates cross-check and the flow's latest configuration. A note records that Genesys documents session and turn resources as kept for approximately, but not before, 10 days. On the right, a gap-aware warehouse feeds three views: the journey map, session replay and the utterance corpus. GENESYS DATA MODEL Flow › session › turn BOT · DIGITALBOT FLOW Genesys Dialog Engine, voice or digital SESSION Bot result · result category REPORTING TURN Bot prompts Caller utterance (STT on voice) Ask action Intent + confidence · slots Ask result, e.g. NoMatchCollection Conversation id Collected for every Architect bot flow Third-party bots do not report here GENESYS PUBLIC API · READ-ONLY Two cursor walks GET /api/v2/flows?type=bot|digitalbot GET /api/v2/analytics/botflows/{id}/sessions 250 PER PAGE · AFTER CURSOR GET …/botflows/{id}/reportingturns GROUPED BY SESSION ID POST /api/v2/analytics/bots/aggregates/query GET /api/v2/flows/{id}/latestconfiguration Sessions and turns: kept approximately, but not before, 10 days (API definition) DASHED: BEST-EFFORT CROSS-CHECK ONE VIEW Bot-session warehouse BOT SANKEY Gap-aware coverage per flow Each session stored once JOURNEY MAPPrompts · asks · results · endings SESSION REPLAYTurn by turn, with confidence bars UTTERANCE CORPUSRanked no-matches · per-ask rates Harvest regularly and the corpus outlives the platform's retention Every session and every reporting turn in the window – full population, not a sample – kept after Genesys ages it out
Bot flows, sessions and reporting turns are read through two cursor walks and kept in one warehouse behind three views.
Read this diagram as text

A three-column diagram, read left to right, showing how Genesys bot data is read through the public API and kept in one bot-session warehouse, with a summary banner beneath.

  1. Left column, "Genesys data model", headed "Flow › session › turn": a nested hierarchy starting with a "Bot · digitalbot flow" box labelled "Genesys Dialog Engine, voice or digital".
  2. Inside the flow sits a "Session" box holding "Bot result · result category", and inside the session a "Reporting turn" box listing "Bot prompts", "Caller utterance (STT on voice)", "Ask action", "Intent + confidence · slots", "Ask result, e.g. NoMatchCollection" and "Conversation id".
  3. Notes under the left column read "Collected for every Architect bot flow" and "Third-party bots do not report here".
  4. An arrow runs from the reporting turn into the middle column, "Genesys public API · read-only", headed "Two cursor walks", which lists five read calls from top to bottom: the bot and digital bot flow list; the bot flow sessions walk, marked "250 per page · after cursor"; the reporting turns walk, marked "Grouped by session ID"; the bot aggregates query; and the flow's latest configuration.
  5. The bot aggregates query is drawn dashed, and a key explains "Dashed: best-effort cross-check".
  6. A highlighted note in the middle column reads "Sessions and turns: kept approximately, but not before, 10 days (API definition)".
  7. An arrow runs from the two cursor walks to the right column, "One view", headed "Bot-session warehouse": a "Bot Sankey" box says "Gap-aware coverage per flow" and "Each session stored once".
  8. Below it, three views are listed: "Journey map" ("Prompts · asks · results · endings"), "Session replay" ("Turn by turn, with confidence bars") and "Utterance corpus" ("Ranked no-matches · per-ask rates"), with the note "Harvest regularly and the corpus outlives the platform's retention".
  9. The banner beneath reads: "Every session and every reporting turn in the window – full population, not a sample – kept after Genesys ages it out".

04

How we engineered the harvest

A run is one bot flow and one UTC window of up to 42 days. The app compares the window with the flow's stored coverage, polls only the gaps plus a six-hour overlap at the tail so late-settling analytics refresh, and walks two cursors against Genesys at 250 items a page: /api/v2/analytics/botflows/{id}/sessions and the matching division-aware reporting turns endpoint. Turns are grouped by session id, each session is projected into its journey path and stored once, keyed by flow and session, so a re-harvest reuses everything already downloaded. A typical window finishes in seconds to a couple of minutes, because there are no per-conversation jobs.

Safety bounds keep it honest. Each uncovered gap's session walk has a cap, 1,200 by default and adjustable up to 20,000, and the turns walk has a derived bound. If either is hit, the run says so and deliberately does not record coverage for that sweep, so nothing is silently missed. A best-effort POST to /api/v2/analytics/bots/aggregates/query corroborates the tiles but is never a dependency. The Genesys client refuses every non-GET request except a short allowlist of read-only analytics queries, runs behind a per-organisation rate limiter and honours Retry-After on 429 responses. The OAuth client needs analytics:botFlowSession:view, analytics:botFlowDivisionAwareReportingTurn:view and architect:flow:view.

05

The journey map: what the bot said and where callers went

The journey map is a zoomable directed graph rendered on the server from the run's stored paths. Ask actions are decision nodes, intent matches, no-matches and no-inputs are result nodes, milestones and nested bot calls get their own nodes, and end states such as Transfer to ACD, Caller hung up and Recognition failure are terminals. The bot's spoken prompts are a layer that is on by default, because for a bot its speech is the product. Reprompt loops are drawn rather than dropped. Node labels show reach as a count and percentage of all sessions, edge width encodes volume, and hovering an edge shows its share of the source node's outflow.

Node labels come only from bounded sets, such as ask names, intent names, result taxonomies and prompt texts, never from raw utterances, so the diagram cannot explode in size. A Summary panel adds client-ready material from the full population: a clean partition of sessions into transferred, hung up in bot, transferred elsewhere and self-served, friction hotspots, top successful outcomes, the last prompt heard before callers abandoned, the busiest journeys and a fidelity score for how faithfully one aggregate diagram reproduces real step order. The map and the summary both export as single-page branded PDFs.

06

Session replay and the utterance corpus

Session replay lists the run's sessions with their end category, final matched intent, channel and step count, and replays the selected one as a conversation: bot bubbles for each prompt, a caller bubble with the utterance or no input, a colour-coded ask result, an intent confidence bar, green at 80% and above, amber from 55%, red below, and slot chips with their own confidence. The header carries the Genesys conversation id so any session can be cross-checked elsewhere, and a map node filter narrows the list to sessions that passed through that node.

The Utterances page aggregates every stored turn. Tiles show sessions, turns, utterances captured, match rate, no match and no input. The no-match table merges phrases case-insensitively, ranks them by frequency and links up to three sessions to hear each in context. On voice that includes the literal misheard forms, the pna for P&A, which are exactly what to add as grammar entries. Per-ask cards show result mixes and match rates, low-confidence matches under 70% are listed for mis-route checks, and an intents table gives average and minimum confidence. Match rate excludes agent requests and guardrail results, because a caller asking for a person is an escalation, not a recognition failure. The whole corpus exports as a five-sheet Excel workbook.

From one session replay to the tuning backlog An illustrative example. On the left, a session replay shows two reporting turns: the bot asks what the caller needs, the caller says pna please and the ask result is No match; the bot reprompts, the caller says parts and availability, and the intent matches with a confidence bar. On the right, the utterance corpus ranks the phrases the bot did not understand, each with the ask it happened at and hear-it-in-context links back into the replay, and the opt-in tuning panel turns the misheard form into a speech-to-text variant to add in Architect. SESSION REPLAY · ILLUSTRATIVE One caller, turn by turn TURN 1 · ASK: MAIN INTENT Bot: What can I help you with today? Caller: “pna please” No match TURN 2 · REPROMPT Bot: Is it about an order, or parts? Caller: “parts and availability” Success Intent: PartsAvailability confidence 62% UTTERANCE CORPUS The no-match backlog UTTERANCE ASK CONTEXT pna pleaseMain intent ▶ ▶ ▶ p and aMain intent ▶ ▶ it's the 7 digit onePart number ▶ MERGED CASE-INSENSITIVELY · RANKED BY FREQUENCY BOT TUNING · OPT-IN PER CLICK stt-variant · high pna p and a click to copy into Architect On voice, the utterance is the speech-to-text the NLU judged – so the misheard form is exactly what to add to the grammar
Illustrative: a no-match in one replayed session joins the ranked corpus, and tuning turns the misheard form into a grammar entry.
Read this diagram as text

An illustrative two-panel diagram: on the left a replayed bot session turn by turn, on the right the utterance corpus and tuning panel that the no-match feeds into.

  1. Left panel, "Session replay · illustrative", headed "One caller, turn by turn".
  2. Turn 1, "Ask: main intent": the bot says "What can I help you with today?", the caller says "pna please", and the result is marked "No match" (highlighted as a problem).
  3. Turn 2, "Reprompt": the bot says "Is it about an order, or parts?", the caller says "parts and availability", and the result is marked "Success".
  4. Below the turns, "Intent: PartsAvailability" is shown with a partly filled bar labelled "confidence 62%".
  5. A dotted arrow links the "pna please" row in the right panel back to the "No match" result in turn 1, showing where the phrase can be heard in context.
  6. Right panel, "Utterance corpus", headed "The no-match backlog": a table with columns Utterance, Ask and Context lists "pna please" at "Main intent" with three context links, "p and a" at "Main intent" with two, and "it's the 7 digit one" at "Part number" with one.
  7. Under the table: "Merged case-insensitively · ranked by frequency", with an arrow down to a "Bot tuning · opt-in per click" panel.
  8. The tuning panel shows "stt-variant · high" with the chips "pna" and "p and a" and the note "click to copy into Architect".
  9. The banner beneath reads: "On voice, the utterance is the speech-to-text the NLU judged – so the misheard form is exactly what to add to the grammar".

07

Tuning, privacy and the team behind it

A Bot tuning panel on the Utterances page can hand a large language model the run's per-ask digest, including every no-match utterance with its count, and returns prioritised recommendations with click-to-copy additions ready for Architect: grammar additions, STT variants, intent gaps, reprompt wording, escalation handling and format hints. It is opt-in per click, cached per run, and sends nothing until a user presses Generate. A separate analyst narrative uses aggregate metrics only, never caller speech. The warehouse itself holds caller utterances and slot values, so we advise clients to treat it, and its exports, with the same care as call recordings.

Bot Flow Diagnostics was built the QVCCS way. Our Senior Business Consultant and Business Analyst defined the questions bot owners ask after go-live; the Solution Architect designed the path projection and the match-rate semantics; Senior Developers built the harvest and renderers under the Senior Platform Practice Lead's engineering standards; and our Systems Integration Tester proved the caps, coverage rules and failure paths against a deterministic synthetic bot, so the demo exercises every feature without customer data. For clients, it closes the loop from what callers said to what to change in Architect next.

How it compares

Bot insight in native Genesys Cloud CX and in QVCCS Bot Flow Diagnostics

Native facts are as documented in the Genesys Cloud Resource Center and the Platform API definition.

AspectNative Genesys Cloud CXQVCCS Bot Flow Diagnostics
ScopeOptimization dashboard for Dialog Engine bot flows and digital bot flows; metrics for core bot flows, not invoked subflows.Architect bot and digital bot flows; nested bot calls appear as their own nodes on the map.
Turn-level historyUtterance History and Bot Conversation Library retain the last 10 days; API resources kept approximately, but not before, 10 days.Gap-aware warehouse keeps every harvested session and turn; windows of up to 42 days per run.
Aggregate metricsBot Performance Detail view: entries, durations, turns and intents, custom ranges up to 6 weeks.KPIs computed from the full harvested population, with a best-effort aggregates cross-check.
Journey viewJourney Flows show milestones and flow outcomes that influence containment.Journey map of prompts, asks, results, milestones and endings, with reach and branch shares.
One sessionBot Conversation Library: session id, bot result and reason.Turn-by-turn replay with prompts, utterances, ask results, confidence bars and slot values.
Action-level issuesTop Bot Flow Actions with Issues: no match, no input and other outcomes by action.Per-ask result mix, match rate and low-confidence matches, exported to a five-sheet workbook.
New intentsIntent Miner mines intents, utterances, topics and phrases from chat and call recordings.Ranked no-match corpus and opt-in tuning recommendations for the bot's existing asks.

Bot Flow Diagnostics complements the native bot insights; use Intent Miner and the Optimization dashboard alongside it.

The takeaways

  • Every bot session and reporting turn in a window, kept after Genesys ages it out.
  • A full-population journey map shows which ask leaks and where callers end.
  • Turn-by-turn replay puts every failure in its context, with confidence and slots.
  • A ranked no-match corpus turns misheard speech into concrete grammar entries.
  • Read-only harvesting with explicit caps, so nothing is ever silently missed.

Bot Flow Diagnostics is part of the QVCCS App Suite, included with every Managed Professional Services tier and built by the same certified team that designs, builds and supports Genesys Cloud CX solutions.

Read the user guide Managed Professional Services

Last reviewed

Questions about what you have read?

Clients, partners and people introduced to us can reach the specialists behind our applications and articles directly.

Who to contact