A Genesys Cloud CX compliance scanner: it harvests Speech & Text Analytics conversation
transcripts into a private cache and pattern-scans every phrase for payment-card (PCI) and personal (PII)
data that STA did not redact — surfacing the real exposure, with the honest caveat that regex can
only see structured data, which in practice lives in digital channels, not voice.
Continuously integrated and updated. This is a published sample of the user guide. The App Suite changes frequently, so this page may not reflect the latest features, screens and behaviour. The current guide is available in the application through the QVCCS apps portal, included with every Managed Professional Services tier.
At a glance
PII/PCI exposure scanner
harvest → scan
voice + digital channels
STA-redaction aware
live streaming
xls / pdf export
read-only against Genesys
Security at a glance
PII/PCI Finder
Read-only
Sign-in
OAuth client credentials you supply; held on the server for your session only and never sent to the browser
Stores
A private cache of harvested transcripts, which can contain unredacted PII/PCI; delete selected days or clear the whole cache in the app
PII/PCI Finder answers a question Genesys Cloud cannot answer natively: “Which of our recorded
conversations contain payment-card or personal data in the clear — data that Speech & Text
Analytics (STA) failed to redact?” Genesys keeps no per-conversation “contains unredacted PII”
flag and its transcript search is keyword-only, so PII/PCI Finder downloads the STA transcripts itself and
scans every phrase, listing only the conversations with genuine findings. Clean or fully-redacted
conversations are hidden.
Compliance / security reviewers — periodic PCI-DSS and data-protection audits of the
contact centre, with per-conversation evidence and highlighted transcripts.
CX engineers — verifying that STA redaction policies actually work after flow or policy
changes, and finding where they leak.
Operations — the Redacted by STA mode gives the inverse view: which conversations
STA did redact, and what kinds of data it caught.
How “unredacted” is decided
When STA redacts a value it replaces it with a bracketed placeholder — Genesys writes snake_case
tokens such as [card_number], [card_expiry_date] or [person].
PII/PCI Finder recognises a curated list of these placeholders as “already redacted”, runs its own detectors
over the remaining text, and treats any detector match outside a placeholder as exposed. A
detector hit overlapping a placeholder is discarded; placeholders themselves are counted separately
(shown green in the transcript viewer).
The two-step model
① Harvest (slow, rate-limited) downloads every transcript in the selected days once, into an
immutable per-conversation JSON cache. ② Scan then runs detection purely against that local
cache — instant, repeatable, and free to re-run with different scopes (PCI-only, PII-only, low-signal
on/off, Redacted-by-STA). A completed conversation’s transcript never changes, so nothing is ever
re-downloaded; overlapping harvests fill only the gaps.
Read results with the blind spot in mind. The detectors are regular
expressions: they catch structured values (digit runs, name@domain, dashed SSNs).
Voice transcripts are speech-to-text — spoken values arrive as words — so voice yields almost
no pattern-detectable findings, and scans default to voice-only. A result of 0 findings is not
proof of no exposure; see §4, “The voice blind spot”, before drawing conclusions.
Read-only posture: every Genesys call is a read — the OAuth token grant, the
analytics conversation query (a POST, but a read query), user-name lookups and transcript downloads.
PII/PCI Finder cannot modify anything in your Genesys org.
2Quick start
Create an OAuth client — Genesys Cloud Menu > IT and Integrations > OAuth → Add client,
grant type Client Credentials, with a role granting at least
analytics:conversationDetail:view, conversation:communication:view,
speechAndTextAnalytics:data:view and (for agent names)
directory:user:view. For exposure auditing, omitrecording:recording:viewSensitiveData — see §5.
Sign in — pick the org’s region, paste the client ID and secret. Credentials are held on the server (for silent token refresh) and never sent to the browser.
Harvest — select days on the coverage calendar (or set Date A → Date B) and press
① Harvest to cache. Tick “Also scan email / chat / message” first if you want digital
channels — that is where the machine-detectable findings live (§4). The harvest keeps running if you close the tab and even auto-resumes after a service restart.
Scan — press ② Scan cached data. Choose the scope (PCI & PII / PCI / PII /
Redacted by STA) and toggles first. Scans are cache-only: nothing downloads, so re-running with a
different scope is instant.
Review — click any flagged conversation ID to read its transcript with findings
highlighted (PCI red, PII amber, STA redactions green), then export XLS or PDF.
UTC everywhere. Dates and times in the range pickers, the calendar, the
results table and all API payloads are UTC. The default time window is 00:00 → 23:59.
3UI walkthrough
Everything lives on one Audit page. The header shows the connected region as a pill, a 📖 Guide button (this page) and
Log out. Top to bottom: cache insights, the multi-select coverage calendar, the date/time range
and action buttons, scan options, the live-streaming toggle, progress, and the results table.
Cache insights panel
Computed entirely from the cache — no Genesys call — and reflecting the current
calendar selection (or the whole cache when nothing is selected):
KPI tiles — cached conversations, days covered, cache size, voice share, high-signal
findings, STA redactions.
Channel mix — voice / email / chat proportions of the cached set.
Top redactions — the most frequent STA placeholder labels (what Genesys is catching).
Trend sparklines — per-day Volume and Findings lines with a range picker
(Year / 6 mo / Qtr / Last mo / This mo), month gridlines and labels, a findings-class filter
(All / PII / PCI — shared with the calendar’s Findings heatmap), and a hover readout showing that
day’s counts with a finding-type breakdown.
Multi-select coverage calendar
Month calendars from ~3 months back (or the oldest cached day) to today. The selection drives
the actions: Harvest, Scan and Delete apply to the selected days; the Date A → Date B range below
is used only when nothing is selected.
Selection gestures — click to toggle a day; drag to paint; Shift-click to
extend a span; click a weekday header to toggle that weekday across the month; keyboard: arrows
move (±1 / ±7), Space toggles, Shift+arrow extends, Home/End jump.
Colour-by modes — Coverage (cached vs empty, with amber for suspect partial
days), Volume heatmap, Findings heatmap (with the All/PII/PCI class toggle), and
vs Genesys, which fetches true per-day totals for the selected days (in 3-day chunks with
live progress) and shades attainment: fully / partially / not harvested.
Suspect detection — a cached day well below the typical volume for its weekday
(< 0.6 × the same-weekday median) is flagged as a likely partial harvest (e.g. one killed by a
restart). The Suspect preset selects them all for one-click re-harvest; vs Genesys
remains the authoritative check.
Saved selections — name and save the current day-set (“Save selection…”); saved windows
reload or delete from chips. Remembered in your browser, not shared.
Delete selected from cache / Clear entire cache — deletes from the app’s cache only (confirm guarded); Genesys is untouched and days can always be re-harvested.
Harvest ① and Scan ②
① Harvest to cache — downloads all transcripts for the selection (no detection). It warns
before duplicating a harvest already running for the org (offering to reattach instead), keeps running if you close the tab, auto-resumes after a service restart, and on
completion nudges you to run ②. The calendar refreshes so newly-harvested days turn green.
The digital-channels checkbox applies here too — an unticked harvest downloads voice only.
② Scan cached data — runs detection over the cache only:
un-harvested conversations are counted as “not harvested”, never downloaded mid-scan.
Scan scope radio — PCI & PII / PCI only / PII only /
Redacted by STA. The last inverts the audit: it flags conversations that contain
Genesys redaction placeholders, with per-label counts — proof of what STA is catching.
Toggles — “Also scan email / chat / message” (off = voice calls only; skipped
digital conversations are counted and reported) and “Also flag phone numbers & IPs”
(include low-signal detectors; noisier).
List all (no scan) — a plain conversation listing with channel, direction and handler
columns. Uses the Date A → B range (not the calendar selection).
Progress — per-day enumeration ticker, a download-based ETA (cached items are ~instant,
so the ETA tracks the true cost: cache misses), a running findings-by-type tally, a current-date
marker, and Cancel (the scan stops at the next conversation).
Reattach — reopening the app (even after logging back in with the same org credentials)
finds any running job for the org and re-binds the progress UI.
Results & the transcript viewer
Flagged table — conversation ID, date and time (UTC), channel, direction, agent/flow
handler (agent names resolved via the Users API, else flow name), and finding counts by category
(PCI / PII badges) with a per-type tooltip. In Redacted-by-STA mode the last column shows
placeholder counts instead.
Honest-zero banner — a completed exposure scan that flags 0 conversations shows a warning
explaining exactly what that does and does not mean (see §4); a voice-only scan additionally
reports how many digital conversations were skipped.
Transcript modal — click a conversation ID: speaker-labelled phrases (Customer / Agent /
System) with findings highlighted (PCI red, PII amber; tooltips carry card brand, confidence and a
reasoning note) and STA redactions in green. A summary bar counts unredacted PCI / PII and Genesys
redactions for the current scope. Rendering is DOM-node based — transcript text is never injected
as HTML.
Live streaming
Toggle Live streaming on to continuously ingest and flag new conversations. Every 30 seconds the app re-checks a window that overlaps the previous one by 10 minutes, downloads new transcripts, and flags findings under the current scope/signal/channel settings.
Self-healing — the app remembers how far it has got for each org; each window overlaps the last; conversations whose STA transcript isn’t ready yet are retried for
30 minutes rather than being cached as empty; catch-up after downtime is clamped to 6 h (older
gaps → run a manual harvest).
A right-rail console streams per-conversation events (✓ clean, ⚑ flagged, ✕ error, cycle
summaries). Flagged rows land in the same results table. Live mode auto-stops after 5 minutes
without a status poll (tab closed); manual scan controls are disabled while live mode is on.
Exports
Export XLS — Excel 2003 SpreadsheetML built entirely client-side (title + range + the current table), saved with the scan type and date range in the file name.
Export PDF — a formatted A4-landscape report opens in a new window and triggers the print
dialog; choose “Save as PDF”. Allow pop-ups for the site.
4Detection model
The engine runs a fixed set of regex detectors over every
transcript phrase, with checksum and context gates to keep false positives down. Each detector carries
a signal level: high-signal findings flag a conversation in the bulk scan by default;
low-signal detectors (noisy in contact-centre text) only flag when “Also flag phone numbers
& IPs” is ticked — but are always highlighted when you open a transcript.
Detector
Class
Signal
Gate — what must be true to match
Payment card number
PCI
high (context can demote to low)
12–19 digit run (spaces/dashes allowed) that passes the Luhn checksum and matches a real card scheme’s issuer range + length (Visa, Mastercard, Amex, Discover, Diners Club, JCB, UnionPay, Maestro). Nearby words score it: payment context (card, CVV, expiry…) raises confidence; counter-context (order, tracking, phone, facebook…) demotes to low signal. Broad-IIN schemes (Maestro, UnionPay) additionally need payment context or 4-4-4-4 grouping.
Spoken card number
PCI
same scoring
Consecutive number-words in voice STT (“four five six… double seven”, incl. double/triple/treble) are rebuilt into digits and pushed through the identical card gate; matches are tagged spoken. Runs only when PCI is in scope.
Card security code (CVV)
PCI
high
3–4 digits only when introduced by a CVV-ish keyword (cvv, cvc, “security code”, “card code”).
Card expiry date
PCI
low
MM/YY or MM/YYYY — too date-like to flag alone.
US Social Security Number
PII
high
Dashed ddd-dd-dddd form only.
Email address
PII
high
Standard name@domain.tld pattern.
IBAN (bank account)
PII
low
ISO 13616 mod-97 checksum must verify (stops all-caps prose matching). Classed PII, not PCI — PCI covers cardholder data only.
Phone number
PII
low
NANP-style pattern with digit-boundary guards so it can’t slice a phone shape out of a longer number.
IPv4 address
PII
low
Strict dotted quad (0–255 octets).
No UK-specific detectors. There is no detector for UK National Insurance numbers, or for UK sort codes and bank account numbers given on their own rather than inside an IBAN, so UK organisations should not rely on PII/PCI Finder to find them.
Context window — the scanner hands each phrase its neighbours as context (an agent asking
for “the card number” right before the customer reads it out), so card scoring isn’t blind to the
conversation around a match.
STA placeholders — a curated list of [bracketed] tokens (card_number,
card_expiry_date, cvv, ssn, account_number, iban, phone, email, person, name, address,
date_of_birth, pin, passport, driver’s licence, …) in snake_case or space-separated form is
recognised as “already redacted”. Any detector hit overlapping one is dropped; curation stops
email/STT artefacts like [external] or [cid:…] from being miscounted.
The voice blind spot — what this scanner cannot see
Regex detectors are effectively blind to voice. Voice transcripts are
speech-to-text: a spoken card number, email address or SSN arrives as words (“four five six…”,
“name at gmail dot com”), not as the digit runs and @-symbols the patterns need. Only the
spoken-card pass reconstructs number-words — and only for card numbers. On a month of data from a production organisation
(~28,000 cached conversations) 2,406 of 2,407 high-signal findings were email addresses inside
digital transcripts; the single PCI finding was one spoken card number in a voice call that STA had
missed. Barely ~2% of voice transcripts contain even one digit character. Meanwhile scans
default to voice-only — the channel where regex finds nothing — so a scan can honestly report
0 findings while exposure exists: in digital channels you didn’t scan, or spoken in voice where
no pattern can see it. Read “0 flagged” as “nothing machine-detectable in the scanned
channels”, never as “no exposure”.
Why voice is structurally poor ground for pattern matching here:
STT writes words, not symbols — numbers and addresses are usually transcribed as words,
which defeats every detector except the spoken-card pass.
STA already redacts much of the spoken risk — on the same cache, Genesys placeholders
appear 5,000+ times across 1,100+ conversations (card numbers, expiry dates, names — including
spoken ones). Long digit runs are rare in voice partly because STA got there first; use the
Redacted by STA scan mode to see that side of the ledger.
Names, addresses and free-text PII need ML/NLP — exactly what STA itself does. No regex will ever catch a spoken name and street address; PII/PCI Finder does not attempt to.
What to do about it: tick “Also scan email / chat / message” before both
① Harvest and ② Scan — digital channels are where structured, detectable data lives (a voice-only
harvest never downloads them, so a later digital scan would just count them “not harvested”). The
voice-only default is a deliberate scope choice (voice is the primary risk channel in most PCI audits and
digital inflates run time); the app backs it up with the honest-zero banner whenever a voice-only scan
returns 0. Keep using STA redaction policies as the primary control for spoken data — PII/PCI Finder audits
the leftovers it can see, it does not replace STA.
Positive signals still matter: when the spoken-card pass does fire it has the
full IIN + Luhn + context gate behind it, so a spoken-card finding is a genuine STA miss worth acting
on — that is exactly how the one real residual card above was found.
5Genesys endpoints & permissions
All Genesys access happens on the server with the session’s bearer token, against your org’s Genesys region (14 public regions built in). HTTP
429 responses honour Retry-After with up to 3 retries.
Endpoint
Used for
POST
login.<region>/oauth/token
Client-credentials sign-in (token cached with a 60 s expiry safety margin; refreshed silently mid-scan)
POST
/api/v2/analytics/conversations/details/query
Enumerate conversations — one UTC day per query (the API caps intervals at 7 days and a busy day can silently exhaust page caps), paged 100/page, ordered by conversationStart, de-duplicated by conversation ID
GET
/api/v2/users?id=…
Agent userId → display name (batched ×50, best-effort — failures fall back to participant/flow names)
GET
/api/v2/conversations/{id}
Conversation detail → communication IDs for the transcript endpoint. Often skipped: the scanner probes whether the analytics voice sessionId works directly as the communicationId; when it does, this call is bypassed (~half the rate-limited calls saved)
Pre-signed transcript download URLs (404 = no transcript for that communication)
GET
pre-signed URL
The transcript JSON itself (no Authorization header — the URL is self-authorising)
OAuth client permissions:analytics:conversationDetail:view (enumeration) ·
conversation:communication:view (detail fallback) ·
speechAndTextAnalytics:data:view (transcript URLs) ·
directory:user:view (optional — without it, agent columns fall back to participant or
flow names).
The client’s permission posture defines what “exposed” means. A client
withoutrecording:recording:viewSensitiveData receives transcripts as an ordinary
user sees them — findings are true exposure gaps. A client with it receives raw unmasked text —
everything STA would normally hide gets flagged. For exposure auditing use a client without
View Sensitive Data, and use the same client for all harvesting: the cache is keyed by
conversation only, so mixing postures would blend masked and unmasked text in one cache.
6Storage, privacy & security
No database — the app keeps a private cache: one entry per conversation transcript (including “no transcript” results so they aren’t re-queried), resumable harvest records, and live-mode progress per org.
The transcript cache is sensitive at rest. It can contain unredacted PII/PCI — that is the point of the tool. Access to it is restricted and writes are atomic. Apply a retention policy: delete days via the UI, or clear the whole cache.
Credentials — the client ID/secret are held on the server for your session only, so an expired token can be renewed silently mid-scan; they are never sent to the browser, which holds only a secure session cookie. Sessions end after a period of inactivity.
Resumable-harvest records keep the client secret encrypted, and each record is deleted the instant its harvest settles.
Auth model — there is no separate app login: authentication is the Genesys OAuth client-credentials grant, and every data request requires that session. The guide itself is public by design.
7Technology
Web application — harvest and scan jobs run in the background on the server — enumerate one UTC day at a time → download → detect → tally — using a fast path that skips per-conversation detail calls when possible. Finished jobs are tidied away about 30 minutes after completion.
Browser interface — a single-page app that polls job progress; the XLS and PDF exports are generated entirely in your browser.
Detection engine — the detectors with their Luhn/IIN/IBAN gates, spoken-number reconstruction, STA-placeholder recognition, and the segmenter the transcript viewer renders from (§4).
Cache & resilience — a read-through transcript cache with a per-day coverage index (volume, channel mix, findings, redaction labels — behind the calendar, KPIs and trends); live mode (30-second cycles, overlapping windows, STA-lag retries, idle auto-stop); and encrypted records of in-flight harvests — the auto-resume mechanism (§6).
Scheduled pre-ingestion — QVCCS can schedule the transcript cache to fill overnight (for example, yesterday’s or the last 7 days’ voice transcripts) so interactive scans run from the cache. Scheduled ingestion is voice-only — to cover email/chat/message, run ① Harvest in the UI with the digital checkbox ticked. It uses the same no-View-Sensitive-Data client as interactive scans: the cache is keyed by conversation, not by permission posture (§5).
8Troubleshooting
“Login failed (HTTP 400/401)”
Wrong client ID/secret — or the right credentials in the wrong region: a client only authenticates against its own org’s region. If Genesys could not be reached, try again shortly.
My voice-only scan reports zero findings.
Working as designed, and the app says so in the results banner: voice STT defeats pattern
matching, and the digital channels (where detectable findings live) were skipped. Tick “Also scan
email / chat / message”, re-run ① Harvest (a voice-only harvest never downloaded them), then ②
Scan. And remember: 0 is “nothing machine-detectable”, not “no exposure” — §4.
A digital-inclusive scan reports lots of “not harvested”.
② Scan is cache-only. The earlier harvest ran voice-only, so the digital conversations were
never downloaded. Re-run ① Harvest with the digital checkbox ticked, then scan again.
“Session expired; please log in again.” / sudden 401s
Your session ended after a period of inactivity, or the service was restarted. Log back in. Running harvests
resume on their own after a restart; plain scans do not — re-run them (the cache makes the
re-run fast).
A scan is slow / seems stuck on “enumerating”.
Enumeration queries one UTC day at a time and large orgs rate-limit; the app backs off on HTTP 429. The heavy phase is downloading cache-misses — the ETA is based on those. Narrow the day-set, or ask QVCCS to schedule overnight pre-ingestion.
Rows say “no transcript”.
STA never produced a transcript for that conversation: transcription wasn’t enabled for the
queue/flow, the conversation is outside the org’s transcript retention, or it is so recent STA
hasn’t finished. These are an unscannable blind spot, tallied separately. Live mode retries recent
empties for 30 minutes before caching them as settled.
“Scan job not found” while polling.
Finished jobs are tidied away after ~30 minutes, and job tracking is reset when the service restarts. Exported results are unaffected; re-run the scan if needed.
The results banner says “truncated”.
A single day exceeded the per-day enumeration cap (20,000 conversations by default), so
enumeration stopped early for that day. Split the day into smaller time ranges and run them
separately.
“A harvest is already running for this org …”
The duplicate guard found a running harvest — possibly started by a colleague or a previous
session (jobs are tagged by region + client ID). Choose OK to attach and watch it rather than
double-downloading; harvesting is idempotent either way.
A day looks harvested but its volume/findings seem low.
Check the Suspect marker — a cached day well under its weekday’s median volume is
probably a partial harvest (e.g. one that was interrupted part-way). Select it (the
Suspect preset does this) and re-run ①: only the gaps download. vs Genesys is the
authoritative comparison.
A “card number” finding looks like an order number.
Findings survive Luhn and real issuer-range checks, but context matters: matches near
words like “order” or “tracking” are demoted to low signal (hidden unless the low-signal toggle is
on). Open the transcript — each highlight carries its brand, confidence and reasoning note.
Live streaming stopped by itself.
Live mode auto-stops after 5 minutes without a status poll — i.e. the tab was closed or asleep. Toggle it back on; it backfills the gap (up to 6 h; older gaps → run a
manual harvest).
Export PDF does nothing.
The report opens in a new window and calls print — allow pop-ups for this site.
Is the cached data itself sensitive?
Yes. The transcript cache can contain unredacted PII/PCI — treat it like call recordings: restricted access and a retention policy (§6). Delete days via the UI or
clear the whole cache; Genesys is never modified.
QVCCS PII/PCI Finder — unredacted PII/PCI exposure auditing for Genesys Cloud CX. Strictly read-only against the Genesys APIs; the transcript cache is sensitive at rest.
Continuously integrated and updated. This is a published sample of the user guide. The App Suite changes frequently, so this page may not reflect the latest features, screens and behaviour. The current guide is available in the application through the QVCCS apps portal, included with every Managed Professional Services tier.
Included at every tier
Use PII/PCI Finder with QVCCS Managed Professional Services.
The App Suite comes with every Bronze, Silver, Gold and Diamond subscription, used by our engineers and your administrators alike – backed by the certified QVCCS bench.