QVCCS innovation · Audit and governance
Finding what redaction missed: scanning Genesys Cloud transcripts for PII and card data
Genesys Cloud CX gives you strong tools to keep card data out of recordings and to redact sensitive data from transcripts. Our PII/PCI Finder answers the auditor’s follow-up question: across a quarter of conversations, where did something slip through? It harvests Speech and Text Analytics transcripts, scans every phrase with checksum-gated detectors, and is candid about what pattern matching cannot see.
Did you know that Genesys Cloud shows unredacted transcripts only to users with the Recording > Recording > View Sensitive Data permission, that no role includes it by default, and that an administrator must grant it manually?
Did you know Genesys states that Automated Sensitive Data Masking is not PCI DSS compliant? It is a safety net for when agents forget Secure Pause; only Secure Pause and Secure Call Flows are validated by a Qualified Security Assessor as Level 1 compliant.
Did you know that, according to Genesys, the headers and descriptions of AI-generated transcript outlines are not redacted, so users with access to outlines may see information that is redacted in the transcript itself?
01
The auditor’s question on a Tuesday afternoon
The data-protection review is booked for Thursday. The auditor has one request that sounds easy: show me every conversation from last quarter where a payment card number, a card security code or an email address appears in a transcript in the clear. You know your controls. Agents use Secure Pause, payments go through a secure flow, and redaction is switched on. But controls are not evidence, and “we think it is fine” is not an answer anyone wants to give in a PCI DSS or GDPR conversation.
The difficulty is scale. A transcript viewer is built for reading one conversation carefully, not for proving a negative across tens of thousands. That gap is what our PII/PCI Finder was built to close. It does not replace a single Genesys control. It audits the leftovers – the values that reached a transcript without being caught – and lists only the conversations with genuine findings, so a reviewer spends their afternoon on evidence rather than on scrolling.
02
What Genesys Cloud CX does natively – and does well
Prevention comes first, and Genesys is clear about it. Secure Pause lets an agent temporarily stop recording while card details are given. A secure flow masks the audio path and the data a caller enters in an automated IVR, so neither the recording nor the agent hears it; Genesys notes that diagnostic protocol captures should not be enabled when secure flows are in use. Genesys Cloud is Service Provider Level 1 compliant with PCI DSS version 4.0, and an org with the PCI DSS setting enabled disables DTMF logging and media capture by the Edge.
Redaction is the second layer. With Speech and Text Analytics, administrators can switch on Sensitive Data Redaction for Payment Cards and for Personal Information in Organization Settings. Detected values become bracketed placeholders such as [card number], [card expiry date], [email] or [name], and the corresponding audio plays back as silence. Genesys describes this as best effort, available only when speech or text analytics is enabled for the interaction, with national identification numbers covered for Australia, Canada, the UK and the US.
Licensing is explicit in the documentation: automatic redaction lists a Genesys Cloud CX 1 WEM Add-on II, Genesys Cloud CX 2 WEM Add-on I, Genesys Cloud CX 3 or Genesys Cloud CX 4 licence, plus the Speech and Text Analytics Upgrade Add-On. Our app depends on the same thing – it reads STA transcripts – so it sits on top of that investment rather than beside it. The question it answers is the one only an after-the-fact audit can: did all of that work, conversation by conversation?
03
Harvest once, scan as many times as you like
The app separates the slow part from the fast part. Step one, Harvest, enumerates conversations through /api/v2/analytics/conversations/details/query – one UTC day per query, paged and de-duplicated, so a busy day cannot quietly exhaust a page cap. For each conversation it asks /api/v2/speechandtextanalytics/conversations/{conversationId}/communications/{communicationId}/transcripturls for pre-signed download links and fetches the transcript JSON into an immutable per-conversation cache. Agent names are resolved through the Users API in batches of fifty.
We spent real engineering effort on the rate-limited path. The transcript endpoint needs a communication id, which normally means a call to /api/v2/conversations/{conversationId} first. The harvester probes whether the analytics voice session id works directly as the communication id; when it does, that detail call is skipped – roughly half of the rate-limited calls saved. HTTP 429 responses honour Retry-After with up to three retries, a long harvest keeps running if the tab closes, and an interrupted harvest resumes on its own after a restart.
Step two, Scan, runs detection against the local cache only. Nothing downloads, so re-running with a different scope – PCI only, PII only, low-signal detectors on or off – is instant. A coverage calendar shows which days are cached, flags suspect days whose volume falls below 0.6 times the median for that weekday, and offers a “vs Genesys” mode that fetches true per-day totals so partial harvests cannot hide.
Read this diagram as text
Three columns show five read-only Genesys Cloud calls feeding a two-step harvest and scan inside the app, which produces one view of flagged conversations.
- Left panel, "Genesys Cloud Platform API: five read-only calls", in order: enumerate conversations one UTC day at a time with the analytics conversation details query; get communication ids from the conversation, often skipped; get pre-signed transcript URLs from Speech and Text Analytics; fetch the transcript JSON from the pre-signed URL, which is self-authorising; and resolve agent display names from users, batched 50 at a time.
- All five calls join and arrow into the middle panel, "In the app: two steps".
- A fast path tries the voice session id as the communication id, which skips the detail call, and leads to step one.
- Step one, "Harvest · cache", stores immutable JSON per conversation.
- Step two, "Scan · detect", checks placeholders first, then Luhn, issuer and context, using the cache only so re-runs are instant.
- Notes in the middle mention HTTP 429 with Retry-After and that the app is read-only against Genesys.
- The scan feeds the right panel, "One view: only genuine findings": flagged conversations (ID, UTC time, channel, handler), PCI and PII counts per type with a tooltip, a transcript viewer, a Redacted by STA mode showing what Genesys caught by label, and XLS workbook and PDF report exports.
- In the transcript viewer, red marks PCI, amber marks PII and green marks STA redactions.
- A final note says clean conversations are hidden.
04
A detection model built not to cry wolf
A contact centre is full of long numbers – order references, tracking codes, phone numbers – so a naive digit-run search is useless. The payment card detector requires a run of 12 to 19 digits, spaces or dashes allowed, that passes the Luhn checksum and also matches a real card scheme’s issuer range and length: Visa, Mastercard, Amex, Discover, Diners Club, JCB, UnionPay or Maestro. Nearby words then score it. Payment context such as card, CVV or expiry raises confidence; counter-context such as order, tracking or phone demotes it to low signal. Broad-range schemes need payment context or 4-4-4-4 grouping as well.
The other detectors are gated just as carefully. A card security code only counts when introduced by a keyword such as CVV, CVC or security code. Expiry dates in MM/YY form are low signal because they look like any date. US Social Security numbers are matched in dashed form only, email addresses by the standard pattern, and IBANs must pass the ISO 13616 mod-97 check – classed as PII, not PCI, because PCI covers cardholder data. Phone numbers and IPv4 addresses are low signal and flag a conversation only when you ask for them. There is also a gap UK organisations must know about: the current detectors do not look for UK National Insurance numbers, or for UK sort codes and bank account numbers given on their own rather than inside an IBAN, so the app should not be relied on to find them.
Crucially, the detector knows what Genesys already redacted. A curated list of STA placeholders is recognised as already handled; any detector hit overlapping one is discarded, and placeholders are counted separately. A phrase reading “my card is [card number]” is a success story, not a finding. A phrase where card digits sit outside any placeholder – which we only ever show masked, as •••• •••• •••• •••• – is what reaches the results table – shown in red in the transcript viewer, with the card brand, confidence and a reasoning note on hover.
Interactive explainer: with JavaScript on, see the Luhn checksum worked step by step on test card numbers.
Read this diagram as text
Two bands show the payment card detector as six gates in sequence, then the other PCI and PII detectors with their signal level.
- Top band, "Payment card detector · PCI: six gates before anything is called a card number", runs left to right.
- Gate 1, Phrase: the phrase plus its neighbours. Gate 2, STA first: skip anything inside a placeholder such as [card number]. Gate 3, Digit run: 12 to 19 digits, with spaces or dashes.
- Gate 4, Luhn: the checksum must pass. Gate 5, Issuer: issuer range and length must match. Gate 6, Context: payment words raise the score, order and tracking words lower it.
- A masked example, "my card is" followed by sixteen masked digits outside any placeholder, passes every gate and becomes a high-signal PCI finding.
- Bottom band, "Other detectors": high signal flags a conversation, low signal is shown only on request.
- PCI detectors: CVV or CVC after a keyword (high), and card expiry as MM/YY or MM/YYYY (low).
- PII detectors: US SSN in dashed form only (high), email as name@domain.tld (high), IBAN with mod-97 checksum (low), phone by NANP pattern (low), and IPv4 as a strict dotted quad (low).
05
The voice blind spot, stated plainly
Here is the part we insisted the app says out loud. Regular expressions see structured values: digit runs, an @ sign, dashed numbers. Voice transcripts are speech-to-text, so a spoken email arrives as “name at example dot com” and a spoken card number as words. Apart from the spoken-card pass, pattern detectors are effectively blind to voice – and names, addresses and free-text personal data need the kind of language model that Speech and Text Analytics itself uses. The app does not pretend otherwise.
So the results screen carries an honest-zero banner. A scan that flags nothing tells you exactly what that means: nothing machine-detectable in the channels you scanned, never “no exposure”. Scans default to voice, and the banner reports how many digital conversations were skipped. In our own testing on a month of production data, almost every high-signal finding was an email address in a digital transcript. The advice in the guide follows: tick “Also scan email / chat / message” before both Harvest and Scan, and keep STA redaction as the primary control for spoken data.
Read this diagram as text
Two panels show three stacked layers of control for sensitive data, and what pattern matching can and cannot see in voice and digital transcripts.
- Left panel, "Layers of control: prevent, redact, then audit", runs top to bottom.
- Layer 1, "Prevent capture · native": Secure Pause and Secure Call Flows, QSA-validated as Level 1 PCI DSS compliant.
- What still reaches a transcript passes to layer 2, "Redact · native STA": sensitive data redaction, best effort, using bracketed placeholders, with audio played as silence.
- What redaction missed passes to layer 3, "Audit · QVCCS PII/PCI Finder": it finds the leftovers it can see and is an audit aid, not a PCI DSS control.
- Right panel, "What patterns can see: voice writes words", compares voice speech-to-text with digital chat and email.
- In voice, spoken digits such as "four five six … double seven" are rebuilt by the spoken-card pass, but a spoken email such as "name at example dot com" has words and no @, so it is invisible.
- In digital channels, a masked email is structured and detectable, and a masked digit run is gated and scored.
- Notes say names, addresses and free text need language models, and zero flagged means nothing machine-detectable, not "no exposure".
06
Permission posture decides what “exposed” means
One design decision matters more than any regex. The OAuth client the app uses should not hold View Sensitive Data. A client without it receives transcripts as an ordinary user sees them, so every finding is a real exposure gap. A client with it receives unmasked text, and everything STA would normally hide gets flagged. Because the cache is keyed by conversation, the guide asks you to use the same client for every harvest so masked and unmasked text never mix. A Redacted by STA scan mode inverts the audit to show what Genesys caught, by placeholder label.
The cache itself is sensitive – that is the point of the tool – so it is created with owner-only file permissions, written atomically, and should live on an encrypted volume under a retention policy; days can be deleted from the calendar at any time, and Genesys is never modified. Every Genesys call is a read, credentials stay server-side, and transcripts render as text nodes, never injected as HTML. In live streaming mode, every 30 seconds the app re-checks a window that overlaps the previous one by 10 minutes, and retries transcripts that are not ready for up to 30 minutes, for continuous monitoring.
This is the kind of build our method exists for. A Senior Business Consultant captured the auditor’s questions as acceptance criteria; a Solution Architect set the security and data-protection design in the Design stage; a Senior Developer built the detection engine; and a Systems Integration Tester ran negative tests against order numbers, tracking codes and placeholder-adjacent text before a design authority review. Findings export to XLS or a formatted PDF report, ready for the review on Thursday.
How it compares
Native controls and the QVCCS PII/PCI Finder
The native features prevent and redact. The QVCCS app audits the result. They are layers, not alternatives.
| Aspect | Native Genesys Cloud CX | QVCCS PII/PCI Finder |
|---|---|---|
| Preventing capture | Secure Pause and Secure Call Flows, validated by a QSA as Level 1 PCI DSS compliant. | Does not prevent capture; audits transcripts afterwards and never replaces these controls. |
| Redaction | Best-effort STA redaction of payment card and personal information entities into bracketed placeholders, when analytics is enabled. | Recognises those placeholders, counts them, and flags detector matches found outside them. |
| Searching transcripts | Keyword search within a transcript in the interaction view. | Scans every phrase of every harvested conversation in a date range and lists only those with findings. |
| Card detection | Card number, expiry and CVV entities detected by STA. | 12–19 digits, Luhn checksum, real issuer range and length, context scoring, spoken-digit reconstruction. |
| Who sees unredacted data | Users with View Sensitive Data, which no role includes by default. | Recommends an OAuth client without it, so findings are true exposure gaps. |
| Coverage gaps | Redaction applies only where speech or text analytics is enabled; national IDs for Australia, Canada, the UK and the US. | Conversations with no transcript are tallied as an unscannable blind spot; an honest-zero banner explains empty results. No detectors for UK National Insurance numbers, sort codes or bank account numbers. |
| Licensing | Automatic redaction lists Genesys Cloud CX licence options plus the Speech and Text Analytics Upgrade Add-On. | Reads STA transcripts, so it relies on STA being in place. |
Native behaviour is as described in the Genesys Cloud Resource Center at the time of writing. The QVCCS app is an audit aid; it is not a PCI DSS control.
The takeaways
- Evidence, not assumption: every harvested conversation scanned, only genuine findings listed.
- Checksum- and issuer-gated card detection with context scoring keeps order numbers out of the results.
- STA-aware: Genesys placeholders are recognised, counted and never miscounted as exposure.
- Candid about limits: voice is a blind spot for pattern matching, and UK National Insurance numbers, sort codes and bank account numbers are not detected.
- Read-only against Genesys, with a sensitive cache you control and can purge.
PII/PCI Finder is part of the QVCCS App Suite, included with every Managed Professional Services tier and built by the same certified team that designs, builds and supports Genesys Cloud CX solutions.
Sources
- Enable automatic redaction of sensitive informationhelp.genesys.cloud
- Work with a voice transcripthelp.genesys.cloud
- PCI DSS compliance – Genesys Cloud Resource Centerhelp.genesys.cloud
- Is Automated Sensitive Data Masking PCI DSS compliant?help.genesys.cloud
- Secure flows overviewhelp.genesys.cloud
- Speech and Text Analytics APIs – Genesys Cloud Developer Centerdeveloper.genesys.cloud