QVCCS innovation · Audit and governance

Finding what redaction missed: scanning Genesys Cloud transcripts for PII and card data

Genesys Cloud CX gives you strong tools to keep card data out of recordings and to redact sensitive data from transcripts. Our PII/PCI Finder answers the auditor’s follow-up question: across a quarter of conversations, where did something slip through? It harvests Speech and Text Analytics transcripts, scans every phrase with checksum-gated detectors, and is candid about what pattern matching cannot see.

Updated QVCCS Innovation teamPII/PCI Finder user guide →

Find what redaction missed An illustrative transcript with three customer phrases. In a voice phrase the card number has been replaced by the Speech and Text Analytics placeholder [card number], highlighted green as redacted. In a chat phrase masked card digits sit outside any placeholder and are highlighted red as unredacted card data, and in an email phrase a masked email address is highlighted amber as unredacted personal data. On the right, three result tiles: redacted by STA, unredacted PCI and unredacted PII. QVCCS INNOVATION · COMPLIANCE Find what redaction missed TRANSCRIPT · ILLUSTRATIVE, MASKED AGENT Can I take the long card number? CUSTOMER · VOICE It is [card number] CUSTOMER · CHAT Card •••• •••• •••• •••• CUSTOMER · EMAIL Reach me at •••••@••••• Redacted by STA Counted, not flagged Unredacted PCI Luhn and issuer checked Unredacted PII Email, SSN, IBAN
  • Did you know that Genesys Cloud shows unredacted transcripts only to users with the Recording > Recording > View Sensitive Data permission, that no role includes it by default, and that an administrator must grant it manually?

  • Did you know Genesys states that Automated Sensitive Data Masking is not PCI DSS compliant? It is a safety net for when agents forget Secure Pause; only Secure Pause and Secure Call Flows are validated by a Qualified Security Assessor as Level 1 compliant.

  • Did you know that, according to Genesys, the headers and descriptions of AI-generated transcript outlines are not redacted, so users with access to outlines may see information that is redacted in the transcript itself?

01

The auditor’s question on a Tuesday afternoon

The data-protection review is booked for Thursday. The auditor has one request that sounds easy: show me every conversation from last quarter where a payment card number, a card security code or an email address appears in a transcript in the clear. You know your controls. Agents use Secure Pause, payments go through a secure flow, and redaction is switched on. But controls are not evidence, and “we think it is fine” is not an answer anyone wants to give in a PCI DSS or GDPR conversation.

The difficulty is scale. A transcript viewer is built for reading one conversation carefully, not for proving a negative across tens of thousands. That gap is what our PII/PCI Finder was built to close. It does not replace a single Genesys control. It audits the leftovers – the values that reached a transcript without being caught – and lists only the conversations with genuine findings, so a reviewer spends their afternoon on evidence rather than on scrolling.

02

What Genesys Cloud CX does natively – and does well

Prevention comes first, and Genesys is clear about it. Secure Pause lets an agent temporarily stop recording while card details are given. A secure flow masks the audio path and the data a caller enters in an automated IVR, so neither the recording nor the agent hears it; Genesys notes that diagnostic protocol captures should not be enabled when secure flows are in use. Genesys Cloud is Service Provider Level 1 compliant with PCI DSS version 4.0, and an org with the PCI DSS setting enabled disables DTMF logging and media capture by the Edge.

Redaction is the second layer. With Speech and Text Analytics, administrators can switch on Sensitive Data Redaction for Payment Cards and for Personal Information in Organization Settings. Detected values become bracketed placeholders such as [card number], [card expiry date], [email] or [name], and the corresponding audio plays back as silence. Genesys describes this as best effort, available only when speech or text analytics is enabled for the interaction, with national identification numbers covered for Australia, Canada, the UK and the US.

Licensing is explicit in the documentation: automatic redaction lists a Genesys Cloud CX 1 WEM Add-on II, Genesys Cloud CX 2 WEM Add-on I, Genesys Cloud CX 3 or Genesys Cloud CX 4 licence, plus the Speech and Text Analytics Upgrade Add-On. Our app depends on the same thing – it reads STA transcripts – so it sits on top of that investment rather than beside it. The question it answers is the one only an after-the-fact audit can: did all of that work, conversation by conversation?

03

Harvest once, scan as many times as you like

The app separates the slow part from the fast part. Step one, Harvest, enumerates conversations through /api/v2/analytics/conversations/details/query – one UTC day per query, paged and de-duplicated, so a busy day cannot quietly exhaust a page cap. For each conversation it asks /api/v2/speechandtextanalytics/conversations/{conversationId}/communications/{communicationId}/transcripturls for pre-signed download links and fetches the transcript JSON into an immutable per-conversation cache. Agent names are resolved through the Users API in batches of fifty.

We spent real engineering effort on the rate-limited path. The transcript endpoint needs a communication id, which normally means a call to /api/v2/conversations/{conversationId} first. The harvester probes whether the analytics voice session id works directly as the communication id; when it does, that detail call is skipped – roughly half of the rate-limited calls saved. HTTP 429 responses honour Retry-After with up to three retries, a long harvest keeps running if the tab closes, and an interrupted harvest resumes on its own after a restart.

Step two, Scan, runs detection against the local cache only. Nothing downloads, so re-running with a different scope – PCI only, PII only, low-signal detectors on or off – is instant. A coverage calendar shows which days are cached, flags suspect days whose volume falls below 0.6 times the median for that weekday, and offers a “vs Genesys” mode that fetches true per-day totals so partial harvests cannot hide.

Genesys Cloud endpoints converge on one cached, scannable view On the left, five read-only calls: the analytics conversation details query enumerates conversations one UTC day at a time, the conversations endpoint supplies communication ids when needed, the Speech and Text Analytics transcript URLs endpoint returns pre-signed links, the pre-signed URL returns the transcript JSON, and the Users API resolves agent names. In the middle, a fast path tries the voice session id as the communication id, transcripts land in an immutable per-conversation cache in step one, and step two scans the cache with placeholder, Luhn, issuer and context checks. On the right, the single view: flagged conversations with PCI and PII counts, a highlighted transcript viewer, a Redacted by STA mode and XLS or PDF exports. GENESYS CLOUD PLATFORM API Five read-only calls Enumerate conversations · one UTC dayPOST /api/v2/analytics/conversations/details/query Communication ids · often skippedGET /api/v2/conversations/{conversationId} Pre-signed transcript URLsGET /api/v2/speechandtextanalytics/conversations/{conversationId}/communications/{commId}/transcripturls Transcript JSONGET pre-signed URL (self-authorising) Agent display names · batched ×50GET /api/v2/users?id=… IN THE APP Two steps FAST PATH Voice session id tried as communication id skips the detail call ① HARVEST · CACHE Immutable JSON per conversation ② SCAN · DETECT Placeholders first Luhn · issuer · context cache only, instant re-runs HTTP 429 · RETRY-AFTER READ-ONLY AGAINST GENESYS ONE VIEW Only genuine findings Flagged conversationsID · UTC time · channel · handler PCI and PII countsPer type, with a tooltip Transcript viewerPCI red · PII amber · STA green Redacted by STA modeWhat Genesys caught, by label ExportsXLS workbook · PDF report Clean conversations are hidden
Analytics, conversations, users and Speech and Text Analytics endpoints converge on one cached, scannable view of flagged conversations.
Read this diagram as text

Three columns show five read-only Genesys Cloud calls feeding a two-step harvest and scan inside the app, which produces one view of flagged conversations.

  1. Left panel, "Genesys Cloud Platform API: five read-only calls", in order: enumerate conversations one UTC day at a time with the analytics conversation details query; get communication ids from the conversation, often skipped; get pre-signed transcript URLs from Speech and Text Analytics; fetch the transcript JSON from the pre-signed URL, which is self-authorising; and resolve agent display names from users, batched 50 at a time.
  2. All five calls join and arrow into the middle panel, "In the app: two steps".
  3. A fast path tries the voice session id as the communication id, which skips the detail call, and leads to step one.
  4. Step one, "Harvest · cache", stores immutable JSON per conversation.
  5. Step two, "Scan · detect", checks placeholders first, then Luhn, issuer and context, using the cache only so re-runs are instant.
  6. Notes in the middle mention HTTP 429 with Retry-After and that the app is read-only against Genesys.
  7. The scan feeds the right panel, "One view: only genuine findings": flagged conversations (ID, UTC time, channel, handler), PCI and PII counts per type with a tooltip, a transcript viewer, a Redacted by STA mode showing what Genesys caught by label, and XLS workbook and PDF report exports.
  8. In the transcript viewer, red marks PCI, amber marks PII and green marks STA redactions.
  9. A final note says clean conversations are hidden.

04

A detection model built not to cry wolf

A contact centre is full of long numbers – order references, tracking codes, phone numbers – so a naive digit-run search is useless. The payment card detector requires a run of 12 to 19 digits, spaces or dashes allowed, that passes the Luhn checksum and also matches a real card scheme’s issuer range and length: Visa, Mastercard, Amex, Discover, Diners Club, JCB, UnionPay or Maestro. Nearby words then score it. Payment context such as card, CVV or expiry raises confidence; counter-context such as order, tracking or phone demotes it to low signal. Broad-range schemes need payment context or 4-4-4-4 grouping as well.

The other detectors are gated just as carefully. A card security code only counts when introduced by a keyword such as CVV, CVC or security code. Expiry dates in MM/YY form are low signal because they look like any date. US Social Security numbers are matched in dashed form only, email addresses by the standard pattern, and IBANs must pass the ISO 13616 mod-97 check – classed as PII, not PCI, because PCI covers cardholder data. Phone numbers and IPv4 addresses are low signal and flag a conversation only when you ask for them. There is also a gap UK organisations must know about: the current detectors do not look for UK National Insurance numbers, or for UK sort codes and bank account numbers given on their own rather than inside an IBAN, so the app should not be relied on to find them.

Crucially, the detector knows what Genesys already redacted. A curated list of STA placeholders is recognised as already handled; any detector hit overlapping one is discarded, and placeholders are counted separately. A phrase reading “my card is [card number]” is a success story, not a finding. A phrase where card digits sit outside any placeholder – which we only ever show masked, as •••• •••• •••• •••• – is what reaches the results table – shown in red in the transcript viewer, with the card brand, confidence and a reasoning note on hover.

Interactive explainer: with JavaScript on, see the Luhn checksum worked step by step on test card numbers.

The PII/PCI Finder detection model Top: the payment card detector as six gates in sequence – take the phrase with its neighbours for context, skip anything inside a Speech and Text Analytics placeholder, find a run of 12 to 19 digits, require the Luhn checksum to pass, require a real card scheme issuer range and length, then score nearby words so payment context raises confidence and order or tracking context demotes it. A masked example shows a finding that passes every gate. Bottom: the other detectors with their class and signal – CVV after a keyword, card expiry, dashed US SSN, email, IBAN with mod-97 checksum, phone and IPv4. PAYMENT CARD DETECTOR · PCI Six gates before anything is called a card number 1 · PHRASEPhrase plus itsneighbours 2 · STA FIRSTSkip anythingin [card number] 3 · DIGIT RUN12–19 digits,spaces or dashes 4 · LUHNChecksummust pass 5 · ISSUERIssuer rangeand length match 6 · CONTEXTPayment words ↑Order, tracking ↓ Masked example: “my card is •••• •••• •••• ••••” outside any placeholder → passes every gate → high-signal PCI finding OTHER DETECTORS · HIGH SIGNAL FLAGS A CONVERSATION, LOW SIGNAL ONLY ON REQUEST PCI · HIGHCVV / CVCafter a keyword PCI · LOWCard expiryMM/YY, MM/YYYY PII · HIGHUS SSNdashed form only PII · HIGHEmailname@domain.tld PII · LOWIBANmod-97 checksum PII · LOWPhoneNANP pattern PII · LOWIPv4strict dotted quad
Each phrase is checked against STA placeholders first, then card candidates must pass Luhn, issuer range and context scoring. Examples are masked.
Read this diagram as text

Two bands show the payment card detector as six gates in sequence, then the other PCI and PII detectors with their signal level.

  1. Top band, "Payment card detector · PCI: six gates before anything is called a card number", runs left to right.
  2. Gate 1, Phrase: the phrase plus its neighbours. Gate 2, STA first: skip anything inside a placeholder such as [card number]. Gate 3, Digit run: 12 to 19 digits, with spaces or dashes.
  3. Gate 4, Luhn: the checksum must pass. Gate 5, Issuer: issuer range and length must match. Gate 6, Context: payment words raise the score, order and tracking words lower it.
  4. A masked example, "my card is" followed by sixteen masked digits outside any placeholder, passes every gate and becomes a high-signal PCI finding.
  5. Bottom band, "Other detectors": high signal flags a conversation, low signal is shown only on request.
  6. PCI detectors: CVV or CVC after a keyword (high), and card expiry as MM/YY or MM/YYYY (low).
  7. PII detectors: US SSN in dashed form only (high), email as name@domain.tld (high), IBAN with mod-97 checksum (low), phone by NANP pattern (low), and IPv4 as a strict dotted quad (low).

05

The voice blind spot, stated plainly

Here is the part we insisted the app says out loud. Regular expressions see structured values: digit runs, an @ sign, dashed numbers. Voice transcripts are speech-to-text, so a spoken email arrives as “name at example dot com” and a spoken card number as words. Apart from the spoken-card pass, pattern detectors are effectively blind to voice – and names, addresses and free-text personal data need the kind of language model that Speech and Text Analytics itself uses. The app does not pretend otherwise.

So the results screen carries an honest-zero banner. A scan that flags nothing tells you exactly what that means: nothing machine-detectable in the channels you scanned, never “no exposure”. Scans default to voice, and the banner reports how many digital conversations were skipped. In our own testing on a month of production data, almost every high-signal finding was an email address in a digital transcript. The advice in the guide follows: tick “Also scan email / chat / message” before both Harvest and Scan, and keep STA redaction as the primary control for spoken data.

Defence in depth and the voice blind spot On the left, three layers of control: prevention with Secure Pause and Secure Call Flows, validated as Level 1 PCI DSS compliant; best-effort Speech and Text Analytics redaction into placeholders; and the QVCCS PII/PCI Finder auditing what is left, which is not itself a PCI control. On the right, what pattern matching can see: in voice transcripts values arrive as words, so only spoken card digits are rebuilt and a spoken email is invisible; in digital transcripts masked emails and digit runs are structured and detectable. Names, addresses and free text need language models, and a zero result means nothing machine-detectable in the channels scanned. LAYERS OF CONTROL Prevent, redact, then audit 1 · PREVENT CAPTURE · NATIVE Secure Pause · Secure Call Flows QSA-validated as Level 1 PCI DSS compliant WHAT STILL REACHES A TRANSCRIPT 2 · REDACT · NATIVE STA Sensitive data redaction Best effort · bracketed placeholders · audio plays as silence WHAT REDACTION MISSED 3 · AUDIT · QVCCS PII/PCI FINDER Finds the leftovers it can see An audit aid, not a PCI DSS control WHAT PATTERNS CAN SEE Voice writes words VOICE · SPEECH-TO-TEXT DIGITAL · CHAT · EMAIL “four five six … double seven” Spoken-card pass rebuilds “name at example dot com” Words, no @: invisible •••••@•••••.com Structured: detectable •••• •••• •••• •••• Digit run: gated and scored Names, addresses and free text need language models Zero flagged = nothing machine-detectable, not “no exposure”
Defence in depth: Secure Pause and secure flows prevent capture, STA redaction masks transcripts, and the PII/PCI Finder audits what is left.
Read this diagram as text

Two panels show three stacked layers of control for sensitive data, and what pattern matching can and cannot see in voice and digital transcripts.

  1. Left panel, "Layers of control: prevent, redact, then audit", runs top to bottom.
  2. Layer 1, "Prevent capture · native": Secure Pause and Secure Call Flows, QSA-validated as Level 1 PCI DSS compliant.
  3. What still reaches a transcript passes to layer 2, "Redact · native STA": sensitive data redaction, best effort, using bracketed placeholders, with audio played as silence.
  4. What redaction missed passes to layer 3, "Audit · QVCCS PII/PCI Finder": it finds the leftovers it can see and is an audit aid, not a PCI DSS control.
  5. Right panel, "What patterns can see: voice writes words", compares voice speech-to-text with digital chat and email.
  6. In voice, spoken digits such as "four five six … double seven" are rebuilt by the spoken-card pass, but a spoken email such as "name at example dot com" has words and no @, so it is invisible.
  7. In digital channels, a masked email is structured and detectable, and a masked digit run is gated and scored.
  8. Notes say names, addresses and free text need language models, and zero flagged means nothing machine-detectable, not "no exposure".

06

Permission posture decides what “exposed” means

One design decision matters more than any regex. The OAuth client the app uses should not hold View Sensitive Data. A client without it receives transcripts as an ordinary user sees them, so every finding is a real exposure gap. A client with it receives unmasked text, and everything STA would normally hide gets flagged. Because the cache is keyed by conversation, the guide asks you to use the same client for every harvest so masked and unmasked text never mix. A Redacted by STA scan mode inverts the audit to show what Genesys caught, by placeholder label.

The cache itself is sensitive – that is the point of the tool – so it is created with owner-only file permissions, written atomically, and should live on an encrypted volume under a retention policy; days can be deleted from the calendar at any time, and Genesys is never modified. Every Genesys call is a read, credentials stay server-side, and transcripts render as text nodes, never injected as HTML. In live streaming mode, every 30 seconds the app re-checks a window that overlaps the previous one by 10 minutes, and retries transcripts that are not ready for up to 30 minutes, for continuous monitoring.

This is the kind of build our method exists for. A Senior Business Consultant captured the auditor’s questions as acceptance criteria; a Solution Architect set the security and data-protection design in the Design stage; a Senior Developer built the detection engine; and a Systems Integration Tester ran negative tests against order numbers, tracking codes and placeholder-adjacent text before a design authority review. Findings export to XLS or a formatted PDF report, ready for the review on Thursday.

How it compares

Native controls and the QVCCS PII/PCI Finder

The native features prevent and redact. The QVCCS app audits the result. They are layers, not alternatives.

AspectNative Genesys Cloud CXQVCCS PII/PCI Finder
Preventing captureSecure Pause and Secure Call Flows, validated by a QSA as Level 1 PCI DSS compliant.Does not prevent capture; audits transcripts afterwards and never replaces these controls.
RedactionBest-effort STA redaction of payment card and personal information entities into bracketed placeholders, when analytics is enabled.Recognises those placeholders, counts them, and flags detector matches found outside them.
Searching transcriptsKeyword search within a transcript in the interaction view.Scans every phrase of every harvested conversation in a date range and lists only those with findings.
Card detectionCard number, expiry and CVV entities detected by STA.12–19 digits, Luhn checksum, real issuer range and length, context scoring, spoken-digit reconstruction.
Who sees unredacted dataUsers with View Sensitive Data, which no role includes by default.Recommends an OAuth client without it, so findings are true exposure gaps.
Coverage gapsRedaction applies only where speech or text analytics is enabled; national IDs for Australia, Canada, the UK and the US.Conversations with no transcript are tallied as an unscannable blind spot; an honest-zero banner explains empty results. No detectors for UK National Insurance numbers, sort codes or bank account numbers.
LicensingAutomatic redaction lists Genesys Cloud CX licence options plus the Speech and Text Analytics Upgrade Add-On.Reads STA transcripts, so it relies on STA being in place.

Native behaviour is as described in the Genesys Cloud Resource Center at the time of writing. The QVCCS app is an audit aid; it is not a PCI DSS control.

The takeaways

  • Evidence, not assumption: every harvested conversation scanned, only genuine findings listed.
  • Checksum- and issuer-gated card detection with context scoring keeps order numbers out of the results.
  • STA-aware: Genesys placeholders are recognised, counted and never miscounted as exposure.
  • Candid about limits: voice is a blind spot for pattern matching, and UK National Insurance numbers, sort codes and bank account numbers are not detected.
  • Read-only against Genesys, with a sensitive cache you control and can purge.

PII/PCI Finder is part of the QVCCS App Suite, included with every Managed Professional Services tier and built by the same certified team that designs, builds and supports Genesys Cloud CX solutions.

Read the user guide Managed Professional Services

Last reviewed

Questions about what you have read?

Clients, partners and people introduced to us can reach the specialists behind our applications and articles directly.

Who to contact