Technical guide
Building a data lake or data warehouse on Genesys Cloud CX
Genesys Cloud CX offers more than a dozen ways to get data out, from synchronous analytics queries to nightly jobs, event streams and a native lakehouse export. This guide sets out what each route returns and when to use it, a reference architecture for landing and modelling the data, how to handle rate limits and API responses gracefully, and the API families QVCCS works with across its App Suite.
In short
- Conversation details jobs remain the dependable source for complete history; the data lakehouse export path adds near-real-time Parquet files; EventBridge and notifications serve events as they happen.
- Land raw responses first, then normalise into conversations, participants, sessions, segments, metrics and attributes, replacing each conversation whole so late updates and duplicates are harmless.
- Design for the platform's rules: data availability dates, cursor paging, 429 responses with Retry-After, channel lifetimes, signed URLs that expire quickly and retention that changes with interaction age.
- Recordings, transcripts and participant data carry personal data, so minimise what you extract, control who can query it and make Genesys erasure and retention travel downstream.
- The same API families that feed a warehouse also expose trunks, SIP traces, flow executions, bot sessions, data actions, licences and configuration, which QVCCS uses across its App Suite.
01
What a Genesys Cloud data platform must deliver
Genesys Cloud CX is very good at running a contact centre and showing what is happening in it, but most organisations eventually want its data somewhere else: joined with CRM, order and billing data, kept for longer than in-platform views allow, or modelled for analysts and data scientists. That is a data engineering problem with contact-centre-specific traps. A single customer contact is a conversation with several participants, each with sessions made of timed segments, with metrics emitted at points along the way. Records keep changing after the customer hangs up, as wrap-up codes, evaluations and late segments arrive.
A good platform therefore needs four properties. It must be complete, so no conversation is silently missed. It must be correct, so totals reconcile with Genesys Cloud performance views. It must be timely enough for each use, which may mean minutes for operations and a day for board reporting. And it must be governed, because transcripts, recordings and participant data hold personal data. Every design decision in the rest of this guide serves one of those four properties.
02
Every route Genesys provides for getting data out
Genesys publishes a data integration decision grid that sorts its options by timeliness, sensitivity to message loss and volume. The table below extends that grid with the newer data lakehouse export path and the content and diagnostic APIs a warehouse often needs. Most production designs combine three routes: a batch route for completeness, a streaming or near-real-time route for freshness, and reference-data snapshots that turn identifiers into names.
Two newer options deserve particular attention. The Genesys Cloud data lakehouse export path, activated through a service order and enabled under Lakehouse Settings in Analytics Settings, publishes flat Parquet files every five to ten minutes for schemas such as conversations, segments, participant attributes, conversation and flow metrics, user presence, agent routing status, flow execution history and flow outcome events, each with documented primary keys. Files are retained for 72 hours, so a consumer must collect them continuously. Genesys describes a future lakehouse platform with query-ready governed tables. Separately, the Genesys Cloud Analytics add-on (A3S, Analytics-as-a-Service) is a fully hosted warehouse populated from the analytics APIs, for organisations that would rather buy than build.
| Route | Endpoints or feature | Freshness and limits | Best used for |
|---|---|---|---|
| Conversation detail query | POST /; GET / | Up to date; interval up to seven days without a filter or 32 days with one; start dates up to 558 days back; numbered pages | Recent conversations, look-ups and filling the latest day |
| Conversation details jobs | POST /, then / and /; / | Served from a data lake updated nightly; cursor paging, default 1,000 and maximum 10,000 per page | Backfill and the nightly warehouse load; includes participant attributes |
| User details query and jobs | /; / | Same query and job split as conversations | Presence and routing status timelines |
| Aggregate queries and jobs | / and / for conversations, users, flows, flow executions, actions, bots, evaluations, journeys and summaries | Metrics bucketed by granularity and grouped by dimensions | KPI series and reconciliation totals |
| Observations | /; users and flows observations | Current state only | Live tiles, not history |
| Notifications (WebSocket) | / and / | Channels last 24 hours; 1,000 topics per connection; events are not replayed after a drop | Responsive interfaces and conversation topics |
| Amazon EventBridge integration | Partner event source in your AWS account; up to ten integrations per organisation | Latency typically a few hundred milliseconds, sometimes seconds; retries for up to four days | Server-side streaming and analytics detail events |
| Data lakehouse export path | GET /; /; POST / | Parquet files every five to ten minutes; 72-hour window | Near-real-time bulk feed with keyed, flat schemas |
| Analytics add-on (A3S) | Genesys-hosted warehouse fed from the analytics APIs | Hosted service with SQL access and BI templates | Buying the warehouse rather than building it |
| Recording bulk export | AWS S3 recording bulk actions integration; / | By recording policy or on demand; two concurrent jobs per organisation | Recordings, screen recordings, attachments and metadata |
| Transcripts | GET …/ | Pre-signed URLs to transcript JSON | Text analytics and leakage scanning |
| Flow execution history | POST /; / | Off by default; ten-day retention; 200 instances per query | Step-level flow diagnostics |
| Configuration and reference data | /, /, /, /, plus the Authorization API for roles | Current state only | Dimensions, snapshotted to keep history |
| Audit, usage and operational events | /; /; / | Asynchronous query executions | Change history, API consumption and platform error events |
03
Access, regions, SDKs and configuration as code
Every extraction service should authenticate with its own OAuth client using the client credentials grant, assigned a role that holds only the permissions it needs, such as analytics:conversationDetail:view for detail data or analytics:datawarehouse:view for the lakehouse export, and scoped to the divisions it should see. Separate clients per pipeline make API usage attributable and let one feed be revoked without stopping the others. Secrets belong in your secrets manager. Each Genesys region has its own API host: an organisation in the London region (eu-west-2) calls api.euw2.pure.cloud, and a client must target the region where its organisation lives.
Genesys publishes Platform API SDKs for Python (the PureCloudPlatformClientV2 package), JavaScript (purecloud-platform-client-v2), Java, .NET and Go, plus a command-line tool for scripted calls. The SDKs map one-to-one onto the documented endpoints, which keeps code reviewable against the API reference. For configuration, CX as Code, the Genesys Terraform provider, can export and manage organisation resources declaratively, and Archy manages Architect flows as YAML. We use both for the reference-data side of a platform: an export gives a versioned, diffable snapshot of how the organisation was configured on a given day.
04
Batch extraction with conversation details jobs
The conversation details job is the workhorse of most warehouses. You submit the same query body as the synchronous endpoint, receive a job identifier, poll until the state is FULFILLED, then walk the results with a cursor until no cursor is returned. Two behaviours matter. Jobs are served from a data lake updated nightly, so the dataAvailabilityDate in each result, also exposed by the availability endpoint, tells you how far the data is complete; Genesys notes that if the nightly batch fails, availability does not move and the latest day must come from the query endpoint. And by default a job returns every conversation whose lifetime overlaps the interval, unlike queries, which match conversations by the UTC day they started; the startOfDayIntervalMatching flag switches jobs to query behaviour.
Jobs also return participant attributes, truncated to 1,024 characters per key and value, which the query endpoint omits, and they keep access to older data, whereas the detail query endpoint stops returning interactions older than about 1.5 years. Each poll and each results page is an API call, so the loop belongs inside the response handling described in the rate limits section below. The sample shows the job lifecycle with the Python SDK.
import os, time
import PureCloudPlatformClientV2 as gc
# London region; credentials come from your secrets store, never from code
gc.configuration.host = gc.PureCloudRegionHosts.eu_west_2.get_api_host()
client = gc.api_client.ApiClient().get_client_credentials_token(
os.environ["GENESYS_CLOUD_CLIENT_ID"], os.environ["GENESYS_CLOUD_CLIENT_SECRET"])
analytics = gc.AnalyticsApi(client)
# 1. Only request what the analytics data lake already holds
ready = analytics.get_analytics_conversations_details_jobs_availability().data_availability_date
print("complete up to", ready)
# 2. Submit the job: by default it returns every conversation whose lifetime overlaps the interval
query = gc.AsyncConversationQuery()
query.interval = "2026-10-07T00:00:00Z/2026-10-08T00:00:00Z"
job_id = analytics.post_analytics_conversations_details_jobs(query).job_id
# 3. Poll until the job is fulfilled
while True:
state = analytics.get_analytics_conversations_details_job(job_id).state
if state == "FULFILLED":
break
if state in ("FAILED", "CANCELLED", "EXPIRED"):
raise RuntimeError(f"job {job_id} ended {state}")
time.sleep(15)
# 4. Walk the cursor; a page may hold fewer rows than page_size, so stop only when no cursor returns
cursor = None
while True:
kwargs = {"page_size": 1000}
if cursor:
kwargs["cursor"] = cursor
page = analytics.get_analytics_conversations_details_job_results(job_id, **kwargs)
for conversation in page.conversations or []:
land_raw(conversation.to_dict()) # your writer: bronze, partitioned by load date
cursor = page.cursor
if not cursor:
break05
Near-real-time: lakehouse export, EventBridge and notifications
Where the nightly cycle is too slow, there are three options with different guarantees. The lakehouse export path is the closest to a native bulk feed: poll the metadata endpoint, track file identifiers you have processed, because Genesys keeps no record of what you downloaded, and fetch new files individually or in bulk. Signed download URLs are valid for only 15 seconds after creation, so request them just before downloading. Overlapping the dateStart of each poll with the previous one avoids gaps at the cost of a few repeat listings.
The Amazon EventBridge integration publishes the topics you choose to a partner event source in your own AWS account, where rules route events to targets such as Lambda, Kinesis, SQS or SNS, with dead-letter queues for failures. Genesys recommends it for server-side integrations, and analytics detail events are available only through EventBridge. WebSocket notifications suit responsive interfaces: channels last 24 hours, each user and application can hold 20 channels, and a connection carries up to 1,000 topics, which can be combined by prefix. Conversation events are available only through notifications. Neither stream replaces a batch reconciliation: events are best treated as early, provisional data.
GET /api/v2/analytics/dataextraction/downloads/metadata?dataSchema=segments&dateStart=2026-10-09T06:00:00Z&pageSize=200
# Response (abridged): one entry per Parquet file, with a stable id and a 72-hour expiry
# { "entities": [ { "id": "…", "dataSchema": "segments", "dateCreated": "…", "dateExpires": "…" } ],
# "nextUri": "…", "enabledDataSchemas": [ "conversations", "segments", "participant-attributes", … ] }
POST /api/v2/analytics/dataextraction/downloads/bulk
{ "files": [ "<file id>", "<file id>" ] }
# Returns a signed URL per file: start each download within its short validity window[
{ "id": "v2.telephony.providers.edges.trunks.<trunkId>.metrics" },
{ "id": "v2.users.<userId>?presence&routingStatus" },
{ "id": "v2.routing.queues.<queueId>.conversations" }
]06
Rate limits and graceful handling of API responses
Genesys Cloud protects the platform with limits that can apply per access token, per user, per application or per organisation. Each limit belongs to a namespace and is documented on the Developer Center Limits page, and API Explorer shows the limits that apply to each operation. For a data platform the ones that bite are 300 requests per minute for each user-based token, a separate configurable allowance for client credentials tokens, 300 token creations per minute, and the analytics limits: a details query page size of 100, 50 concurrent details queries, and five aggregates job requests per second. An organisation's effective values are readable through GET /api/v2/organizations/limits/namespaces/{namespaceName}, and Customer Care can consider a case for raising a configurable limit; some limits are hard and cannot change.
Graceful handling starts with treating every status code as a decision, not an error to print. A 429 identifies a breached limit: Genesys returns a Retry-After header in seconds and, in most payloads, a limit object naming the limiter, and nothing more should be sent on that client until the wait has passed. A 502, 503 or 504 is retryable on Genesys's regimented schedule of 3 seconds, rising to 9 after five minutes and 27 after ten, deliberately ignoring Retry-After; Genesys says retryable conditions lasting longer than thirty seconds are unexpected and should be reported. A 400, 403 or 404 is about the request itself and is never retried blindly. Error bodies carry status, code, message and contextId, and every response carries an ININ-Correlation-Id header, which we log against each failure because Customer Care can trace it.
The Java and .NET SDKs include an optional retry configuration, off by default, that follows Genesys's backoff logic. The Python and JavaScript SDKs document none, so we wrap every call, as the sample shows. Retrying is the last line of defence rather than the control mechanism: each of our pipelines also throttles itself below its limits with a client-side token bucket, staggers job submissions and spreads reference-data refreshes, so a 429 is an event worth investigating rather than part of normal running.
| Response | What Genesys means | What our clients do |
|---|---|---|
| 200, 201, 202 | Success; 202 means accepted for asynchronous processing | Process the body; for jobs and query executions, poll the returned identifier |
| 400 | The request is malformed or breaks a rule, such as a missing field or an interval that is too long | Log status, code, message and contextId; fix the request; never retry unchanged |
| 401 | The token is missing, expired or revoked | Acquire a new token once and repeat the call; alert if the new token is refused too |
| 403 | The OAuth client's role or divisions do not allow the operation | Correct the role or division scope; never retry |
| 404 | The resource does not exist, or is not visible to this client | Record and skip, for example a deleted user still referenced in history |
| 429 | A rate limit was breached; Retry-After gives the wait in seconds | Pause that client for the Retry-After period, log the limit object, then resume |
| 502, 503, 504 | A transient platform condition | Retry at 3, 9 then 27 seconds; raise an incident if it persists |
import logging, time
from PureCloudPlatformClientV2.rest import ApiException
log = logging.getLogger("genesys")
def backoff(elapsed):
# Genesys's schedule: 3 s for five minutes, then 9 s, then 27 s after ten minutes
return 3 if elapsed < 300 else 9 if elapsed < 600 else 27
def call(fn, *args, give_up_after=900, **kwargs):
started = time.monotonic()
while True:
try:
return fn(*args, **kwargs)
except ApiException as e:
headers = e.headers or {}
elapsed = time.monotonic() - started
log.warning("genesys status=%s correlation=%s body=%s",
e.status, headers.get("ININ-Correlation-Id"), e.body)
if e.status == 429 and headers.get("Retry-After"):
wait = int(headers["Retry-After"])
elif e.status in (502, 503, 504):
wait = backoff(elapsed)
else:
raise # 400, 403, 404: fix the request; 401: refresh the token upstream
if elapsed + wait > give_up_after:
raise
time.sleep(wait)
# Usage, inside the cursor loop shown earlier:
# page = call(analytics.get_analytics_conversations_details_job_results,
# job_id, cursor=cursor, page_size=1000)07
Reference architecture: bronze, silver and gold
We keep the warehouse itself platform-neutral, because the Genesys-side design matters more than the product underneath. The pattern that works is a layered one, with raw data preserved so that any model can be rebuilt. Incremental extraction runs on overlapping windows, for example re-reading the last few days each night, and every load is idempotent: a conversation is replaced as a whole, keyed by conversation ID, so a late wrap-up, an extra segment or a duplicate event never produces double counts. Participant attributes are best stored as key-value rows rather than columns, because keys appear and disappear as flows change; schema drift then becomes something to report on, not a failed load.
Completeness is proved rather than assumed. Each window is reconciled against aggregate queries for the same interval, conversation counts and key metrics are compared with Genesys Cloud performance views on sample days, and pipeline monitoring checks row volumes as well as job success, because a job that succeeds with half the expected rows is worse than one that fails. Time is stored in UTC with local-time conversion at the gold layer. Dimensions keep history, so a renamed queue does not rewrite last year's reports.
- Bronze: raw JSON responses, Parquet files and events, landed unchanged in object storage, partitioned by load date and source, with the request that produced them.
- Silver: normalised tables for conversations, participants, sessions, segments, session metrics, participant attributes, flow outcomes, user status intervals and evaluations, keyed by Genesys identifiers.
- Reference snapshots: users, queues, skills, wrap-up codes, flows, divisions, sites, trunks and DIDs captured daily, giving slowly changing dimensions.
- Gold: modelled marts and KPIs defined once in a metrics dictionary, in a columnar warehouse, for BI tools and data science.
- Control tables: extract windows, job identifiers, file identifiers processed, availability dates and reconciliation results.
-- Silver load: replace each conversation as a whole, never patch individual rows.
-- staging.segments_flat holds the latest extract, one row per segment.
BEGIN;
DELETE FROM silver.segments
WHERE conversation_id IN (SELECT DISTINCT conversation_id FROM staging.segments_flat);
INSERT INTO silver.segments (
conversation_id, participant_id, session_id, purpose, media_type,
segment_type, segment_start, segment_end, queue_id, wrap_up_code,
disconnect_type, extract_batch_id)
SELECT conversation_id, participant_id, session_id, purpose, media_type,
segment_type, segment_start, segment_end, queue_id, wrap_up_code,
disconnect_type, extract_batch_id
FROM staging.segments_flat;
COMMIT;
-- Gold: talk time per queue and day, joined to the queue name as it was at the time
SELECT d.queue_name, CAST(m.emit_date AS DATE) AS day,
COUNT(DISTINCT m.conversation_id) AS conversations,
SUM(m.metric_value) / 1000.0 AS talk_seconds
FROM silver.session_metrics m
JOIN (SELECT DISTINCT conversation_id, session_id, queue_id
FROM silver.segments WHERE queue_id IS NOT NULL) s
ON s.conversation_id = m.conversation_id AND s.session_id = m.session_id
JOIN gold.dim_queue d
ON d.queue_id = s.queue_id
AND m.emit_date >= d.valid_from AND m.emit_date < d.valid_to
WHERE m.metric_name = 'tTalk'
GROUP BY d.queue_name, CAST(m.emit_date AS DATE);08
Recordings, transcripts and personal data
Content is where data platforms create risk. The AWS S3 recording bulk actions integration exports recordings, screen recordings, attachments and metadata to your bucket, either automatically through a recording policy or on demand through the recording bulk job API, which runs at most two concurrent jobs per organisation. Accessing recordings that contain PCI DSS or PII data where redaction is enabled requires the recording:recording:viewSensitiveData permission. Speech and text analytics transcripts are fetched through pre-signed transcript URLs and include words with timings plus sentiment and topic events.
Genesys can redact payment card and personal information from transcripts and recordings when speech or text analytics processes the interaction, but describes redaction as best effort and recommends Secure Pause or secure call flows as the primary PCI DSS controls. That is why QVCCS built PII/PCI Finder, which scans transcripts for card numbers and personal data that slipped through. In the warehouse, apply data minimisation under UK GDPR: extract only the attributes a use case needs, tokenise identifiers, restrict who can query content, and make retention and erasure follow Genesys, including requests made through the GDPR API.
09
The art of the possible across the App Suite
A warehouse is one consumer of Genesys Cloud data; the same API families support operational and diagnostic tools. Over many months QVCCS has built its App Suite against them, from SIP trunk health to transcript scanning and licence attribution, and each app has taught us something about limits, edge cases and data shape that now informs our pipeline designs. The table maps each data domain to the Genesys API family and to the QVCCS apps that work in it. Endpoints are as Genesys documents them; where the table names a topic, it is a notifications or EventBridge topic.
| Domain | Genesys API family | What it yields | QVCCS app |
|---|---|---|---|
| SIP trunks | / and /; topic v2. | Trunk state and live metrics per trunk | Trunk Monitor |
| SIP traces | GET /; POST / | SIP message metadata by call or conversation, and PCAP files on request | SIP Trace Analyser |
| Live conversations | Topic v2.; / | Conversations in progress and queue state | Live Call Map |
| Conversation detail | / and /; / | Participants, sessions, segments, metrics and media quality such as MOS | Conversation Detail, Conversation Analyser, Voice Analysis |
| Participant data | Conversation details jobs, participants[].attributes | Attribute keys, values, trends and schema drift | Participant Data Analytics |
| Queues and membership | / and / | Queue configuration and members | Config Validator, CX Pulse |
| Users and presence | Topics v2. and .routingStatus; / | Presence and routing status as they change | CX Pulse |
| Flow outcomes and milestones | /; /; / | Outcome success and failure, milestone counts, IVR paths | IVR Sankey, Journey Analyser |
| Flow execution | / and / | Ordered step trace of each flow execution | Flow Journey, Conversation Analyser |
| Bot flows | / and /; / | Sessions, turns, intents and utterances | Bot Flow Diagnostics |
| Data actions | /; / | Action inventory and execution counts and durations | Data Actions Diagnostics |
| Transcripts | …/ | Transcript JSON with timings, sentiment and topics | PII/ Finder |
| PII/PCI leakage | Detail query to select conversations, then transcripts | Unredacted card numbers and personal data in text | PII/ Finder |
| Recordings | /; / | Media and recording metadata | Download Recordings |
| Configuration and as-built | /, /, /, /; / | Inventory, validation and flow documentation | Config Validator, Flow Mapper, Flow Narrator |
| Permissions and licences | Authorization API roles and permissions endpoints; / and / | Who holds which permission and licence, and why | Permissions Auditor, Licence Optimiser |
| DIDs and numbers | / and /; / | Number-to-flow mapping and conflicts | DID Conflict Checker, Number Inventory |
| Knowledge | / | Knowledge base inventory and snapshots | Knowledge Manager |
| IP ranges | GET / | Published address ranges by region and service | IP Ranges |
| Audit | /; topic v2. | Who changed what, and when | Roles Audit, Role Compliance |
10
Genesys Cloud API best practices we apply
Handling responses well is half of using the Platform API responsibly; the other half is not making calls you do not need. The practices below come from Genesys's own Rate Limits, API Tips and Change Management pages, and from building and running the QVCCS App Suite against production organisations. None is exotic, but together they decide whether an integration runs quietly for years or becomes the reason an organisation hits its limits at the busiest hour of the week.
Genesys's Change Management Policy says breaking changes are generally announced 60 days ahead, with 30 days as the minimum, and resource deprecations at least 90 days ahead; a reduced rate limit counts as a breaking change. Deprecated routes are flagged in API Explorer and changes appear on the Developer Center announcements calendar. We review that calendar against every integration we support, and the results feed the release impact assessments in our support model.
- One OAuth client per integration, with a least-privilege role scoped to the divisions it needs, and its secret held in a secrets manager and rotated.
- Reuse each access token until it expires or a 401 says otherwise; never request a token per call, since token creation is itself limited.
- Call the API host for the region where the organisation lives, as Genesys requires.
- Prefer push to polling: notifications or EventBridge for presence, conversation state and anything else that changes often, as Genesys's rate limit guidance advises.
- Prefer bulk and asynchronous endpoints: details and aggregates jobs, user search, fetching users by ID in bulk, and recording jobs.
- Cache reference data such as users, queues, skills and wrap-up codes, refreshing it on change events or on a schedule.
- Page deliberately: the default page size is 25, so request the largest page each endpoint allows and follow nextUri or the cursor to the end.
- Treat expand parameters as best effort, as Genesys documents, and use the specific endpoint when the expanded data is essential.
- Log the ININ-Correlation-Id of every failed call, and measure consumption per OAuth client with the Usage API, because API requests count against the organisation's monthly fair-use allocation.
- Test every change against a non-production organisation, and keep configuration in version control with CX as Code.
11
How QVCCS approaches a Genesys Cloud data platform
We start in Discovery with the questions the data must answer, the freshness each needs, volumes, retention and the security constraints, and record the choice of route for each dataset in architecture decision records. The Interface Control Document for each feed fixes fields, masking, schedules, identifiers, error handling and ownership, and a metrics dictionary defines every KPI against its Genesys metric. A Solution Architect owns the design, a Business Analyst works with your analysts on definitions, and a Senior Developer and Software Developers build the extractors, streams and transformations as peer-reviewed code under the engineering standards of our Senior Platform Practice Lead.
Our Systems Integration Tester proves completeness by reconciling warehouse figures with Genesys Cloud views across sample days, and by introducing throttling, expired tokens, failed jobs, late updates and duplicate events on purpose. We muster that team from our own bench, and it stays backed by the whole practice. After go-live, support and maintenance are aligned to your Genesys Cloud CX consumption and support model: we work inside your provider's service delivery model, or provide SLA-based support, up to 24x7x365 where you contract it with us, including release impact assessments for API and schema changes.
Questions
Common questions
Should we use the data lakehouse export path or conversation details jobs?
Often both. The export path delivers flat Parquet files every five to ten minutes with documented primary keys, but files are kept for only 72 hours and it needs commercial activation. Conversation details jobs reach back across the full history and are the natural source for backfills and nightly reconciliation. A common design takes the export path for freshness and jobs for completeness, with both landing in the same bronze layer.
Why do conversations from yesterday appear in today's extract?
Because Genesys conversation records keep changing after they end, and because jobs by default return every conversation whose lifetime overlaps the interval, including ones that started earlier. Wrap-up codes, evaluations and late segments can arrive later too. Re-read an overlapping window and replace each conversation as a whole by its ID, and the repeats become harmless updates rather than duplicates.
Can we stream everything through EventBridge and skip batch extraction?
We advise against it. EventBridge gives reliable, retried delivery into AWS and is the only route for analytics detail events, but events describe moments in a conversation rather than the final, complete record, and conversation events are available only through notifications. Treat streams as early data for operations and alerting, and reconcile them against batch extraction for anything reported as fact.
How do we stop personal data spreading into the warehouse?
Decide per dataset which attributes are needed, and exclude or tokenise the rest at extraction. Keep transcripts and recordings in separately secured storage with narrower access than metrics. Apply the same retention as Genesys and propagate erasure requests. Remember that redaction is best effort, so scan transcript content for leakage, as PII/PCI Finder does, rather than assuming it is clean.
Sources
- Genesys Cloud data lakehouse (Genesys Cloud Resource Center)help.genesys.cloud
- Data Extraction API (Genesys Cloud Developer Center)developer.genesys.cloud
- Conversation Detail job (Genesys Cloud Developer Center)developer.genesys.cloud
- Analytics Integration Guide (Genesys Cloud Developer Center)developer.genesys.cloud
- Notifications overview (Genesys Cloud Developer Center)developer.genesys.cloud
- Amazon EventBridge integration (Genesys Cloud Developer Center)developer.genesys.cloud
- Rate limits (Genesys Cloud Developer Center)developer.genesys.cloud
- Retention period for analytics data and recording (Genesys Cloud Resource Center)help.genesys.cloud
- Limits (Genesys Cloud Developer Center)developer.genesys.cloud
- API tips and tricks (Genesys Cloud Developer Center)developer.genesys.cloud
- Platform API change management policy (Genesys Cloud Developer Center)developer.genesys.cloud
- Analytics Add-on overview (Genesys Cloud Resource Center)help.genesys.cloud