A provider of ethics-and-compliance programs to Fortune 500 employees — 500+ courses, 70+ languages, 30M+ learners a year — had its data trapped across Postgres, MongoDB and third-party systems. We consolidated it into a secure, GDPR-compliant AWS data lake and made learning effectiveness measurable.
The client delivers compliance programs through a white-labeled portal, tracking course completions in Postgres, portal activity in MongoDB, and employee feedback in third-party systems. Consolidating it — and evaluating learning effectiveness across 70+ languages — was the core challenge.
Course completions, portal activity and feedback lived in Postgres, MongoDB and third-party tools like Salesforce and QuestionPro — with no unified view.
30M+ learners a year generate enormous volumes of data; consolidating it efficiently meant cutting heavy manual processing.
The client needed to evaluate learning outcomes by analyzing feedback and completions — not just store the data.
With 70+ languages served, region-level data had to be managed through a single generalized structure.
Comparative visualization required securely sharing data across regions — while respecting data-residency rules.
As a compliance company itself, the client demanded a GDPR-compliant solution with tightly controlled access.
We implemented a streamlined migration and a data lake on AWS that handles complex, nested formats, upholds rigorous data-quality standards, and runs on frequent processing schedules — with a secure pipeline that consolidates datasets and enables cross-region comparison.
Data pulled from Postgres and MongoDB via AWS DMS, plus third-party systems like Salesforce and QuestionPro, then cleaned, translated and partitioned with AWS Glue ETL jobs over intricate, deeply nested structures.
A centralized S3 data lake with encrypted files and multi-format support, governed by Lake Formation for fine-grained access and authorized-user-only sharing.
Multiple QuickSight dashboards built and embedded directly into the product portals — secure, performant, multi-tenant, with language and sentiment analysis layered in.
GDPR via separate US & EU implementations — only anonymized data combined across regions for benchmarking · CI/CD with multi-level monitoring & validation
Across 70+ languages, the platform detects language and generates sentiment, providing an analytic description of each client's dashboards.
Frequent source-data changes are immediately reflected in the final dashboards, with end-to-end automation, CI/CD, and multi-level monitoring that alerts on any data, ETL or infrastructure failure.
Handles a high influx of data at speed — currently up to 4TB — with each component able to scale independently.
Separate implementations in the US and Europe, combining only anonymized data across regions for benchmarking purposes.
We build governed, GDPR-aware data lakes that unify your silos and surface insight in dashboards your teams can actually use. Let's map yours.
Book a consultation →