Case Studies · Travel Tech

Streaming analytics for 3.3 billion records a month.

A global hotel-booking SaaS connecting hotels and online travel aggregators was ingesting up to half a gigabyte every 30 seconds — overwhelming its database with IOPS and connection drops. We replaced it with a near real-time, pay-per-use streaming pipeline on AWS.

Industry
Travel Tech / Hospitality
Technology
Kinesis · Glue · Athena
Duration
4 months · 4-person team
Scale
3.3B records / month
The opportunity

A database buckling under real-time load.

The client's interface enables two-way data exchange between hotels and OTAs, with ingestion ranging from a few megabytes to half a gigabyte every 30 seconds — roughly 3.3 billion records a month. The resulting IOPS and frequent connection disruptions were hammering database performance and cost.

Extreme ingestion

Up to 0.5 GB every 30 seconds — ~3.3 billion records monthly — far beyond what the existing database was built to absorb.

High IOPS & drops

Exceptionally high input/output operations and frequent connection disruptions degraded performance across the application.

Runaway cost

Those same pressures drove database costs up, with capacity provisioned for peaks that sat idle the rest of the time.

The solution

A near real-time streaming pipeline, billed by usage.

We built a near real-time solution on AWS using cloud-native services that expands effortlessly as data volume grows and is billed solely on actual usage — no charges for idle capacity. The application streams into a data lake via Kinesis Firehose, with transactional tables kept fresh from SQL Server and Spark-based reporting jobs on Glue and Athena.

  • Five-minute data-availability window after ingestion
  • Reverse-ETL back into client applications, plus real-time API integration
  • Threshold alerting so stakeholders hear about breaches immediately
stream — firehose monitor
$ devotica stream status --live
ingest 512 MB / 30s · firehose
datalake partitioned · s3
availability < 5 min
report job 25M rows · 4m 52s
spend (mo) < $1,000
# resilient to spikes · no record throttling
Solution architecture

From datastream to dashboard.

Ingestion

Backend ApplicationHotels ⇄ OTAs
Kinesis FirehoseStreaming datastream
SQL ServerDimension tables

Data Lake & Processing

Amazon S3Data lake
AWS GlueSpark reporting jobs
Lake FormationSecurity

Analytics, Alerting & Reverse-ETL

Amazon AthenaAnalytics
Amazon QuickSightBI dashboards
Zabbix · SendGridThreshold alerts
API / Reverse-ETLBack into apps

Provisioned with Terraform · CI/CD via Jenkins · one-click deployment · validated to 5× current volume (~100M)

Highlights

Resilient, real-time and cost-efficient.

<5 min
Data availability after ingestion
<$1K
Monthly cloud spend
Headroom validated (~100M)
25M
Rows reported in <5 min

Resilient to spikes

The pipeline absorbs sudden volume surges with no data loss or record throttling.

Seamless deployment

End-to-end automation enables effortless one-click deployment across environments.

Pay-per-use

Cost-optimized per service and billed on actual usage — keeping monthly spend below $1,000.

Technology stack
Amazon Kinesis FirehoseAWS GlueAmazon S3AWS LambdaAWS DMSAmazon AthenaAmazon QuickSightAWS Lake FormationZabbixSendGridTerraformJenkins
Your real-time data

Is a database the wrong tool for your firehose?

When ingestion outgrows your database, a streaming lake bills by usage and scales on demand. We'll design one that fits your volume — and your budget.

Book a consultation →