Context

Sector: Commercial ticketing. Role: Software Engineer, DevOps. Environment: Mission-critical ticketing services subjected to massive concurrency bursts during on-sale events and multi-year migration backfills.

Challenge

  • Stabilize streaming and storage layers under burst-traffic load that overwhelmed the prior architecture.
  • Reduce scale-time and infrastructure cost without sacrificing reliability during on-sale events.
  • Diagnose and remediate cross-system bottlenecks across Kafka and downstream storage.

Architecture

Compute & Scaling

  • Migrated mission-critical ticketing services to lower-cost spot-instance compute with automated scale-out.
  • Reduced node scale-time from 15 minutes to approximately 2 minutes (largely consumer group rebalance).

Streaming & Storage

  • Diagnosed downstream DynamoDB write bottleneck causing Kafka Streams backpressure during on-sale bursts.
  • Prototyped Cassandra and ScyllaDB alternative storage backends; designs subsequently adopted by the team.

Outcomes

  • Sustained 42M messages per minute at peak during on-sale events and multi-year migration backfills.
  • Reduced node scale-time from 15 minutes to approximately 2 minutes.
  • Reduced infrastructure cost by 40% through migration to lower-cost spot-instance compute.
  • Eliminated downstream storage bottleneck through prototyped alternative storage architecture.