Data platform engineer, teaching what actually works

Your data platform bill is climbing and nobody can explain why. Learn to fix it.

I tune production Spark, cut cloud data platform cost, and lead migrations off CDH and Teradata. Now I teach the same diagnostic loops in video courses and a guided roadmap - so the next slow job or climbing bill is a diagnosis, not a guess.

Free guide

Not sure where to start? Follow the roadmap.

The Data Engineer Roadmap - the full path from your first SQL query to owning a platform you can defend on cost, reliability, and scale. Free and public, with the deep phases pointing at a course.

View the roadmap →

Recent writing

All writing →

Or work with me directly

How I work →

Apache Spark performance tuning

Jobs that run slower than they should, without anyone knowing why.

  • Skew and shuffle pathologies after joins and aggregations
  • Broadcast join thresholds and partition sizing
  • Killing expensive Python UDFs and cache misuse
  • Reading query plans to find the number that actually matters

Cloud platform cost optimization

A bill that keeps climbing and nobody can fully explain.

  • Idle compute and always-on clusters running around the clock
  • Ephemeral job clusters, auto-termination, and spot instances with fallback
  • Vectorized engines and autoscaling where they pay for themselves, not everywhere
  • Delta table maintenance and small-file cleanup

Platform migrations

Moving off CDH or Teradata without breaking trust.

  • Migrations onto modern cloud data platforms
  • Validation-first approach so the numbers provably match
  • The boring patterns that survive the next platform change
  • De-risking the cutover before it reaches production

Platform architecture and reliability

A modern platform one promotion away from being a legacy monolith.

  • Data contracts, SLAs, and pipeline trust
  • Delta and Iceberg maintenance, retention, and metadata hygiene
  • Catching silent degradation before a missed deadline finds it
  • Boring, durable design over clever, fragile design

Got a pipeline that costs too much or runs too slow?

That is the work. Tell me the symptom - a platform bill nobody can explain, a Spark job that keeps creeping, a migration you are not sure will land. If it is not something I can help with, I will tell you straight.