Independent data platform consultant
Your Databricks bill is climbing and nobody can explain why. That's the work.
If your Spark job is slower than it should be, your Databricks bill keeps creeping, or your modern data platform is one promotion away from being a legacy monolith, I fix the parts that are actually costing you.
I tune production Spark, kill Databricks cost, and lead migrations off CDH, Teradata, and Hadoop onto Snowflake and Databricks. Based in Spain, working remotely.
What I fix
Apache Spark performance tuning
Jobs that run slower than they should, without anyone knowing why.
- Skew and shuffle pathologies after joins and aggregations
- Broadcast join thresholds and partition sizing
- Killing expensive Python UDFs and cache misuse
- Reading query plans to find the number that actually matters
Databricks cost optimization
A bill that keeps climbing and nobody can fully explain.
- Idle compute and all-purpose clusters running around the clock
- Job clusters, auto-termination, and spot instances with fallback
- Photon and autoscaling where they pay for themselves, not everywhere
- Delta table maintenance and small-file cleanup
Platform migrations
Moving off CDH, Teradata, or Hadoop without breaking trust.
- Migrations onto Snowflake and Databricks
- Validation-first approach so the numbers provably match
- The boring patterns that survive the next platform change
- De-risking the cutover before it reaches production
Platform architecture and reliability
A modern platform one promotion away from being a legacy monolith.
- Data contracts, SLAs, and pipeline trust
- Delta and Iceberg maintenance, retention, and metadata hygiene
- Catching silent degradation before a missed deadline finds it
- Boring, durable design over clever, fragile design
Recent writing
All writing →Your Iceberg metadata directory is growing faster than your data
Every commit writes a snapshot and nothing cleans them up by default. Streaming ingest can leave tens of thousands of files in one directory.
Read →You can be productive in Spark for months before you understand what it is doing
I showed a mentee a query plan for the first time. He asked what an Exchange was. Forty minutes later we were still on that one question.
Read →Your Delta table is Iceberg-compatible. Your readers just don't know it yet
Iceberg compatibility on a Delta table is not instant. The metadata sync lags, and downstream readers can silently consume a stale snapshot.
Read →Got a pipeline that costs too much or runs too slow?
That is the work. Tell me the symptom - a Databricks bill nobody can explain, a Spark job that keeps creeping, a migration you are not sure will land. If it is not something I can help with, I will tell you straight.