Client and company code is private, so these are described by problem and outcome.
MySQL to ClickHouse migration
Problem: Analytics queries on 10M+ records were too slow for real-time diagnostics.
What I did: Designed a two-node ClickHouse cluster and migrated from MySQL with a blue-green strategy and no downtime.
100x faster queries and real-time diagnostics.
ClickHouseMySQLMigrationSQL
Infrastructure cost & data architecture redesign
Problem: Redundant compute, storage and transfer layers, with ~20x event growth ahead.
What I did: Mapped data flows end to end, built a cost model around a 'cost per million events' KPI, and benchmarked direct ClickHouse ingestion, managed queues, self-hosted HA Nomad, optimized MySQL and DuckLake/Parquet on object storage.
A prioritized modernization roadmap built to absorb ~20x ingestion without linear cost growth.
GCPBigQueryClickHouseFinOpsDuckLake
Identity resolution & graph validation
Problem: No objective way to trust a production identity-resolution system.
What I did: Defined KPIs (identifiable and identified sessions, hard candidate and identification rates) and compared BFS and DSU (union-find) engines across 2-day to 7-month windows.
Verified output consistency before rollout.
GraphsSQLPythonData quality
Server-side event tracking in Go
Problem: Client-side tracking lost conversions and could not feed several ad platforms reliably.
What I did: Built an event collection service in Go, containerized and run on Nomad, that processes events and routes them to multiple destinations such as Meta and Google Ads.
The foundation of the company's tracking and attribution pipeline.
GoDockerNomadETL
Multi-agent AI orchestrator
Problem: Shopify merchants needed analytical answers beyond what standard tools offer.
What I did: Designed and built an orchestrator from scratch that processes complex user queries, with OpenTelemetry tracing across the whole system.
Insight queries standard tooling could not answer, fully observable in production.
PythonOpenTelemetryLLM agents
Production incident detection & recovery
Problem: Recurring pipeline failures were being found by customers, not by us.
What I did: Centralized monitoring over infrastructure, application and data-quality signals, plus restart, rollback and database-recovery runbooks with post-recovery validation.
Recurring failures turned into alerts and repeatable procedures; MTTR down 65%.
GrafanaPrometheusGraylogNomad