Client and company code is private, so these are described by problem and outcome.
Production incident detection & recovery
Problem: Recurring pipeline failures were being found by customers, not by us.
What I did: Centralized monitoring over infrastructure, application and data-quality signals, plus restart, rollback and database-recovery runbooks with post-recovery validation.
Recurring failures turned into alerts and repeatable procedures; MTTR down 65%.
GrafanaPrometheusGraylogNomad
Infrastructure cost & data architecture redesign
Problem: Redundant compute, storage and transfer layers, with ~20x event growth ahead.
What I did: Mapped data flows end to end, built a cost model around a 'cost per million events' KPI, and benchmarked direct ClickHouse ingestion, managed queues, self-hosted HA Nomad, optimized MySQL and DuckLake/Parquet on object storage.
A prioritized modernization roadmap built to absorb ~20x ingestion without linear cost growth.
GCPBigQueryClickHouseFinOpsDuckLake
MySQL to ClickHouse migration
Problem: Analytics queries on 10M+ records were too slow for real-time diagnostics.
What I did: Designed a two-node ClickHouse cluster and migrated from MySQL with a blue-green strategy and no downtime.
100x faster queries and real-time diagnostics.
ClickHouseMySQLMigrationSQL
GCP cost reduction
Problem: Cloud spend was growing faster than traffic.
What I did: Found obsolete and cost-heavy services to replace or remove, and rewrote hot production paths in Go to save memory.
50% lower GCP cost.
GCPCloud RunPub/SubGo
CI/CD pipeline for Spring Boot & Angular
Problem: One-hour manual releases.
What I did: Jenkins, SonarQube and Git pipeline with quality gates and test-driven development.
15-minute automated deploys, 85% code coverage.
JenkinsSonarQubeDocker