GCP cost reduction
Problem: Cloud spend was growing faster than traffic.
What I did: Found obsolete and cost-heavy services to replace or remove, and rewrote hot production paths in Go to save memory.
50% lower GCP cost.
Open to remote roles · Tunis, UTC+1 · EU work authorization
Cloud and SRE engineer who operates production on Google Cloud: Pub/Sub, Cloud Run, Cloud Functions and BigQuery. I led a 50% cost reduction and built the FinOps model that turns cloud spend into an engineering metric.
The Quantic Factory · Paris, remote
The Quantic Factory · Paris, remote
Client and company code is private, so these are described by problem and outcome.
Problem: Cloud spend was growing faster than traffic.
What I did: Found obsolete and cost-heavy services to replace or remove, and rewrote hot production paths in Go to save memory.
50% lower GCP cost.
Problem: Redundant compute, storage and transfer layers, with ~20x event growth ahead.
What I did: Mapped data flows end to end, built a cost model around a 'cost per million events' KPI, and benchmarked direct ClickHouse ingestion, managed queues, self-hosted HA Nomad, optimized MySQL and DuckLake/Parquet on object storage.
A prioritized modernization roadmap built to absorb ~20x ingestion without linear cost growth.
Problem: Analytics queries on 10M+ records were too slow for real-time diagnostics.
What I did: Designed a two-node ClickHouse cluster and migrated from MySQL with a blue-green strategy and no downtime.
100x faster queries and real-time diagnostics.
Problem: Recurring pipeline failures were being found by customers, not by us.
What I did: Centralized monitoring over infrastructure, application and data-quality signals, plus restart, rollback and database-recovery runbooks with post-recovery validation.
Recurring failures turned into alerts and repeatable procedures; MTTR down 65%.
Problem: Shopify merchants needed analytical answers beyond what standard tools offer.
What I did: Designed and built an orchestrator from scratch that processes complex user queries, with OpenTelemetry tracing across the whole system.
Insight queries standard tooling could not answer, fully observable in production.
Problem: One-hour manual releases.
What I did: Jenkins, SonarQube and Git pipeline with quality gates and test-driven development.
15-minute automated deploys, 85% code coverage.
I work remotely from Tunis with managers and a team based in Paris. Tunisia is on UTC+1 all year, so I overlap France's working day almost entirely.
I took on-call and overnight pager duty for production systems, so I know that remote reliability means alerts, runbooks and clear handoffs, not being in the room.
Post-mortems, runbooks, architecture decision notes and benchmark reports. I leave a trail someone else can pick up without a call.
Promoted to Technical Project Lead to run five-plus workstreams and coordinate priorities and deadlines, all remotely.
Italian citizenship: no visa or sponsorship needed to work in the EU.
GCP (Pub/Sub, Cloud Run, Cloud Functions, BigQuery, IAM, Cloud Monitoring) · Azure · Docker · Kubernetes · HashiCorp Nomad · Ansible
On-call · Incident response · Runbooks · Fluentd · Elasticsearch · Graylog · Grafana · Prometheus · OpenTelemetry
ClickHouse · BigQuery · MySQL · PostgreSQL · DuckDB / DuckLake · Parquet & object storage · Message queues · ETL · Data quality
Go · Python · SQL · Bash · Rust · JavaScript · Java
Jenkins · GitHub Actions · SonarQube · Git · TDD
Looking for a remote cloud engineer (gcp) role with a European team.
khalilrezgui0@gmail.com