Viewing asSRE / PlatformData EngineeringGCP / CloudGo / Backend

Open to remote roles · Tunis, UTC+1 · EU work authorization

I run GCP workloads reliably and at half the cost.

Cloud and SRE engineer who operates production on Google Cloud: Pub/Sub, Cloud Run, Cloud Functions and BigQuery. I led a 50% cost reduction and built the FinOps model that turns cloud spend into an engineering metric.

50%GCP cost reduction
99.9%uptime on GCP and Nomad
65%lower mean time to resolution
20xgrowth the new architecture absorbs

01 Experience

Jan 2026 — Present

Technical Project Lead

The Quantic Factory · Paris, remote

  • Own reliability and delivery for 5+ production data workstreams: CRM, tracking, monitoring, attribution and identity resolution.
  • Led a company-wide FinOps and data architecture redesign across GCP, BigQuery, MySQL, ClickHouse, Hetzner and Scaleway, with a roadmap for ~20x ingestion growth without proportional cost.
  • Built post-production validation in SQL and ClickHouse that catches event-volume anomalies and identifier gaps before customers notice.
Jan 2024 — Jan 2026

DataOps & Site Reliability Engineer

The Quantic Factory · Paris, remote

  • On-call and overnight pager duty: resolved live incidents such as MySQL crashes and pipeline bugs by rollback or in-incident debugging.
  • Built the observability stack (Fluentd, Elasticsearch, Graylog, Grafana, Prometheus) across 100+ daily jobs, cutting MTTR by 65%.
  • Migrated critical infrastructure from MySQL to ClickHouse: 100x faster queries on 10M+ records.
  • Led a team optimizing GCP infrastructure for a 50% cost reduction.
  • Shipped a Python AI agent to production with 85% prediction accuracy.

02 Selected work

Client and company code is private, so these are described by problem and outcome.

GCP cost reduction

Problem: Cloud spend was growing faster than traffic.

What I did: Found obsolete and cost-heavy services to replace or remove, and rewrote hot production paths in Go to save memory.

50% lower GCP cost.

GCPCloud RunPub/SubGo

Infrastructure cost & data architecture redesign

Problem: Redundant compute, storage and transfer layers, with ~20x event growth ahead.

What I did: Mapped data flows end to end, built a cost model around a 'cost per million events' KPI, and benchmarked direct ClickHouse ingestion, managed queues, self-hosted HA Nomad, optimized MySQL and DuckLake/Parquet on object storage.

A prioritized modernization roadmap built to absorb ~20x ingestion without linear cost growth.

GCPBigQueryClickHouseFinOpsDuckLake

MySQL to ClickHouse migration

Problem: Analytics queries on 10M+ records were too slow for real-time diagnostics.

What I did: Designed a two-node ClickHouse cluster and migrated from MySQL with a blue-green strategy and no downtime.

100x faster queries and real-time diagnostics.

ClickHouseMySQLMigrationSQL

Production incident detection & recovery

Problem: Recurring pipeline failures were being found by customers, not by us.

What I did: Centralized monitoring over infrastructure, application and data-quality signals, plus restart, rollback and database-recovery runbooks with post-recovery validation.

Recurring failures turned into alerts and repeatable procedures; MTTR down 65%.

GrafanaPrometheusGraylogNomad

Multi-agent AI orchestrator

Problem: Shopify merchants needed analytical answers beyond what standard tools offer.

What I did: Designed and built an orchestrator from scratch that processes complex user queries, with OpenTelemetry tracing across the whole system.

Insight queries standard tooling could not answer, fully observable in production.

PythonOpenTelemetryLLM agents

CI/CD pipeline for Spring Boot & Angular

Problem: One-hour manual releases.

What I did: Jenkins, SonarQube and Git pipeline with quality gates and test-driven development.

15-minute automated deploys, 85% code coverage.

JenkinsSonarQubeDocker

03 Public code

More on GitHub

04 Why I work well remotely

I already do it, across a border.

I work remotely from Tunis with managers and a team based in Paris. Tunisia is on UTC+1 all year, so I overlap France's working day almost entirely.

Carrying a pager from a distance.

I took on-call and overnight pager duty for production systems, so I know that remote reliability means alerts, runbooks and clear handoffs, not being in the room.

Written by default.

Post-mortems, runbooks, architecture decision notes and benchmark reports. I leave a trail someone else can pick up without a call.

Trusted with ownership.

Promoted to Technical Project Lead to run five-plus workstreams and coordinate priorities and deadlines, all remotely.

Easy to hire in Europe.

Italian citizenship: no visa or sponsorship needed to work in the EU.

05 Toolbox

Cloud & infrastructure

GCP (Pub/Sub, Cloud Run, Cloud Functions, BigQuery, IAM, Cloud Monitoring) · Azure · Docker · Kubernetes · HashiCorp Nomad · Ansible

Reliability & observability

On-call · Incident response · Runbooks · Fluentd · Elasticsearch · Graylog · Grafana · Prometheus · OpenTelemetry

Data

ClickHouse · BigQuery · MySQL · PostgreSQL · DuckDB / DuckLake · Parquet & object storage · Message queues · ETL · Data quality

Code

Go · Python · SQL · Bash · Rust · JavaScript · Java

Delivery

Jenkins · GitHub Actions · SonarQube · Git · TDD

Languages: English (fluent) · French (fluent) · Arabic (native)
Education: Engineering degree in Computer Science (Cloud Computing & IT Architecture), Esprit School of Engineering, 2024 · Azure Fundamentals, 2023

Let's talk

Looking for a remote cloud engineer (gcp) role with a European team.

khalilrezgui0@gmail.com