Viewing asSRE / PlatformData EngineeringGCP / CloudGo / Backend

Open to remote roles · Tunis, UTC+1 · EU work authorization

I keep production reliable, observable and cheap to run.

SRE and DataOps engineer with two years of on-call experience on GCP and HashiCorp Nomad. I cut incident resolution time, migrate databases without downtime, and design infrastructure that costs less per event.

99.9%uptime on GCP and Nomad
65%lower mean time to resolution
100xfaster queries after ClickHouse migration
50%lower GCP cost

01 Experience

Jan 2026 — Present

Technical Project Lead

The Quantic Factory · Paris, remote

  • Own reliability and delivery for 5+ production data workstreams: CRM, tracking, monitoring, attribution and identity resolution.
  • Led a company-wide FinOps and data architecture redesign across GCP, BigQuery, MySQL, ClickHouse, Hetzner and Scaleway, with a roadmap for ~20x ingestion growth without proportional cost.
  • Built post-production validation in SQL and ClickHouse that catches event-volume anomalies and identifier gaps before customers notice.
Jan 2024 — Jan 2026

DataOps & Site Reliability Engineer

The Quantic Factory · Paris, remote

  • On-call and overnight pager duty: resolved live incidents such as MySQL crashes and pipeline bugs by rollback or in-incident debugging.
  • Built the observability stack (Fluentd, Elasticsearch, Graylog, Grafana, Prometheus) across 100+ daily jobs, cutting MTTR by 65%.
  • Migrated critical infrastructure from MySQL to ClickHouse: 100x faster queries on 10M+ records.
  • Led a team optimizing GCP infrastructure for a 50% cost reduction.
  • Shipped a Python AI agent to production with 85% prediction accuracy.

02 Selected work

Client and company code is private, so these are described by problem and outcome.

Production incident detection & recovery

Problem: Recurring pipeline failures were being found by customers, not by us.

What I did: Centralized monitoring over infrastructure, application and data-quality signals, plus restart, rollback and database-recovery runbooks with post-recovery validation.

Recurring failures turned into alerts and repeatable procedures; MTTR down 65%.

GrafanaPrometheusGraylogNomad

Infrastructure cost & data architecture redesign

Problem: Redundant compute, storage and transfer layers, with ~20x event growth ahead.

What I did: Mapped data flows end to end, built a cost model around a 'cost per million events' KPI, and benchmarked direct ClickHouse ingestion, managed queues, self-hosted HA Nomad, optimized MySQL and DuckLake/Parquet on object storage.

A prioritized modernization roadmap built to absorb ~20x ingestion without linear cost growth.

GCPBigQueryClickHouseFinOpsDuckLake

MySQL to ClickHouse migration

Problem: Analytics queries on 10M+ records were too slow for real-time diagnostics.

What I did: Designed a two-node ClickHouse cluster and migrated from MySQL with a blue-green strategy and no downtime.

100x faster queries and real-time diagnostics.

ClickHouseMySQLMigrationSQL

GCP cost reduction

Problem: Cloud spend was growing faster than traffic.

What I did: Found obsolete and cost-heavy services to replace or remove, and rewrote hot production paths in Go to save memory.

50% lower GCP cost.

GCPCloud RunPub/SubGo

CI/CD pipeline for Spring Boot & Angular

Problem: One-hour manual releases.

What I did: Jenkins, SonarQube and Git pipeline with quality gates and test-driven development.

15-minute automated deploys, 85% code coverage.

JenkinsSonarQubeDocker

03 Public code

More on GitHub

04 Why I work well remotely

I already do it, across a border.

I work remotely from Tunis with managers and a team based in Paris. Tunisia is on UTC+1 all year, so I overlap France's working day almost entirely.

Carrying a pager from a distance.

I took on-call and overnight pager duty for production systems, so I know that remote reliability means alerts, runbooks and clear handoffs, not being in the room.

Written by default.

Post-mortems, runbooks, architecture decision notes and benchmark reports. I leave a trail someone else can pick up without a call.

Trusted with ownership.

Promoted to Technical Project Lead to run five-plus workstreams and coordinate priorities and deadlines, all remotely.

Easy to hire in Europe.

Italian citizenship: no visa or sponsorship needed to work in the EU.

05 Toolbox

Reliability & observability

On-call · Incident response · Runbooks · Fluentd · Elasticsearch · Graylog · Grafana · Prometheus · OpenTelemetry

Cloud & infrastructure

GCP (Pub/Sub, Cloud Run, Cloud Functions, BigQuery, IAM, Cloud Monitoring) · Azure · Docker · Kubernetes · HashiCorp Nomad · Ansible

Data

ClickHouse · BigQuery · MySQL · PostgreSQL · DuckDB / DuckLake · Parquet & object storage · Message queues · ETL · Data quality

Code

Go · Python · SQL · Bash · Rust · JavaScript · Java

Delivery

Jenkins · GitHub Actions · SonarQube · Git · TDD

Languages: English (fluent) · French (fluent) · Arabic (native)
Education: Engineering degree in Computer Science (Cloud Computing & IT Architecture), Esprit School of Engineering, 2024 · Azure Fundamentals, 2023

Let's talk

Looking for a remote site reliability & platform engineer role with a European team.

khalilrezgui0@gmail.com