Skip to content

Summary

Site Reliability and Platform Engineer with 9+ years building and running infrastructure where downtime is very visible: live sports streaming with peaks of 20–30M concurrent users across three continents, a banking core kept at 99.99% through a fintech-to-regulated-bank transition, and cloud platforms built from scratch. Deep in AWS and GCP, Kubernetes, Terraform, GitOps and observability.

20–30Mconcurrent viewers at peak
99.99%banking-core availability
9+ yrsinfrastructure & operations

Experience

  1. Senior Site Reliability Engineer

    NBCUniversal · Peacock · New York · remote

    Nov 2022 — present3 yrs 11 mo

    • Lead infrastructure for Peacock's live sports events (e.g. NFL), handling traffic peaks of 20–30M concurrent users.
    • Design and run multi-region deployments across Europe, Africa and the Americas, optimizing latency and replication.
    • Architected the network and security layers that integrate AWS and GCP for secure, high-speed data replication.
    • Run the observability stack (Prometheus, Thanos, Loki, Grafana) that product teams rely on during live events.
    AWSGCPMulti-regionPrometheusThanosLokiGrafana
  2. Author & maintainer

    Open source · github.com/andresgarcia29

    Oct 2025 — present1 yr

    • ark-cli: one AWS SSO login puts every EKS cluster across all accounts and regions into kubeconfig in seconds.
    • harness-creator, harness-daemon and harness-ui: an agentic engineering harness where agents propose and deterministic gates verify, a Go daemon that observes coding agents live, and its dashboard.
    GoTypeScriptShellClaude CodeMCP
  3. DevOps Lead

    Covalto · Mexico City · remote

    Nov 2021 — Nov 20221 yr

    • Led the technical migration during the acquisition of Banco Finterra, taking a fintech to a fully regulated national bank.
    • Designed the architecture and documentation to meet CNBV banking security and data sovereignty standards.
    • Standardized deployments with GitOps and built health checkers in Go and Python, keeping critical banking cores at 99.99% availability.
    GitOpsGoPythonCNBV
  4. Senior DevOps Engineer

    GoExpedi · Houston, TX · remote

    Aug 2020 — Nov 20211 yr 3 mo

    • Built the cloud platform from scratch: from basic Docker deployments to a multi-account AWS strategy managed with Terraform.
    • Secured B2B e-commerce platforms for global enterprise clients.
    • Broke legacy monoliths into microservices and established a high-performance CI/CD culture.
    AWSTerraformDockerMicroservicesCI/CD
  5. DevOps Engineer

    Rever · San Francisco, CA · remote

    Dec 2018 — Aug 20201 yr 8 mo

    • Built the company's Data Warehouse and Data Lake from scratch, enabling real-time analytics.
    • Implemented GDPR standards for data privacy and security.
    • Helped scale engineering from 2 to 30+ engineers while keeping DevOps best practices.
    AWSData lakeGDPR

Earlier · 2016 – 2018

  • Data Scientist · Digitt Apr 2018 — Dec 2018
  • Back-end Developer · Agave Lab Jan 2018 — May 2018
  • Machine Learning Developer · Tecnológico de Monterrey Oct 2017 — Mar 2018
  • Full Stack Developer · Kukumi May 2017 — Dec 2017
  • Full Stack Developer · MAXCOOL Jul 2016 — Aug 2017

Selected work & open source

Company-wide AI gateway and MCP rollout

Governed access to LLMs for every team: per-team cost tracking and rate limits behind one gateway, plus MCP servers so assistants reach internal tools securely.

ark-cli ↗

v0.13.1Go

One AWS SSO login → every EKS cluster across all your accounts and regions in your kubeconfig, in seconds. Scans in parallel, writes kubeconfig natively, retries AWS throttling with backoff.

52 commits32 releaseslast push

$brew tap andresgarcia29/agm && brew install ark --cask
ark k8s filtering clusters and switching context

harness-creator ↗

v0.62.5Shell

Claude Code plugin that turns a multi-repo workspace into an agentic engineering harness. Agents propose; deterministic gates verify.

417 commits88 releaseslast push

$/plugin marketplace add andresgarcia29/harness-creator

harness-daemon ↗

v0.60.0Go

Single Go binary that shows what your coding agents are doing, waiting on, deciding and spending, live, across machines over SSH.

131 commits34 releaseslast push

$brew install andresgarcia29/agm/harness && harness init
Live agent terminals: one waiting on a decision with one-click answers, one fixing failing tests

harness-ui ↗

TypeScript

React dashboard for a fleet of coding agents: every session, gate and dollar at a glance.

72 commitslast push

harness-ui overview with agents, tasks waiting on you and daily spend

Activity · live from GitHub

busiest day
223 on Jul 25
longest streak
37 days
current streak
2 days
active days
202 / 367

8,529 contributions in the last 12 months, public and private

Refreshed every 6h by a scheduled workflow · updated

How I work

  1. Boring on purpose. The 3 a.m. version of me has to understand it.
  2. If a script can check it, a script checks it. People and models bring judgment, not diffs.
  3. Observability before scale. You can't fix what you can't see while 25 million people are watching.
  4. Everything in git. Infra, policy, runbooks and decisions. The audit trail is a feature.
  5. Small blast radius. Canary first, roll back first, diagnose second.
  6. Write it down. Documentation is the on-call engineer you don't have to page.

Let's talk.

Open to remote Senior SRE / Platform Engineering roles, full-time or contract. Email is fastest.

esc