Västerås, Sweden

Max Agahi

Senior DevOps Engineer

Max Agahi

Nineteen years keeping production systems running — from 400 Cisco switches across 35 physical sites, to Kubernetes platforms and Terraform-managed AWS estates.

What I do

I keep cloud platforms running, and make them cheaper and safer to operate. Day to day that means Kubernetes, AWS, Terraform and the CI/CD around them, plus the observability that tells you whether any of it is actually working.

I came to cloud the long way: workstations and structured cabling first, then enterprise networks and on-premises systems. So when a platform problem turns out to be a routing problem, a storage problem, or something physical, I have usually seen it before.

I have run a team of fifteen and mentored six people into their next roles. More recently I have led programmes rather than people — an AWS Control Tower migration, a tested disaster-recovery capability, a multi-region Kubernetes fleet upgrade. Either way I take the incident rather than the ticket: response, root cause, and the remediation that stops the repeat.

Experience

  1. Quinyx A.B.

    Four years, promoted into the senior role after seventeen months.

    1. Senior DevOps Engineer

      Jan 2024 – Present
      • Lead the migration of the AWS estate into a new Control Tower organisation — landing zone, organisational units, governed regions, service control policies and billing foundations — sequenced across five waves.
      • Delivered a tested disaster-recovery capability for the platform against a 24-hour recovery-time and 3-hour recovery-point objective, including cross-region replicas and global databases for tier-one data, in five and a half months.
      • Upgrade a seven-cluster, two-region Kubernetes fleet across successive minor versions, keeping every cluster inside its support window and off extended-support charges.
      • Converging that fleet onto Flux-managed GitOps with common chart versions, tracked as 42 sub-tasks across five workstreams, with target versions verified against live cluster state rather than against what the manifests claim.
      • Own continuous cost optimisation and FinOps practice across the cloud estate, including the spend-control layer built from nothing — 33 daily budgets across 11 cost signals, plus anomaly detection.
      • Replaced an unsustainable cross-region backup mechanism with a tag-driven export pipeline meeting a 3-hour recovery point and a 35-day retention bar, with reusable Terraform modules and an auditor that catches coverage drift.
      • Authored an organisation-wide commit and pull-request convention and made it a required merge gate across every production-connected repository, application code and infrastructure alike.
      • Consolidating infrastructure Terraform out of two application repositories into one canonical repository, after a sweep found 146 module directories; tracked as 140 individually scoped migrations.
      • Automated the production-to-lower-environment database refresh as a self-healing Argo Workflows pipeline, removing a recurring manual retry from the on-call rotation.
      • Built an MCP server that provisions database-access connections, folders and grants through GraphQL, so an AI coding agent can do work that was previously manual.
      • Onboarded four engineers onto the platform, three at junior level and one senior, covering the infrastructure they needed to ship and operate their own services.
    2. DevOps Engineer

      Aug 2022 – Dec 2023
      • Built the Terraform codebase that provisions and manages the company's Kubernetes clusters, giving every engineering team one infrastructure-as-code path to them.
      • Brought previously unmanaged VPC and security-group coverage under Terraform, closing foundational gaps in the estate's infrastructure as code.
      • Implemented VPC endpoints so internal traffic to cloud services stopped crossing the public internet, and established the reusable module pattern for it — delivered in five weeks.
      • Upgraded the production Kubernetes platform and validated every workload against removed APIs ahead of successive releases.
      • Cut production log noise across eleven services, reducing pipeline load and storage cost and making severity filtering usable during an incident.
      • Hardened the platform with restrictive Pod Security Standards, auditing and updating pod specifications across every service.
      • Built Terraform modules for gateway and routing-table management, replacing hand-managed network resources with reviewable code.
  2. NOC Operator

    Trustly Group A.B.

    • Monitored availability and performance of infrastructure, systems and products.
    • Raised incidents, wrote postmortems and problem cases, and drove them with other teams to resolution.
    • Wrote scripts and tools for the NOC team's day-to-day work.
    • Created and revised NOC process and system documentation.
  3. Relocated to Sweden. Passed the AWS Cloud Practitioner exam and started a research master’s in electronics during the move.

  4. DevOps Engineer Remote

    SmartCat Innovation and Technology Group

    • Built and maintained CI/CD pipelines to testing, staging and production with Jenkins, Git, JFrog Artifactory, Docker, Ansible and Bash.
    • Installed and maintained servers, containers and backups.
    • Added visibility across the pipeline and the infrastructure with Zabbix.
  5. Greater Tehran Electricity Distribution Company

    Twelve years and five titles, at a utility running roughly 2,500 workstations, 100 servers and 200 virtual machines across 35 sites.

    1. Senior Systems and Network Engineer

      Dec 2013 – Sep 2019
      • Kept the electricity-meter billing system running for three years without an unplanned outage — 1.5 million meters across Tehran, against twenty years of accumulated reading history.
      • Grew a team to 15 direct reports on promotion into the senior role, including three technicians on shift rotation for out-of-hours cover.
      • Mentored six of them into their next roles — network administration, Windows domain and systems administration, and technician positions.
      • Built a fully automated CI/CD pipeline on Jenkins, Git and Ansible, replacing hand deployment across 20 internal applications.
      • Wrote the Ansible playbooks that turned a six-hour manual server build into a repeatable run.
      • Helped design the company's new data centre, from capacity estimation through to choosing the servers and network hardware that went into it.
      • Deployed Cisco ASA firewalls as a high-availability pair, standardising rules that had drifted site by site.
      • Recovered the internal mail and messaging store after a third-party backup system corrupted its SAN volume, rebuilding the partition tables by hand.
      • Moved 10 internal services onto Docker, retiring the servers each one had been pinned to.
      • Led the network side of five new substation and office builds, from cabling and racks through to switching and monitoring.
      • Ran remote-access VPN for around 100 staff.
      • Set up centralised logging, so incidents could be reconstructed rather than guessed at.
      • Introduced configuration backup for network devices, making a failed switch a swap rather than a rebuild from memory.
    2. Systems and Network Engineer

      Dec 2011 – Dec 2013
      • Re-addressed the estate onto a class A plan that separated data, telephony, physical security and management traffic, and introduced EIGRP routing across it.
      • Introduced offsite tape backup. Until then, backups were written to the same server disk as the data they were meant to protect.
      • Consolidated 120 physical servers onto VMware ESXi, cutting rack space, power draw and hardware spend.
      • Deployed Zabbix across roughly 800 devices — switches, routers, UPS units, radio links and servers — so capacity and link problems surfaced before users reported them.
      • Segmented the operational-technology network away from the office network, so a compromised desktop could not reach control systems.
      • Introduced redundant inter-site links, and tested the failover rather than assuming it.
      • Standardised Cisco IOS across 400+ switches, ending a long tail of divergent and unpatched firmware.
    3. Junior Systems and Network Engineer

      Dec 2010 – Dec 2011
      • Took over day-to-day operation of the metropolitan network linking 35 sites, having come up through the helpdesk that depended on it.
      • Inherited an undocumented network and produced the first accurate topology and IP address plan for it.
      • Replaced manual switch configuration with templated, reviewed configs.
      • Automated the nightly backup verification that had been a manual checklist.
    4. IT Technician, then Junior IT Technician

      Dec 2007 – Dec 2010
      • Supported around 500 workstations and their users across a regional office and the six branch offices under it, as one of four technicians covering that region.
      • Standardised workstation builds with network installation, unattended configuration and group policy, cutting repeat tickets.
      • Rolled out Cisco and Alcatel IP telephony and Polycom video conferencing to 35 sites, retiring analogue lines.
      • Wrote the helpdesk's first runbooks, replacing knowledge that lived only in people's heads.
      • Took on server-room work early — UPS load checks, electrical distribution panels, rack installation and structured cabling — which is how the move into systems work started.

Side projects

  1. Python
    Rust

    yadgar

    A persistent memory engine for AI coding agents. It decays what you stop touching, promotes what recurs, and filters recall to the git branch you are on, pairing every memory with a curated wiki searched through the same ranking pipeline. Works with any MCP client.

    A rewrite as gRPC modules is underway at yadgarhq — each logic service paired with its own storage twin, behind one bilingual gateway.

  2. Python

    karyab

    Job-hunting automation for the terminal. It fetches postings from several sources, ranks them against a profile, and builds a daily shortlist. It also tracks what you submitted, and nudges you to follow up when a posting goes quiet.

  3. Python

    ostad

    Turns a certification syllabus into a tutor that teaches and then examines you, across learn, drill, mock-exam and review modes.

    Every graded question traces back to a primary source, enforced in CI. No braindumps and no copyrighted prep material, by design.

  4. Python

    ccpm

    A personal plugin marketplace for Claude Code, currently carrying seven plugins.

  5. Terraform

    terraform-modules

    Reusable Terraform modules, with pre-commit hooks wired for formatting, validation and generated documentation.

Skills

Public cloud
AWSGoogle CloudAzure
Infrastructure as code
TerraformOpenTofuCloudFormationAnsible
CI/CD
ArgoCDArgo WorkflowsFluxGitHub ActionsSpinnakerJenkins
Containers
KubernetesDockerPodmanOpenStack
Observability
GrafanaOpenSearchElasticsearchPrometheusZabbix
Operating systems
LinuxNixOSRHELDebianUbuntuAlpineWindows Server
Virtualisation
KVMVMwareProxmox
Networking
Switching and routingDNSNTPDHCPLDAPCisco IOSMikroTik
Network security
Cisco ASAFortinetpfSenseOPNsenseJuniper
VPN and TLS
WireGuardOpenVPNLet's Encryptcert-managerDNSSEC
Secrets management
Vault1PasswordSOPSsealed-secrets
Security scanning
TrivySnykSonarQube CloudGitleaks
Cloud security
IAM policySecurity groupsWAFGuardDuty
Storage
SANHPENetApp
Languages
BashPythonGo
Version control
GitGitHub

Education

Master by Research, Electronics
Mid Sweden University, Sundsvall, 2021

First term completed. Left the programme to take a role in industry.

Bachelor of Computer Engineering
Azad University, 2006

Certifications

Certified Ethical Hacker (CEH v0.9)
EC-Council

AWS Certified Cloud Practitioner
Amazon Web Services, 2021 · expired 2024

The badge is issued in my legal name, Mohammadmahmoud Agahi.

IELTS Academic — band 7.5
expired

Band 7.5 corresponds to CEFR C1.

Training

CCNA and CCNP coursework
Private institutes, Tehran

Coursework, not Cisco certification.

Get in touch