Daniel Calvo

Dani smiling beside his black cat

Hello! I'm Dani, and I do SRE.

I thrive making sense of difficult problems and environments, figuring out what's the most impactful thing we can do together, and then taking ownership of what we deliver.

If you are looking for someone with the right mix of soft & hard skills to make your platform reliable, usable and affordable, I'd be happy to chat!

Senior Site Reliability Engineer at RingCentral Events (acquisition)

August 2023 — April 2026 Remote Spain

  • As the sole remaining SRE at acquisition, helped onboard a new team of three SREs and a manager, transferring platform knowledge, architectural history, operational practices and broader engineering culture to enable independent support of the product.

    This kicked off with a visit to RingCentral's office in which I presented to the new folks as much as I could: the repos we had, how we released, how the infra was set up, what the monitoring looked like, how new projects were created, how on-call and postmortems were handled, areas in which we were lacking and a bunch more. I wish I still had the slides. The presentation was well received and no one fell asleep, so that was a win!

    I also tried to close it on a positive note -- things weren't so great in the past, but now hopefully we would be able to improve the situation.

    That was a good starting point, and afterwards most follow-ups were async, with an occasional conversation after a stand-up meeting. Before long everyone was up to speed and getting it done. Good times.

  • Investigated our Datadog and Cloudflare usage and lowered it to support contract renewals with lower spend, reducing our spend by 30% and 25% respectively.

    For Datadog, I remember I dug deep into their billing docs to understand how much allocation we got with certain products we had contracted (e.g., hosts gave you free metrics). We then stopped ingesting some logs entirely, sent others to S3, dropped a bunch of metrics from being indexed and learned that a few things were already included in our allotments. Then we went over every single item we had contracted to make sure it was right-sized -- I put together a huge table on Confluence we all reviewed and agreed on before renewal.

    Cloudflare's billing was a lot more opaque: they had no cost dashboard. I got Codex/Gemini to map our usage to certain websites and endpoints and discovered about 70% of our metered Cloudflare usage eaten up by load tests against staging 🫪. We also had some redirect rules matching way more than they should.

    We adjusted our contracted usage to match decreased load testing and a more conservative volume of production traffic, and fixed all our misconfigured rules. The procurement team further negotiated with these usage numbers and got us a final renewal quote 25% below the initial quote.

  • Built dashboards to identify release pipeline bottlenecks and flaky tests, enabling fixes by owning teams, and reduced monolith release time from 45 to 35 minutes with pre-deployment image caching and capacity provisioning.

    This was a timeboxed effort to reduce our monolith release time. After getting everything dashboarded and investigated, the easiest improvements were at release time. Right after the build step, we cached the new image on all prod/stg machines with a DaemonSet, did some overprovisioning so new nodes wouldn't be needed during deployment rollout, and changed the deployment maxSurge to 100%. These adjustments across both environments saved us 10 minutes overall.

    I wish I could've spent more time on this. Making systems go fast is always fun.

  • Owned our migration from JFrog to AWS CodeArtifact end-to-end, migrating ~40 repositories and handling project planning, package synchronization tooling, CI changes, developer coordination and onboarding documentation.

    JFrog cost us ~100x more than AWS CodeArtifact for package hosting -- so we set out to drop JFrog as a vendor.

    I listed our repos and package types, created a task per repo and started off writing some Go to sync packages from JFrog to AWS, change lockfiles, and index packages on GitLab. Then I started going through our Yarn/npm repos.

    Simple pipelines were straightforward: run a command to get a CodeArtifact token, change .yarnrc, change yarn.lock, update the docs and create a PR. Others were a lot more involved: we had monorepos with up to 30 separate pipelines across parent, child and scheduled ones -- I really got my PhD in GitLab pipelines with this project.

    With this project I also had my first proper exploration of AI-assisted dev, right when Opus 4.5 came out in late 2025. It was really useful to write the package sync tooling and help me understand complicated pipelines. I also tried creating skills with a lot of reference context and safety guardrails so it could migrate the remaining repos using my previous work as examples. Having AI do the changes was hit or miss: the FE pipelines did not follow a clear standard and the models weren't as sharp back then. It was faster and cleaner to just make changes by hand.

    I coordinated reviews with dev teams and left them instructions on how to use CodeArtifact. I migrated about 40 repos by the time I was laid off. The migration was left unfinished.

  • Created a short, actionable first responder runbook for approximately 30 on-call engineers, clarifying what to do when paged: incident triage, declaration, escalation, communications and postmortem responsibilities.

    This came up when a principal engineer in our team was running a first responder workshop with new hires. I attended the first training round and noticed we had a bunch of reference material on Confluence, but no actual guide to follow when getting paged at 3 a.m.

    I volunteered to write this and it turned out quite handy. It gave you a script you could follow with confidence: when to acknowledge the page, what customer journeys warranted raising an incident, how to actually raise the incident, how to escalate it to the owning team, what communication you had to do as the first responder and so on.

    I feel it helped folks feel more sure-footed and safe when being on call, and it also helped people to not jump straight into troubleshooting and actually follow the process. I ended up using the guide myself several weeks later and it did help a lot!

Senior Site Reliability Engineer at Hopin

January 2021 — July 2023 Remote Spain

  • Worked on Hopin's migration from Heroku to AWS EKS, delivering features on infra releases, cluster access & networking as well as troubleshooting during the transition.

    This was our cross-team effort to move our monolith from Heroku to k8s on AWS. When I joined, the project was already underway, so I jumped in to help with infrastructure work and troubleshooting that needed doing. I worked on our GitHub Actions to add multi-account cluster deployments, added read-only cluster access for developers, and sorted out our IAM and security group practices so services could reach what they needed and have the necessary permissions.

    There was also plenty of troubleshooting along the way: slow disk I/O, Cloudflare handshake errors, image pull errors, cert-manager/external-dns/ingress issues, and a few more I'm likely forgetting.

    I have very fond memories of this project. Things moved quickly, there was a new challenge every day and everyone worked well together. In retrospect, we built way more complexity than was needed, assuming growth would continue and bigger teams would maintain things. This growth didn't happen, and this influenced how I later approached the second version of our k8s setup.

  • Created a golden path for new Kubernetes services, documenting and templating service delivery at Hopin. It described our usage of Docker, Terraform, AWS, Kubernetes, GitLab, Datadog, Cloudflare and others. This was adopted as the starting point for all new services, standardizing deployments and simplifying maintenance.

    We moved the monolith that served hopin.com to Kubernetes, but there wasn't a documented way for teams to deploy new services to our platform. I pitched a sample app to my manager: an example that teams could pick up and use as a base to deploy a service to production. He agreed, so I got to work.

    The app covered local dev with Docker/docker-compose, the Terraform for AWS resources and permissions, GitLab pipelines, k8s manifests, Cloudflare configuration, and an overview of monitoring on Datadog. I also included some troubleshooting guidance, as many engineers were new to Kubernetes.

    The initiative was a success: about 15 services went live in the next year and all of them adopted the sample app as their starting point, with teams later adjusting their release processes to match their needs. This consistency across projects (they all used the same template) really did help later with rightsizing, cluster migrations, and ad hoc troubleshooting.

  • Drove our AWS cost categorization & reduction effort that saved Hopin about $1M a year.

    This project had two parts: first the finance folks wanted to know where the spend was going, and later we wanted to see how we could lower it.

    I worked with finance on what AWS tags were needed to categorize spend, then used the AWS Terraform provider default_tags to apply them. Well, trying to apply tags uncovered tons of TF config drift, so addressing that was a side quest in and of itself.

    Once we could categorize spend, we set out to lower it. We removed unused environments and resources, reduced oversized k8s resource requests, and reviewed database capacity and backup retention. I also reviewed the S3 usage patterns on our recordings and moved all of them to S3 Glacier IR -- that lowered our S3 spend by about 70%.

    I eventually presented the database backup and S3 spend improvements at the company all-hands. S3 and DB savings alone amounted to $500K a year. What started out as a tagging task ended up as a larger project between engineering and finance. This was cool!

  • Built the second iteration of Hopin's Kubernetes platform, favoring simplicity, community maintained components, easier to follow GitOps, and up-to-date Kubernetes and cluster components, later working with developers to migrate our production services with zero migration-related outages.

    The first iteration of our k8s platform, built in 2021, was built for hypergrowth: custom Terraform modules, complicated GitOps, and lots of nested Kustomize. By the time we needed to do cluster upgrades, we were six k8s versions behind and I was the only remaining SRE. Upgrading this would be difficult. We needed something we could maintain with the people we had.

    I built a replacement using community Terraform modules, simpler config with a saner Kustomize structure (just a base and per-cluster patches), and all the components updated. We also moved from Terragrunt to Atlantis and adopted Argo CD and Karpenter. We were using cdk8s to generate our monolith YAML, which I replaced with plain YAML, envsubst and two variables.

    For the actual migration, we adjusted our IAM and pipelines to deploy to both clusters for a while. We tested in staging, and eventually shifted production traffic gradually through Cloudflare while keeping an eye on Datadog for issues. The migration completed with zero migration-related outages. I was happy with how we managed to simplify the platform and migrate all services with no outages.

  • Other accomplishments include: simplifying TF modules, creating log based SLOs, writing Prometheus exporters and onboarding junior SREs.

    There are some other small hits here: if you wanted to host some static content, we had an in-house Terraform module for S3, another one for CloudFront, and then you had to figure out the TF for Cloudflare. This was a terrible experience. I took a page out of the configuration best practices chapter of the SRE Workbook and created a new Terraform module that required only two inputs: bucket name and DNS entry. These were the only things developers cared about; the rest could be left as defaults and tweaked as needed.

    For SLOs, I explored using Cloudflare logs sent to S3, using Logstash and DogStatsD to push metrics to Datadog, as doing this natively within DD was too expensive. I had the log-to-metrics pipeline all ready to go, but this was shelved after the second layoff wave. Good learning experience though.

    We sometimes bumped into the AWS account limits for ECR image counts and Parameter Store entries, so I wrote exporters to track those, as, surprisingly, AWS did not make those metrics available. These kept being useful years later when ephemeral environments would fail to clean up entirely.

    I fondly remember being the team buddy for newly hired SREs, helping them with their first tasks and getting them up to speed. They went on to do awesome in no time.

DevOps Engineer at trivago

July 2018 — December 2020 Palma de Mallorca

  • Delivered several release processes using Saltstack, Kubernetes, Jenkins and Nomad

  • Initiated, implemented & delivered "The monitoring project" using Prometheus, Grafana & ELK, making product telemetry available for Ops, Dev and Product.

  • Created, documented & shared ownership of all release processes from scratch using Jenkins pipelines

  • Gave internal workshops on Cloud native tech, encouraging everyone to adopt new technologies.

  • Did my best to live up to the DevOps principles & deliver value together with everyone in the team :-)

System Administrator & DevOps Engineer at trivago

October 2013 — June 2018 Palma de Mallorca

  • Maintained release processes in Shell, Python & Groovy as part of the operations team, as well as our infrastructure set up in Saltstack.

  • Helped set up & maintain an in-house Openstack cluster, encouraging development teams to get involved in their release processes.

  • Was the key person for troubleshooting production issues, following up on issues with all involved parties (dev, sys, product) until their resolution.

  • Onboarded & mentored new ops team members on tools, processes, workflows and memes.

System Administrator at Genexies Mobile

July 2012 — August 2013 Madrid

  • Administered production and testing environments on an internet business, being responsible with the Sysadmin team for various highly available services

  • Helped develop and implement the Service Support project, a set of monitored application metrics that indicated the health of all business critical applications, developed in Bash.

  • Documented and implemented various policies such as backup, disaster recovery and alarm response

  • Worked with external system administrators on the troubleshooting of technical issues between Genexies infrastructure and client's, such as Telefonica

Certifications

CKA · Terraform Associate · AWS Solutions Architect Associate · LPIC 1 & 2

Education · Technologist degree in System & Network Administration

January 2008 — January 2011 SENAI, Florianópolis