Platform / SRE
Platform and SRE work across many Kubernetes clusters
One Kubernetes cluster is a project. Twenty of them is an operations problem that will not be solved by another dashboard. Most of the platform and SRE work I have been hired for starts at that second number: many clusters, many pipelines, a reliability target that is already promised to someone, and a team that is tired of being the human API for YAML.
At Virtasant I worked on CI and platform infrastructure where the daily pipeline volume was large — on the order of a thousand executions — and the cluster estate was measured in tens, not one. Earlier delivery work at Rackspace and in consulting engagements was the same shape at a smaller scale: EKS and AKS clusters that had to be created more than once, upgraded more than once, and explained to a client who was not going to read the control-plane logs. The lessons below are from that work. They are patterns, not a promise that your estate will hit someone else’s percentage.
A platform team is not a ticket desk
The failure I walk into most often is a “DevOps team” that applies other people’s manifests. Every service request is a chat message. Every cluster is a snowflake with a name that made sense in the quarter it was born. The team is busy, the backlog is moral, and nobody can say what a new service receives on day one.
Platform work, as I want the word used, is the opposite. A new workload gets a golden path: a namespace or a cluster baseline, identity, ingress, secrets, logs, metrics, and a deploy path that already exists. The platform team’s job is to make that path boring and to keep it boring while Kubernetes versions move. The application team’s job is to ship the service. When those two jobs collapse into one queue, you do not have a platform. You have a bottleneck with kubectl.
SRE sits on top of that path, not instead of it. The useful SRE questions are about the user: what does down mean, what is the error budget, what is the page that fires when the budget is burning, and what toil are we doing every month that a change to the golden path would delete. I have written a lot of runbooks. The ones that mattered were the ones that shortened recovery because the failure mode was already named. A runbook for a failure you have never seen is a wish.
Upgrades are the product
Clusters rot on a schedule whether or not you plan for it. Kubernetes minor versions do not wait for a convenient quarter, and managed offerings (EKS, AKS, GKE) will eventually force the conversation. Treating upgrades as a project that starts when a version is already unsupported is how you get a freeze, a hero, and a weekend.
The version of this that worked for me was an upgrade train. Clusters are grouped. A non-production group moves first. The diff is the same diff everywhere: control plane, node image, the two or three add-ons that always break (CNI, ingress, cluster-autoscaler, the thing that talks to IAM). You do not “upgrade Kubernetes”. You upgrade a known set of components, with a rollback that is written before you start, and you do it often enough that the procedure is still true.
Pod Security is part of the same train. A cluster that is “upgraded” but still admits privileged pods from any namespace was not upgraded. It was patched. I would rather slip a date and land a baseline that matches the policy than publish a version number that the security review will reject in the same month.
GitOps is how the estate stays one estate
Argo CD and Flux are not religions. They are how I stop twenty clusters from becoming twenty opinions. The cluster should be a projection of git. If someone hotfixes a live deployment, the next reconcile should either revert it or make the hotfix impossible to forget because the UI is shouting drift. I have used both tools in client work. The choice matters less than whether application teams can ship without asking the platform team to type.
Helm charts are the packaging, not the platform. A chart that encodes one team’s entire history — every port, every annotation, every cloud-specific annotation behind a boolean — will be forked. I want a small chart for the boring deployment shape and an escape hatch that is still reviewed. Terraform (or Pulumi, when the generation problem is real) creates the cluster and the account boundaries. GitOps creates the workloads. Mixing those two loops in one apply, so that a namespace change and a VPC change land in the same blast radius, is how a Friday deploy takes down a network.
What actually moves reliability
People ask for a Prometheus stack because it looks like maturity. The stack is fine — Prometheus, Alertmanager, Grafana, and a disciplined label set have been the default in the environments I have run. The maturity is the rule that every page maps to an action. An alert that says “CPU high” trains the on-call to mute the channel. An alert that says “the error budget for this service is burning because the readiness probe is failing in one zone” is a page.
The other half of reliability is toil with a calendar. Secret rotation that depends on a person remembering a spreadsheet will be late. Compliance findings that are re-discovered by a scanner every week, with no owner and no pull request, will be explained in a slide instead of closed. I have spent real time on both: automating rotation for large sets of service accounts, and turning scanner output into changes that stick because they live in the baseline. Mean time to recovery drops when the failure is in a runbook and the change that prevents it is in git. It does not drop because the status page got a new color.
Disaster recovery is a design, not a document in the appendix. Two regions only help if identity, data, and the deploy path exist in both, and if somebody has failed the primary on purpose. A 99.95% target is a constraint on that design. It is not a number you paste into a proposal after the architecture is already a single availability zone with a hopeful backup.
What I ask before adding another cluster
Who is on call for it, and do they have a path that is not me? What is the baseline commit it will be created from? Which pipeline deploys into it, and is that pipeline the same one as production? What is the upgrade group it joins? If those answers are “we will figure it out”, the cluster is a liability with a bill.
I am glad to build the next cluster. I would rather spend the first week making the existing ones the same kind of thing. That is slower in a status meeting and faster in the quarter when a CVE lands on the node image and you have one change to roll instead of twenty snowflakes.
If you want that kind of platform — golden path, upgrade train, GitOps, pages that mean something — I take freelance engagements. Email ah.zahran@outlook.com. The companion piece on the operating model is Everything as Code.