Ayman Zahran

Everything as Code

Everything as Code is an operating model, not a tooling preference

Ayman Zahran · 28 September 2026 · about 8 minutes

The footer of this site says git commit -m "Everything as Code" because that is the standard I want the work held to. It is not a product name and it is not a request to put the company on one vendor’s platform. It is a claim about how a change should travel from an idea to production: someone can read it, someone else can review it, and a machine can apply it again without a hero in a terminal.

I have spent the last several years doing that work across AWS, Azure, Google Cloud, VMware, and Kubernetes — as a freelancer and, before that, inside delivery teams at places such as Virtasant, Rackspace, and client engagements where the cluster count and the pipeline count were both large enough that memory was not a control plane. The pattern that survived those environments was not “we standardized on Terraform” or “we standardized on Ansible”. It was “the real system is whatever git says, and the language is allowed to change when the layer changes”.

What “everything” actually includes

Infrastructure as code is the part people already nod at: networks, accounts, clusters, databases, IAM. That is necessary and it is not enough. A platform that is only half-declared still drifts. The rest of the system has to be able to sit next to it.

Pipelines belong in the same review. If the Jenkins job, the GitHub Actions workflow, or the Azure Pipelines YAML is edited in a UI that nobody diffs, the infrastructure pull request is theatre. The deploy path is part of production. Policy belongs there too. A Kyverno rule, an OPA check, a pipeline step that fails a privileged pod, or even a small script that rejects a Terraform plan with a public bucket is more honest than a wiki page titled “security standards”. Observability config belongs there when the team is large enough to forget which dashboard is canonical. Prometheus rules, Grafana dashboards as JSON, and alert routes checked into git age better than a folder of screenshots.

Runbooks are the awkward one. A runbook that is only prose will rot, and a runbook that pretends a human incident is a shell script will get someone hurt. I want the procedure in the same repository as the system it describes, with commands that match the current interface, and with an explicit line that says which steps are judgment. The win is not automation for its own sake. The win is that the next on-call engineer is not reconstructing the system from Slack.

Polyglot on purpose

My MBA thesis looked at this as a transformation strategy rather than a tooling argument. The short version: use the language that fits the layer. HCL is a good language for cloud resources because the plan is the point. YAML is a reasonable language for Kubernetes objects and for many pipeline schemas because the objects are data. A general-purpose language earns its place when the abstraction leaks — when you are generating many similar resources, when you need real tests, or when the provider’s HCL forces copy-paste that a loop would have prevented. Pulumi, CDK, and CDKTF are tools for that leak. They are not a personality.

Bash stays at the edges: glue in a container entrypoint, a migration that will run twice, a break-glass command written down so it is not invented at 2 a.m. I do not want the platform’s source of truth to be a 2,000-line shell script with no plan and no review. I also do not want a team ashamed of a twenty-line script that does one job cleanly.

The capstone vehicle for that thesis was a designed Cloud and DevOps practice I called EchoOps. It was a design and a financial model, not an employer and not a product you can buy. I mention it only because the conclusion is still how I take work: polyglot programming plus Everything as Code is a way to change how infrastructure is delivered. Picking a single blessed tool and calling the rest shadow IT is how platforms become shelfware.

Where the phrase goes wrong

The failure mode I see most often is encoding a broken process. If every change needs six approvals in a ticket tool and then a pull request that restates the ticket, you have added git without removing the queue. The pull request has to be the place the decision is made, or it is a printout.

The second failure is a monorepo nobody can review. “Everything” does not mean one repository the size of the company, owned by a platform team that becomes a ticket desk. It means each thing that matters has a diff. Boundaries still matter: a platform repository for the cluster baseline, an application repository for the service, a module registry with a small interface. Reviewability is the constraint. If a reviewer cannot tell what will change in production, the code is not finished, regardless of which language it is in.

The third failure is a console that is still the real control plane. Terraform applied once, then “just a small click” for the rest of the quarter, is how state files become fan fiction. Drift is not a moral failing. It is what happens when the path of least resistance is the UI. The fix is boring: make the coded path faster than the click, and make the click detectable. A scheduled plan, a config connector, or a policy that denies console changes on production accounts will do more than a slide about culture.

The fourth failure is generated code that nobody reads. AI assistants are now good at drafting Terraform, Helm, and pipeline YAML. I use that. I do not merge it because it compiled. Generated infrastructure is still infrastructure. It needs an owner, a plan you looked at, and a module boundary so the next change is not a rewrite of a thousand lines of machine prose.

A bar I can actually hold

When I join an engagement I am not asking for a green-field rewrite. I am asking whether a production change can be explained in a pull request this month. If it cannot, the first deliverable is usually a thin slice: one environment, one service, one pipeline, brought under review, with the console path closed behind it. Then we widen.

Modules and charts should have a small interface. If every caller passes forty variables, the abstraction is a copy of the resource and the team will fork it. Policy should sit next to the code that can violate it, and it should fail closed in continuous integration before it fails in production. Documentation should point at the repository. A diagram that disagrees with the module is a bug.

None of this removes judgment. Incident command, a conversation with a client about what “down” means, and the choice to delete a feature instead of automating it are still human. Everything as Code is a constraint on the parts of the system that should be repeatable. It is not a claim that the job is typing.

If that is the kind of platform work you want done — reviewable, polyglot where it helps, and honest about what stays a decision — write to me at ah.zahran@outlook.com. Related notes: platform and SRE across clusters and how this site is shipped.