GitOps for nineteen services across four environments
What should have happened
The repository describes what should be running in dev, staging, beta and production, and the cluster follows the repository.
What actually happened
Several developers working on the same project with configuration that had drifted apart, and nothing in the cluster flagging it. Changing a variable meant running the whole pipeline and shipping the application again.
What I did
Nineteen ArgoCD Applications reconciling the cluster against the repository. Each service's configuration became a sha256 hash in a pod template annotation, so changing a variable changes the template and triggers a rollout — without going through a build. With maxUnavailable at zero and a readiness probe, no old pod leaves before a new one answers. And when someone edits the cluster by hand, the application shows up OutOfSync instead of becoming a surprise weeks later.
0downtime when applying configuration
What This Unlocked
- The cluster's history became the git history. Who changed what, when, in which commit, and with what reasoning in the pull request.
- Rolling back became a revert, instead of someone rebuilding from memory what used to be in place.
- Drift stopped being invisible. Editing the cluster directly makes the application show up OutOfSync.
- Secrets left the repository: the cluster pulls credentials from an external store, and what is versioned is the reference, not the value.
- Configuration left the application repository. Environment variables, even the ones that are not secrets, moved to the platform repository — so having the code no longer means having production configuration by default.
- Each environment ended up isolated in three layers: its own namespace, its own secret store and its own path inside it. Staging cannot reach production credentials, not even through a misconfiguration.
- The cluster became rebuildable from the repository, because the repository stopped being documentation and became the source.
What I Would Do Differently
I left the resync on its default, which polls git every three minutes. That is dead time between the merge and the change taking effect, and during it the cluster is not what the repository says. Today I would wire the Azure DevOps webhook to trigger the sync on commit, and keep polling as a fallback.