Developer infrastructure
Concept studyMaking deploys boring across forty repositories
A single release path for a growing engineering organisation, where the previous approach made every deploy a coordinated event involving three teams.
- Engagement
- Embedded engineer
- Duration
- Ongoing, eight months to first restructure
- Year
- 2025
- Client
- Not applicable
The challenge
What made this difficult.
Services had accumulated their own pipelines, their own deployment triggers, and their own rollback procedures. A change touching a shared library required a coordinated release across teams, and the cost of that coordination had made teams batch unrelated changes into larger, riskier deploys. The bottleneck was not code review — it was the release path itself.
Constraints
Non-negotiables we designed around.
No big-bang migration
Incremental migration mattered more than a clean end state. Forty services could not be frozen while the platform changed.
Different runtimes already in production
The constraint applied to containerised services and to scheduled batch jobs with entirely different failure characteristics.
Existing ownership boundaries
Team boundaries had to keep matching on-call responsibility. A pipeline change that obscured ownership would be rejected.
Audit requirements on release events
Every production change needed an attributable record, and that record had to survive platform changes.
Approach
How we would build it.
Started with one service, end to end
Rather than designing the target platform on paper, we migrated a single representative service completely — build, test, canary, promote, rollback — and used the friction as the requirements document for everything else.
Made provenance a build property
Every artefact carries the source revision, the build inputs, and the change record with it. Promotion becomes a reference to a verified artefact rather than a rebuild, which is what removed the coordinated-release requirement.
Progressive rollout as the default, not an option
Canary analysis runs on every production change with automatic halt thresholds. Small teams stopped needing a manual go/no-go decision they had no context to make well.
Rollback designed before rollout
Because rollout is a shift of traffic between immutable artefacts, rollback is a traffic operation. We rehearsed it under failure conditions rather than assuming it worked.
Migrated on an agreed service queue
Teams opted in as they had capacity. A thin compatibility layer meant both patterns ran in parallel, so the migration never became a coordination exercise.
Architecture
How the pieces fit together.
Unified release path. Promotion shifts traffic between immutable artefacts, so rollback is a traffic operation.
Source
- Monorepo
- Service modules
- Shared libraries
- Infrastructure definitions
Build
- Hermetic builds
- Content-addressed artefacts
- SBOM generation
- Vulnerability scanning
Verify
- Unit and integration
- Contract tests
- Environment gates
- Artefact attestation
Release
- Progressive rollout
- Automatic halt thresholds
- Traffic shifting
- Instant rollback
Operate
- Unified telemetry
- Change correlation
- Ownership metadata
- Incident annotations
Stack
What it would run on.
Language
- Go
- TypeScript
- Python
- Bash
Platform
- Container images
- Declarative infrastructure
- OCI registries
- Secret management
Release
- Progressive delivery controller
- Service mesh traffic management
- Policy-as-code
- CI with attestation
Expected outcomes
What success would look like.
Coordinated releases
Eliminated
Design goal: shared-library changes no longer require a multi-team release window.
Rollback
Traffic shift
Because promotion references immutable artefacts rather than rebuilding, rollback no longer depends on a reversible build.
Adoption
Opt-in queue
Both release patterns ran in parallel throughout, so migration was never a freeze-window negotiation.
What we learned
The conclusions we would carry forward.
Migrating one service completely was worth more than designing the whole platform — the failure modes showed up immediately.
Making builds hermetic and artefacts immutable is what turned rollback from a build problem into a routing problem.
An agreed migration queue kept platform work from turning into a coordination tax on every team.
Related
