Automation 5 min read

What DevOps at a 1,000-Person Company Actually Taught Me About Automation

Enterprise credibility — what DevOps at a 1,000-person company actually taught me about automation

What DevOps at a 1,000-Person Company Actually Taught Me About Automation

There's a version of enterprise DevOps that exists in blog posts and conference talks. It involves elegant GitOps pipelines, thoughtful platform engineering teams, and mature SRE practices. Decisions are data-driven. The on-call rotation is humane. Everyone agrees on the right way to do things.

That's not what I experienced.

I'm not saying my company was dysfunctional — it wasn't. But the gap between "enterprise DevOps as described" and "enterprise DevOps as practiced" is worth being honest about, because the lessons from the gap are more useful than the lessons from the ideal version.

Here are the things I actually learned.

1. Scale Exposes What Sloppiness Hides

At small scale, a lot of bad practices are survivable. Your deployment process is undocumented? Fine — everyone on the team was there when you set it up. Your monitoring configuration is a pile of manually-created dashboards? Fine — there are three of you and you know what's important.

At 1,000 employees, with 12 people on the DevOps team and 40+ applications in production, none of that survives. The undocumented deployment process becomes a 3am incident when the one person who knew it took PTO. The manual monitoring dashboards become a liability when the application it was watching gets decommissioned and nobody removes the alerts.

Scale is a forcing function for rigor. Everything you got away with at small scale — the undocumented runbooks, the one-off scripts nobody committed to version control, the configuration managed "by hand" — all of it becomes a problem.

This was the first lesson: rigor isn't a luxury that large teams can afford and small teams can skip. It's a necessity that large teams can't avoid and small teams just haven't paid the bill for yet.

2. Automation That Nobody Trusts Is Worse Than No Automation

Early in my tenure, I inherited an automated deployment pipeline that was technically capable but had a history of spurious failures. Not consistent failures — random ones. A deployment would work 19 times, then fail for no reproducible reason on the 20th.

The engineering team had responded to this by treating the automation as untrustworthy. They'd run it, and if it failed, they'd manually complete the steps it failed on rather than investigating why it failed. The automation was still running, technically — but it had become a starting ritual rather than a reliable system.

This is a specific failure mode: automation that lacks reliability erodes trust, and lost trust is very hard to rebuild. Engineers who have been burned by unreliable automation will route around it even after you fix it. You have to over-communicate every fix and demonstrate reliability over an extended period before they'll trust it again.

The pipeline was fixable. The trust problem took six months after the fix to resolve.

The lesson: reliability is not optional. An unreliable automation is actively harmful — it trains the team to treat automation as untrustworthy.

3. The People Problem Is Bigger Than the Technical Problem

You can build a perfectly designed automation system and have it fail because the humans who interact with it don't understand it, don't trust it, or don't know it exists.

At scale, this becomes a major category of work that nobody talks about in the technical communities. Change management. Documentation. Training. Communication. Getting engineers from different teams to agree on standards. Convincing management to give you runway to build the right thing instead of the fast thing.

I spent roughly 30% of my time on explicitly non-technical work: presenting to leadership, writing internal documentation, running training sessions on new tooling, facilitating cross-team standards discussions. I resented it at first. Later I understood that those 30% of hours were the multiplier on the technical work.

An automation nobody knows about helps nobody. A deployment pipeline nobody trusts gets routed around. Documentation nobody reads might as well not exist.

4. Toil Is a Lagging Indicator

"Toil" — the term from SRE for manual, repetitive, automatable work — doesn't announce itself. It accumulates gradually and becomes normalized before anyone names it.

The ticket triage process I eventually automated had been running manually for years before I arrived. It was the way things were done. Nobody had sat down and said "this is toil, we should automate it" — because nobody had done the audit.

Doing that audit — actually measuring how much time different categories of work consumed — was the trigger for most of the automation I built. Not inspiration, not a flash of insight. A spreadsheet tracking where the team's hours went.

The spreadsheet showed that triage was consuming 15% of our total capacity. That number made the automation project easy to justify. It also showed three other categories consuming 8-10% each that became the next wave of automation work.

If you don't have a way to measure where your time goes, you can't prioritize what to automate. The audit is not optional.

5. Small Companies Have the Advantage

This last one surprised me, and it took a while after leaving the corporate role to understand it.

Large organizations have resources but they also have inertia. A decision about tooling that would take me an afternoon at my MSP takes three meetings, a security review, and a procurement process in an enterprise.

The engineer who can move fast, apply rigor without bureaucracy, and automate without a platform team is more capable per capita than an enterprise team — if they're disciplined about it.

The mistake small operators make is thinking enterprise practices are too heavyweight for them. Some are. Many aren't. The rigorous parts — documentation, error handling, monitoring, version control — scale down without losing their value.

The lesson I brought from enterprise to solo MSP operation is not "enterprise practices are gold standard." It's "here is what rigor actually looks like, and here is how to apply it at the scale you're working at."

That's what I tried to put in the book.


Matt Fitzgerald is the founder of Fitzgerald Tech Solutions, former DevOps lead at a 1,000-person organization, and author of Live Life Automated. The Operator's Edge newsletter goes out every Monday.

Comments

No comments yet — be the first.

Leave a comment

Links aren't allowed. Short thanks are posted right away; everything else is reviewed first.

related

More from the blog.

Weekly · Free · Unsubscribe any time

The Operator's Brief

Every week: something I've automated, a tool I've found useful, and whatever I'm thinking about. Short, practical, no sales pitch.

✓

You're in.

Check your inbox to confirm — first issue arrives next week.