Hire DevOps developers
DevOps is the practice of making deployment and operation routine, and hiring for it means separating people who automate delivery from people who only administer servers.
What DevOps actually is
DevOps describes a set of practices aimed at making the path from a developer's change to running software fast, repeatable and safe. In job-market terms it covers continuous integration and deployment, infrastructure as code, containerisation, monitoring and the operational work of keeping systems running. The title is used loosely enough that establishing what you actually mean is the first task in any search.
The useful distinction is between automating delivery and administering infrastructure. Both are legitimate and they are different jobs. Someone who builds pipelines, defines infrastructure in code and improves how quickly and safely a team can ship is doing the first. Someone who manages servers, handles access and responds to alerts is doing the second. Job specifications routinely describe one and interview for the other.
The measure that matters is not which tools someone knows but what they have made possible. A team that deploys once a quarter with a change-freeze weekend operates differently from one that deploys several times a day without ceremony, and moving from the first to the second is the actual work. Tool familiarity is a means to that, and it is frequently mistaken for the thing itself.
The part that separates seniors from mid-levels
The pipeline is the core artefact, and a good one does more than run tests. It builds once and promotes the same artefact through environments rather than rebuilding per stage, because rebuilding means the thing you tested is not the thing you shipped. It fails fast, gives useful output, and is fast enough that developers do not work around it. A pipeline that takes forty minutes will be bypassed, and the bypass becomes the real process.
Infrastructure as code is the second pillar, and the discipline it requires is usually underestimated. State management, drift between what is declared and what exists, and the blast radius of a change are all real operational concerns. An engineer who has recovered corrupted state, or who has had a plan reveal a destroy on something important, has learned things that cannot be read.
Observability is the third, and it is where the difference between competence and seniority is most visible. Collecting metrics is easy. Knowing which ones indicate that users are affected, setting alerts that fire on those rather than on every anomaly, and keeping the on-call burden sustainable is much harder. Alert fatigue is an engineering failure, and engineers who have been woken up by their own bad alerts design differently.
Where DevOps is used
The label “DevOps developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Delivery acceleration
Teams whose deployments are slow, manual or frightening, where the goal is to make shipping unremarkable.
Cloud migration and modernisation
Moving workloads into the cloud and rebuilding the deployment path around them.
Platform engineering
Building internal tooling so product teams can deploy and operate without needing a specialist each time.
Reliability work
Improving uptime and incident response for systems that fail more often than the business can tolerate.
Compliance automation
Regulated environments where audit evidence, access control and change approval must be produced by the system rather than by people.
Cost and capacity management
Right-sizing and governing infrastructure spend as an ongoing engineering concern.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
DevOps is a practice rather than a product, so currency is judged by the delivery outcomes someone has achieved and the tooling they have used recently rather than by any version number.
DevOps itself is not versioned as a single product, so the useful equivalent is the support status of the tools a DevOps engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
| Hashicorp Terraform | 1.16.4 | 2026-09-23 | none published | none published |
| Docker Engine | 29.8.1 | 2026-09-15 | 1 | 2026-12-04 |
| Ansible | 14.4.0 | 2026-09-08 | none published | none published |
| Jenkins | 2.583 | 2026-09-22 | none published | none published |
| GitLab | 19.4.1 | 2026-09-22 | 3 | 2026-12-17 |
| Prometheus | 3.14.0 | 2026-08-17 | 2 | 2027-07-31 |
| Grafana | 13.2.2 | 2026-09-15 | 4 | 2027-05-24 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for DevOps alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- Git and a code host
- The source of truth for both application and infrastructure, with review as the control point.
- GitHub Actions, GitLab CI or similar
- Pipeline execution. The specific tool matters less than how the pipeline is structured.
- Docker
- Containerisation, which is what makes build-once-promote-everywhere practical.
- Terraform or equivalent
- Infrastructure as code, with state management as the operational concern.
- Kubernetes or a managed platform
- Orchestration where the complexity is warranted, which is less often than it is adopted.
- Prometheus and Grafana, or a managed equivalent
- Metrics and dashboards, including the ones that indicate user impact.
- OpenTelemetry
- Tracing across services, which is what makes distributed failures diagnosable.
- A secret manager
- Credentials kept out of repositories and environment files, and rotated.
Related skills that frequently appear on the same specification: Go, Microsoft Azure, AWS, Kubernetes, Terraform.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
What they changed about delivery
The outcome that defines the role, rather than the tools used to reach it.
- Strong answer: Can describe deployment frequency or lead time before and after, and what specifically changed.
- Warning sign: Lists tools without being able to say what improved.
Pipeline design
Reveals whether they understand why a pipeline is structured as it is.
- Strong answer: Builds once and promotes, fails fast, keeps it quick enough that nobody routes around it.
- Warning sign: Rebuilds at each stage, or has a pipeline slow enough that the team bypasses it.
Infrastructure state management
Where infrastructure as code causes real incidents.
- Strong answer: Remote state with locking, understands drift, reviews plans carefully, and has recovered from a state problem.
- Warning sign: Local state files, or applies without reading the plan.
Alerting and on-call
Separates engineers who monitor from engineers who are woken up.
- Strong answer: Alerts on user-visible symptoms, keeps the pager quiet, and has deleted alerts that were not actionable.
- Warning sign: Alerts on every metric threshold and treats noise as normal.
Secret management
A common and serious failure with a straightforward correct answer.
- Strong answer: Secrets in a manager, injected at runtime, rotated, never in the repository.
- Warning sign: Credentials in environment files committed to source control.
An incident they handled
Operational judgement is learned in incidents.
- Strong answer: Describes detection, diagnosis, mitigation and what changed afterwards, without blaming individuals.
- Warning sign: Has never been on call, or describes the cause as human error and stops there.
When not to use Kubernetes
The clearest test of judgement over enthusiasm in this field.
- Strong answer: Can name cases where a managed platform or plain containers would serve better, and has recommended against it.
- Warning sign: Treats it as the default for every workload regardless of team size.
Warning signs in a DevOps codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
Rebuilding at each stage
- What you see: A pipeline that builds separately for test, staging and production.
- What it costs: The artefact tested is not the artefact deployed, which defeats the purpose of testing.
- The fix: Build once, tag it, promote the same artefact. Configuration varies per environment; the build does not.
Manual production changes
- What you see: Fixes applied directly to running infrastructure to resolve an incident.
- What it costs: Drift between declared and actual state, and a subsequent deployment that silently reverts the fix.
- The fix: Fix forward through the pipeline. Where an emergency change is unavoidable, reconcile it into code the same day.
Alert fatigue
- What you see: Alerts firing constantly, with an unwritten understanding that most can be ignored.
- What it costs: The one that matters is missed, and on-call becomes something people leave jobs over.
- The fix: Alert on user-visible symptoms with a clear action. Delete anything that has never required one.
Secrets in the repository
- What you see: Credentials in environment files, configuration or pipeline definitions.
- What it costs: Exposure to everyone with repository access, permanently, including in history after removal.
- The fix: Use a secret manager, inject at runtime, rotate anything that has ever been committed.
Kubernetes for two services
- What you see: A cluster adopted for a workload a managed container service would run.
- What it costs: A large operational surface that someone must maintain, for capability the team does not use.
- The fix: Start with the simplest platform that meets the requirement. Move up when a specific need forces it.
Pipelines nobody trusts
- What you see: A test suite with known intermittent failures that people re-run until green.
- What it costs: The signal is gone. Real failures are indistinguishable from flakes and get ignored.
- The fix: Treat a flaky test as a broken test. Quarantine and fix it, because tolerating one teaches the team to ignore all of them.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a DevOps engineer.
- Junior
- Maintains existing pipelines and infrastructure definitions. Needs review on anything touching production.
- Mid-level
- Owns the delivery path for a service, including its pipeline, infrastructure and alerts. Participates in on-call.
- Senior
- Owns the platform: deployment strategy, infrastructure architecture, observability and incident process. Can argue against complexity the team does not need.
- Staff
- Owns delivery across teams, the security and compliance baseline, the cost model, and the internal tooling that lets product teams work without a specialist in the loop.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
Pipeline build or rebuild
- Usual team: One engineer.
- What governs it: Scoped by service count and test suite reliability. Flaky tests are usually the real project.
Infrastructure as code adoption
- Usual team: One engineer.
- What governs it: Importing an existing estate is slower than greenfield and reveals configuration nobody documented.
Observability implementation
- Usual team: One engineer plus the owning teams.
- What governs it: Instrumentation is straightforward; deciding what to alert on requires the teams and is the harder half.
Platform engineering
- Usual team: Two engineers.
- What governs it: Ongoing rather than a project. Judged by whether product teams need a specialist to deploy.
Incident response improvement
- Usual team: One senior engineer.
- What governs it: Scoped by current on-call pain. Usually starts by deleting alerts rather than adding them.
Migration work you may actually be hiring for
A large share of DevOps work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Manual deployment to an automated pipeline
- Why teams do it: Repeatability, speed and the ability to roll back without improvisation.
- What to watch: Automate the existing process first, then improve it. Teams that redesign the process and automate it simultaneously cannot tell which change caused a problem.
Hand-managed infrastructure to infrastructure as code
- Why teams do it: Reproducibility, reviewable change and real disaster recovery.
- What to watch: Import rather than recreate, and expect to find configuration nobody remembers making. Do it one workload at a time and freeze console access to that workload once it is under code.
Virtual machines to containers
- Why teams do it: Consistent environments and the ability to build once and run anywhere.
- What to watch: Stateful workloads and anything depending on the host are the hard cases. Move stateless services first and leave databases to a managed service rather than containerising them for symmetry.
Threshold alerts on everything to symptom-based alerting
- Why teams do it: A pager that means something, and an on-call rotation people will stay in.
- What to watch: Delete before you add. Start from the user-visible symptoms that constitute an outage and work back, rather than trying to classify the alerts you already have.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a DevOps engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- Whether the emphasis is delivery automation, cloud infrastructure, or operations, since these are different jobs.
- How deployment works today and how long it takes from approval to live.
- Whether infrastructure is defined in code, and how much of the estate is not.
- Whether the person joins an on-call rotation and what its current state is.
- What the team size is, since below a certain size a full-time role is hard to justify.
- Whether compliance requirements mean the pipeline must produce audit evidence.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
DevOps work is counted by the US Bureau of Labor Statistics under Network and Computer Systems Administrators. That classification is broader than the technology itself, so treat the figures as the shape of the market a DevOps engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 314,340 people in this occupation, with a median annual wage of $99,130.
The spread matters more than the midpoint. The 90th percentile is about 2.5 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a DevOps engineer can sit at $62,640 and $155,050 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for DevOps frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Network and Computer Systems Administrators | 314,340 | $78,010 | $99,130 | $126,640 | $155,050 |
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
| Computer Network Architects | 179,740 | $104,620 | $134,050 | $168,200 | $202,680 |
| Computer and Information Systems Managers | 670,570 | $138,060 | $175,140 | $220,730 | $297,510 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Network and Computer Systems Administrators, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.7. San Jose sits at the top with a median of $133,360; Pittsburgh sits at the bottom with $80,440. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 2,460 | $133,360 | +35% | 1.07 |
| San Francisco, CA | 3,910 | $129,680 | +31% | 0.81 |
| Washington, D.C. | 9,920 | $125,430 | +27% | 1.57 |
| Baltimore, MD | 4,620 | $122,950 | +24% | 1.69 |
| New York, NY | 17,690 | $119,390 | +20% | 0.92 |
| Boston, MA | 6,760 | $116,770 | +18% | 1.24 |
| Los Angeles, CA | 8,770 | $106,290 | +7% | 0.69 |
| San Diego, CA | 2,510 | $105,170 | +6% | 0.81 |
| Denver, CO | 4,670 | $105,090 | +6% | 1.44 |
| Austin, TX | 5,030 | $104,520 | +5% | 1.92 |
| Seattle, WA | 5,530 | $104,440 | +5% | 1.31 |
| Raleigh, NC | 2,730 | $103,380 | +4% | 1.82 |
| Dallas-Fort Worth, TX | 11,740 | $103,260 | +4% | 1.43 |
| Chicago, IL | 6,950 | $103,170 | +4% | 0.76 |
| Portland, OR | 2,820 | $102,950 | +4% | 1.15 |
| Minneapolis-St. Paul, MN | 3,200 | $102,790 | +4% | 0.81 |
| Tampa, FL | 4,720 | $101,560 | +2% | 1.61 |
| Houston, TX | 6,330 | $101,430 | +2% | 0.95 |
| Atlanta, GA | 6,140 | $101,000 | +2% | 1.05 |
| Philadelphia, PA | 4,050 | $100,460 | +1% | 0.69 |
| Salt Lake City, UT | 1,130 | $99,110 | 0% | 0.68 |
| Miami, FL | 6,340 | $97,180 | -2% | 1.11 |
| Detroit, MI | 3,010 | $96,400 | -3% | 0.78 |
| Phoenix, AZ | 4,370 | $93,730 | -5% | 0.91 |
| Kansas City, MO | 2,760 | $91,980 | -7% | 1.25 |
| Orlando, FL | 4,260 | $91,800 | -7% | 1.49 |
| Charlotte, NC | 4,180 | $89,990 | -9% | 1.52 |
| Pittsburgh, PA | 1,570 | $80,440 | -19% | 0.70 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. Washington, D.C., Baltimore, Austin, Raleigh, Tampa, Charlotte each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Tool familiarity without delivery outcomes. Ask what deployment frequency and lead time were before and after their work. The answer separates the field quickly.
System administration presented as DevOps. Both are valid, and they are different. Decide which you need and interview for it explicitly.
Complexity enthusiasm. Ask when they would not use Kubernetes. Someone with no answer will build something your team cannot run.
No on-call experience. Design judgement about reliability comes from being woken up by your own systems.
Hiring DevOps developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of DevOps engineer who will be available.
- New York, NY $119,390 median
- Seattle, WA $104,440 median
- San Jose, CA $133,360 median
- Washington, D.C. $125,430 median
- San Francisco, CA $129,680 median
- Dallas-Fort Worth, TX $103,260 median
- Los Angeles, CA $106,290 median
- Boston, MA $116,770 median
- Chicago, IL $103,170 median
- Atlanta, GA $101,000 median
- Austin, TX $104,520 median
- Phoenix, AZ $93,730 median
- Philadelphia, PA $100,460 median
- Minneapolis-St. Paul, MN $102,790 median
- Denver, CO $105,090 median
- Detroit, MI $96,400 median
- Houston, TX $101,430 median
- Charlotte, NC $89,990 median
- San Diego, CA $105,170 median
- Salt Lake City, UT $99,110 median
- Miami, FL $97,180 median
- Portland, OR $102,950 median
- Baltimore, MD $122,950 median
- Tampa, FL $101,560 median
- Orlando, FL $91,800 median
- Raleigh, NC $103,380 median
- Kansas City, MO $91,980 median
- Pittsburgh, PA $80,440 median
Frequently asked questions
Do we need a dedicated DevOps engineer?
Below roughly ten engineers, usually not full time, and the work is better handled by a strong developer with operational interest plus periodic specialist help. Above that, deployment and infrastructure become a real drag on everyone else's time and a dedicated person pays for themselves. The clearest signal is whether shipping has become something people schedule around.
What should a DevOps engineer actually deliver?
Deployments that are frequent, boring and reversible; infrastructure defined in code and reproducible; alerts that fire when users are affected and rarely otherwise; and developers who can ship without asking them. If after six months the team still cannot deploy without their involvement, the role has become a bottleneck rather than a solution.
Is DevOps a role or a practice?
It began as a practice and the market uses it as a title, so arguing about it is unproductive. What matters is being specific in the brief: pipelines and delivery automation, infrastructure and cloud architecture, and on-call operations are three different emphases, and a specification that implies all three equally will produce a confusing shortlist.
Should we adopt Kubernetes?
Only for a reason you can state. It solves real problems around orchestration, scaling and portability at a real cost in operational surface. For a handful of services, a managed container platform does the job with a fraction of the burden. The question to ask is who will operate the cluster at three in the morning, and whether that person exists.
How do we know if our deployment process is bad?
Two symptoms. First, whether people avoid deploying near the end of the week, which indicates deployments are risky rather than routine. Second, how long it takes from a change being approved to being live. If the answer is days, the process is the constraint on everything else the team does.
Is DevOps just cloud work?
Largely overlapping in practice, and not identical. Cloud expertise is about a specific platform's services and cost model. DevOps is about how changes reach production safely, which applies on any infrastructure. Most roles want both, and it is worth knowing which one the work actually weighs toward.
How do we reduce on-call burden?
Start by deleting alerts rather than adding automation. Most painful on-call rotations are painful because of alerts that do not require action, which trains people to ignore the pager. After that, the levers are fixing the top recurring causes properly rather than repeatedly mitigating them, and making sure alerts point to something the responder can actually do.
What does a DevOps engineer need from us on day one?
Access, which sounds trivial and is the usual blocker. Cloud accounts, repositories, the pipeline tool, the monitoring system and production credentials all need to exist before they can do anything. Teams that spend a new engineer's first two weeks on access requests have paid full rate for nothing.