Hire QA Automation developers
QA automation is software engineering aimed at confidence, and hiring for it means valuing judgement about what to test over the ability to write many tests.
What QA Automation actually is
QA automation is the practice of building and maintaining automated checks that tell a team whether software still works. In practice it spans test strategy, writing tests at several levels, building the infrastructure they run on, and the harder organisational work of making the results trusted enough that people act on them.
The role is often framed as writing tests, which undersells it and produces poor hires. The valuable skill is deciding what deserves a test, at which level, and what does not. A suite of three thousand tests that takes an hour and fails intermittently is worse than two hundred fast reliable ones, because the first teaches everyone to ignore failures and the second does not.
The distinction from manual testing matters when writing a brief. Exploratory testing, where a person probes software looking for problems nobody anticipated, is a genuine and different skill that automation does not replace. Automation is for regression: confirming that what worked yesterday still works. Organisations that automate everything and stop exploring find the bugs that automation was never going to catch.
The part that separates seniors from mid-levels
The test pyramid is the organising idea, and misapplying it is the most common structural problem. Many fast unit tests, fewer integration tests, and a small number of end-to-end tests over the flows that actually matter. Teams that invert this, with hundreds of browser tests and few unit tests, get slow suites that fail unpredictably and never tell anyone where the problem is. An engineer who can explain where a given check belongs is worth more than one who can write any of them.
Flakiness is the second concern and it is the thing that destroys the value of a suite. A test that fails occasionally without a real defect trains everyone to re-run rather than investigate, at which point the suite has stopped being a signal. Most flakiness comes from timing assumptions: waiting a fixed period rather than for a condition, or tests sharing state. Engineers who treat a flaky test as a broken test, to be fixed or removed, are describing the only sustainable position.
The third is test data and environment control. Tests that depend on data somebody set up by hand, or on a shared environment other people are changing, fail for reasons unrelated to the code. Building tests that create what they need and clean up after themselves is unglamorous and it is what makes a suite reliable enough to run on every change.
Where QA Automation is used
The label “QA Automation developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Continuous delivery
Teams deploying frequently, where automated confidence is what makes that safe.
Regulated software
Financial, healthcare and safety-critical systems where evidence of testing is a requirement.
Complex web applications
Products where manual regression testing has become impractical.
Mobile applications
Device and version fragmentation making manual coverage impossible.
API and integration testing
Service estates where contract testing between components prevents a large class of failure.
Legacy systems
Characterisation tests around code nobody understands, enabling safe change.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
Test tooling moves quickly and browser automation libraries in particular release frequently, so currency is judged by the tools a team uses now rather than by any long-lived version.
QA Automation itself is not versioned as a single product, so the useful equivalent is the support status of the tools a QA automation engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Node.js | 26.10.0 | 2026-09-22 | 6 | 2029-04-30 |
| Python | 3.14.7 | 2026-08-05 | 5 | 2030-10-31 |
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
| Docker Engine | 29.8.1 | 2026-09-15 | 1 | 2026-12-04 |
| GitLab | 19.4.1 | 2026-09-22 | 3 | 2026-12-17 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for QA Automation alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- Playwright or Cypress
- Browser automation. Playwright has become the common default for new work.
- A unit testing framework
- Whatever the application's language uses, where most tests should live.
- Testcontainers
- Real dependencies in containers for integration tests instead of mocks.
- Contract testing
- Verifying service boundaries without full end-to-end runs.
- A CI pipeline
- Where tests run. Speed and reliability here determine whether they are trusted.
- Appium or platform tools
- Mobile automation, a distinct and harder discipline.
- k6 or similar
- Load and performance testing, usually a separate specialism.
- Test reporting
- Making failures legible, including flakiness tracking over time.
Related skills that frequently appear on the same specification: Python, Java, JavaScript, TypeScript, DevOps.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
What they choose not to test
The most informative question, because it tests judgement rather than capability.
- Strong answer: Can explain what is not worth automating, and has deleted tests that were not earning their maintenance.
- Warning sign: Believes more coverage is always better, or quotes a coverage percentage as a goal.
Handling flaky tests
Flakiness is what kills suites.
- Strong answer: Treats a flaky test as broken, fixes the timing or state assumption, and quarantines rather than tolerating.
- Warning sign: Adds a retry, or accepts that some tests fail sometimes.
Where a given check belongs
Tests whether they understand the pyramid rather than reciting it.
- Strong answer: Pushes checks to the lowest level that can catch the problem, and reserves end-to-end tests for revenue-critical paths.
- Warning sign: Writes browser tests for logic a unit test would cover.
Test data and isolation
The most common source of unreliable suites.
- Strong answer: Tests create their own data and clean up, and do not depend on shared environment state.
- Warning sign: Relies on a seeded shared environment that somebody maintains by hand.
Suite speed
A slow suite is a suite people route around.
- Strong answer: Knows how long the suite takes, parallelises, and treats duration as a metric to defend.
- Warning sign: Has no idea how long it takes, or accepts an hour without concern.
Working with developers
Automation that sits outside the team does not get maintained.
- Strong answer: Tests live with the application code, developers contribute, and failures are the team's responsibility.
- Warning sign: Describes a separate suite owned solely by QA, which is the arrangement that reliably decays.
A bug automation missed
Honest about the limits of the practice.
- Strong answer: Can describe one, and what changed as a result, including where exploratory testing was the answer.
- Warning sign: Believes comprehensive automation removes the need for exploratory testing.
Warning signs in a QA Automation codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
An inverted pyramid
- What you see: Hundreds of browser tests and few unit tests.
- What it costs: A slow, brittle suite that takes hours and does not indicate where a failure originated.
- The fix: Push checks down to the lowest level that catches the problem. Keep end-to-end tests few and aimed at the flows that earn money.
Tolerated flakiness
- What you see: Tests known to fail occasionally, with re-running as the accepted response.
- What it costs: Everyone learns to ignore failures, so a real one is missed. The suite has become decoration.
- The fix: Treat a flaky test as a broken test. Quarantine it immediately and fix or delete it.
Fixed waits
- What you see: Tests pausing for a set number of seconds.
- What it costs: Slow when the wait is too long and flaky when it is too short, which varies by machine and load.
- The fix: Wait for a condition, never for a duration. Modern tools do this by default and the pattern usually indicates older habits.
Shared mutable test data
- What you see: Tests depending on records that exist in a shared environment.
- What it costs: Tests interfere with each other and fail for reasons unrelated to the code, especially in parallel.
- The fix: Each test creates what it needs and removes it. This is the prerequisite for parallel execution.
Coverage as the target
- What you see: A percentage treated as a goal in its own right.
- What it costs: Tests written to touch lines rather than to verify behaviour, which provides confidence without substance.
- The fix: Measure whether tests catch real regressions. Coverage is a diagnostic, not an objective.
A separate QA-owned suite
- What you see: Tests in their own repository, maintained only by QA, run after development finishes.
- What it costs: It drifts from the application, breaks constantly, and becomes a bottleneck rather than a safety net.
- The fix: Tests live with the code and the whole team owns failures. This single change fixes more than any tooling decision.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a QA automation engineer.
- Junior
- Writes tests within an established framework. Needs review on level selection and on flakiness.
- Mid-level
- Owns test coverage for a feature area, builds reliable tests, and keeps them fast and isolated.
- Senior
- Owns test strategy, the infrastructure, the pipeline integration, and the judgement about what is not worth automating.
- Staff
- Owns quality practice across teams, the relationship between automated and exploratory testing, and the standards that make results trusted.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
Building a suite from nothing
- Usual team: One engineer alongside the development team.
- What governs it: Start with the few flows that would be most damaging to break. Comprehensive coverage attempted first never finishes.
Stabilising an unreliable suite
- Usual team: One senior engineer, time-boxed.
- What governs it: Usually higher value than adding tests. Begins with measuring which tests fail intermittently.
Pipeline integration
- Usual team: One engineer with DevOps overlap.
- What governs it: Scoped by suite duration. A slow suite must be parallelised before it can gate anything.
Characterisation tests on legacy code
- Usual team: One engineer with a developer who knows the system.
- What governs it: Enables safe change. The constraint is understanding what the current behaviour actually is.
Mobile automation
- Usual team: One engineer with mobile experience.
- What governs it: Substantially harder than web. Device management is a real ongoing cost.
Migration work you may actually be hiring for
A large share of QA Automation work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Manual regression testing to automation
- Why teams do it: Manual regression does not scale with release frequency and is the first thing cut under deadline.
- What to watch: Automate the highest-risk flows first rather than working through a test plan in order. Keep exploratory testing; automation replaces regression checking, not investigation.
An inverted pyramid to a balanced suite
- Why teams do it: Faster feedback and failures that indicate where the problem is.
- What to watch: Add lower-level tests before removing browser tests, so coverage is never lost in the transition. Delete the end-to-end tests only once something else covers the case.
A separate QA-owned suite to tests alongside the application
- Why teams do it: Tests that stay in step with the code and failures the whole team owns.
- What to watch: The organisational change is harder than the technical one. It only works if developers are genuinely accountable for failures rather than forwarding them.
An older browser automation tool to Playwright
- Why teams do it: Better reliability, parallel execution and cross-browser coverage.
- What to watch: Rewriting a large suite is rarely justified on tooling grounds alone. Write new tests in the new tool and migrate the old ones as they need attention anyway.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a QA automation engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- What exists today: no tests, an unreliable suite, or a healthy one needing extension.
- Whether tests live with the application code and who owns failures.
- How long the suite takes and whether it gates deployment.
- Whether the scope is web, API, mobile or all three, since mobile is substantially harder.
- Whether exploratory testing will continue, because automation does not replace it.
- Whether the person is expected to build infrastructure as well as write tests.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
QA Automation work is counted by the US Bureau of Labor Statistics under Software Quality Assurance Analysts and Testers. That classification is broader than the technology itself, so treat the figures as the shape of the market a QA automation engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 186,740 people in this occupation, with a median annual wage of $104,300.
The spread matters more than the midpoint. The 90th percentile is about 2.7 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a QA automation engineer can sit at $61,440 and $167,010 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for QA Automation frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Software QA Analysts and Testers | 186,740 | $80,310 | $104,300 | $133,180 | $167,010 |
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Software QA Analysts and Testers, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 2.0. San Jose sits at the top with a median of $166,530; Charlotte sits at the bottom with $81,990. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 6,480 | $166,530 | +60% | 4.76 |
| San Francisco, CA | 5,880 | $134,630 | +29% | 2.06 |
| New York, NY | 11,460 | $130,540 | +25% | 1.00 |
| Baltimore, MD | 2,670 | $129,880 | +25% | 1.64 |
| Seattle, WA | 6,650 | $129,690 | +24% | 2.65 |
| Washington, D.C. | 8,460 | $126,350 | +21% | 2.25 |
| Denver, CO | 3,620 | $123,690 | +19% | 1.87 |
| Boston, MA | 6,050 | $122,210 | +17% | 1.86 |
| Los Angeles, CA | 7,200 | $122,210 | +17% | 0.96 |
| San Diego, CA | 2,280 | $120,210 | +15% | 1.24 |
| Raleigh, NC | 1,760 | $110,800 | +6% | 1.97 |
| Portland, OR | 1,550 | $109,760 | +5% | 1.07 |
| Atlanta, GA | 5,270 | $105,800 | +1% | 1.52 |
| Chicago, IL | 5,010 | $105,580 | +1% | 0.93 |
| Orlando, FL | 1,940 | $104,250 | 0% | 1.15 |
| Dallas-Fort Worth, TX | 8,840 | $103,210 | -1% | 1.82 |
| Austin, TX | 3,340 | $103,070 | -1% | 2.15 |
| Philadelphia, PA | 2,650 | $102,740 | -1% | 0.76 |
| Detroit, MI | 2,140 | $102,390 | -2% | 0.93 |
| Minneapolis-St. Paul, MN | 2,190 | $102,130 | -2% | 0.93 |
| Phoenix, AZ | 3,160 | $100,740 | -3% | 1.11 |
| Miami, FL | 1,960 | $100,050 | -4% | 0.58 |
| Tampa, FL | 1,930 | $98,870 | -5% | 1.11 |
| Houston, TX | 2,420 | $98,690 | -5% | 0.61 |
| Kansas City, MO | 1,430 | $94,970 | -9% | 1.09 |
| Salt Lake City, UT | 1,610 | $82,940 | -20% | 1.62 |
| Pittsburgh, PA | 830 | $82,900 | -21% | 0.62 |
| Charlotte, NC | 3,900 | $81,990 | -21% | 2.39 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Baltimore, Seattle, Washington, D.C., Denver each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Test writing without engineering practice. Test code is code. Ask about how they structure and maintain it, because unmaintainable tests get deleted.
No judgement about what not to test. Ask what they have deleted. Volume is easy and discrimination is what you are paying for.
Tolerance for flakiness. The single most important cultural question. A tolerant answer predicts a suite nobody trusts.
Isolation from the development team. Separately owned suites decay. Confirm the arrangement before hiring into it.
Hiring QA Automation developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of QA automation engineer who will be available.
- New York, NY $130,540 median
- Seattle, WA $129,690 median
- San Jose, CA $166,530 median
- Washington, D.C. $126,350 median
- San Francisco, CA $134,630 median
- Dallas-Fort Worth, TX $103,210 median
- Los Angeles, CA $122,210 median
- Boston, MA $122,210 median
- Chicago, IL $105,580 median
- Atlanta, GA $105,800 median
- Austin, TX $103,070 median
- Phoenix, AZ $100,740 median
- Philadelphia, PA $102,740 median
- Minneapolis-St. Paul, MN $102,130 median
- Denver, CO $123,690 median
- Detroit, MI $102,390 median
- Houston, TX $98,690 median
- Charlotte, NC $81,990 median
- San Diego, CA $120,210 median
- Salt Lake City, UT $82,940 median
- Miami, FL $100,050 median
- Portland, OR $109,760 median
- Baltimore, MD $129,880 median
- Tampa, FL $98,870 median
- Orlando, FL $104,250 median
- Raleigh, NC $110,800 median
- Kansas City, MO $94,970 median
- Pittsburgh, PA $82,900 median
Frequently asked questions
How much of our testing should be automated?
Regression testing, which is checking that what worked still works, should be automated as far as it is worth maintaining. Exploratory testing, where a person looks for problems nobody thought of, should not be, because automation can only check what somebody anticipated. Organisations that automate everything and stop exploring find the class of bug automation was never going to catch.
What is a reasonable amount of test coverage?
Coverage is a diagnostic rather than a target, and setting a percentage goal produces tests written to touch lines rather than to verify behaviour. The better question is whether your tests catch real regressions before customers do. A team with modest coverage over the paths that matter is in a better position than one with high coverage and an unreliable suite.
Why does nobody trust our test suite?
Almost always flakiness. Once tests fail intermittently without real defects, people learn to re-run rather than investigate, and from that point the suite provides no signal. Stabilising an unreliable suite is nearly always more valuable than adding to it, and it is usually the work that gets skipped.
Should QA be a separate team?
Separate ownership of the test suite is the arrangement that reliably decays: the tests drift from the application, break constantly, and become a bottleneck. Quality engineers embedded in teams, with tests living alongside the application code and the whole team responsible for failures, works considerably better. The specialism is real; the separation is what causes problems.
Playwright or Cypress?
Playwright has become the common default for new work, with better cross-browser support, parallel execution and a more capable automation model. Cypress has a strong developer experience and a large installed base. Either is a reasonable choice, and for an existing suite the cost of switching rarely justifies itself on tooling grounds alone.
How long should our test suite take?
Short enough that developers run it before pushing, which in practice means minutes rather than tens of minutes for the fast layers. Longer end-to-end suites can run separately. The real threshold is the point at which people start routing around it, and once that happens the suite is no longer protecting anything.
Do we need automation engineers or can developers write the tests?
Developers should write most tests, particularly at the unit and integration levels, and in healthy teams they do. A dedicated automation engineer earns their place by owning strategy, building the infrastructure, keeping the suite fast and reliable, and bringing judgement about what belongs where. That is a different contribution from writing more tests.
What should we automate first?
The few flows whose failure would be most damaging: signing in, paying, and whatever your core transaction is. A handful of reliable tests over those is worth more than broad shallow coverage, and it gives you something trustworthy to build on. Teams that start by trying to cover everything usually produce a suite nobody believes.