Hire Data Engineering developers
Data engineering builds the pipelines everything else depends on, and hiring for it means valuing reliability and data quality over familiarity with any particular tool.
What Data Engineering actually is
Data engineering is the discipline of moving data from where it is produced to where it is used, reliably and in a shape people can work with. In practice that means building and operating pipelines, modelling data for analysis, managing warehouses and lakes, and being responsible when the numbers in a report are wrong. It sits between the systems that generate data and the analysts and models that consume it.
The role is often confused with data science, and the distinction is worth being clear about. A data scientist asks questions of data and builds models. A data engineer makes sure the data exists, arrives on time, and means what it claims to mean. Organisations that hire the second when they needed the first, or more commonly the reverse, end up with expensive people blocked on work that is not theirs.
What has changed in recent years is that the warehouse became powerful enough to do the transformation itself. The older pattern transformed data before loading it; the current one loads raw data and transforms it inside the warehouse with SQL. That shift means modern data engineering is closer to software engineering practice, with version control, testing and review applied to transformation logic that used to live in undocumented scripts.
The part that separates seniors from mid-levels
Idempotency is the foundation and the thing to interview for first. Pipelines fail, get retried, and get re-run over historical periods. A pipeline that produces different results when run twice will eventually produce wrong numbers that nobody can explain, and it will do so quietly. Engineers who design for re-runnability talk about partitioned writes, deterministic transformations and the difference between appending and replacing a window of data.
Data quality is the second, and the difference between a competent data engineer and a valuable one. Bad data is worse than missing data, because missing data stops a report and bad data produces a confident wrong answer that somebody acts on. Testing at the boundary, asserting expectations about row counts, uniqueness, ranges and referential integrity, and failing loudly rather than continuing is the practice that separates them.
The third is understanding that a pipeline is a production system with users. Someone depends on a dashboard being correct by nine in the morning. That means monitoring, alerting, defined freshness expectations and a process when something breaks, exactly as for any service. Data teams that treat pipelines as scripts rather than as services are the ones where the business quietly stops trusting the numbers.
Where Data Engineering is used
The label “Data Engineering developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Analytics platforms
Building the warehouse and models that reporting and business intelligence run on.
Machine learning infrastructure
Feature pipelines and training data, where reproducibility matters more than anywhere else.
Product data
Event collection and behavioural data feeding product decisions and customer-facing features.
Financial and regulatory reporting
Where correctness and auditability are the requirement and the tolerance for error is zero.
Systems integration
Moving data between operational systems, often the least glamorous and most depended-upon work.
Migration and consolidation
Moving off legacy warehouses or consolidating several sources into one platform.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
Data engineering is a discipline rather than a product, so currency is judged by the platforms and patterns someone has worked with recently; the table below shows support status across common components of the stack.
Data Engineering itself is not versioned as a single product, so the useful equivalent is the support status of the tools a data engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Python | 3.14.7 | 2026-08-05 | 5 | 2030-10-31 |
| PostgreSQL | 18.6 | 2026-08-11 | 5 | 2030-11-14 |
| Apache Kafka | 4.3.1 | 2026-06-23 | 1 | 2024-11-06 |
| Elasticsearch | 9.5.4 | 2026-09-15 | 1 | 2027-07-15 |
| MongoDB Server | 8.3.11 | 2026-09-11 | 3 | 2029-10-31 |
| Redis | 8.10.2 | 2026-09-17 | 5 | 2030-09-01 |
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for Data Engineering alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- SQL
- The core skill. Modern transformation is mostly SQL executed in the warehouse.
- Python
- Orchestration, extraction and anything SQL cannot express.
- A warehouse
- Snowflake, BigQuery, Redshift or similar, where cost and performance characteristics differ substantially.
- dbt
- Transformation with version control, testing, documentation and lineage. Close to standard in current practice.
- Airflow, Dagster or Prefect
- Orchestration and scheduling, with dependency management and retry behaviour.
- Kafka or a managed stream
- Event streaming where near-real-time data is genuinely required.
- Object storage
- The landing zone for raw data, and increasingly the storage layer for open table formats.
- Data quality testing
- Assertions on freshness, uniqueness, ranges and relationships, run as part of the pipeline.
Related skills that frequently appear on the same specification: Machine Learning, Python, SQL, Data Science, AWS.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
Idempotency and re-runs
The single most important property of a pipeline and the most common omission.
- Strong answer: Designs so a re-run produces the same result, uses partitioned replacement, and can explain backfilling a historical window safely.
- Warning sign: Appends on every run, so a retry silently duplicates data.
Data quality testing
Separates engineers who move data from engineers who are accountable for it.
- Strong answer: Tests uniqueness, freshness, ranges and relationships, and fails the pipeline rather than passing bad data downstream.
- Warning sign: Relies on consumers noticing that numbers look wrong.
A pipeline failure they handled
Operational reality, and where judgement is formed.
- Strong answer: Describes detection, the effect on downstream consumers, the fix and the prevention.
- Warning sign: Has never been responsible when a report was wrong.
Warehouse cost
Warehouse spend is usually a significant line item and is driven by engineering decisions.
- Strong answer: Understands what drives cost in their platform, has optimised a model, and can attach a number to it.
- Warning sign: Has never seen the warehouse bill.
Modelling for consumers
Technically correct data that analysts cannot use has not solved the problem.
- Strong answer: Models for how the data will be queried, documents meaning, and talks to the people who use it.
- Warning sign: Exposes raw source structures and expects analysts to work it out.
Batch versus streaming
Streaming is frequently adopted where batch would do, at considerable cost in complexity.
- Strong answer: Can articulate when latency genuinely justifies streaming, and defaults to batch otherwise.
- Warning sign: Proposes streaming by default, or cannot explain what it costs operationally.
Schema changes upstream
Source systems change without warning and pipelines break.
- Strong answer: Detects schema drift, fails explicitly rather than silently dropping fields, and has a relationship with the source owners.
- Warning sign: Discovers upstream changes when a report is wrong.
Warning signs in a Data Engineering codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
Non-idempotent pipelines
- What you see: Runs that append unconditionally, so a retry duplicates data.
- What it costs: Silently inflated numbers, discovered weeks later when somebody questions a total.
- The fix: Replace a defined partition rather than appending. A re-run must produce the same result as a first run.
No tests on data
- What you see: Pipelines that validate code and never validate what flows through them.
- What it costs: Bad data reaches dashboards and decisions with full confidence attached.
- The fix: Assert uniqueness, freshness, ranges and referential integrity. Fail the pipeline rather than publishing.
Transformation logic nobody can review
- What you see: Business rules embedded in scheduled scripts or notebooks outside version control.
- What it costs: Nobody knows why a number is calculated as it is, and the person who did has left.
- The fix: Version-controlled, reviewed transformation with documented lineage.
Streaming where batch would do
- What you see: A real-time pipeline feeding a dashboard people look at each morning.
- What it costs: Substantially more operational complexity and cost for latency nobody needs.
- The fix: Start with batch. Move to streaming when a consumer has a genuine latency requirement they can state.
Silent schema drift handling
- What you see: Pipelines that ignore unexpected columns or coerce types quietly.
- What it costs: Fields silently dropped or corrupted, producing wrong results with no error anywhere.
- The fix: Fail explicitly on unexpected schema changes. A broken pipeline is much cheaper than wrong data.
Ungoverned warehouse cost
- What you see: Models re-processing full history on every run, and unused tables refreshed indefinitely.
- What it costs: A bill that grows without anyone being able to explain the increase.
- The fix: Process incrementally where possible, retire unused models, and attribute cost to the teams that own them.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a data engineer.
- Junior
- Builds transformations within an established framework. Needs review on idempotency and testing.
- Mid-level
- Owns a pipeline end to end including orchestration, tests and monitoring. Responds when it breaks.
- Senior
- Owns platform architecture, the modelling approach, data quality standards and warehouse cost. Works directly with data consumers on what they need.
- Staff
- Owns data contracts with source systems, governance and lineage across the organisation, and the platform strategy including build-versus-buy decisions.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
Warehouse and pipeline build
- Usual team: One to two data engineers plus a stakeholder per source.
- What governs it: Source data quality is the variable that determines the timeline and it is always worse than described.
Legacy pipeline migration
- Usual team: Two engineers, one who knows the existing system.
- What governs it: Undocumented business logic in old scripts is the hard part. Expect archaeology.
Data quality programme
- Usual team: One engineer plus data owners.
- What governs it: Technically straightforward; the work is agreeing what correct means with the people who use the data.
Warehouse cost reduction
- Usual team: One engineer, time-boxed.
- What governs it: Well-defined and usually high return on any platform that has grown without review.
Streaming implementation
- Usual team: One to two senior engineers.
- What governs it: Only justified by a stated latency requirement. Considerably harder to operate than batch.
Migration work you may actually be hiring for
A large share of Data Engineering work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Transform before load to load then transform in the warehouse
- Why teams do it: Raw data is preserved, transformations become reviewable SQL, and reprocessing history is possible.
- What to watch: Warehouse cost becomes an engineering concern rather than an infrastructure one. Set up cost visibility at the same time, not afterwards.
Scheduled scripts to an orchestrator
- Why teams do it: Dependency management, retries, visibility and alerting instead of cron and hope.
- What to watch: Migrate one pipeline at a time. The hard part is usually discovering what the old scripts actually did, since the business logic in them is rarely documented.
An on-premises warehouse to a cloud warehouse
- Why teams do it: Elastic compute, separation of storage and compute, and no capacity planning.
- What to watch: Cost shifts from fixed to variable, which is an advantage only if somebody watches it. Dialect differences in SQL are the other main source of work.
Untested transformations to tested, version-controlled models
- Why teams do it: Data quality becomes something the pipeline enforces rather than something consumers discover.
- What to watch: Start with the models that feed the most important reports. Adding tests everywhere at once produces noise that gets ignored.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a data engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- Which warehouse or platform is in use, since cost and optimisation differ substantially between them.
- What the source systems are and who owns them, because that relationship determines pipeline stability.
- Whether transformation is version-controlled and tested today, or lives in scripts nobody reviews.
- Whether anyone is genuinely blocked on latency, which is the only good reason to require streaming experience.
- Who consumes the data and whether they currently trust it.
- Whether the person will be on call when a pipeline fails before a morning report.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
Data Engineering work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a data engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.
The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a data engineer can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for Data Engineering frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Data Scientists | 262,440 | $85,660 | $120,230 | $158,880 | $199,130 |
| Database Architects | 67,140 | $109,370 | $139,500 | $169,290 | $204,000 |
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
| Database Administrators | 69,990 | $79,610 | $104,620 | $135,460 | $163,320 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 6,060 | $185,080 | +54% | 3.16 |
| San Francisco, CA | 10,460 | $170,110 | +41% | 2.61 |
| Seattle, WA | 8,370 | $164,740 | +37% | 2.38 |
| New York, NY | 23,160 | $135,980 | +13% | 1.45 |
| Baltimore, MD | 1,090 | $134,320 | +12% | 0.48 |
| Charlotte, NC | 4,420 | $132,460 | +10% | 1.93 |
| Washington, D.C. | 9,260 | $132,200 | +10% | 1.75 |
| Boston, MA | 7,930 | $132,040 | +10% | 1.74 |
| San Diego, CA | 2,830 | $130,990 | +9% | 1.09 |
| Minneapolis-St. Paul, MN | 3,250 | $129,780 | +8% | 0.99 |
| Los Angeles, CA | 9,850 | $129,740 | +8% | 0.93 |
| Portland, OR | 1,700 | $129,600 | +8% | 0.83 |
| Dallas-Fort Worth, TX | 10,120 | $127,750 | +6% | 1.48 |
| Miami, FL | 3,040 | $127,450 | +6% | 0.64 |
| Austin, TX | 3,730 | $127,360 | +6% | 1.71 |
| Raleigh, NC | 1,990 | $120,710 | 0% | 1.59 |
| Salt Lake City, UT | 2,970 | $114,990 | -4% | 2.13 |
| Phoenix, AZ | 3,480 | $114,540 | -5% | 0.87 |
| Denver, CO | 4,510 | $112,520 | -6% | 1.66 |
| Tampa, FL | 1,730 | $109,990 | -9% | 0.71 |
| Philadelphia, PA | 6,480 | $109,910 | -9% | 1.32 |
| Atlanta, GA | 6,820 | $108,940 | -9% | 1.40 |
| Chicago, IL | 7,940 | $107,640 | -10% | 1.04 |
| Houston, TX | 4,060 | $106,750 | -11% | 0.73 |
| Orlando, FL | 1,440 | $106,590 | -11% | 0.61 |
| Detroit, MI | 3,810 | $103,330 | -14% | 1.18 |
| Kansas City, MO | 1,260 | $99,870 | -17% | 0.68 |
| Pittsburgh, PA | 2,270 | $96,670 | -20% | 1.21 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Data science background without engineering practice. Common and consequential. Test version control, testing, orchestration and on-call behaviour explicitly.
Tool familiarity without reliability thinking. Ask how a pipeline behaves when re-run. This separates the field faster than any tool question.
No accountability experience. Ask about a time the numbers were wrong. Someone who has never owned that has not felt the constraint.
Building without talking to consumers. Ask how they decided what to model. Technically clean data nobody can use is a common and expensive outcome.
Hiring Data Engineering developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of data engineer who will be available.
- New York, NY $135,980 median
- Seattle, WA $164,740 median
- San Jose, CA $185,080 median
- Washington, D.C. $132,200 median
- San Francisco, CA $170,110 median
- Dallas-Fort Worth, TX $127,750 median
- Los Angeles, CA $129,740 median
- Boston, MA $132,040 median
- Chicago, IL $107,640 median
- Atlanta, GA $108,940 median
- Austin, TX $127,360 median
- Phoenix, AZ $114,540 median
- Philadelphia, PA $109,910 median
- Minneapolis-St. Paul, MN $129,780 median
- Denver, CO $112,520 median
- Detroit, MI $103,330 median
- Houston, TX $106,750 median
- Charlotte, NC $132,460 median
- San Diego, CA $130,990 median
- Salt Lake City, UT $114,990 median
- Miami, FL $127,450 median
- Portland, OR $129,600 median
- Baltimore, MD $134,320 median
- Tampa, FL $109,990 median
- Orlando, FL $106,590 median
- Raleigh, NC $120,710 median
- Kansas City, MO $99,870 median
- Pittsburgh, PA $96,670 median
Frequently asked questions
What is the difference between a data engineer and a data scientist?
A data engineer builds and operates the systems that make data available, correct and timely. A data scientist analyses that data and builds models from it. They are complementary and frequently confused in job specifications. Hiring a data scientist into a role that is actually pipeline work is a common and expensive mistake, and the person usually leaves.
Do we need a data engineer or would an analyst do?
If the question is mostly about producing reports from data that already arrives reliably, an analyst with strong SQL will go further than you expect. If data does not arrive reliably, or arrives wrong, or lives in several systems that disagree, that is engineering work. The clearest signal is whether people currently spend more time fixing data than analysing it.
Should we build a warehouse or use the application database for reporting?
The application database works until reporting queries start competing with the application for resources, or until you need to combine several sources. Both happen sooner than teams expect. A separate analytical store removes the contention and allows a data model designed for questions rather than for transactions, which is a different shape entirely.
How do we know if our data is trustworthy?
Ask whether the pipelines test anything. Most do not. If nothing asserts that a table has the expected number of rows, that identifiers are unique, that values are in range and that data is fresh, then nobody finds out about a problem until somebody questions a number, which means the wrong numbers before it went unquestioned.
Do we need real-time data?
Usually less than people say. Real-time is attractive and it carries substantial operational complexity and cost. The test is whether a consumer can state a decision they would make differently with fresher data. Where there is a genuine answer, streaming is worth it. Where the answer is that it would be nice, batch is the better engineering decision.
Why is our warehouse bill increasing?
Usually models reprocessing full history on every run, tables that are refreshed but no longer used, and inefficient transformations that were written when the data was small. Warehouse cost is driven by engineering decisions, so it responds well to engineering attention. A focused review typically finds substantial savings on any platform that has grown without governance.
What should we build first?
Get raw data landing reliably and reproducibly before you model anything. Teams that start with the dashboard build transformations on a foundation that is still shifting, and end up rebuilding. Reliable ingestion is unglamorous and it is the thing everything else depends on.
How do we stop pipelines breaking when source systems change?
Partly with technology and mostly with relationships. Technically, fail loudly on unexpected schema rather than coercing silently, so you find out immediately. Organisationally, the durable fix is that the teams owning source systems know their data has consumers, which is what a data contract is really for.