FuturByte

Hire Data Engineering developers

Data engineering builds the pipelines everything else depends on, and hiring for it means valuing reliability and data quality over familiarity with any particular tool.

What Data Engineering actually is

Data engineering is the discipline of moving data from where it is produced to where it is used, reliably and in a shape people can work with. In practice that means building and operating pipelines, modelling data for analysis, managing warehouses and lakes, and being responsible when the numbers in a report are wrong. It sits between the systems that generate data and the analysts and models that consume it.

The role is often confused with data science, and the distinction is worth being clear about. A data scientist asks questions of data and builds models. A data engineer makes sure the data exists, arrives on time, and means what it claims to mean. Organisations that hire the second when they needed the first, or more commonly the reverse, end up with expensive people blocked on work that is not theirs.

What has changed in recent years is that the warehouse became powerful enough to do the transformation itself. The older pattern transformed data before loading it; the current one loads raw data and transforms it inside the warehouse with SQL. That shift means modern data engineering is closer to software engineering practice, with version control, testing and review applied to transformation logic that used to live in undocumented scripts.

The part that separates seniors from mid-levels

Idempotency is the foundation and the thing to interview for first. Pipelines fail, get retried, and get re-run over historical periods. A pipeline that produces different results when run twice will eventually produce wrong numbers that nobody can explain, and it will do so quietly. Engineers who design for re-runnability talk about partitioned writes, deterministic transformations and the difference between appending and replacing a window of data.

Data quality is the second, and the difference between a competent data engineer and a valuable one. Bad data is worse than missing data, because missing data stops a report and bad data produces a confident wrong answer that somebody acts on. Testing at the boundary, asserting expectations about row counts, uniqueness, ranges and referential integrity, and failing loudly rather than continuing is the practice that separates them.

The third is understanding that a pipeline is a production system with users. Someone depends on a dashboard being correct by nine in the morning. That means monitoring, alerting, defined freshness expectations and a process when something breaks, exactly as for any service. Data teams that treat pipelines as scripts rather than as services are the ones where the business quietly stops trusting the numbers.

Where Data Engineering is used

The label “Data Engineering developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.

Analytics platforms

Building the warehouse and models that reporting and business intelligence run on.

Machine learning infrastructure

Feature pipelines and training data, where reproducibility matters more than anywhere else.

Product data

Event collection and behavioural data feeding product decisions and customer-facing features.

Financial and regulatory reporting

Where correctness and auditability are the requirement and the tolerance for error is zero.

Systems integration

Moving data between operational systems, often the least glamorous and most depended-upon work.

Migration and consolidation

Moving off legacy warehouses or consolidating several sources into one platform.

If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.

Support status of the tools in this stack

Data engineering is a discipline rather than a product, so currency is judged by the platforms and patterns someone has worked with recently; the table below shows support status across common components of the stack.

Data Engineering itself is not versioned as a single product, so the useful equivalent is the support status of the tools a data engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.

Release and support status across the Data Engineering toolchain
ToolLatest releaseRelease dateMaintained linesFurthest end-of-life date
Python3.14.72026-08-0552030-10-31
PostgreSQL18.62026-08-1152030-11-14
Apache Kafka4.3.12026-06-2312024-11-06
Elasticsearch9.5.42026-09-1512027-07-15
MongoDB Server8.3.112026-09-1132029-10-31
Redis8.10.22026-09-1752030-09-01
Kubernetes1.37.12026-09-2342027-10-28

Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.

The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.

The toolchain around it

Nobody hires for Data Engineering alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.

SQL
The core skill. Modern transformation is mostly SQL executed in the warehouse.
Python
Orchestration, extraction and anything SQL cannot express.
A warehouse
Snowflake, BigQuery, Redshift or similar, where cost and performance characteristics differ substantially.
dbt
Transformation with version control, testing, documentation and lineage. Close to standard in current practice.
Airflow, Dagster or Prefect
Orchestration and scheduling, with dependency management and retry behaviour.
Kafka or a managed stream
Event streaming where near-real-time data is genuinely required.
Object storage
The landing zone for raw data, and increasingly the storage layer for open table formats.
Data quality testing
Assertions on freshness, uniqueness, ranges and relationships, run as part of the pipeline.

Related skills that frequently appear on the same specification: Machine Learning, Python, SQL, Data Science, AWS.

What to test in an interview

These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.

Idempotency and re-runs

The single most important property of a pipeline and the most common omission.

Data quality testing

Separates engineers who move data from engineers who are accountable for it.

A pipeline failure they handled

Operational reality, and where judgement is formed.

Warehouse cost

Warehouse spend is usually a significant line item and is driven by engineering decisions.

Modelling for consumers

Technically correct data that analysts cannot use has not solved the problem.

Batch versus streaming

Streaming is frequently adopted where batch would do, at considerable cost in complexity.

Schema changes upstream

Source systems change without warning and pipelines break.

Warning signs in a Data Engineering codebase

The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.

Non-idempotent pipelines

No tests on data

Transformation logic nobody can review

Streaming where batch would do

Silent schema drift handling

Ungoverned warehouse cost

What each level can own

Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a data engineer.

Junior
Builds transformations within an established framework. Needs review on idempotency and testing.
Mid-level
Owns a pipeline end to end including orchestration, tests and monitoring. Responds when it breaks.
Senior
Owns platform architecture, the modelling approach, data quality standards and warehouse cost. Works directly with data consumers on what they need.
Staff
Owns data contracts with source systems, governance and lineage across the organisation, and the platform strategy including build-versus-buy decisions.

How the work is usually scoped

Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.

Warehouse and pipeline build

Legacy pipeline migration

Data quality programme

Warehouse cost reduction

Streaming implementation

Migration work you may actually be hiring for

A large share of Data Engineering work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.

Transform before load to load then transform in the warehouse

Scheduled scripts to an orchestrator

An on-premises warehouse to a cloud warehouse

Untested transformations to tested, version-controlled models

Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.

What a good brief for this role contains

Most of the time lost in hiring a data engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.

If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.

What the US market pays for this work

Data Engineering work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a data engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.

US annual wages, Data Scientists, May 2025
US annual wages, Data Scientists, May 2025$120,230Median$67,240$199,13010th pct90th pctMiddle half $86K to $159K

The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a data engineer can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.

Related classifications are worth reading alongside it, because teams hiring for Data Engineering frequently end up recruiting against these titles too:

US national wages, May 2025
OccupationEmployed25th percentileMedian75th percentile90th percentile
Data Scientists262,440$85,660$120,230$158,880$199,130
Database Architects67,140$109,370$139,500$169,290$204,000
Software Developers1,687,890$105,210$135,980$171,980$214,670
Database Administrators69,990$79,610$104,620$135,460$163,320

Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.

These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.

How US metro markets compare for this role

The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.

Median wage for data scientists, by US metro area
Median wage for data scientists, by US metro areaSan Jose, CA: $185,080San Jose, CASan Jose, CA$185,080San Francisco, CA: $170,110San Francisco, CASan Francisco, CA$170,110Seattle, WA: $164,740Seattle, WASeattle, WA$164,740New York, NY: $135,980New York, NYNew York, NY$135,980Baltimore, MD: $134,320Baltimore, MDBaltimore, MD$134,320Charlotte, NC: $132,460Charlotte, NCCharlotte, NC$132,460Washington, D.C.: $132,200Washington, D.C.Washington, D.C.$132,200Boston, MA: $132,040Boston, MABoston, MA$132,040San Diego, CA: $130,990San Diego, CASan Diego, CA$130,990Minneapolis-St. Paul, MN: $129,780Minneapolis-St. Paul, MNMinneapolis-St. Paul, MN$129,780Los Angeles, CA: $129,740Los Angeles, CALos Angeles, CA$129,740Portland, OR: $129,600Portland, ORPortland, OR$129,600Dallas-Fort Worth, TX: $127,750Dallas-Fort Worth, TXDallas-Fort Worth, TX$127,750Miami, FL: $127,450Miami, FLMiami, FL$127,450Austin, TX: $127,360Austin, TXAustin, TX$127,360Raleigh, NC: $120,710Raleigh, NCRaleigh, NC$120,710Salt Lake City, UT: $114,990Salt Lake City, UTSalt Lake City, UT$114,990Phoenix, AZ: $114,540Phoenix, AZPhoenix, AZ$114,540Denver, CO: $112,520Denver, CODenver, CO$112,520Tampa, FL: $109,990Tampa, FLTampa, FL$109,990Philadelphia, PA: $109,910Philadelphia, PAPhiladelphia, PA$109,910Atlanta, GA: $108,940Atlanta, GAAtlanta, GA$108,940Chicago, IL: $107,640Chicago, ILChicago, IL$107,640Houston, TX: $106,750Houston, TXHouston, TX$106,750Orlando, FL: $106,590Orlando, FLOrlando, FL$106,590Detroit, MI: $103,330Detroit, MIDetroit, MI$103,330Kansas City, MO: $99,870Kansas City, MOKansas City, MO$99,870Pittsburgh, PA: $96,670Pittsburgh, PAPittsburgh, PA$96,670
Data Scientists by metro area, May 2025, ranked by median wage
Metro areaEmployedMedian wagevs US medianLocation quotient
San Jose, CA6,060$185,080+54%3.16
San Francisco, CA10,460$170,110+41%2.61
Seattle, WA8,370$164,740+37%2.38
New York, NY23,160$135,980+13%1.45
Baltimore, MD1,090$134,320+12%0.48
Charlotte, NC4,420$132,460+10%1.93
Washington, D.C.9,260$132,200+10%1.75
Boston, MA7,930$132,040+10%1.74
San Diego, CA2,830$130,990+9%1.09
Minneapolis-St. Paul, MN3,250$129,780+8%0.99
Los Angeles, CA9,850$129,740+8%0.93
Portland, OR1,700$129,600+8%0.83
Dallas-Fort Worth, TX10,120$127,750+6%1.48
Miami, FL3,040$127,450+6%0.64
Austin, TX3,730$127,360+6%1.71
Raleigh, NC1,990$120,7100%1.59
Salt Lake City, UT2,970$114,990-4%2.13
Phoenix, AZ3,480$114,540-5%0.87
Denver, CO4,510$112,520-6%1.66
Tampa, FL1,730$109,990-9%0.71
Philadelphia, PA6,480$109,910-9%1.32
Atlanta, GA6,820$108,940-9%1.40
Chicago, IL7,940$107,640-10%1.04
Houston, TX4,060$106,750-11%0.73
Orlando, FL1,440$106,590-11%0.61
Detroit, MI3,810$103,330-14%1.18
Kansas City, MO1,260$99,870-17%0.68
Pittsburgh, PA2,270$96,670-20%1.21

Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.

The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.

The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.

Hiring risks worth naming

Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.

Data science background without engineering practice. Common and consequential. Test version control, testing, orchestration and on-call behaviour explicitly.

Tool familiarity without reliability thinking. Ask how a pipeline behaves when re-run. This separates the field faster than any tool question.

No accountability experience. Ask about a time the numbers were wrong. Someone who has never owned that has not felt the constraint.

Building without talking to consumers. Ask how they decided what to model. Technically clean data nobody can use is a common and expensive outcome.

Hiring Data Engineering developers by metro area

Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of data engineer who will be available.

Frequently asked questions

What is the difference between a data engineer and a data scientist?

A data engineer builds and operates the systems that make data available, correct and timely. A data scientist analyses that data and builds models from it. They are complementary and frequently confused in job specifications. Hiring a data scientist into a role that is actually pipeline work is a common and expensive mistake, and the person usually leaves.

Do we need a data engineer or would an analyst do?

If the question is mostly about producing reports from data that already arrives reliably, an analyst with strong SQL will go further than you expect. If data does not arrive reliably, or arrives wrong, or lives in several systems that disagree, that is engineering work. The clearest signal is whether people currently spend more time fixing data than analysing it.

Should we build a warehouse or use the application database for reporting?

The application database works until reporting queries start competing with the application for resources, or until you need to combine several sources. Both happen sooner than teams expect. A separate analytical store removes the contention and allows a data model designed for questions rather than for transactions, which is a different shape entirely.

How do we know if our data is trustworthy?

Ask whether the pipelines test anything. Most do not. If nothing asserts that a table has the expected number of rows, that identifiers are unique, that values are in range and that data is fresh, then nobody finds out about a problem until somebody questions a number, which means the wrong numbers before it went unquestioned.

Do we need real-time data?

Usually less than people say. Real-time is attractive and it carries substantial operational complexity and cost. The test is whether a consumer can state a decision they would make differently with fresher data. Where there is a genuine answer, streaming is worth it. Where the answer is that it would be nice, batch is the better engineering decision.

Why is our warehouse bill increasing?

Usually models reprocessing full history on every run, tables that are refreshed but no longer used, and inefficient transformations that were written when the data was small. Warehouse cost is driven by engineering decisions, so it responds well to engineering attention. A focused review typically finds substantial savings on any platform that has grown without governance.

What should we build first?

Get raw data landing reliably and reproducibly before you model anything. Teams that start with the dashboard build transformations on a foundation that is still shifting, and end up rebuilding. Reliable ingestion is unglamorous and it is the thing everything else depends on.

How do we stop pipelines breaking when source systems change?

Partly with technology and mostly with relationships. Technically, fail loudly on unexpected schema rather than coercing silently, so you find out immediately. Organisationally, the durable fix is that the teams owning source systems know their data has consumers, which is what a data contract is really for.