FuturByte

Hire Data Science developers

Data science turns data into decisions, and hiring for it means testing statistical judgement and communication rather than familiarity with libraries.

What Data Science actually is

Data science is the practice of drawing reliable conclusions from data and communicating them so that somebody can act. In commercial settings it spans exploratory analysis, experiment design and evaluation, statistical modelling, forecasting and the presentation of all of it to people who will make decisions on the basis of it.

The distinction from machine learning engineering matters when writing a job specification. A data scientist is judged on the quality of the conclusion and whether it changed a decision. A machine learning engineer is judged on a system that runs. Many organisations advertise for one and interview for the other, and the result is a hire who is competent and mismatched.

The part that is consistently underweighted is communication. An analysis nobody acts on has produced nothing, however rigorous. The data scientists who are most valuable commercially are usually those who can explain uncertainty to a non-technical audience without either overstating confidence or hedging so heavily that no decision is possible. That skill is rarer than the technical ability and much harder to teach.

The part that separates seniors from mid-levels

Experiment design is where real competence is most visible. Understanding what a control group is for, why sample size must be determined in advance, what happens when you check results repeatedly until they become significant, and why a statistically significant result can be commercially irrelevant are the foundations. The single most common failure in commercial data science is stopping a test when it looks good, which reliably produces conclusions that do not replicate.

Causation is the second area. Most business questions are causal: did this change work, would this intervention help. Most available data is observational, which supports correlation and not causation without careful design. Data scientists who reach for randomisation where possible, and who understand confounding and selection effects where it is not, produce conclusions that survive contact with reality. Those who do not produce confident recommendations that quietly fail.

The third is knowing when a simple answer is sufficient. A great deal of commercially valuable data science is careful descriptive work: understanding what actually happened, segmenting properly, and presenting it clearly. Reaching for a complex model where a well-constructed summary would answer the question is a common way to spend three weeks producing something less useful than three days would have.

Where Data Science is used

The label “Data Science developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.

Product analytics

Understanding user behaviour, measuring feature impact and designing the experiments that inform the roadmap.

Marketing and growth

Attribution, campaign measurement, segmentation and lifetime value modelling.

Pricing and revenue

Elasticity, discount effectiveness and forecasting, where conclusions translate directly into money.

Risk and fraud

Scoring and detection, frequently alongside machine learning engineering.

Operations

Demand forecasting, capacity planning and supply chain analysis.

Executive reporting

Turning organisational data into the small number of numbers leadership actually uses.

If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.

Support status of the tools in this stack

Data science tooling is a fast-moving ecosystem and library compatibility causes regular friction, so what matters is the currency of the stack a team runs rather than any single version.

Data Science itself is not versioned as a single product, so the useful equivalent is the support status of the tools a data scientist works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.

Release and support status across the Data Science toolchain
ToolLatest releaseRelease dateMaintained linesFurthest end-of-life date
Python3.14.72026-08-0552030-10-31
NumPy2.5.32026-09-0642028-06-22
PostgreSQL18.62026-08-1152030-11-14
Kubernetes1.37.12026-09-2342027-10-28
Redis8.10.22026-09-1752030-09-01
Docker Engine29.8.12026-09-1512026-12-04

Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.

The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.

The toolchain around it

Nobody hires for Data Science alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.

SQL
The most used skill by a wide margin. Most data science is querying before it is anything else.
Python with pandas or Polars
Analysis and manipulation, with notebooks for exploration.
statsmodels or R
Statistical modelling and inference, where the emphasis is on understanding rather than prediction.
scikit-learn
Predictive modelling where prediction rather than explanation is the goal.
An experimentation platform
Assignment, tracking and analysis for controlled experiments.
A visualisation tool
Whatever the organisation uses for dashboards, plus the ability to make a clear chart from scratch.
A warehouse
Where the data actually lives, and the constraint on what can be asked.
Version control
Underused in this discipline and a reasonable signal of engineering maturity.

Related skills that frequently appear on the same specification: Machine Learning, AI Engineering, Python, SQL, Data Engineering.

What to test in an interview

These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.

Experiment design

The highest-value skill in commercial data science and where weak candidates are exposed.

Causation versus correlation

Most business questions are causal and most data is not.

Communicating uncertainty

Determines whether the work changes any decision.

An analysis that changed a decision

The outcome the role exists for.

Data quality

Real data is messy and conclusions inherit its problems.

When not to model

Tests judgement, which is what separates useful from impressive.

Engineering practice

Analyses that cannot be reproduced cannot be trusted or updated.

Warning signs in a Data Science codebase

The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.

Stopping a test when it looks good

Correlational findings presented as causal

Analysis nobody can reproduce

Complexity where description would do

Ignoring how the data was collected

Dashboards nobody uses

What each level can own

Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a data scientist.

Junior
Runs analyses on defined questions with prepared data. Needs review on statistical choices and on framing conclusions.
Mid-level
Owns a question end to end, designs experiments, and presents conclusions to the team that will act on them.
Senior
Shapes which questions are worth asking, owns the experimentation standard, and is trusted to tell leadership that a favoured idea did not work.
Staff
Owns the analytical agenda across the organisation, the standards for measurement and experimentation, and the relationship between data and decision-making at senior level.

How the work is usually scoped

Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.

Experimentation programme

A specific business question

Forecasting

Analytics foundation

Embedded in a product team

Migration work you may actually be hiring for

A large share of Data Science work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.

Ad hoc analysis to a structured experimentation programme

Notebooks to version-controlled reproducible analysis

A central analytics queue to embedded data scientists

Dashboards as the default output to decision-focused analysis

Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.

What a good brief for this role contains

Most of the time lost in hiring a data scientist is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.

If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.

What the US market pays for this work

Data Science work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a data scientist is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.

US annual wages, Data Scientists, May 2025
US annual wages, Data Scientists, May 2025$120,230Median$67,240$199,13010th pct90th pctMiddle half $86K to $159K

The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a data scientist can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.

Related classifications are worth reading alongside it, because teams hiring for Data Science frequently end up recruiting against these titles too:

US national wages, May 2025
OccupationEmployed25th percentileMedian75th percentile90th percentile
Data Scientists262,440$85,660$120,230$158,880$199,130
Computer and Information Research Scientists37,200$103,570$140,300$188,700$230,630
Software Developers1,687,890$105,210$135,980$171,980$214,670

Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.

These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.

How US metro markets compare for this role

The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.

Median wage for data scientists, by US metro area
Median wage for data scientists, by US metro areaSan Jose, CA: $185,080San Jose, CASan Jose, CA$185,080San Francisco, CA: $170,110San Francisco, CASan Francisco, CA$170,110Seattle, WA: $164,740Seattle, WASeattle, WA$164,740New York, NY: $135,980New York, NYNew York, NY$135,980Baltimore, MD: $134,320Baltimore, MDBaltimore, MD$134,320Charlotte, NC: $132,460Charlotte, NCCharlotte, NC$132,460Washington, D.C.: $132,200Washington, D.C.Washington, D.C.$132,200Boston, MA: $132,040Boston, MABoston, MA$132,040San Diego, CA: $130,990San Diego, CASan Diego, CA$130,990Minneapolis-St. Paul, MN: $129,780Minneapolis-St. Paul, MNMinneapolis-St. Paul, MN$129,780Los Angeles, CA: $129,740Los Angeles, CALos Angeles, CA$129,740Portland, OR: $129,600Portland, ORPortland, OR$129,600Dallas-Fort Worth, TX: $127,750Dallas-Fort Worth, TXDallas-Fort Worth, TX$127,750Miami, FL: $127,450Miami, FLMiami, FL$127,450Austin, TX: $127,360Austin, TXAustin, TX$127,360Raleigh, NC: $120,710Raleigh, NCRaleigh, NC$120,710Salt Lake City, UT: $114,990Salt Lake City, UTSalt Lake City, UT$114,990Phoenix, AZ: $114,540Phoenix, AZPhoenix, AZ$114,540Denver, CO: $112,520Denver, CODenver, CO$112,520Tampa, FL: $109,990Tampa, FLTampa, FL$109,990Philadelphia, PA: $109,910Philadelphia, PAPhiladelphia, PA$109,910Atlanta, GA: $108,940Atlanta, GAAtlanta, GA$108,940Chicago, IL: $107,640Chicago, ILChicago, IL$107,640Houston, TX: $106,750Houston, TXHouston, TX$106,750Orlando, FL: $106,590Orlando, FLOrlando, FL$106,590Detroit, MI: $103,330Detroit, MIDetroit, MI$103,330Kansas City, MO: $99,870Kansas City, MOKansas City, MO$99,870Pittsburgh, PA: $96,670Pittsburgh, PAPittsburgh, PA$96,670
Data Scientists by metro area, May 2025, ranked by median wage
Metro areaEmployedMedian wagevs US medianLocation quotient
San Jose, CA6,060$185,080+54%3.16
San Francisco, CA10,460$170,110+41%2.61
Seattle, WA8,370$164,740+37%2.38
New York, NY23,160$135,980+13%1.45
Baltimore, MD1,090$134,320+12%0.48
Charlotte, NC4,420$132,460+10%1.93
Washington, D.C.9,260$132,200+10%1.75
Boston, MA7,930$132,040+10%1.74
San Diego, CA2,830$130,990+9%1.09
Minneapolis-St. Paul, MN3,250$129,780+8%0.99
Los Angeles, CA9,850$129,740+8%0.93
Portland, OR1,700$129,600+8%0.83
Dallas-Fort Worth, TX10,120$127,750+6%1.48
Miami, FL3,040$127,450+6%0.64
Austin, TX3,730$127,360+6%1.71
Raleigh, NC1,990$120,7100%1.59
Salt Lake City, UT2,970$114,990-4%2.13
Phoenix, AZ3,480$114,540-5%0.87
Denver, CO4,510$112,520-6%1.66
Tampa, FL1,730$109,990-9%0.71
Philadelphia, PA6,480$109,910-9%1.32
Atlanta, GA6,820$108,940-9%1.40
Chicago, IL7,940$107,640-10%1.04
Houston, TX4,060$106,750-11%0.73
Orlando, FL1,440$106,590-11%0.61
Detroit, MI3,810$103,330-14%1.18
Kansas City, MO1,260$99,870-17%0.68
Pittsburgh, PA2,270$96,670-20%1.21

Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.

The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.

The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.

Hiring risks worth naming

Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.

Hired for analysis when engineering was needed. The most common mismatch. If the data does not yet exist reliably, that is data engineering.

Technical strength without communication. Ask them to explain a result to a non-technical audience. Analysis nobody acts on has produced nothing.

Weak experimental discipline. Ask about stopping rules. This single question separates the field more reliably than any other.

No engineering practice. Ask how somebody else would reproduce their work. Irreproducible analysis cannot be trusted.

Hiring Data Science developers by metro area

Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of data scientist who will be available.

Frequently asked questions

Do we need a data scientist or a data engineer?

If your data is unreliable, scattered or does not exist in usable form, you need a data engineer first. A data scientist hired into that situation spends their time building pipelines, does it less well than a specialist would, and usually leaves. If the data is there and the question is what it means, that is data science.

When should we hire our first data scientist?

When there are decisions being made on intuition that data could inform, and when the data to inform them exists and is accessible. Hiring earlier than that produces someone who spends a year on infrastructure. A useful test is whether you can name three decisions you would make differently with better analysis.

Should our data scientist be central or embedded in product teams?

Embedded works better in most organisations, because the value comes from proximity to the decisions. Central teams tend to become a request queue producing analyses that arrive after the decision was made. A central team makes sense for shared standards and infrastructure once there are several data scientists.

How do we know if the analysis is any good?

Ask what would have changed their mind, and ask how confident they are and why. Strong data scientists state their uncertainty plainly and can describe the evidence that would overturn their conclusion. Analysis that arrives with no acknowledged uncertainty is usually either trivial or overconfident.

What is the most common mistake in commercial experimentation?

Stopping a test when the result looks favourable. Checking repeatedly and concluding when significance appears dramatically inflates the false positive rate, so the finding does not replicate and the change does not deliver what was promised. Deciding the sample size in advance, or using a method built for continuous monitoring, is the fix.

Does a data scientist need to code well?

Well enough that someone else can run and check their work, which is a lower bar than software engineering and a higher one than many meet. The failure mode is analysis that exists only in a notebook with manual steps, which means conclusions cannot be reproduced, checked or updated. That is a real problem regardless of how good the statistics were.

Python or R?

Python if the work sits alongside engineering, which is most commercial contexts, because the path from analysis to production is shorter. R where the emphasis is statistical inference and the community around it. Both are capable and the choice usually follows the surrounding team rather than the work.

Why did our A/B test result not hold up?

Most often because the test was stopped early, the sample was too small, or the metric was chosen after seeing the data. All three produce findings that look convincing and do not replicate. Fixing the process is more valuable than re-running the test, because otherwise the next result will have the same problem.