Hire Data Science developers
Data science turns data into decisions, and hiring for it means testing statistical judgement and communication rather than familiarity with libraries.
What Data Science actually is
Data science is the practice of drawing reliable conclusions from data and communicating them so that somebody can act. In commercial settings it spans exploratory analysis, experiment design and evaluation, statistical modelling, forecasting and the presentation of all of it to people who will make decisions on the basis of it.
The distinction from machine learning engineering matters when writing a job specification. A data scientist is judged on the quality of the conclusion and whether it changed a decision. A machine learning engineer is judged on a system that runs. Many organisations advertise for one and interview for the other, and the result is a hire who is competent and mismatched.
The part that is consistently underweighted is communication. An analysis nobody acts on has produced nothing, however rigorous. The data scientists who are most valuable commercially are usually those who can explain uncertainty to a non-technical audience without either overstating confidence or hedging so heavily that no decision is possible. That skill is rarer than the technical ability and much harder to teach.
The part that separates seniors from mid-levels
Experiment design is where real competence is most visible. Understanding what a control group is for, why sample size must be determined in advance, what happens when you check results repeatedly until they become significant, and why a statistically significant result can be commercially irrelevant are the foundations. The single most common failure in commercial data science is stopping a test when it looks good, which reliably produces conclusions that do not replicate.
Causation is the second area. Most business questions are causal: did this change work, would this intervention help. Most available data is observational, which supports correlation and not causation without careful design. Data scientists who reach for randomisation where possible, and who understand confounding and selection effects where it is not, produce conclusions that survive contact with reality. Those who do not produce confident recommendations that quietly fail.
The third is knowing when a simple answer is sufficient. A great deal of commercially valuable data science is careful descriptive work: understanding what actually happened, segmenting properly, and presenting it clearly. Reaching for a complex model where a well-constructed summary would answer the question is a common way to spend three weeks producing something less useful than three days would have.
Where Data Science is used
The label “Data Science developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Product analytics
Understanding user behaviour, measuring feature impact and designing the experiments that inform the roadmap.
Marketing and growth
Attribution, campaign measurement, segmentation and lifetime value modelling.
Pricing and revenue
Elasticity, discount effectiveness and forecasting, where conclusions translate directly into money.
Risk and fraud
Scoring and detection, frequently alongside machine learning engineering.
Operations
Demand forecasting, capacity planning and supply chain analysis.
Executive reporting
Turning organisational data into the small number of numbers leadership actually uses.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
Data science tooling is a fast-moving ecosystem and library compatibility causes regular friction, so what matters is the currency of the stack a team runs rather than any single version.
Data Science itself is not versioned as a single product, so the useful equivalent is the support status of the tools a data scientist works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Python | 3.14.7 | 2026-08-05 | 5 | 2030-10-31 |
| NumPy | 2.5.3 | 2026-09-06 | 4 | 2028-06-22 |
| PostgreSQL | 18.6 | 2026-08-11 | 5 | 2030-11-14 |
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
| Redis | 8.10.2 | 2026-09-17 | 5 | 2030-09-01 |
| Docker Engine | 29.8.1 | 2026-09-15 | 1 | 2026-12-04 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for Data Science alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- SQL
- The most used skill by a wide margin. Most data science is querying before it is anything else.
- Python with pandas or Polars
- Analysis and manipulation, with notebooks for exploration.
- statsmodels or R
- Statistical modelling and inference, where the emphasis is on understanding rather than prediction.
- scikit-learn
- Predictive modelling where prediction rather than explanation is the goal.
- An experimentation platform
- Assignment, tracking and analysis for controlled experiments.
- A visualisation tool
- Whatever the organisation uses for dashboards, plus the ability to make a clear chart from scratch.
- A warehouse
- Where the data actually lives, and the constraint on what can be asked.
- Version control
- Underused in this discipline and a reasonable signal of engineering maturity.
Related skills that frequently appear on the same specification: Machine Learning, AI Engineering, Python, SQL, Data Engineering.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
Experiment design
The highest-value skill in commercial data science and where weak candidates are exposed.
- Strong answer: Determines sample size in advance, understands the cost of repeated checking, and distinguishes statistical from practical significance.
- Warning sign: Would stop a test early because the result looked good.
Causation versus correlation
Most business questions are causal and most data is not.
- Strong answer: Reaches for randomisation, understands confounding, and is explicit about what an observational result can and cannot support.
- Warning sign: Presents correlational findings as causal recommendations.
Communicating uncertainty
Determines whether the work changes any decision.
- Strong answer: Explains confidence in plain language, states what would change their mind, and does not hide behind hedging.
- Warning sign: Either presents point estimates as certainties, or qualifies so heavily that no decision follows.
An analysis that changed a decision
The outcome the role exists for.
- Strong answer: Describes the question, the work, the recommendation and what was actually done.
- Warning sign: Describes analyses produced but cannot name anything that changed as a result.
Data quality
Real data is messy and conclusions inherit its problems.
- Strong answer: Validates assumptions, checks for selection effects and survivorship bias, and has caught an error before publishing.
- Warning sign: Takes the data at face value.
When not to model
Tests judgement, which is what separates useful from impressive.
- Strong answer: Recognises when descriptive work answers the question and says so.
- Warning sign: Reaches for a model regardless of the question.
Engineering practice
Analyses that cannot be reproduced cannot be trusted or updated.
- Strong answer: Uses version control, writes code somebody else can run, and can reproduce a past result.
- Warning sign: Work exists only in notebooks nobody else can execute.
Warning signs in a Data Science codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
Stopping a test when it looks good
- What you see: Monitoring an experiment continuously and concluding when significance appears.
- What it costs: Conclusions that do not replicate, and decisions made on noise.
- The fix: Fix the sample size in advance, or use a method designed for sequential testing. This is the most common serious error in the field.
Correlational findings presented as causal
- What you see: Recommendations to change something based on observational association.
- What it costs: Interventions that do not work, and credibility lost when they do not.
- The fix: Randomise where it is possible at all. Where it is not, state the assumptions the causal claim rests on and name the confounders you could not rule out, so the reader can judge the strength of the conclusion.
Analysis nobody can reproduce
- What you see: Results in notebooks with manual steps and no version control.
- What it costs: Findings that cannot be checked or updated, and disagreement that cannot be resolved.
- The fix: Version control the analysis, make it runnable end to end, and record the data used.
Complexity where description would do
- What you see: A model built for a question a clear summary would answer.
- What it costs: Weeks spent producing something harder to explain and no more useful.
- The fix: Answer the question the simplest way that is honest, then escalate if the answer is insufficient.
Ignoring how the data was collected
- What you see: Conclusions drawn without considering who is missing from the dataset.
- What it costs: Survivorship and selection effects producing confident conclusions about a population that was never observed.
- The fix: Ask what the data would look like if the hypothesis were false, and who is absent from it.
Dashboards nobody uses
- What you see: Extensive reporting infrastructure with no identified decision attached.
- What it costs: Maintenance burden and analyst time spent on output that changes nothing.
- The fix: Start from a decision somebody makes. Build the smallest thing that informs it.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a data scientist.
- Junior
- Runs analyses on defined questions with prepared data. Needs review on statistical choices and on framing conclusions.
- Mid-level
- Owns a question end to end, designs experiments, and presents conclusions to the team that will act on them.
- Senior
- Shapes which questions are worth asking, owns the experimentation standard, and is trusted to tell leadership that a favoured idea did not work.
- Staff
- Owns the analytical agenda across the organisation, the standards for measurement and experimentation, and the relationship between data and decision-making at senior level.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
Experimentation programme
- Usual team: One data scientist plus engineering support.
- What governs it: The infrastructure for reliable assignment and tracking is usually the prerequisite and is often missing.
A specific business question
- Usual team: One data scientist, time-boxed.
- What governs it: Scoped by data availability rather than by analytical difficulty, always.
Forecasting
- Usual team: One data scientist.
- What governs it: Scoped by data history and quality. Establishing a simple baseline first is essential and often skipped.
Analytics foundation
- Usual team: One data scientist plus a data engineer.
- What governs it: Most organisations wanting data science discover this is the actual first project.
Embedded in a product team
- Usual team: One data scientist alongside engineers and a product manager.
- What governs it: Usually more effective than a central team, because proximity to decisions is what makes the work land.
Migration work you may actually be hiring for
A large share of Data Science work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Ad hoc analysis to a structured experimentation programme
- Why teams do it: Findings that replicate, and decisions that can be defended.
- What to watch: The infrastructure for reliable assignment and tracking is the prerequisite. Running experiments on unreliable instrumentation produces confident conclusions from noise.
Notebooks to version-controlled reproducible analysis
- Why teams do it: Conclusions that can be checked, repeated and updated when the data changes.
- What to watch: Does not require full software engineering practice. Version control, a runnable script and recorded data sources cover most of the benefit.
A central analytics queue to embedded data scientists
- Why teams do it: Proximity to decisions, which is what determines whether analysis changes anything.
- What to watch: Embedded analysts can drift apart in methods. Keep shared standards for measurement and experimentation even when people sit in product teams.
Dashboards as the default output to decision-focused analysis
- Why teams do it: Dashboards accumulate and are rarely retired, and most are not attached to a decision.
- What to watch: Audit which dashboards are actually opened before building more. The answer is usually uncomfortable and clarifying.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a data scientist is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- Which decisions this person is expected to inform, since that determines everything else.
- Whether the data exists and is reliable, or whether building that is part of the job.
- Whether experimentation infrastructure exists, because analysis without it is limited to observation.
- Whether the role is embedded in a product team or central.
- Whether modelling for production is in scope, since that is machine learning engineering.
- Who the audience is for the work, because communication skill should be tested against it.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
Data Science work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a data scientist is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.
The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a data scientist can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for Data Science frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Data Scientists | 262,440 | $85,660 | $120,230 | $158,880 | $199,130 |
| Computer and Information Research Scientists | 37,200 | $103,570 | $140,300 | $188,700 | $230,630 |
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 6,060 | $185,080 | +54% | 3.16 |
| San Francisco, CA | 10,460 | $170,110 | +41% | 2.61 |
| Seattle, WA | 8,370 | $164,740 | +37% | 2.38 |
| New York, NY | 23,160 | $135,980 | +13% | 1.45 |
| Baltimore, MD | 1,090 | $134,320 | +12% | 0.48 |
| Charlotte, NC | 4,420 | $132,460 | +10% | 1.93 |
| Washington, D.C. | 9,260 | $132,200 | +10% | 1.75 |
| Boston, MA | 7,930 | $132,040 | +10% | 1.74 |
| San Diego, CA | 2,830 | $130,990 | +9% | 1.09 |
| Minneapolis-St. Paul, MN | 3,250 | $129,780 | +8% | 0.99 |
| Los Angeles, CA | 9,850 | $129,740 | +8% | 0.93 |
| Portland, OR | 1,700 | $129,600 | +8% | 0.83 |
| Dallas-Fort Worth, TX | 10,120 | $127,750 | +6% | 1.48 |
| Miami, FL | 3,040 | $127,450 | +6% | 0.64 |
| Austin, TX | 3,730 | $127,360 | +6% | 1.71 |
| Raleigh, NC | 1,990 | $120,710 | 0% | 1.59 |
| Salt Lake City, UT | 2,970 | $114,990 | -4% | 2.13 |
| Phoenix, AZ | 3,480 | $114,540 | -5% | 0.87 |
| Denver, CO | 4,510 | $112,520 | -6% | 1.66 |
| Tampa, FL | 1,730 | $109,990 | -9% | 0.71 |
| Philadelphia, PA | 6,480 | $109,910 | -9% | 1.32 |
| Atlanta, GA | 6,820 | $108,940 | -9% | 1.40 |
| Chicago, IL | 7,940 | $107,640 | -10% | 1.04 |
| Houston, TX | 4,060 | $106,750 | -11% | 0.73 |
| Orlando, FL | 1,440 | $106,590 | -11% | 0.61 |
| Detroit, MI | 3,810 | $103,330 | -14% | 1.18 |
| Kansas City, MO | 1,260 | $99,870 | -17% | 0.68 |
| Pittsburgh, PA | 2,270 | $96,670 | -20% | 1.21 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Hired for analysis when engineering was needed. The most common mismatch. If the data does not yet exist reliably, that is data engineering.
Technical strength without communication. Ask them to explain a result to a non-technical audience. Analysis nobody acts on has produced nothing.
Weak experimental discipline. Ask about stopping rules. This single question separates the field more reliably than any other.
No engineering practice. Ask how somebody else would reproduce their work. Irreproducible analysis cannot be trusted.
Hiring Data Science developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of data scientist who will be available.
- New York, NY $135,980 median
- Seattle, WA $164,740 median
- San Jose, CA $185,080 median
- Washington, D.C. $132,200 median
- San Francisco, CA $170,110 median
- Dallas-Fort Worth, TX $127,750 median
- Los Angeles, CA $129,740 median
- Boston, MA $132,040 median
- Chicago, IL $107,640 median
- Atlanta, GA $108,940 median
- Austin, TX $127,360 median
- Phoenix, AZ $114,540 median
- Philadelphia, PA $109,910 median
- Minneapolis-St. Paul, MN $129,780 median
- Denver, CO $112,520 median
- Detroit, MI $103,330 median
- Houston, TX $106,750 median
- Charlotte, NC $132,460 median
- San Diego, CA $130,990 median
- Salt Lake City, UT $114,990 median
- Miami, FL $127,450 median
- Portland, OR $129,600 median
- Baltimore, MD $134,320 median
- Tampa, FL $109,990 median
- Orlando, FL $106,590 median
- Raleigh, NC $120,710 median
- Kansas City, MO $99,870 median
- Pittsburgh, PA $96,670 median
Frequently asked questions
Do we need a data scientist or a data engineer?
If your data is unreliable, scattered or does not exist in usable form, you need a data engineer first. A data scientist hired into that situation spends their time building pipelines, does it less well than a specialist would, and usually leaves. If the data is there and the question is what it means, that is data science.
When should we hire our first data scientist?
When there are decisions being made on intuition that data could inform, and when the data to inform them exists and is accessible. Hiring earlier than that produces someone who spends a year on infrastructure. A useful test is whether you can name three decisions you would make differently with better analysis.
Should our data scientist be central or embedded in product teams?
Embedded works better in most organisations, because the value comes from proximity to the decisions. Central teams tend to become a request queue producing analyses that arrive after the decision was made. A central team makes sense for shared standards and infrastructure once there are several data scientists.
How do we know if the analysis is any good?
Ask what would have changed their mind, and ask how confident they are and why. Strong data scientists state their uncertainty plainly and can describe the evidence that would overturn their conclusion. Analysis that arrives with no acknowledged uncertainty is usually either trivial or overconfident.
What is the most common mistake in commercial experimentation?
Stopping a test when the result looks favourable. Checking repeatedly and concluding when significance appears dramatically inflates the false positive rate, so the finding does not replicate and the change does not deliver what was promised. Deciding the sample size in advance, or using a method built for continuous monitoring, is the fix.
Does a data scientist need to code well?
Well enough that someone else can run and check their work, which is a lower bar than software engineering and a higher one than many meet. The failure mode is analysis that exists only in a notebook with manual steps, which means conclusions cannot be reproduced, checked or updated. That is a real problem regardless of how good the statistics were.
Python or R?
Python if the work sits alongside engineering, which is most commercial contexts, because the path from analysis to production is shorter. R where the emphasis is statistical inference and the community around it. Both are capable and the choice usually follows the surrounding team rather than the work.
Why did our A/B test result not hold up?
Most often because the test was stopped early, the sample was too small, or the metric was chosen after seeing the data. All three produce findings that look convincing and do not replicate. Fixing the process is more valuable than re-running the test, because otherwise the next result will have the same problem.