FuturByte

Hire Machine Learning developers

Machine learning engineering is mostly software and data engineering with modelling attached, and hiring for it means testing evaluation discipline rather than algorithm recall.

What Machine Learning actually is

Machine learning engineering covers building systems that learn patterns from data and then serve predictions in production. In practice the role spans data preparation, feature engineering, model training, evaluation, deployment and the ongoing monitoring of something whose behaviour changes as the world changes. Most of the work is not modelling.

That last point is the one job specifications most often get wrong. A production machine learning system is a piece of software with a statistical component, and the proportion of effort that goes into pipelines, reproducibility, serving infrastructure and monitoring far exceeds the proportion spent choosing and tuning models. Hiring someone who is strong at modelling and weak at engineering produces notebooks that never ship.

The distinction from data science is worth being explicit about. A data scientist is typically judged on insight and on models that answer a question. A machine learning engineer is judged on systems that run reliably and produce predictions users depend on. Both are legitimate; they attract different people, and a specification that blends them produces a shortlist where nobody is quite right.

The part that separates seniors from mid-levels

Evaluation is where competence is most visible and most often absent. A model's reported accuracy is meaningless without knowing how the data was split, whether the split respected time, whether the classes were balanced, and whether any information from the future leaked into training. Data leakage is the classic failure: a model that performs superbly in testing and poorly in production, because a feature encoded the answer. Candidates who have been caught by this once describe evaluation very differently from those who have not.

The training and serving boundary is the second area. Features computed one way during training and another way at serving time produce silent degradation that no test catches, because both paths work in isolation. This is among the most common production failures in machine learning systems and the reason feature consistency has become an explicit engineering concern rather than an incidental one.

The third is that models decay. The world changes, the input distribution shifts, and a model that was accurate at launch quietly becomes less so. Unlike ordinary software, nothing errors. Monitoring prediction distributions, tracking performance against outcomes when they eventually arrive, and having a retraining path are what separate a system somebody owns from a model somebody deployed.

Where Machine Learning is used

The label “Machine Learning developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.

Recommendation and ranking

Product, content and search ranking, where the feedback loop between predictions and behaviour is itself a modelling problem.

Fraud and risk

Financial services and marketplaces, where class imbalance is severe and the cost of the two error types is very different.

Demand and forecasting

Inventory, pricing and capacity planning, where time-aware evaluation is essential and often mishandled.

Computer vision

Inspection, medical imaging and document processing, where data labelling is usually the dominant cost.

Natural language processing

Classification, extraction and search, increasingly overlapping with large language model work.

Operational machine learning

Predictive maintenance and anomaly detection on sensor data, where the engineering matters more than the model.

If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.

Support status of the tools in this stack

Machine learning tooling moves quickly and library compatibility is a frequent source of friction, so what matters is the currency of the stack a team runs rather than any single version.

Machine Learning itself is not versioned as a single product, so the useful equivalent is the support status of the tools a machine learning engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.

Release and support status across the Machine Learning toolchain
ToolLatest releaseRelease dateMaintained linesFurthest end-of-life date
Python3.14.72026-08-0552030-10-31
NumPy2.5.32026-09-0642028-06-22
Kubernetes1.37.12026-09-2342027-10-28
PostgreSQL18.62026-08-1152030-11-14
Redis8.10.22026-09-1752030-09-01
Docker Engine29.8.12026-09-1512026-12-04

Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.

The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.

The toolchain around it

Nobody hires for Machine Learning alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.

Python
The language, used as an interface to libraries implemented in faster languages.
scikit-learn
Classical models, and still the right answer far more often than deep learning.
PyTorch
Deep learning, dominant in research and increasingly in production.
SQL and a warehouse
Where training data actually comes from. Underrated and central.
MLflow or Weights and Biases
Experiment tracking, so results are reproducible and comparable.
A feature store or disciplined equivalent
Consistency between training and serving, the main source of silent failure.
Docker and a serving layer
Deployment, whether batch scoring or real-time inference.
Monitoring for drift
Watching input and prediction distributions, since degradation raises no errors.

Related skills that frequently appear on the same specification: AI Engineering, Python, Data Engineering, Data Science, AWS.

What to test in an interview

These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.

How they evaluate a model

The single most informative question, and where weak candidates are exposed immediately.

Data leakage

The most common reason a model performs well in testing and badly in production.

Training and serving consistency

Where production machine learning fails silently.

Monitoring a deployed model

Separates people who deploy models from people who own systems.

When not to use machine learning

Tests judgement, and rules are frequently the better answer.

Working with imperfect labels

Real labels are scarce, noisy and expensive, unlike benchmark datasets.

A model that failed in production

Operational judgement in this field comes from failures.

Warning signs in a Machine Learning codebase

The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.

Data leakage

Accuracy on an imbalanced problem

Training and serving skew

Notebooks in production

No monitoring after deployment

Deep learning by default

What each level can own

Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a machine learning engineer.

Junior
Trains and evaluates models on prepared data. Needs review on evaluation design and leakage.
Mid-level
Owns a model end to end including its pipeline, deployment and monitoring. Designs evaluation that reflects the business cost.
Senior
Owns system architecture, the feature and training infrastructure, the monitoring and retraining strategy, and can say when machine learning is not the answer.
Staff
Owns the platform across teams, the standards for evaluation and deployment, and the judgement about which problems merit models at all.

How the work is usually scoped

Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.

First model into production

Research to production

Platform build

Model improvement

Monitoring and retraining

Migration work you may actually be hiring for

A large share of Machine Learning work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.

Notebooks to production pipelines

Separate training and serving feature code to shared transformations

Manual retraining to an automated pipeline

Deep learning on tabular data to gradient-boosted trees

Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.

What a good brief for this role contains

Most of the time lost in hiring a machine learning engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.

If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.

What the US market pays for this work

Machine Learning work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a machine learning engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.

US annual wages, Data Scientists, May 2025
US annual wages, Data Scientists, May 2025$120,230Median$67,240$199,13010th pct90th pctMiddle half $86K to $159K

The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a machine learning engineer can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.

Related classifications are worth reading alongside it, because teams hiring for Machine Learning frequently end up recruiting against these titles too:

US national wages, May 2025
OccupationEmployed25th percentileMedian75th percentile90th percentile
Data Scientists262,440$85,660$120,230$158,880$199,130
Software Developers1,687,890$105,210$135,980$171,980$214,670
Computer and Information Research Scientists37,200$103,570$140,300$188,700$230,630

Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.

These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.

How US metro markets compare for this role

The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.

Median wage for data scientists, by US metro area
Median wage for data scientists, by US metro areaSan Jose, CA: $185,080San Jose, CASan Jose, CA$185,080San Francisco, CA: $170,110San Francisco, CASan Francisco, CA$170,110Seattle, WA: $164,740Seattle, WASeattle, WA$164,740New York, NY: $135,980New York, NYNew York, NY$135,980Baltimore, MD: $134,320Baltimore, MDBaltimore, MD$134,320Charlotte, NC: $132,460Charlotte, NCCharlotte, NC$132,460Washington, D.C.: $132,200Washington, D.C.Washington, D.C.$132,200Boston, MA: $132,040Boston, MABoston, MA$132,040San Diego, CA: $130,990San Diego, CASan Diego, CA$130,990Minneapolis-St. Paul, MN: $129,780Minneapolis-St. Paul, MNMinneapolis-St. Paul, MN$129,780Los Angeles, CA: $129,740Los Angeles, CALos Angeles, CA$129,740Portland, OR: $129,600Portland, ORPortland, OR$129,600Dallas-Fort Worth, TX: $127,750Dallas-Fort Worth, TXDallas-Fort Worth, TX$127,750Miami, FL: $127,450Miami, FLMiami, FL$127,450Austin, TX: $127,360Austin, TXAustin, TX$127,360Raleigh, NC: $120,710Raleigh, NCRaleigh, NC$120,710Salt Lake City, UT: $114,990Salt Lake City, UTSalt Lake City, UT$114,990Phoenix, AZ: $114,540Phoenix, AZPhoenix, AZ$114,540Denver, CO: $112,520Denver, CODenver, CO$112,520Tampa, FL: $109,990Tampa, FLTampa, FL$109,990Philadelphia, PA: $109,910Philadelphia, PAPhiladelphia, PA$109,910Atlanta, GA: $108,940Atlanta, GAAtlanta, GA$108,940Chicago, IL: $107,640Chicago, ILChicago, IL$107,640Houston, TX: $106,750Houston, TXHouston, TX$106,750Orlando, FL: $106,590Orlando, FLOrlando, FL$106,590Detroit, MI: $103,330Detroit, MIDetroit, MI$103,330Kansas City, MO: $99,870Kansas City, MOKansas City, MO$99,870Pittsburgh, PA: $96,670Pittsburgh, PAPittsburgh, PA$96,670
Data Scientists by metro area, May 2025, ranked by median wage
Metro areaEmployedMedian wagevs US medianLocation quotient
San Jose, CA6,060$185,080+54%3.16
San Francisco, CA10,460$170,110+41%2.61
Seattle, WA8,370$164,740+37%2.38
New York, NY23,160$135,980+13%1.45
Baltimore, MD1,090$134,320+12%0.48
Charlotte, NC4,420$132,460+10%1.93
Washington, D.C.9,260$132,200+10%1.75
Boston, MA7,930$132,040+10%1.74
San Diego, CA2,830$130,990+9%1.09
Minneapolis-St. Paul, MN3,250$129,780+8%0.99
Los Angeles, CA9,850$129,740+8%0.93
Portland, OR1,700$129,600+8%0.83
Dallas-Fort Worth, TX10,120$127,750+6%1.48
Miami, FL3,040$127,450+6%0.64
Austin, TX3,730$127,360+6%1.71
Raleigh, NC1,990$120,7100%1.59
Salt Lake City, UT2,970$114,990-4%2.13
Phoenix, AZ3,480$114,540-5%0.87
Denver, CO4,510$112,520-6%1.66
Tampa, FL1,730$109,990-9%0.71
Philadelphia, PA6,480$109,910-9%1.32
Atlanta, GA6,820$108,940-9%1.40
Chicago, IL7,940$107,640-10%1.04
Houston, TX4,060$106,750-11%0.73
Orlando, FL1,440$106,590-11%0.61
Detroit, MI3,810$103,330-14%1.18
Kansas City, MO1,260$99,870-17%0.68
Pittsburgh, PA2,270$96,670-20%1.21

Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.

The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.

The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.

Hiring risks worth naming

Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.

Research background without engineering practice. The most common gap. Test version control, testing, deployment and on-call thinking directly.

Benchmark experience only. Public datasets are clean. Ask about label noise, imbalance and the cost of acquiring data.

No production ownership. Ask about a model that degraded. Deployment is where the real lessons are.

Modelling enthusiasm over problem fit. Ask when rules would beat a model. Someone with no answer will build one where it is not needed.

Hiring Machine Learning developers by metro area

Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of machine learning engineer who will be available.

Frequently asked questions

Do we need a machine learning engineer or a data scientist?

A data scientist if the output is insight and analysis that informs decisions. A machine learning engineer if the output is a system that produces predictions users or processes depend on. The second role is mostly software engineering, and hiring a strong researcher into it typically produces excellent prototypes that never reach production.

What do we need before we can do machine learning at all?

Data that exists, is accessible, and is labelled or has an outcome you can learn from. Most organisations that want to start discover that this is the actual project. It is entirely normal for the first six months of a machine learning initiative to be data engineering, and treating that as a failure rather than as the work is a common mistake.

Why did our model work in testing and fail in production?

The two usual causes are data leakage, where a feature encoded information that would not be available at prediction time, and training-serving skew, where features are computed differently in the two paths. Both produce exactly this symptom and neither raises an error. They are the first two things to check.

How do we know if a model is still working?

You have to measure it deliberately, because nothing will fail. Monitor the distribution of inputs and predictions for drift, and compare predictions against outcomes once those outcomes arrive. A model with no monitoring is a model whose current quality nobody knows, and quality decays as the world changes.

Should we build or buy?

Buy when the problem is general, such as translation, transcription, or standard document extraction, because a provider with vastly more data will beat anything you build. Build when the value comes from your own data and your specific problem. The mistake in both directions is common: building commodity capability, or buying something that needed your data to be useful.

How much data do we need?

It depends far more on the problem than on a number anyone can quote. What matters more is quality, label accuracy and whether the data represents the situations the model will meet. A smaller, well-labelled, representative dataset routinely beats a larger noisy one, and effort spent on labelling quality is often better spent than effort on modelling.

Is machine learning the right approach for our problem?

Often not, and a good engineer will tell you. If the rules are known and stable, write the rules: they are cheaper, explainable and easier to change. Machine learning earns its place where the pattern is real, complex and present in data, but not expressible as rules anyone can write down.

How long before a model delivers value?

The modelling is usually the short part. Getting data pipelines reliable, deploying, and building enough monitoring to trust the output is where the time goes. For a first model in an organisation with no existing infrastructure, plan in quarters rather than weeks, and be suspicious of any estimate that does not account for the data.