Hire Machine Learning developers
Machine learning engineering is mostly software and data engineering with modelling attached, and hiring for it means testing evaluation discipline rather than algorithm recall.
What Machine Learning actually is
Machine learning engineering covers building systems that learn patterns from data and then serve predictions in production. In practice the role spans data preparation, feature engineering, model training, evaluation, deployment and the ongoing monitoring of something whose behaviour changes as the world changes. Most of the work is not modelling.
That last point is the one job specifications most often get wrong. A production machine learning system is a piece of software with a statistical component, and the proportion of effort that goes into pipelines, reproducibility, serving infrastructure and monitoring far exceeds the proportion spent choosing and tuning models. Hiring someone who is strong at modelling and weak at engineering produces notebooks that never ship.
The distinction from data science is worth being explicit about. A data scientist is typically judged on insight and on models that answer a question. A machine learning engineer is judged on systems that run reliably and produce predictions users depend on. Both are legitimate; they attract different people, and a specification that blends them produces a shortlist where nobody is quite right.
The part that separates seniors from mid-levels
Evaluation is where competence is most visible and most often absent. A model's reported accuracy is meaningless without knowing how the data was split, whether the split respected time, whether the classes were balanced, and whether any information from the future leaked into training. Data leakage is the classic failure: a model that performs superbly in testing and poorly in production, because a feature encoded the answer. Candidates who have been caught by this once describe evaluation very differently from those who have not.
The training and serving boundary is the second area. Features computed one way during training and another way at serving time produce silent degradation that no test catches, because both paths work in isolation. This is among the most common production failures in machine learning systems and the reason feature consistency has become an explicit engineering concern rather than an incidental one.
The third is that models decay. The world changes, the input distribution shifts, and a model that was accurate at launch quietly becomes less so. Unlike ordinary software, nothing errors. Monitoring prediction distributions, tracking performance against outcomes when they eventually arrive, and having a retraining path are what separate a system somebody owns from a model somebody deployed.
Where Machine Learning is used
The label “Machine Learning developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Recommendation and ranking
Product, content and search ranking, where the feedback loop between predictions and behaviour is itself a modelling problem.
Fraud and risk
Financial services and marketplaces, where class imbalance is severe and the cost of the two error types is very different.
Demand and forecasting
Inventory, pricing and capacity planning, where time-aware evaluation is essential and often mishandled.
Computer vision
Inspection, medical imaging and document processing, where data labelling is usually the dominant cost.
Natural language processing
Classification, extraction and search, increasingly overlapping with large language model work.
Operational machine learning
Predictive maintenance and anomaly detection on sensor data, where the engineering matters more than the model.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
Machine learning tooling moves quickly and library compatibility is a frequent source of friction, so what matters is the currency of the stack a team runs rather than any single version.
Machine Learning itself is not versioned as a single product, so the useful equivalent is the support status of the tools a machine learning engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Python | 3.14.7 | 2026-08-05 | 5 | 2030-10-31 |
| NumPy | 2.5.3 | 2026-09-06 | 4 | 2028-06-22 |
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
| PostgreSQL | 18.6 | 2026-08-11 | 5 | 2030-11-14 |
| Redis | 8.10.2 | 2026-09-17 | 5 | 2030-09-01 |
| Docker Engine | 29.8.1 | 2026-09-15 | 1 | 2026-12-04 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for Machine Learning alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- Python
- The language, used as an interface to libraries implemented in faster languages.
- scikit-learn
- Classical models, and still the right answer far more often than deep learning.
- PyTorch
- Deep learning, dominant in research and increasingly in production.
- SQL and a warehouse
- Where training data actually comes from. Underrated and central.
- MLflow or Weights and Biases
- Experiment tracking, so results are reproducible and comparable.
- A feature store or disciplined equivalent
- Consistency between training and serving, the main source of silent failure.
- Docker and a serving layer
- Deployment, whether batch scoring or real-time inference.
- Monitoring for drift
- Watching input and prediction distributions, since degradation raises no errors.
Related skills that frequently appear on the same specification: AI Engineering, Python, Data Engineering, Data Science, AWS.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
How they evaluate a model
The single most informative question, and where weak candidates are exposed immediately.
- Strong answer: Chooses metrics that match the business cost of each error type, splits appropriately including by time, and holds out a genuine test set.
- Warning sign: Quotes accuracy on an imbalanced problem, or tunes against the test set.
Data leakage
The most common reason a model performs well in testing and badly in production.
- Strong answer: Can describe a leak they found, and checks for it systematically rather than hoping.
- Warning sign: Unfamiliar with the concept, or has never had a model underperform after deployment.
Training and serving consistency
Where production machine learning fails silently.
- Strong answer: Computes features the same way in both paths, and has diagnosed a mismatch.
- Warning sign: Has never considered that the two paths could differ.
Monitoring a deployed model
Separates people who deploy models from people who own systems.
- Strong answer: Monitors input and prediction distributions, tracks outcomes when they arrive, and has a retraining trigger.
- Warning sign: Deploys and moves on, treating the model as finished.
When not to use machine learning
Tests judgement, and rules are frequently the better answer.
- Strong answer: Can describe problems solved better with rules, heuristics or a simpler statistical approach.
- Warning sign: Treats machine learning as the answer to every prediction problem.
Working with imperfect labels
Real labels are scarce, noisy and expensive, unlike benchmark datasets.
- Strong answer: Has dealt with label noise, class imbalance and the cost of acquiring labels.
- Warning sign: Experience only on clean public datasets.
A model that failed in production
Operational judgement in this field comes from failures.
- Strong answer: Describes what degraded, how it was detected, and what changed afterwards.
- Warning sign: Has never had a model in production.
Warning signs in a Machine Learning codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
Data leakage
- What you see: A feature that encodes information unavailable at prediction time.
- What it costs: Excellent test results and poor production performance, with the cause often taking weeks to find.
- The fix: Construct features from the perspective of the moment of prediction. Split by time where the problem is temporal.
Accuracy on an imbalanced problem
- What you see: A fraud model reported at high accuracy when the positive class is rare.
- What it costs: A model that predicts the majority class and looks excellent while being useless.
- The fix: Use metrics matched to the cost of each error type, and state the base rate alongside any figure.
Training and serving skew
- What you see: Feature computation implemented separately in the training pipeline and the serving path.
- What it costs: Predictions quietly worse than evaluation suggested, with no error to alert anyone.
- The fix: Share the transformation code, or use a feature store. Verify with the same input through both paths.
Notebooks in production
- What you see: Training or scoring code run from a notebook on a schedule.
- What it costs: Hidden state, no tests, no reproducibility, and failures nobody else can debug.
- The fix: Move to tested modules with orchestration. Keep notebooks for exploration.
No monitoring after deployment
- What you see: A model deployed with infrastructure alerts only.
- What it costs: Silent decay over months while the system reports itself healthy.
- The fix: Monitor input and prediction distributions, and compare against outcomes when they become available.
Deep learning by default
- What you see: A neural network applied to tabular data where a gradient-boosted tree would do better.
- What it costs: More compute, more complexity, longer iteration, and frequently worse results.
- The fix: Start simple and establish a baseline. Escalate only when the simpler approach is demonstrably insufficient.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a machine learning engineer.
- Junior
- Trains and evaluates models on prepared data. Needs review on evaluation design and leakage.
- Mid-level
- Owns a model end to end including its pipeline, deployment and monitoring. Designs evaluation that reflects the business cost.
- Senior
- Owns system architecture, the feature and training infrastructure, the monitoring and retraining strategy, and can say when machine learning is not the answer.
- Staff
- Owns the platform across teams, the standards for evaluation and deployment, and the judgement about which problems merit models at all.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
First model into production
- Usual team: One engineer plus a domain expert.
- What governs it: Getting data and deployment right takes far longer than modelling. Budget accordingly and expect the data to be worse than described.
Research to production
- Usual team: One engineer paired with whoever built the prototype.
- What governs it: Scoped by how much hidden state the prototype carries. Reproducing its numbers first is the essential step.
Platform build
- Usual team: Two engineers.
- What governs it: Justified once several models exist. Premature for a team with one.
Model improvement
- Usual team: One engineer, time-boxed.
- What governs it: Usually better data rather than a better model. Establish that honestly before spending on modelling.
Monitoring and retraining
- Usual team: One engineer.
- What governs it: Frequently the highest-return work available on an existing system, and the most often skipped.
Migration work you may actually be hiring for
A large share of Machine Learning work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Notebooks to production pipelines
- Why teams do it: Anything serving predictions must be reproducible, tested and debuggable by someone other than its author.
- What to watch: Reproduce the original numbers with the rewritten code before switching anything. Hidden state from out-of-order execution means exported notebook code frequently does not do what the notebook did.
Separate training and serving feature code to shared transformations
- Why teams do it: Training-serving skew is among the most common and least visible production failures.
- What to watch: Verify by passing identical inputs through both paths and comparing outputs. The discrepancies are usually small, specific and quietly damaging.
Manual retraining to an automated pipeline
- Why teams do it: Models decay, and retraining that depends on someone remembering will not happen.
- What to watch: Automated retraining needs automated evaluation gates, otherwise you automate the deployment of a worse model. Build the gate before the automation.
Deep learning on tabular data to gradient-boosted trees
- Why teams do it: Frequently better results with far less compute and much faster iteration.
- What to watch: Worth testing as a baseline before assuming the opposite. Teams are often surprised, and the simpler model is usually easier to explain to stakeholders too.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a machine learning engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- What decision the model will inform, and what happens to the prediction once it exists.
- Whether the data exists, is accessible and is labelled, since this is usually the real constraint.
- Whether the work is prototyping or production, because these attract different people.
- Whether existing pipelines and infrastructure exist, or whether that is part of the job.
- What the cost of each type of error is, because this determines the evaluation metric.
- Who owns the model once deployed, including monitoring and retraining.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
Machine Learning work is counted by the US Bureau of Labor Statistics under Data Scientists. That classification is broader than the technology itself, so treat the figures as the shape of the market a machine learning engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 262,440 people in this occupation, with a median annual wage of $120,230.
The spread matters more than the midpoint. The 90th percentile is about 3.0 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a machine learning engineer can sit at $67,240 and $199,130 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for Machine Learning frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Data Scientists | 262,440 | $85,660 | $120,230 | $158,880 | $199,130 |
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
| Computer and Information Research Scientists | 37,200 | $103,570 | $140,300 | $188,700 | $230,630 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Data Scientists, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.9. San Jose sits at the top with a median of $185,080; Pittsburgh sits at the bottom with $96,670. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 6,060 | $185,080 | +54% | 3.16 |
| San Francisco, CA | 10,460 | $170,110 | +41% | 2.61 |
| Seattle, WA | 8,370 | $164,740 | +37% | 2.38 |
| New York, NY | 23,160 | $135,980 | +13% | 1.45 |
| Baltimore, MD | 1,090 | $134,320 | +12% | 0.48 |
| Charlotte, NC | 4,420 | $132,460 | +10% | 1.93 |
| Washington, D.C. | 9,260 | $132,200 | +10% | 1.75 |
| Boston, MA | 7,930 | $132,040 | +10% | 1.74 |
| San Diego, CA | 2,830 | $130,990 | +9% | 1.09 |
| Minneapolis-St. Paul, MN | 3,250 | $129,780 | +8% | 0.99 |
| Los Angeles, CA | 9,850 | $129,740 | +8% | 0.93 |
| Portland, OR | 1,700 | $129,600 | +8% | 0.83 |
| Dallas-Fort Worth, TX | 10,120 | $127,750 | +6% | 1.48 |
| Miami, FL | 3,040 | $127,450 | +6% | 0.64 |
| Austin, TX | 3,730 | $127,360 | +6% | 1.71 |
| Raleigh, NC | 1,990 | $120,710 | 0% | 1.59 |
| Salt Lake City, UT | 2,970 | $114,990 | -4% | 2.13 |
| Phoenix, AZ | 3,480 | $114,540 | -5% | 0.87 |
| Denver, CO | 4,510 | $112,520 | -6% | 1.66 |
| Tampa, FL | 1,730 | $109,990 | -9% | 0.71 |
| Philadelphia, PA | 6,480 | $109,910 | -9% | 1.32 |
| Atlanta, GA | 6,820 | $108,940 | -9% | 1.40 |
| Chicago, IL | 7,940 | $107,640 | -10% | 1.04 |
| Houston, TX | 4,060 | $106,750 | -11% | 0.73 |
| Orlando, FL | 1,440 | $106,590 | -11% | 0.61 |
| Detroit, MI | 3,810 | $103,330 | -14% | 1.18 |
| Kansas City, MO | 1,260 | $99,870 | -17% | 0.68 |
| Pittsburgh, PA | 2,270 | $96,670 | -20% | 1.21 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Charlotte, Washington, D.C., Boston each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Research background without engineering practice. The most common gap. Test version control, testing, deployment and on-call thinking directly.
Benchmark experience only. Public datasets are clean. Ask about label noise, imbalance and the cost of acquiring data.
No production ownership. Ask about a model that degraded. Deployment is where the real lessons are.
Modelling enthusiasm over problem fit. Ask when rules would beat a model. Someone with no answer will build one where it is not needed.
Hiring Machine Learning developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of machine learning engineer who will be available.
- New York, NY $135,980 median
- Seattle, WA $164,740 median
- San Jose, CA $185,080 median
- Washington, D.C. $132,200 median
- San Francisco, CA $170,110 median
- Dallas-Fort Worth, TX $127,750 median
- Los Angeles, CA $129,740 median
- Boston, MA $132,040 median
- Chicago, IL $107,640 median
- Atlanta, GA $108,940 median
- Austin, TX $127,360 median
- Phoenix, AZ $114,540 median
- Philadelphia, PA $109,910 median
- Minneapolis-St. Paul, MN $129,780 median
- Denver, CO $112,520 median
- Detroit, MI $103,330 median
- Houston, TX $106,750 median
- Charlotte, NC $132,460 median
- San Diego, CA $130,990 median
- Salt Lake City, UT $114,990 median
- Miami, FL $127,450 median
- Portland, OR $129,600 median
- Baltimore, MD $134,320 median
- Tampa, FL $109,990 median
- Orlando, FL $106,590 median
- Raleigh, NC $120,710 median
- Kansas City, MO $99,870 median
- Pittsburgh, PA $96,670 median
Frequently asked questions
Do we need a machine learning engineer or a data scientist?
A data scientist if the output is insight and analysis that informs decisions. A machine learning engineer if the output is a system that produces predictions users or processes depend on. The second role is mostly software engineering, and hiring a strong researcher into it typically produces excellent prototypes that never reach production.
What do we need before we can do machine learning at all?
Data that exists, is accessible, and is labelled or has an outcome you can learn from. Most organisations that want to start discover that this is the actual project. It is entirely normal for the first six months of a machine learning initiative to be data engineering, and treating that as a failure rather than as the work is a common mistake.
Why did our model work in testing and fail in production?
The two usual causes are data leakage, where a feature encoded information that would not be available at prediction time, and training-serving skew, where features are computed differently in the two paths. Both produce exactly this symptom and neither raises an error. They are the first two things to check.
How do we know if a model is still working?
You have to measure it deliberately, because nothing will fail. Monitor the distribution of inputs and predictions for drift, and compare predictions against outcomes once those outcomes arrive. A model with no monitoring is a model whose current quality nobody knows, and quality decays as the world changes.
Should we build or buy?
Buy when the problem is general, such as translation, transcription, or standard document extraction, because a provider with vastly more data will beat anything you build. Build when the value comes from your own data and your specific problem. The mistake in both directions is common: building commodity capability, or buying something that needed your data to be useful.
How much data do we need?
It depends far more on the problem than on a number anyone can quote. What matters more is quality, label accuracy and whether the data represents the situations the model will meet. A smaller, well-labelled, representative dataset routinely beats a larger noisy one, and effort spent on labelling quality is often better spent than effort on modelling.
Is machine learning the right approach for our problem?
Often not, and a good engineer will tell you. If the rules are known and stable, write the rules: they are cheaper, explainable and easier to change. Machine learning earns its place where the pattern is real, complex and present in data, but not expressible as rules anyone can write down.
How long before a model delivers value?
The modelling is usually the short part. Getting data pipelines reliable, deploying, and building enough monitoring to trust the output is where the time goes. For a first model in an organisation with no existing infrastructure, plan in quarters rather than weeks, and be suspicious of any estimate that does not account for the data.