Hire AI Engineering developers
AI engineering builds products on top of large language models, and hiring for it means testing evaluation and cost discipline rather than familiarity with any particular model.
What AI Engineering actually is
AI engineering is the practice of building applications on top of large language models and related foundation models. It is distinct from machine learning engineering in that the model is usually somebody else's: you are integrating, orchestrating, grounding and evaluating rather than training. The skills are closer to software engineering than to research.
The work typically involves designing prompts and structured outputs, retrieving relevant context to ground responses, handling tool use where the model calls into your systems, evaluating quality systematically, and managing the cost and latency of a component that is slower and more expensive than anything else in the request path.
It is a young discipline and the hiring market reflects that. Very few people have long experience, the tooling changes fast, and a great deal of claimed expertise amounts to having used an API. The candidates worth hiring are distinguished less by which models or frameworks they know than by whether they have shipped something that had to work reliably, and can tell you how they knew it did.
The part that separates seniors from mid-levels
Evaluation is the defining problem, and it is genuinely harder than in traditional machine learning. Outputs are open-ended text, correctness is frequently a matter of judgement, and the same input can produce different outputs. Teams that ship reliable systems build evaluation sets from real failures, score against them automatically on every change, and treat prompt modifications as changes requiring regression testing. Teams that do not are changing prompts and hoping, which works until it does not.
Retrieval is the second area and where most quality problems actually live. When a system answers from your documents, the failure is usually that the right document was not retrieved rather than that the model reasoned badly. Chunking strategy, embedding choice, hybrid search combining semantic and keyword matching, and reranking are the levers. Candidates who jump straight to changing the prompt when quality is poor are skipping the more likely cause.
The third is engineering discipline around a component that is unreliable by nature. The model can be slow, can fail, can return malformed output, and can produce something confident and wrong. Production systems need timeouts, retries, structured output validation, fallbacks and a clear decision about what happens when the model is unavailable. Treating the model as a normal dependency that can fail is what separates a product from a demonstration.
Where AI Engineering is used
The label “AI Engineering developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.
Customer support
Assistants answering from documentation and account data, usually with escalation paths and strict grounding requirements.
Document processing
Extraction, classification and summarisation over contracts, invoices and forms, where structured output matters more than fluency.
Internal knowledge tools
Search and question answering across an organisation's own material, where permissions are a first-class concern.
Content and drafting
Assisted writing inside existing products, where the model produces a draft rather than a final answer.
Agentic workflows
Systems where the model calls tools to accomplish multi-step tasks, which is where reliability is hardest.
Code assistance
Developer tooling within an organisation's own codebase and conventions.
If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.
Support status of the tools in this stack
Model providers deprecate and replace models on their own schedules, so any system built on them carries a recurring maintenance obligation independent of its own development.
AI Engineering itself is not versioned as a single product, so the useful equivalent is the support status of the tools a AI engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.
| Tool | Latest release | Release date | Maintained lines | Furthest end-of-life date |
|---|---|---|---|---|
| Python | 3.14.7 | 2026-08-05 | 5 | 2030-10-31 |
| Node.js | 26.10.0 | 2026-09-22 | 6 | 2029-04-30 |
| PostgreSQL | 18.6 | 2026-08-11 | 5 | 2030-11-14 |
| Redis | 8.10.2 | 2026-09-17 | 5 | 2030-09-01 |
| Kubernetes | 1.37.1 | 2026-09-23 | 4 | 2027-10-28 |
Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.
The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.
The toolchain around it
Nobody hires for AI Engineering alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.
- A model provider API
- The foundation model. Most production systems use more than one for cost, fallback or capability reasons.
- Python or TypeScript
- The implementation languages, with TypeScript common where the product is a web application.
- A vector database or pgvector
- Embedding storage and retrieval. Postgres with an extension is sufficient far more often than assumed.
- Hybrid search
- Semantic and keyword retrieval combined, which materially outperforms either alone.
- An evaluation framework
- Automated scoring against a test set, the practice that separates reliable systems from demonstrations.
- Structured output validation
- Schema enforcement on model output, because malformed responses are routine.
- Tracing and observability
- Recording prompts, retrieved context, outputs and cost per request. Essential for debugging.
- A caching layer
- Reducing cost and latency on repeated or similar requests.
Related skills that frequently appear on the same specification: Machine Learning, Node.js, Python, Data Engineering, TypeScript.
What to test in an interview
These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.
How they evaluate quality
The defining question in this field and where most candidates have nothing to say.
- Strong answer: Has an evaluation set built from real failures, scores automatically, and treats prompt changes as changes requiring regression testing.
- Warning sign: Evaluates by trying a few examples manually and forming an impression.
Diagnosing a bad answer
Tests whether they understand where quality problems actually originate.
- Strong answer: Checks retrieval first, inspects what context the model was given, and only then considers the prompt.
- Warning sign: Immediately rewrites the prompt without looking at what was retrieved.
Retrieval design
Where most grounded-system quality is won or lost.
- Strong answer: Has opinions on chunking, uses hybrid search, has tried reranking, and can describe measuring retrieval quality separately.
- Warning sign: Embeds whole documents and retrieves the nearest few without evaluation.
Handling unreliable output
Production systems must cope with malformed, slow or wrong responses.
- Strong answer: Validates against a schema, retries deliberately, sets timeouts, and has a defined fallback.
- Warning sign: Parses model output optimistically and assumes it will be well formed.
Cost and latency management
These are engineering constraints here in a way they are not in ordinary web development.
- Strong answer: Knows their cost per request, routes to cheaper models where adequate, caches, and can describe a reduction they achieved.
- Warning sign: Has never measured cost per request.
Prompt injection and data boundaries
A real security concern, particularly with tool use and retrieved content.
- Strong answer: Treats retrieved and user content as untrusted, constrains tool permissions, and does not rely on instructions to enforce security.
- Warning sign: Believes instructing the model not to do something is a security control.
Something they shipped that had to work
Distinguishes production experience from experimentation, which is the main divide in this market.
- Strong answer: Describes a real system, its failure modes, and what they did about them.
- Warning sign: Experience is entirely demonstrations and prototypes.
Warning signs in a AI Engineering codebase
The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.
No evaluation set
- What you see: Prompt changes assessed by trying a handful of examples by hand.
- What it costs: Improvements in one area silently break another, and nobody can tell whether the system is getting better.
- The fix: Build a set from real failures and score automatically on every change. This is the highest-value practice in the field.
Blaming the prompt for retrieval failures
- What you see: Repeated prompt rewriting while answer quality stays poor.
- What it costs: Time spent on the wrong component while the actual problem persists.
- The fix: Inspect what was retrieved before touching the prompt. Measure retrieval quality as its own metric.
Trusting model output structure
- What you see: Parsing responses without validation.
- What it costs: Runtime failures when the model returns something unexpected, which it periodically will.
- The fix: Enforce a schema, validate, and handle failure explicitly with a retry or a fallback.
Instructions as a security boundary
- What you see: Relying on telling the model not to reveal or do something.
- What it costs: Prompt injection through user input or retrieved content, which is a genuine and demonstrated attack.
- The fix: Enforce permissions outside the model. Constrain what tools can do and filter what content can reach the context.
Unmeasured cost
- What you see: No visibility of spend per request or per feature.
- What it costs: A bill that scales with usage in ways nobody predicted, sometimes dramatically.
- The fix: Track cost per request as a first-class metric. Route to smaller models where they are adequate and cache aggressively.
Agents where a workflow would do
- What you see: A model given open-ended tool access for a task with known steps.
- What it costs: Unpredictable behaviour, difficult debugging and far higher cost than a deterministic implementation.
- The fix: Write the workflow where the steps are known. Reserve open-ended agency for cases where the path genuinely cannot be predetermined.
What each level can own
Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a AI engineer.
- Junior
- Implements features against a model API within an existing structure. Needs review on evaluation and on handling failure.
- Mid-level
- Owns a feature end to end including retrieval, evaluation and cost. Can debug a quality problem systematically.
- Senior
- Owns system architecture, the evaluation strategy, the retrieval design, the cost model and the security posture around untrusted content.
- Staff
- Owns the organisation's approach: which problems merit these systems, provider strategy and portability, data governance, and the standards other teams build against.
How the work is usually scoped
Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.
First production feature
- Usual team: One engineer plus a domain expert for evaluation.
- What governs it: A demonstration takes days; something reliable takes considerably longer. The gap is almost entirely evaluation and failure handling.
Retrieval over internal documents
- Usual team: One engineer plus a content owner.
- What governs it: Document quality and permissions are the real scope. Retrieval is easy to build and hard to make good.
Evaluation infrastructure
- Usual team: One engineer.
- What governs it: Often the highest-return work on an existing system, and almost always the thing that was skipped.
Cost reduction
- Usual team: One engineer, time-boxed.
- What governs it: Model routing, caching and prompt size. Frequently large and quick savings.
Agentic workflow
- Usual team: One to two senior engineers.
- What governs it: Considerably harder to make reliable than it looks in a demonstration. Scope conservatively.
Migration work you may actually be hiring for
A large share of AI Engineering work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.
Manual quality checking to an automated evaluation set
- Why teams do it: Without it, nobody can tell whether a change improved the system or broke something else.
- What to watch: Build the set from real failures rather than invented examples. Start small; fifty genuine cases scored consistently is worth more than a thousand synthetic ones.
Naive retrieval to hybrid search with reranking
- Why teams do it: Most grounded-answer quality problems are retrieval problems, and semantic search alone misses exact terms.
- What to watch: Measure retrieval quality separately from answer quality, otherwise you cannot tell which component you improved.
One large model for everything to routing by request
- Why teams do it: Cost and latency. Many requests are handled perfectly well by a smaller, cheaper model.
- What to watch: Route on measured quality against the evaluation set rather than on intuition, and keep the fallback to the larger model for cases the smaller one fails.
Free-form text output to validated structured output
- Why teams do it: Anything consumed by code needs a guaranteed shape rather than a usually-correct one.
- What to watch: Validation must include a failure path. A retry with a corrective message handles most cases; silently accepting malformed output does not.
Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.
What a good brief for this role contains
Most of the time lost in hiring a AI engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.
- What the system does and what happens to its output, since a draft for a human and an automated decision carry very different reliability requirements.
- Whether an evaluation set exists, because if not, building one is the first task.
- Whether retrieval over your own content is involved, and what state that content is in.
- Whether tool use or agentic behaviour is in scope, which is substantially harder.
- What the cost and latency constraints are, since both are real engineering limits here.
- Which provider is in use and whether portability matters to the organisation.
If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.
What the US market pays for this work
AI Engineering work is counted by the US Bureau of Labor Statistics under Software Developers. That classification is broader than the technology itself, so treat the figures as the shape of the market a AI engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 1,687,890 people in this occupation, with a median annual wage of $135,980.
The spread matters more than the midpoint. The 90th percentile is about 2.6 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a AI engineer can sit at $82,460 and $214,670 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.
Related classifications are worth reading alongside it, because teams hiring for AI Engineering frequently end up recruiting against these titles too:
| Occupation | Employed | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|---|
| Software Developers | 1,687,890 | $105,210 | $135,980 | $171,980 | $214,670 |
| Data Scientists | 262,440 | $85,660 | $120,230 | $158,880 | $199,130 |
| Computer and Information Research Scientists | 37,200 | $103,570 | $140,300 | $188,700 | $230,630 |
Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.
These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.
How US metro markets compare for this role
The same job is priced very differently across the country. Ranked by median annual wage for Software Developers, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.7. San Jose sits at the top with a median of $213,110; Pittsburgh sits at the bottom with $124,500. A budget built from a national median will be wrong in both of those markets, in opposite directions.
| Metro area | Employed | Median wage | vs US median | Location quotient |
|---|---|---|---|---|
| San Jose, CA | 87,350 | $213,110 | +57% | 7.09 |
| San Francisco, CA | 69,030 | $186,640 | +37% | 2.68 |
| Seattle, WA | 92,770 | $167,280 | +23% | 4.10 |
| New York, NY | 121,000 | $166,830 | +23% | 1.17 |
| Boston, MA | 42,310 | $166,090 | +22% | 1.44 |
| San Diego, CA | 20,610 | $163,270 | +20% | 1.23 |
| Los Angeles, CA | 55,540 | $160,920 | +18% | 0.82 |
| Portland, OR | 18,260 | $156,000 | +15% | 1.39 |
| Washington, D.C. | 69,060 | $154,930 | +14% | 2.03 |
| Baltimore, MD | 16,850 | $138,900 | +2% | 1.14 |
| Denver, CO | 27,010 | $137,610 | +1% | 1.55 |
| Charlotte, NC | 20,820 | $135,920 | 0% | 1.41 |
| Chicago, IL | 40,370 | $134,380 | -1% | 0.82 |
| Austin, TX | 31,960 | $134,120 | -1% | 2.28 |
| Dallas-Fort Worth, TX | 67,030 | $133,290 | -2% | 1.52 |
| Philadelphia, PA | 28,480 | $133,040 | -2% | 0.91 |
| Atlanta, GA | 36,300 | $132,960 | -2% | 1.16 |
| Raleigh, NC | 12,580 | $132,770 | -2% | 1.56 |
| Miami, FL | 18,900 | $132,650 | -2% | 0.62 |
| Phoenix, AZ | 29,380 | $131,750 | -3% | 1.14 |
| Minneapolis-St. Paul, MN | 27,410 | $130,920 | -4% | 1.29 |
| Detroit, MI | 24,870 | $130,760 | -4% | 1.20 |
| Tampa, FL | 14,230 | $130,450 | -4% | 0.91 |
| Orlando, FL | 13,440 | $129,620 | -5% | 0.88 |
| Salt Lake City, UT | 19,040 | $129,600 | -5% | 2.12 |
| Houston, TX | 22,940 | $129,440 | -5% | 0.64 |
| Kansas City, MO | 12,160 | $124,990 | -8% | 1.02 |
| Pittsburgh, PA | 10,320 | $124,500 | -8% | 0.85 |
Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.
The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Washington, D.C., Denver, Austin each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.
The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.
Hiring risks worth naming
Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.
Demonstration experience presented as production experience. The central risk in this market. Ask what they shipped, who used it, and how they knew it worked.
No evaluation discipline. Ask how they would know a prompt change improved things. Weak answers here predict a system nobody can improve confidently.
No cost awareness. These systems can be expensive in ways that scale with success. Ask for a cost per request figure.
Security naivety about untrusted content. Prompt injection is real. Ask how they would stop retrieved content from influencing tool use.
Hiring AI Engineering developers by metro area
Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of AI engineer who will be available.
- New York, NY $166,830 median
- Seattle, WA $167,280 median
- San Jose, CA $213,110 median
- Washington, D.C. $154,930 median
- San Francisco, CA $186,640 median
- Dallas-Fort Worth, TX $133,290 median
- Los Angeles, CA $160,920 median
- Boston, MA $166,090 median
- Chicago, IL $134,380 median
- Atlanta, GA $132,960 median
- Austin, TX $134,120 median
- Phoenix, AZ $131,750 median
- Philadelphia, PA $133,040 median
- Minneapolis-St. Paul, MN $130,920 median
- Denver, CO $137,610 median
- Detroit, MI $130,760 median
- Houston, TX $129,440 median
- Charlotte, NC $135,920 median
- San Diego, CA $163,270 median
- Salt Lake City, UT $129,600 median
- Miami, FL $132,650 median
- Portland, OR $156,000 median
- Baltimore, MD $138,900 median
- Tampa, FL $130,450 median
- Orlando, FL $129,620 median
- Raleigh, NC $132,770 median
- Kansas City, MO $124,990 median
- Pittsburgh, PA $124,500 median
Frequently asked questions
Do we need an AI engineer or can our existing developers do this?
Competent developers can integrate a model API, and that is the easy part. What takes specific experience is evaluation, retrieval quality, handling unreliable output and controlling cost. A sensible arrangement for many teams is one person with real experience setting the patterns and evaluation infrastructure, then the existing team building within them.
How do we know if one of these systems is actually working?
You need an evaluation set: real inputs with expected outputs or quality criteria, scored automatically whenever anything changes. Without it, quality is a matter of impression and every prompt change is a gamble. This is the single practice that most separates teams shipping reliable systems from teams shipping demonstrations.
Why does our assistant give wrong answers about our own documents?
Usually retrieval rather than reasoning. The model can only use what it was given, and if the relevant passage was not retrieved it will answer from general knowledge, confidently. Check what context was actually supplied before changing the prompt. Chunking strategy and hybrid search are typically where the fix lies.
Should we fine-tune a model?
Less often than people assume. Retrieval handles knowledge, and prompting handles most behaviour. Fine-tuning earns its place for consistent formatting, a specific style, or a narrow classification task where you have good labelled examples. It is usually the wrong first move and a reasonable later optimisation.
How do we control the cost?
Measure cost per request first, because most teams cannot answer what a feature costs. After that the levers are routing simpler requests to cheaper models, caching repeated and similar requests, and reducing context size, which is often the largest single contributor. Cost here scales with usage, so success makes it worse rather than better.
Are we locked into one model provider?
Partially, and it is manageable if you plan for it. Keep provider-specific code behind an interface, keep your evaluation set portable, and test periodically against an alternative. Prompts do not transfer perfectly between models, so the realistic position is that switching is a project rather than a configuration change, and knowing its size is the useful thing.
Is prompt injection a real risk?
Yes, and it should be designed for rather than instructed against. Any content the model sees, including retrieved documents and user input, can carry instructions. The defence is architectural: constrain what tools can do, enforce permissions outside the model, and treat model output as untrusted input to anything downstream. Telling the model to ignore instructions is not a control.
Should we build agents?
Only where the sequence of steps genuinely cannot be determined in advance. Where the steps are known, writing the workflow gives you something predictable, debuggable and far cheaper. Open-ended agency is appealing in demonstrations and consistently harder to make reliable in production, and a good engineer will push back on it.