FuturByte

Hire AI Engineering developers

AI engineering builds products on top of large language models, and hiring for it means testing evaluation and cost discipline rather than familiarity with any particular model.

What AI Engineering actually is

AI engineering is the practice of building applications on top of large language models and related foundation models. It is distinct from machine learning engineering in that the model is usually somebody else's: you are integrating, orchestrating, grounding and evaluating rather than training. The skills are closer to software engineering than to research.

The work typically involves designing prompts and structured outputs, retrieving relevant context to ground responses, handling tool use where the model calls into your systems, evaluating quality systematically, and managing the cost and latency of a component that is slower and more expensive than anything else in the request path.

It is a young discipline and the hiring market reflects that. Very few people have long experience, the tooling changes fast, and a great deal of claimed expertise amounts to having used an API. The candidates worth hiring are distinguished less by which models or frameworks they know than by whether they have shipped something that had to work reliably, and can tell you how they knew it did.

The part that separates seniors from mid-levels

Evaluation is the defining problem, and it is genuinely harder than in traditional machine learning. Outputs are open-ended text, correctness is frequently a matter of judgement, and the same input can produce different outputs. Teams that ship reliable systems build evaluation sets from real failures, score against them automatically on every change, and treat prompt modifications as changes requiring regression testing. Teams that do not are changing prompts and hoping, which works until it does not.

Retrieval is the second area and where most quality problems actually live. When a system answers from your documents, the failure is usually that the right document was not retrieved rather than that the model reasoned badly. Chunking strategy, embedding choice, hybrid search combining semantic and keyword matching, and reranking are the levers. Candidates who jump straight to changing the prompt when quality is poor are skipping the more likely cause.

The third is engineering discipline around a component that is unreliable by nature. The model can be slow, can fail, can return malformed output, and can produce something confident and wrong. Production systems need timeouts, retries, structured output validation, fallbacks and a clear decision about what happens when the model is unavailable. Treating the model as a normal dependency that can fail is what separates a product from a demonstration.

Where AI Engineering is used

The label “AI Engineering developer” covers several jobs that share a technology and little else. These are the settings the work usually turns up in, and the one you are hiring into should shape the whole process, because the judgement each demands is different.

Customer support

Assistants answering from documentation and account data, usually with escalation paths and strict grounding requirements.

Document processing

Extraction, classification and summarisation over contracts, invoices and forms, where structured output matters more than fluency.

Internal knowledge tools

Search and question answering across an organisation's own material, where permissions are a first-class concern.

Content and drafting

Assisted writing inside existing products, where the model produces a draft rather than a final answer.

Agentic workflows

Systems where the model calls tools to accomplish multi-step tasks, which is where reliability is hardest.

Code assistance

Developer tooling within an organisation's own codebase and conventions.

If a candidate's experience sits in a different row of that list from the work you have, that is not a reason to reject them, but it is the thing to probe. Ask what would be different about their approach in your setting. Someone who can answer that has transferable judgement. Someone who says it would be much the same has probably not thought about it.

Support status of the tools in this stack

Model providers deprecate and replace models on their own schedules, so any system built on them carries a recurring maintenance obligation independent of its own development.

AI Engineering itself is not versioned as a single product, so the useful equivalent is the support status of the tools a AI engineer works with daily. The table is read from public release data rather than written by hand, so it states what is supported now. It is worth having in front of you during an interview: asking which of these a candidate has upgraded, and what broke, gets you further than asking how many years they have used each.

Release and support status across the AI Engineering toolchain
ToolLatest releaseRelease dateMaintained linesFurthest end-of-life date
Python3.14.72026-08-0552030-10-31
Node.js26.10.02026-09-2262029-04-30
PostgreSQL18.62026-08-1152030-11-14
Redis8.10.22026-09-1752030-09-01
Kubernetes1.37.12026-09-2342027-10-28

Source: endoflife.date public release data, read 2026-09-25. A tool with no published end-of-life dates sets its support boundary by ecosystem practice rather than by policy.

The practical use of this is in judging an estate rather than a person. A team running several of these past their support dates is usually not behind by accident; it is behind because upgrades were never anyone's job. That is worth knowing before you hire, because it tells you whether the first six months will be building new things or paying down what was deferred.

The toolchain around it

Nobody hires for AI Engineering alone. The surrounding tools are where most of the day-to-day work happens, and a gap in any of them costs more time than a gap in the core library. This is the set that turns up most often on real job specifications alongside it.

A model provider API
The foundation model. Most production systems use more than one for cost, fallback or capability reasons.
Python or TypeScript
The implementation languages, with TypeScript common where the product is a web application.
A vector database or pgvector
Embedding storage and retrieval. Postgres with an extension is sufficient far more often than assumed.
Hybrid search
Semantic and keyword retrieval combined, which materially outperforms either alone.
An evaluation framework
Automated scoring against a test set, the practice that separates reliable systems from demonstrations.
Structured output validation
Schema enforcement on model output, because malformed responses are routine.
Tracing and observability
Recording prompts, retrieved context, outputs and cost per request. Essential for debugging.
A caching layer
Reducing cost and latency on repeated or similar requests.

Related skills that frequently appear on the same specification: Machine Learning, Node.js, Python, Data Engineering, TypeScript.

What to test in an interview

These are the topics that separate candidates in practice. Each one is given with why it discriminates, what a strong answer sounds like, and the response that should make you slow down. None of them requires a whiteboard.

How they evaluate quality

The defining question in this field and where most candidates have nothing to say.

Diagnosing a bad answer

Tests whether they understand where quality problems actually originate.

Retrieval design

Where most grounded-system quality is won or lost.

Handling unreliable output

Production systems must cope with malformed, slow or wrong responses.

Cost and latency management

These are engineering constraints here in a way they are not in ordinary web development.

Prompt injection and data boundaries

A real security concern, particularly with tool use and retrieved content.

Something they shipped that had to work

Distinguishes production experience from experimentation, which is the main divide in this market.

Warning signs in a AI Engineering codebase

The fastest way to read a candidate is to ask what they have found wrong in code they inherited. These are the patterns that come up most often, what they cost, and what fixing them looks like. A developer who recognises three or four of these from their own experience is worth more than one who can recite the documentation.

No evaluation set

Blaming the prompt for retrieval failures

Trusting model output structure

Instructions as a security boundary

Unmeasured cost

Agents where a workflow would do

What each level can own

Job titles are not comparable between companies, so it is more useful to describe levels by what a person can be left to own without supervision. These are the boundaries we use when we assess a AI engineer.

Junior
Implements features against a model API within an existing structure. Needs review on evaluation and on handling failure.
Mid-level
Owns a feature end to end including retrieval, evaluation and cost. Can debug a quality problem systematically.
Senior
Owns system architecture, the evaluation strategy, the retrieval design, the cost model and the security posture around untrusted content.
Staff
Owns the organisation's approach: which problems merit these systems, provider strategy and portability, data governance, and the standards other teams build against.

How the work is usually scoped

Team shape follows the kind of work, not the headcount you happen to have budget for. These are the shapes that come up most often and the constraint that actually governs each one.

First production feature

Retrieval over internal documents

Evaluation infrastructure

Cost reduction

Agentic workflow

Migration work you may actually be hiring for

A large share of AI Engineering work is not new development. It is moving an existing system from one state to another while it stays in service. These are the migrations that come up most often, and each one asks for a different kind of experience from the person you hire.

Manual quality checking to an automated evaluation set

Naive retrieval to hybrid search with reranking

One large model for everything to routing by request

Free-form text output to validated structured output

Migration work rewards a different temperament from greenfield work. The useful question in an interview is not whether someone has done the specific migration you face, but whether they have ever run one incrementally: behind a flag, with both paths live, and with a way back. Developers who have only done big-bang cutovers tend to propose them again.

What a good brief for this role contains

Most of the time lost in hiring a AI engineer is lost before anyone is interviewed, in the gap between what the brief says and what the team actually needs. These are the points that, for this technology specifically, change who the right candidate is. A brief that answers them can be matched in days. One that does not produces a shortlist that looks reasonable and converts badly.

If you cannot answer some of these yet, that is normal and it is still worth writing down which ones are open. An unknown that is named can be worked around. An unknown that is papered over in a job specification turns into a rejected shortlist and a restart four weeks later.

What the US market pays for this work

AI Engineering work is counted by the US Bureau of Labor Statistics under Software Developers. That classification is broader than the technology itself, so treat the figures as the shape of the market a AI engineer is hired into rather than as a rate card for the skill. Across the United States the Bureau counts 1,687,890 people in this occupation, with a median annual wage of $135,980.

US annual wages, Software Developers, May 2025
US annual wages, Software Developers, May 2025$135,980Median$82,460$214,67010th pct90th pctMiddle half $105K to $172K

The spread matters more than the midpoint. The 90th percentile is about 2.6 times the 10th, which is a wide band for a single occupation and tells you that the title on its own carries very little pricing information. Two people described as a AI engineer can sit at $82,460 and $214,670 in the same national dataset. When a budget is set from a median without asking which end of that range the work actually needs, the hire that follows is usually the wrong one in one direction or the other.

Related classifications are worth reading alongside it, because teams hiring for AI Engineering frequently end up recruiting against these titles too:

US national wages, May 2025
OccupationEmployed25th percentileMedian75th percentile90th percentile
Software Developers1,687,890$105,210$135,980$171,980$214,670
Data Scientists262,440$85,660$120,230$158,880$199,130
Computer and Information Research Scientists37,200$103,570$140,300$188,700$230,630

Source: BLS Occupational Employment and Wage Statistics, May 2025. Figures cover all US employers and are not FuturByte rates.

These are employer-side wage figures for people on a US payroll. They exclude employer taxes, benefits, recruitment cost and the months a seat sits empty, all of which are real and none of which appear in a salary line. The useful way to read the table is as the cost of the alternative you are comparing against, not as a number to match.

How US metro markets compare for this role

The same job is priced very differently across the country. Ranked by median annual wage for Software Developers, the gap between the highest and lowest of the 28 metro areas covered here is a factor of about 1.7. San Jose sits at the top with a median of $213,110; Pittsburgh sits at the bottom with $124,500. A budget built from a national median will be wrong in both of those markets, in opposite directions.

Median wage for software developers, by US metro area
Median wage for software developers, by US metro areaSan Jose, CA: $213,110San Jose, CASan Jose, CA$213,110San Francisco, CA: $186,640San Francisco, CASan Francisco, CA$186,640Seattle, WA: $167,280Seattle, WASeattle, WA$167,280New York, NY: $166,830New York, NYNew York, NY$166,830Boston, MA: $166,090Boston, MABoston, MA$166,090San Diego, CA: $163,270San Diego, CASan Diego, CA$163,270Los Angeles, CA: $160,920Los Angeles, CALos Angeles, CA$160,920Portland, OR: $156,000Portland, ORPortland, OR$156,000Washington, D.C.: $154,930Washington, D.C.Washington, D.C.$154,930Baltimore, MD: $138,900Baltimore, MDBaltimore, MD$138,900Denver, CO: $137,610Denver, CODenver, CO$137,610Charlotte, NC: $135,920Charlotte, NCCharlotte, NC$135,920Chicago, IL: $134,380Chicago, ILChicago, IL$134,380Austin, TX: $134,120Austin, TXAustin, TX$134,120Dallas-Fort Worth, TX: $133,290Dallas-Fort Worth, TXDallas-Fort Worth, TX$133,290Philadelphia, PA: $133,040Philadelphia, PAPhiladelphia, PA$133,040Atlanta, GA: $132,960Atlanta, GAAtlanta, GA$132,960Raleigh, NC: $132,770Raleigh, NCRaleigh, NC$132,770Miami, FL: $132,650Miami, FLMiami, FL$132,650Phoenix, AZ: $131,750Phoenix, AZPhoenix, AZ$131,750Minneapolis-St. Paul, MN: $130,920Minneapolis-St. Paul, MNMinneapolis-St. Paul, MN$130,920Detroit, MI: $130,760Detroit, MIDetroit, MI$130,760Tampa, FL: $130,450Tampa, FLTampa, FL$130,450Orlando, FL: $129,620Orlando, FLOrlando, FL$129,620Salt Lake City, UT: $129,600Salt Lake City, UTSalt Lake City, UT$129,600Houston, TX: $129,440Houston, TXHouston, TX$129,440Kansas City, MO: $124,990Kansas City, MOKansas City, MO$124,990Pittsburgh, PA: $124,500Pittsburgh, PAPittsburgh, PA$124,500
Software Developers by metro area, May 2025, ranked by median wage
Metro areaEmployedMedian wagevs US medianLocation quotient
San Jose, CA87,350$213,110+57%7.09
San Francisco, CA69,030$186,640+37%2.68
Seattle, WA92,770$167,280+23%4.10
New York, NY121,000$166,830+23%1.17
Boston, MA42,310$166,090+22%1.44
San Diego, CA20,610$163,270+20%1.23
Los Angeles, CA55,540$160,920+18%0.82
Portland, OR18,260$156,000+15%1.39
Washington, D.C.69,060$154,930+14%2.03
Baltimore, MD16,850$138,900+2%1.14
Denver, CO27,010$137,610+1%1.55
Charlotte, NC20,820$135,9200%1.41
Chicago, IL40,370$134,380-1%0.82
Austin, TX31,960$134,120-1%2.28
Dallas-Fort Worth, TX67,030$133,290-2%1.52
Philadelphia, PA28,480$133,040-2%0.91
Atlanta, GA36,300$132,960-2%1.16
Raleigh, NC12,580$132,770-2%1.56
Miami, FL18,900$132,650-2%0.62
Phoenix, AZ29,380$131,750-3%1.14
Minneapolis-St. Paul, MN27,410$130,920-4%1.29
Detroit, MI24,870$130,760-4%1.20
Tampa, FL14,230$130,450-4%0.91
Orlando, FL13,440$129,620-5%0.88
Salt Lake City, UT19,040$129,600-5%2.12
Houston, TX22,940$129,440-5%0.64
Kansas City, MO12,160$124,990-8%1.02
Pittsburgh, PA10,320$124,500-8%0.85

Location quotient compares how concentrated this occupation is in the metro against the national average. A value above 1 means the metro has more of this work than its size would predict.

The location quotient column is the more useful one for hiring. A high median tells you what a role costs; a high quotient tells you whether the people exist. San Jose, San Francisco, Seattle, Washington, D.C., Denver, Austin each have a quotient of 1.5 or above, meaning the work is concentrated there well beyond what the size of the local economy would predict. Those are the markets where a search is likely to be quick and competitive at the same time, and where a counter-offer is most likely to take a candidate off the table late in the process.

The opposite case is worth planning for too. In a metro with a low quotient, the total pool is small even when wages look reasonable, so the realistic options are to widen the search radius, accept a longer time to hire, or bring the capability in from outside the local market entirely. That last option is what most teams are weighing when they come to us.

Hiring risks worth naming

Every one of these has produced a bad hire somewhere. They are written down so that the process tests for them deliberately rather than discovering them in month three.

Demonstration experience presented as production experience. The central risk in this market. Ask what they shipped, who used it, and how they knew it worked.

No evaluation discipline. Ask how they would know a prompt change improved things. Weak answers here predict a system nobody can improve confidently.

No cost awareness. These systems can be expensive in ways that scale with success. Ask for a cost per request figure.

Security naivety about untrusted content. Prompt injection is real. Ask how they would stop retrieved content from influencing tool use.

Hiring AI Engineering developers by metro area

Wages for this occupation vary more between US metro areas than most budget models assume. Each page below sets out the published employment and wage figures for that market, how it compares with the national picture, and what the local industry mix means for the kind of AI engineer who will be available.

Frequently asked questions

Do we need an AI engineer or can our existing developers do this?

Competent developers can integrate a model API, and that is the easy part. What takes specific experience is evaluation, retrieval quality, handling unreliable output and controlling cost. A sensible arrangement for many teams is one person with real experience setting the patterns and evaluation infrastructure, then the existing team building within them.

How do we know if one of these systems is actually working?

You need an evaluation set: real inputs with expected outputs or quality criteria, scored automatically whenever anything changes. Without it, quality is a matter of impression and every prompt change is a gamble. This is the single practice that most separates teams shipping reliable systems from teams shipping demonstrations.

Why does our assistant give wrong answers about our own documents?

Usually retrieval rather than reasoning. The model can only use what it was given, and if the relevant passage was not retrieved it will answer from general knowledge, confidently. Check what context was actually supplied before changing the prompt. Chunking strategy and hybrid search are typically where the fix lies.

Should we fine-tune a model?

Less often than people assume. Retrieval handles knowledge, and prompting handles most behaviour. Fine-tuning earns its place for consistent formatting, a specific style, or a narrow classification task where you have good labelled examples. It is usually the wrong first move and a reasonable later optimisation.

How do we control the cost?

Measure cost per request first, because most teams cannot answer what a feature costs. After that the levers are routing simpler requests to cheaper models, caching repeated and similar requests, and reducing context size, which is often the largest single contributor. Cost here scales with usage, so success makes it worse rather than better.

Are we locked into one model provider?

Partially, and it is manageable if you plan for it. Keep provider-specific code behind an interface, keep your evaluation set portable, and test periodically against an alternative. Prompts do not transfer perfectly between models, so the realistic position is that switching is a project rather than a configuration change, and knowing its size is the useful thing.

Is prompt injection a real risk?

Yes, and it should be designed for rather than instructed against. Any content the model sees, including retrieved documents and user input, can carry instructions. The defence is architectural: constrain what tools can do, enforce permissions outside the model, and treat model output as untrusted input to anything downstream. Telling the model to ignore instructions is not a control.

Should we build agents?

Only where the sequence of steps genuinely cannot be determined in advance. Where the steps are known, writing the workflow gives you something predictable, debuggable and far cheaper. Open-ended agency is appealing in demonstrations and consistently harder to make reliable in production, and a good engineer will push back on it.