How we assess developers
Work somebody has actually done, and a conversation about the decisions in it. No whiteboards, no trivia, no unpaid weekends.
Most technical hiring fails in the same way: a process that tests something other than the job. Whiteboard algorithm puzzles test recall under pressure. Trivia tests what somebody happened to read. Long unpaid take-home exercises test availability, which selects for people without caring responsibilities rather than for people who are good.
What we do instead is look at work somebody has actually done and talk about the decisions in it. That is harder to run at scale, which is why fewer people do it, and it is the only thing we have found that predicts whether someone will be effective on a real codebase.
What we look at
Code they have written and can talk about
Not a puzzle solution. Something real, with the constraints that shaped it. The interesting part is never the code; it is what they would do differently now and why.
Decisions they regret
The single most informative question in technical interviewing. Strong developers have specific answers with the constraint that produced the decision and the cost that followed. Weak ones either have none or blame the framework.
What they have removed
Mid-level developers describe what they added. Senior ones describe a store they deleted, an abstraction they collapsed, a dependency they took out. The ability to make a codebase smaller is the clearest seniority signal there is.
How they read unfamiliar code
Most of the job is reading rather than writing. We give people something they have not seen and watch what they ask. The questions are more informative than the conclusions.
What they choose not to do
Judgement is mostly refusal: not reaching for a store, not adding a dependency, not automating something that does not need it. A developer who can explain what they decided against is telling you they were deciding.
Whether they can say they do not know
Under interview pressure this is genuinely difficult, and its absence is a real risk. Somebody who invents an answer in an interview will invent one in a design review.
What the levels actually mean
Job titles are not comparable between companies, so it is more useful to describe levels by what somebody can be left to own without supervision. These are the boundaries we use, and every technology page states what each level looks like in that specific technology.
- Junior
- Delivers defined pieces of work inside an existing structure, with the approach decided for them. Reviews catch the things experience teaches. Valuable where there is strong direction and capacity to teach, and a poor investment where there is neither.
- Mid-level
- Owns a feature end to end, including its tests, its failure states and its data access. Can be given a problem rather than a solution. Still benefits from review on decisions that will be expensive to reverse.
- Senior
- Owns the shape of an area. Makes the decisions that are costly to change later and can justify them against the alternatives. Raises the level of the people reviewing alongside them, which is the part most often left out of the definition.
- Staff
- Owns cross-cutting concerns: the contracts between areas, the standards others build against, the migration path off whatever the system has outgrown. Measured by what other teams stop having to decide.
Years of experience correlate with these far less than people expect. Somebody who has spent six years maintaining one application carries less transferable judgement than somebody who has spent three across three very different ones. When a CV and a level disagree, the interview should settle it.
The most reliable seniority signal we know is what somebody has removed. Mid-level developers describe what they added. Senior ones describe a store they deleted, an abstraction they collapsed, a dependency they took out. The ability to make a system smaller is harder to fake than any amount of technical vocabulary.
What we do not do
- Whiteboard algorithm puzzles, which test recall under artificial pressure and correlate poorly with the work.
- Unpaid take-home exercises longer than about two hours, which filter for availability rather than ability.
- Trivia about framework APIs, which is what documentation is for.
- Certification counting, which tells you about familiarity rather than judgement.
- Years-of-experience arithmetic, since six years on one application teaches less than three across three.
What we tell you
Every shortlist comes with our reasoning: what we think each person is strong at, where we have reservations, and why we ruled others out. The reservations are the part that matters. A shortlist with no stated weaknesses has not been assessed; it has been assembled. If we cannot articulate a reservation about somebody, we have not looked hard enough.
We also tell you when we are unsure. A candidate we think is probably strong but have not been able to verify in one area is presented that way, with the specific thing we could not confirm, so your interview can test it rather than repeating what we already did.
Frequently asked questions
Why not use algorithm interviews?
Because they test recall under artificial pressure and correlate poorly with performance on a real codebase. They persist because they are easy to standardise and compare, which is a property of the process rather than of the signal. The main thing they reliably select for is recent practice at that specific format.
Do you give take-home exercises?
Only short ones, and we pay for them. Long unpaid exercises filter for people with spare evenings, which is a filter on circumstances rather than on ability. A short live session on real code gives more signal per hour anyway.
How do you assess someone in a technology you do not specialise in?
The questions that discriminate are mostly not technology-specific: state placement, failure handling, what they removed, what they decided against. Where the technology genuinely matters, each of our technology pages sets out what we test and why, and we bring in someone who works in it when we are not confident.
Can we see your assessment before interviewing?
Yes. It comes with the shortlist, including reservations and what we could not verify, so your interview tests the open questions rather than repeating ours.
What if we disagree with your assessment?
Tell us. Either we missed something, in which case we want to know, or the brief and the need have drifted apart, which is worth finding out. Both outcomes are more useful than us defending a shortlist.
Do you test for communication?
Continuously, because the assessment is a conversation rather than a test. Whether someone can explain a decision to a person who was not there is not a soft skill in distributed work; it is most of the job.