Hire a pro

AI / ML Engineer

An engineer who can take a model from a notebook to something a customer depends on. The demo is the easy half. The work is the retrieval, the evaluation set, the fallback behaviour and the cost ceiling that make it safe to leave running.

Shortlist
Within 24 hours
Deployed
Inside 2 weeks
Seniority
4+ years, 2 in production ML
Rate
Flat monthly
Wrong fit
Replaced, not billed twice
What they do

What this person does between Monday and Friday.

Very little of it is model selection. Most of it is the scaffolding that decides whether the model is allowed near a customer.

The evaluation set outlives every model you will swap through it. That is the thing worth paying an engineer to build.
Day to day

Gets a baseline in first

Before any model work, something dumb and cheap that answers the question badly. Without it nobody can say whether the clever version is an improvement.

Day to day

Builds the evaluation set

A few hundred real cases with the right answer written down by somebody who knows the domain. Graded, not eyeballed.

Day to day

Owns retrieval, not just prompting

Most quality problems in a language model feature are retrieval problems wearing a prompt costume. Chunking, embedding choice, reranking.

Day to day

Defines the failure behaviour

What the product does when the model is unsure, wrong, slow or down. Decided up front and written into the code.

Day to day

Watches the bill

Token cost per request, tracked from day one, because the first version of anything is ten times more expensive than it needs to be.

What they know

The tools, grouped by what they are for.

Nobody on the bench knows all of this. We shortlist against what your problem actually needs, and tell you where the gaps are.

Models
ClaudeGPTOpen weightsLocal inference
Retrieval
Vector searchHybrid searchRerankingChunking strategies
Evaluation
Golden setsLLM-as-judgeRegression runsHuman review
Serving
FastAPIQueuesCachingStreamingRate limits
Classic ML
scikit-learnPyTorchFeature storesDrift monitoring
How we vet

Four exercises, all of them from real work.

No algorithm puzzles. Every exercise below is a task this role does in a normal week, and we watch the method more than the answer.

A real retrieval problem, not a puzzle

We give them a messy document set and a question their chunking will get wrong. We are watching whether they diagnose retrieval or start rewriting the prompt.

Most candidates blame the model.

Build an evaluation set in an hour

Given twenty example interactions, produce something that can grade a change. The output matters less than whether they ask what "correct" means before starting.

A third produce something ungradeable.

Cost and latency arithmetic

Rough numbers on the back of an envelope for a feature at ten thousand requests a day. We are looking for an engineer who thinks about the bill without being asked.

Explain a failure to a non-engineer

Ten minutes explaining why a model got something wrong, to somebody in operations. Half of this job is this conversation.

Interview signals

Four questions for when you interview them yourself.

You interview every candidate we put forward, so these are yours to use. They separate somebody who has run this in production from somebody who interviews well.

Ask: how would you know if this got worse after a model upgrade?
A strong answer

Describes a regression suite that runs against a fixed set of cases, with a pass rate they can compare across versions. Mentions that the set has to be built before the upgrade, not after something breaks.

A worrying answer

Says they would test it, or that they would watch for user complaints. Both mean there is no measurement and the first sign of trouble will be a customer.

Ask: the answers are wrong about a quarter of the time. Where do you look first?
A strong answer

Looks at what was retrieved before looking at what was generated. Wants to see the actual chunks that went into the context for a handful of failures.

A worrying answer

Reaches straight for prompt wording or a bigger model. Both can help, neither finds the cause, and one of them triples the bill.

Ask: what happens when the model is down?
A strong answer

Has an answer already, because they have thought about it as a product decision: degrade to a cached response, queue it, or tell the user plainly.

A worrying answer

Has not considered it. This is the difference between a demo and something you can leave running over a holiday weekend.

Ask: tell me about a time the AI was the wrong tool.
A strong answer

Has a real example and is comfortable saying so. A rules engine, a database query or a better form would often have been cheaper and more reliable.

A worrying answer

Cannot think of one. An engineer who has never talked a client out of using a model will not talk you out of it either.

Honestly

When this is the right hire, and when it is not.

Ask for this when

Good fit

  • You have a model working in a notebook and no idea how to make it safe to ship.
  • Quality is inconsistent and nobody can say by how much.
  • The bill is growing faster than the usage.
  • You need somebody who will tell you when the answer is not AI.
Ask for something else when

Poor fit

  • You want original research or a novel architecture. That is a lab, not us.
  • You want a chatbot on your website this week and nothing more.
  • There is no data yet and no plan to get any.
Questions

Before you ask for a shortlist.

Almost never, and you should be suspicious of anyone who suggests it for a normal business problem. Fine-tuning has a place. Training a foundation model does not, unless you have a research budget and a reason nobody else can serve.

Yes. Local or self-hosted inference is a normal request and we staff for it. It costs more in engineering time and usually less in tokens, and the trade is worth making explicitly rather than by default.

Then the first two weeks are finding out, and that is a legitimate use of the time. What we will not do is build a feature nobody can name a metric for, because that is how the 95% of pilots that produce nothing get built.

No, and the distinction matters. A data scientist answers questions about your data. This person ships software that uses a model and stays up. Some people do both well; most are clearly better at one.

Next step

Describe the problem, not the job title.

Thirty minutes, and a shortlist within a day. If a ai / ml engineer is the wrong hire for what you described, you will hear that instead.