Gets a baseline in first
Before any model work, something dumb and cheap that answers the question badly. Without it nobody can say whether the clever version is an improvement.
An engineer who can take a model from a notebook to something a customer depends on. The demo is the easy half. The work is the retrieval, the evaluation set, the fallback behaviour and the cost ceiling that make it safe to leave running.
Very little of it is model selection. Most of it is the scaffolding that decides whether the model is allowed near a customer.
Before any model work, something dumb and cheap that answers the question badly. Without it nobody can say whether the clever version is an improvement.
A few hundred real cases with the right answer written down by somebody who knows the domain. Graded, not eyeballed.
Most quality problems in a language model feature are retrieval problems wearing a prompt costume. Chunking, embedding choice, reranking.
What the product does when the model is unsure, wrong, slow or down. Decided up front and written into the code.
Token cost per request, tracked from day one, because the first version of anything is ten times more expensive than it needs to be.
Nobody on the bench knows all of this. We shortlist against what your problem actually needs, and tell you where the gaps are.
No algorithm puzzles. Every exercise below is a task this role does in a normal week, and we watch the method more than the answer.
We give them a messy document set and a question their chunking will get wrong. We are watching whether they diagnose retrieval or start rewriting the prompt.
Given twenty example interactions, produce something that can grade a change. The output matters less than whether they ask what "correct" means before starting.
Rough numbers on the back of an envelope for a feature at ten thousand requests a day. We are looking for an engineer who thinks about the bill without being asked.
Ten minutes explaining why a model got something wrong, to somebody in operations. Half of this job is this conversation.
You interview every candidate we put forward, so these are yours to use. They separate somebody who has run this in production from somebody who interviews well.
Describes a regression suite that runs against a fixed set of cases, with a pass rate they can compare across versions. Mentions that the set has to be built before the upgrade, not after something breaks.
Says they would test it, or that they would watch for user complaints. Both mean there is no measurement and the first sign of trouble will be a customer.
Looks at what was retrieved before looking at what was generated. Wants to see the actual chunks that went into the context for a handful of failures.
Reaches straight for prompt wording or a bigger model. Both can help, neither finds the cause, and one of them triples the bill.
Has an answer already, because they have thought about it as a product decision: degrade to a cached response, queue it, or tell the user plainly.
Has not considered it. This is the difference between a demo and something you can leave running over a holiday weekend.
Has a real example and is comfortable saying so. A rules engine, a database query or a better form would often have been cheaper and more reliable.
Cannot think of one. An engineer who has never talked a client out of using a model will not talk you out of it either.
Almost never, and you should be suspicious of anyone who suggests it for a normal business problem. Fine-tuning has a place. Training a foundation model does not, unless you have a research budget and a reason nobody else can serve.
Yes. Local or self-hosted inference is a normal request and we staff for it. It costs more in engineering time and usually less in tokens, and the trade is worth making explicitly rather than by default.
Then the first two weeks are finding out, and that is a legitimate use of the time. What we will not do is build a feature nobody can name a metric for, because that is how the 95% of pilots that produce nothing get built.
No, and the distinction matters. A data scientist answers questions about your data. This person ships software that uses a model and stays up. Some people do both well; most are clearly better at one.
Thirty minutes, and a shortlist within a day. If a ai / ml engineer is the wrong hire for what you described, you will hear that instead.
Either one reaches Umer directly. No forms sitting in a queue.