Hire a pro

DevOps Engineer

An engineer whose job is making deployment boring. Pipelines, infrastructure as code, observability that answers questions, and a rollback that somebody has actually rehearsed rather than written down.

Shortlist
Within 24 hours
Deployed
Inside 2 weeks
Seniority
5+ years, on-call experience
Rate
Flat monthly
Wrong fit
Replaced, not billed twice
What they do

The measure is how uneventful a Friday deploy is.

Good platform work is invisible. You notice it only in what stops happening: the manual step, the deploy freeze, the person who is the only one who can release.

The loop should be short enough that a deploy is not an event. If it is an event, it happens rarely, and rare deploys are large and dangerous.
Day to day

Removes manual steps

Every step a human does before a release is a step that gets skipped the one time it mattered.

Day to day

Writes infrastructure down

Terraform or equivalent, in version control, so that the environment can be rebuilt rather than remembered.

Day to day

Makes observability answer questions

Dashboards nobody looks at are decoration. The test is whether you can find out what broke in under five minutes.

Day to day

Practises the rollback

Not documents it. Runs it, on a normal Tuesday, so the first attempt is not during an outage.

Day to day

Manages cost as a first-class concern

Cloud bills grow by default. Somebody has to own the number, and it should not be the finance team finding out in arrears.

What they know

The tools, grouped by what they are for.

Nobody on the bench knows all of this. We shortlist against what your problem actually needs, and tell you where the gaps are.

Cloud
AWSAzureGCPHetznerBare metal
Infrastructure
TerraformPulumiAnsibleCloud-init
Containers
DockerKubernetesECSNomadHelm
Pipelines
GitHub ActionsGitLab CIArgo CDBlue-greenCanary
Observability
PrometheusGrafanaLokiOpenTelemetryPagerDuty
How we vet

Four exercises, all of them from real work.

No algorithm puzzles. Every exercise below is a task this role does in a normal week, and we watch the method more than the answer.

Break a pipeline and fix it

A real failing build with an unhelpful error. We are watching the debugging method rather than the fix.

Rebuild an environment from code

Given a Terraform state and a description, stand up a working environment. Missing pieces are the point of the exercise.

Incident walkthrough

A described outage with the real signals. Diagnose it out loud. This is where experience shows immediately.

Candidates without on-call experience stall here.

Cost review

An itemised cloud bill with obvious waste in it. Find it and say what you would change first.

Interview signals

Four questions for when you interview them yourself.

You interview every candidate we put forward, so these are yours to use. They separate somebody who has run this in production from somebody who interviews well.

Ask: what happens when a deploy goes wrong at 2am?
A strong answer

Describes a rollback they have practised, an alert that fires on the right signal, and a runbook somebody other than them can follow.

A worrying answer

Describes a plan that exists only in documentation. The first real execution of an untested rollback is not the moment you want to discover its gaps.

Ask: do we need Kubernetes?
A strong answer

Asks how many services, how many engineers and what the traffic looks like before answering. Often says no, and can explain what it would cost in operational overhead.

A worrying answer

Yes, always. Kubernetes is excellent and it is also a full-time job. For four services and three engineers it is usually a tax rather than a tool.

Ask: how do you know a deploy was healthy?
A strong answer

Names the specific signals watched afterwards: error rate, latency percentile, and a business metric that would move if something subtle broke.

A worrying answer

Says the pipeline went green. That tells you it deployed, not that it works.

Ask: where are your secrets?
A strong answer

In a secret manager, injected at runtime, rotated, with a story for how access is revoked when somebody leaves.

A worrying answer

In environment variables in the CI settings, or worse in the repository. It works right up until it is the headline.

Honestly

When this is the right hire, and when it is not.

Ask for this when

Good fit

  • Deploys are manual, rare, or depend on one specific person.
  • An outage took hours to diagnose because nothing was instrumented.
  • The cloud bill is growing and nobody can explain why.
Ask for something else when

Poor fit

  • You want somebody to babysit servers. That is managed hosting, and it is cheaper.
  • You have one application on one box and no delivery pain. You do not need this yet.
  • You want Kubernetes because it is on the CV. Ask us and we will talk you out of it.
Questions

Before you ask for a shortlist.

They overlap heavily and the emphasis differs. DevOps tends to lead with delivery speed and automation, SRE with reliability targets and error budgets. Tell us which problem you have and we will match the emphasis rather than the title.

For work they built, during a defined engagement, yes, and that is worth arranging explicitly in the contract. What does not work is an external engineer carrying a pager for a system they do not own.

Not at all, and it is often the cheaper right answer. It means slightly more of the platform is yours to run, which is exactly the work this engineer does.

Next step

Describe the problem, not the job title.

Thirty minutes, and a shortlist within a day. If a devops engineer is the wrong hire for what you described, you will hear that instead.