Skip to content

Free interview kit

DevOps / Site Reliability Engineer interview kit: questions and rubric

For hiring teams interviewing a mid-level DevOps engineer or site reliability engineer, roughly three to seven years in, who runs cloud infrastructure, delivery pipelines and on-call for product teams. It covers a short screen, five competencies with an incident role-play, an infrastructure code review and behavioral questions, anchored ratings, the questions never to ask, and a rule for deciding.

5 competencies · 11 questions · anchored 1-to-5 scale · Updated · Written by the Itya team

How to use this kit

  1. Agree the competencies with the hiring manager before the job is posted.
  2. Give each interviewer one or two competencies, so nothing is asked twice.
  3. Ask every candidate the same core questions; follow up freely.
  4. Rate each competency against the anchors straight after the interview, alone.
  5. Debrief on quotes, not adjectives, then decide with the rule at the end.

First-round screen

Fifteen to twenty minutes, by phone, video or a disclosed AI screen. The aim is to confirm the basics before the panel spends its time.

  1. What draws you to infrastructure and reliability work, and which part of it do you want more of in your next job?
  2. Walk me through the platform you run today: what runs where, how code reaches production, and what pages you at night.
  3. Which of the tools named in the job description (the cloud provider, the infrastructure-as-code tool, the container platform, the CI system) have you run in production, and at what scale?
  4. The job description sets out the location, working pattern, on-call rota and start date. Do those work for you, what is your notice period, and what compensation range are you expecting?

Competency 1

Incident response

Incidents show the habits that matter most in this role: mitigating before diagnosing, keeping people informed under pressure, and learning without blame.

  1. Role-play, 25 minutes: you are on call. At 14:10 an alert fires: checkout errors are at 18% and rising. A config change and an application deploy both went out in the last hour. I play the system and your colleagues; ask for any dashboard, log or person. Walk me through what you do, minute by minute. (Script it in advance: what each dashboard shows, and the cause, such as a config change that shrank the database connection pool.)

    Follow up with

    • Who do you tell, when, and what does your first update say?
    • Rolling back the deploy did not help. What next?
    • What goes into the incident review, and what stays out?
  2. Tell me about the worst incident you have been part of. What was your role, and what changed afterwards?

    Follow up with

    • What did the timeline show that people did not expect?
    • Which follow-up actions actually got done?
RatingWhat it looks like
5 · StrongDeclares and coordinates early, tries the least risky mitigation first, updates people on a steady rhythm, separates mitigation from root cause, and drives a blameless review to completed actions.
3 · MixedMitigates sensibly and communicates when prompted. Finds the cause, but mixes diagnosis with mitigation or needs a nudge to roll back.
1 · WeakWorks alone and silently, changes several things at once, and treats the incident as over when the alert clears. Looks for someone to blame.

Competency 2

Infrastructure as code and automation

Infrastructure that lives only in a console cannot be reviewed, repeated or rebuilt after a disaster. Engineers who manage it as code make every change visible and reversible.

  1. Work sample, 15 minutes: here is a short Terraform or OpenTofu change (use your own tool if different). It renames a database resource, opens a security group to 0.0.0.0/0 on the database port, and hard-codes a cloud access key. Review it as you would for a teammate.

    Follow up with

    • What would the plan show for the renamed database, and why does that matter?
    • Where should the key live instead?
    • How should this change reach production?
  2. Tell me about something you automated that people used to do by hand. How did you decide it was worth it?

    Follow up with

    • What did it break the first time?
    • Who maintains it now?
  3. What would you do if production no longer matches the infrastructure code, because people make console changes during incidents?

    Follow up with

    • How would you find all the drift?
    • What would you allow in an emergency?
RatingWhat it looks like
5 · StrongReads every plan for destructive changes, keeps secrets and broad access out of code, applies only through a reviewed pipeline, and handles drift with a clear path back to code.
3 · MixedWrites and reviews infrastructure code competently and keeps secrets out of it. Needs prompting on state, destructive plans or drift.
1 · WeakMisses that the change replaces the database, and treats console changes and keys in code as normal.

Competency 3

CI/CD and release safety

The pipeline decides how fast and how safely every team ships. A good one makes the safe path the easy path: quick feedback, gradual rollouts and a fast way back.

  1. Our pipeline takes 45 minutes from merge to production, and a bad release last month took two hours to roll back. You have one quarter. What would you look at and change?

    Follow up with

    • What would you measure first?
    • How would you make rollback fast enough that people use it without fear?
    • What would you leave alone?
  2. Tell me about a release that went wrong in a pipeline you owned. What did the pipeline do, and what did you change?

    Follow up with

    • Why did no check catch it?
    • How long did it take to get back to safe?
RatingWhat it looks like
5 · StrongMeasures the pipeline, makes the safe path fast, ships gradually with automatic checks, keeps rollback one tested step, and fixes the kind of failure after each bad release.
3 · MixedBuilds and maintains working pipelines with tests and a rollback path. Needs prompting on progressive rollouts, migration safety or measuring where time goes.
1 · WeakTreats the pipeline as a black box, cannot describe how a rollback works, and fixes failures with manual gates.

Competency 4

Observability and alerting

An on-call engineer is woken by whatever someone chose to measure. Alerts tied to what users feel, with dashboards and traces that answer the next question, decide how fast anyone can respond.

  1. A service has 60 alerts, mostly on CPU, memory and disk, and the on-call engineer is paged about ten times a week, usually for nothing. What would you change?

    Follow up with

    • What would you page on instead?
    • How would you set the thresholds?
    • How do you get the team to agree to delete alerts?
  2. Tell me about a time your monitoring missed a problem that users noticed first. What did you change?

    Follow up with

    • What signal would have caught it?
    • How did you check the new alert fires?
RatingWhat it looks like
5 · StrongPages only on user-facing symptoms tied to SLOs, keeps the rest as dashboards or tickets, closes blind spots after every miss, and tests that alerts fire.
3 · MixedUses metrics, logs and traces competently and knows what an SLO is. Needs prompting to move paging to user-facing symptoms or to test alerts.
1 · WeakAlerts on every resource metric, cannot explain an SLO, and accepts constant noisy pages as normal.

Competency 5

Cost and security

Platform engineers hold the keys and the cloud bill. Least-privilege access, patched images and right-sized resources cost less built in than bolted on after an audit finding or a surprise invoice.

  1. What would you do if the monthly cloud bill rose 40% in two months and finance asked you why?

    Follow up with

    • Where would you look first?
    • Which savings would you not take, and why?
  2. Tell me about a change you made to make access to production safer. What was the risk, and what did you do?

    Follow up with

    • How were secrets stored and rotated afterwards?
    • What would you do in the first hour after a key leaked?
RatingWhat it looks like
5 · StrongBreaks cost down by owner and cause, takes savings that do not hurt reliability, and builds least privilege, short-lived credentials, rotation and pipeline scanning into the platform by default.
3 · MixedFinds the obvious cost drivers and follows least-privilege practice. Needs prompting on ownership, budgets, rotation or the reliability cost of a saving.
1 · WeakCannot explain what drives the cloud bill, and accepts shared admin keys or long-lived secrets as normal.

Never ask

  • Don't ask a candidate's age, or their graduation year to work it out. Why: age says nothing about operations skill, and age-based decisions are unlawful in many places.
  • Don't ask about marital status, children, pregnancy or caregiving duties, even to judge on-call availability. Why: none of it is job-related, and it is put mostly to women. Describe the on-call rota and ask every candidate whether they can work it.
  • Don't ask about religion, caste, community or the origin of a surname, directly or by proxy ('What does your father do?'). Why: in India these questions signal caste and religion, which have no bearing on the job.
  • Don't ask 'Where are you from originally?' or 'What is your native place?'. Why: it invites judgments about region, language and community. If location matters, ask whether the candidate can work from the job's location.
  • Don't ask about health conditions, disabilities or past medical leave, including whether they can handle night pages. Why: the only relevant question is whether the candidate can perform the essential functions of the job, with or without reasonable accommodation.
  • Don't ask about arrests or criminal records in the interview, even though the role holds production access. Why: many US states and cities restrict when employers may ask, and an arrest is not a conviction. If the role needs a background check, run it lawfully and at the stage the law allows.
  • Don't ask for salary history where the law restricts it: many US states and cities do, and the EU Pay Transparency Directive requires member states to. Why: it carries past pay gaps into the new offer. Ask for the expected range; in India, where current pay is commonly asked, still set the offer from the role's band.
  • Don't ask the candidate to describe their current employer's network layout, security controls or unpatched weaknesses. Why: it asks them to breach confidentiality, and a candidate who complies would do the same with yours. Ask about problems and their own decisions in general terms.

Making the decision

Each interviewer rates their competencies alone, against the anchors, before the debrief. The must-haves are incident response, infrastructure as code and automation, and observability and alerting. Any 1 on a must-have is a no, and so is any competency rated 1 by two interviewers. A hire needs an average of 3.5 or higher. The hiring manager decides after the debrief and records the role-play and review evidence behind each rating.

Questions people ask

What is the difference between a DevOps engineer and an SRE interview?
The overlap is large: both own infrastructure, pipelines and production. SRE roles weigh reliability practice more heavily (service level objectives, error budgets, incident command), while many DevOps roles weigh delivery tooling and developer experience. Use one kit for both and move the must-haves to match the job description.
How do you run an incident role-play in an interview?
Write the scenario and the answers in advance: what each dashboard shows, what each colleague says, and the real cause. The interviewer plays the system and answers only what the candidate asks. Use the same script for every candidate, and rate the process (communication, mitigation, hypotheses) more than whether they find the cause in time.
Should DevOps candidates do a hands-on lab or take-home?
A short live review of an infrastructure or pipeline change shows most of what a lab does, without hours of setup. If you use a lab, run it in a disposable sandbox you provide, cap it at an hour or two, and never use your real accounts. Discuss the result in a follow-up interview.
Do cloud or Kubernetes certifications matter when hiring DevOps engineers?
They show study, not production judgment. Treat a certification as a small plus at most, and do not screen candidates out for lacking one. The role-play and the code review in this kit show the skills directly.

A kit is step one.

Itya holds every interviewer to the same rubric, records the interview, and drafts the scorecard with the quotes behind each rating. The AI drafts; your team decides. Free plan, unlimited teammates.