Free interview kit
Data Scientist interview kit: questions and rubric
5 competencies · 11 questions · anchored 1-to-5 scale · Updated · Written by the Itya team
How to use this kit
- Agree the competencies with the hiring manager before the job is posted.
- Give each interviewer one or two competencies, so nothing is asked twice.
- Ask every candidate the same core questions; follow up freely.
- Rate each competency against the anchors straight after the interview, alone.
- Debrief on quotes, not adjectives, then decide with the rule at the end.
First-round screen
Fifteen to twenty minutes, by phone, video or a disclosed AI screen. The aim is to confirm the basics before the panel spends its time.
- What draws you to this role, and which kind of data science work do you want more of: experimentation, modeling, or analysis that shapes product decisions?
- Tell me about a model or experiment you worked on that changed a decision. What was the decision, and what was your part?
- Which of the tools in the job description (Python or R, SQL, an experimentation platform, a notebook or cloud environment) have you used at work, and for what?
- The job description sets out the location, working pattern and start date. Do those work for you, what is your notice period, and what compensation range are you expecting?
Competency 1
Problem framing
Data scientists are often handed a method ('build a model') instead of a problem. The ones who find the decision behind the request, and the simplest approach that serves it, avoid months of work nobody uses.
The head of customer success asks you to 'build a model that predicts which customers will churn'. What do you ask before you touch any data?
Follow up with
- What will the team do differently for a customer the model flags?
- What is the simplest approach you would compare the model against?
- How would you know, three months later, whether the model was worth building?
Tell me about a project where the question you were given was not the one worth answering. How did you find out, and what did you do?
Follow up with
- Who did you check the new framing with?
- What happened to the work already done?
| Rating | What it looks like |
|---|---|
| 5 · Strong | Finds the decision behind every request, defines the target and the success measure up front, compares against a simple baseline, and agrees the framing before building. |
| 3 · Mixed | Clarifies the goal and defines the target on larger requests, but needs prompting to name a baseline or the action that follows a prediction. |
| 1 · Weak | Starts from a method or an algorithm. Does not ask what decision the work serves or how success will be measured. |
Competency 2
Statistics and experimentation
Many product decisions rest on A/B tests, and a badly read test is worse than none because it carries false confidence. A data scientist must design tests that can answer the question and spot the ones that cannot.
Case (20 minutes): here is a one-page summary of a checkout redesign that ran as an A/B test for nine days. Treatment conversion is up 3.1% with p = 0.04. The traffic split is 52/48 although it was set to 50/50. Mobile is up 8%, desktop is down 2%, and average order value is slightly lower. The product manager checked the dashboard daily, stopped the test when it looked good, and wants to ship. What do you tell them?
Follow up with
- What does the 52/48 split tell you, and what would you check first?
- Why does checking daily and stopping early matter?
- Would you ship, ship to mobile only, or rerun? What would you need to see?
Tell me about an experiment you designed. How did you choose the unit of randomization, the primary metric and how long it would run?
Follow up with
- How did you decide the sample size?
- What would you do when a change cannot be randomized at all?
| Rating | What it looks like |
|---|---|
| 5 · Strong | Designs tests that fit the question, checks assignment before reading results, separates planned comparisons from exploration, and turns a messy result into a clear next step. |
| 3 · Mixed | Designs sound standard tests and spots the obvious problems in a result, but needs prompting on sample ratio checks, guardrail metrics or the limits of segment results. |
| 1 · Weak | Reads a p-value as the answer. Misses broken assignment, early stopping and after-the-fact slicing even when prompted. |
Competency 3
Modeling and evaluation
A model is only useful if its evaluation matches how it will be used, and many problems are better served by a rule, a query or a simple regression. Judgment about both matters more than knowing many algorithms.
Work sample (10 minutes): a teammate's fraud model reports 99.4% accuracy on a test set drawn at random from two years of transactions. Fraud is about 0.5% of transactions. Here is the feature list, which includes the account's status at the time the data was exported. Review it as you would before it goes live.
Follow up with
- Which metric would you use instead, and how would you pick the threshold?
- How should the test set be drawn?
- What would you compare the model against?
Tell me about a time you decided not to use machine learning, or used it when a simpler approach would have done.
Follow up with
- What did you compare, and on what measure?
- What would each option have cost to run and maintain?
A product team wants to replace the support-ticket classifier with a prompt to a large language model. How would you decide whether to?
Follow up with
- How would you build the evaluation set?
- What would you compare besides accuracy?
| Rating | What it looks like |
|---|---|
| 5 · Strong | Evaluates a model the way it will be used, catches leakage and imbalance, compares against simple baselines, and recommends the least complex approach that meets the need. |
| 3 · Mixed | Builds sound models with a proper hold-out set, but needs prompting to compare against a baseline, choose a metric tied to the cost of errors, or consider an option without a model. |
| 1 · Weak | Accepts headline metrics, misses leakage and class imbalance, and treats a more complex model as the fix for every problem. |
Competency 4
Data quality and rigor
Models and experiments inherit every flaw in their data: logging changes, wrong labels, samples that leave people out. A data scientist must find these before they reach a decision.
A churn model that worked well at launch has got steadily worse for three months, and nobody has changed its code. What do you investigate?
Follow up with
- How would you tell a pipeline problem from a change in the customers?
- Could the model have changed the outcomes it learned from?
- What would you monitor from now on?
Tell me about a data problem that reached, or nearly reached, a result you shared. How was it found?
Follow up with
- What check would have caught it earlier?
- Who did you tell, and how?
| Rating | What it looks like |
|---|---|
| 5 · Strong | Checks data and labels by routine, asks who is missing from a sample, diagnoses before retraining, monitors inputs and outcomes, and reports problems openly with their impact. |
| 3 · Mixed | Checks data before major work and catches obvious problems, but misses subtler ones such as who is missing from a sample, label delays or upstream changes. |
| 1 · Weak | Trusts the data as delivered. Has no routine checks and fixes problems silently when they surface. |
Competency 5
Communicating uncertainty
Every estimate a data scientist produces has a range, and decision-makers need to know how wide it is without a statistics lesson. Overstating certainty does more damage than an honest 'we do not know yet'.
Role-play (5 minutes): I am the head of sales. Your forecast says next quarter's revenue will most likely be 4.6 million, with an 80% range of 4.2 to 5.0 million. I need one number for the board. What do you tell me?
Follow up with
- Which number should I plan hiring on?
- What would make the forecast change?
Tell me about a time someone treated your result as more certain than it was. What did you do?
Follow up with
- How do you present results now?
| Rating | What it looks like |
|---|---|
| 5 · Strong | Leads with the finding and its range in plain words, ties the uncertainty to the decision, corrects overconfident readings, and says what would change the answer. |
| 3 · Mixed | Gives ranges and caveats accurately, but needs prompting to translate them into what the decision-maker should do. |
| 1 · Weak | States results as certain or buries them in jargon. Decision-makers leave without knowing how far to trust the number. |
Never ask
- Don't ask a candidate's age, or their graduation year to work it out. Why: age says nothing about data science skill, and age-based decisions are unlawful in many places.
- Don't ask about marital status, children, pregnancy or plans to start a family. Why: none of it is job-related, and the question is still put mostly to women, which makes it discriminatory as well as irrelevant.
- Don't ask about religion, caste, community or the origin of a surname, directly or by proxy ('What does your father do?'). Why: in India these questions signal caste and religion, which have no bearing on the job.
- Don't ask 'Where are you from originally?' or 'What is your native place?'. Why: it invites judgments about region and community. If location matters, ask whether the candidate can work from the job's location.
- Don't ask about citizenship, nationality or visa history. Why: the lawful question in most places is whether the candidate is authorized to work in the job's country and whether they will need sponsorship. Ask every candidate the same way.
- Don't ask about health conditions, disabilities or past medical leave. Why: the only relevant question is whether the candidate can perform the essential functions of the job, with or without reasonable accommodation.
- Don't ask for salary history where the law restricts it, as many US states and cities do. Why: it carries past pay gaps into the new offer. Ask for the expected compensation range instead.
- Don't screen on a PhD, a college tier or competition rankings as a proxy for skill. Why: they measure access and spare time more than judgment. The A/B case and the model review show the skill directly.
- Don't set a take-home on your own live data or on a problem your team is working on now. Why: that is unpaid consulting. Use a public or synthetic dataset and cap the time.
Making the decision
The must-haves are problem framing, statistics and experimentation, and modeling and evaluation; for a role that runs no experiments, swap in data quality and rigor before the loop starts. Any 1 on a must-have is a no, and so is any competency rated 1 by two interviewers. A hire needs an average of 3.5 or higher. Each interviewer rates alone before the debrief; the hiring manager decides after it and records the case and work-sample evidence behind each rating.
Questions people ask
- What is the difference between a data scientist and a data analyst interview?
- An analyst interview tests whether someone can answer business questions correctly with SQL and clear reporting. A data scientist interview adds experiment design, modeling and evaluation, and the judgment to say when a model is not needed. Many roles blend the two, so pick competencies from the work the person will do, not the title.
- Should a data scientist interview include live coding?
- A short, realistic exercise is fair: cleaning a small dataset or writing a query in Python, R or SQL, with permission to look things up. Algorithm puzzles mostly reward puzzle practice, not data work. Watch how the candidate checks the result.
- Should we use a take-home assignment for data scientists?
- It can show depth, but it costs the candidate unpaid time and favors people with free evenings. If you use one, provide a public or synthetic dataset, cap it at three hours, say what you will assess, and discuss the choices in a later interview. The 20-minute A/B case gives much of the same evidence live.
- How do we adapt this kit for a senior data scientist or a machine learning engineer?
- For a senior role, widen the framing questions to which problems the team should take on at all, and expect them to set standards for experiments and model reviews. For a machine learning engineer, keep the evaluation and data-quality questions and add deployment, monitoring and system design. Write the changes into the anchors before the loop.
A kit is step one.
Itya holds every interviewer to the same rubric, records the interview, and drafts the scorecard with the quotes behind each rating. The AI drafts; your team decides. Free plan, unlimited teammates.