Skip to content
mlmentorship

Free prep resources

Simulate the loop, not a favorite question.

A useful mock removes predictable cues, preserves real round timing, and produces bounded repair tasks. Run one early enough to improve and one only if the first exposed broad uncertainty.

Subscribe to get the printable workbook by email. Linked examples and blank tracking sheets. Print your saved map below for your own routes and dates.

Simulation rules

  1. Use prompts the candidate has not rehearsed verbatim.
  2. The observer must interrupt, clarify, and push back as a real interviewer would.
  3. Score each round before discussing it; do not coach during the round.
  4. Take 5–10 minute breaks, but do not read notes between rounds.
  5. Turn misses into at most three repair tasks. A list of 20 problems is not a diagnosis.

Quick builder

Generate a concrete prompt set

Select a role to draw one core-library prompt for each typical round. Regenerate to change prompts; replace any question the candidate rehearsed recently.

00:00

Do not open the linked answer during the round. Score after the timer or when you choose to stop.

Simulation packet

Applied Scientist

3h 35m including breaks

Prefer an experienced scientist, MLE, product partner, or interviewer.

  1. 01

    ML breadth 45 min

    Choose three unseen breadth questions. Require mechanism, failure mode, and one alternative on each.

  2. 02

    ML system design 45 min

    Design an ML system under explicit user, latency, cost, and delayed-label constraints.

  3. 03

    Project deep-dive 40 min

    Probe ownership, the hardest decision, one failure, evidence, and what the candidate would change.

  4. 04

    Product & experimentation 35 min

    Define a ship decision, primary metric, guardrails, experiment unit, and validity threats.

  5. 05

    Behavioral / leadership 30 min

    Ask one conflict and one prioritization question; push back on at least two claims.

Simulation packet

Research Scientist

4h 10m including breaks

Prefer an active researcher or Research Engineer who can test derivations, experimental validity, and ownership of claims.

  1. 01

    Math and ML depth 45 min

    Use one derivation and two mechanism questions. Change an assumption after the first answer.

  2. 02

    Paper and research defense 45 min

    Defend one central claim, its evidence, the strongest alternative explanation, and one limitation.

  3. 03

    Research work sample 60 min

    Turn an unfamiliar observation into competing hypotheses, discriminating probes, and a short evidence readout.

  4. 04

    Research implementation 40 min

    Implement or debug a small ML mechanism with tests, numerical checks, and complexity analysis.

  5. 05

    Job-talk follow-up 40 min

    Present one contribution briefly, then probe ownership, failed paths, impact, and the next research question.

Simulation packet

Machine Learning Engineer

3h 40m including breaks

Prefer a strong software engineer or MLE who will inspect correctness and operational detail.

  1. 01

    ML implementation 45 min

    Use an ML primitive, evaluation, or debugging problem. Require examples, tests, complexity, and working code.

  2. 02

    ML system design 45 min

    Design data, training, serving, monitoring, rollback, and iteration for a product ML system.

  3. 03

    Systems / production 45 min

    Diagnose a training or serving bottleneck from incomplete symptoms; require a measurement plan.

  4. 04

    ML breadth 35 min

    Use two mechanism questions plus one model-choice trade-off.

  5. 05

    Behavioral / project 30 min

    Probe a production failure, ownership boundaries, and influence during incident recovery.

Simulation packet

Research Engineer

3h 50m including breaks

Prefer an RE, systems engineer, or researcher who can test both implementation and evidence quality.

  1. 01

    ML implementation 45 min

    Implement an ML primitive or debugging task with tests and complexity analysis.

  2. 02

    Training / inference systems 45 min

    Quantify a workload, identify the bottleneck, select parallelism, and design recovery.

  3. 03

    Research depth 45 min

    Critique a claim, identify alternative explanations, and design discriminating ablations.

  4. 04

    ML breadth 35 min

    Test mechanisms that connect algorithms to training or systems behavior.

  5. 05

    Project deep-dive 40 min

    Probe research-to-production translation, a failed experiment, and the candidate’s personal decisions.

Simulation packet

Staff / Principal ML level overlay

4h 15m including breaks

Use an experienced staff or principal IC who can challenge both technical mechanisms and organization-level claims.

  1. 01

    Multi-team architecture 60 min

    Use the multi-team ML platform case. Change one constraint after the architecture, then probe state, failure, migration, ownership, and adoption.

  2. 02

    Technical strategy and portfolio 45 min

    Give three credible investments under one constraint. Require an order, opportunity cost, quarterly checkpoints, and evidence that would stop or reverse the strategy.

  3. 03

    High-scope project deep-dive 45 min

    Probe problem selection, personal authority, cross-team consequences, one wrong bet, operating results, and who carried the work afterward.

  4. 04

    Influence and recovery 35 min

    Use one disagreement and one stopped or redirected project. Test incentives, evidence, cost of delay, and whether the candidate changed their own view.

  5. 05

    Retained domain depth 40 min

    Choose one mechanism from the candidate’s primary role. Move from strategy to architecture to an algorithm, invariant, experiment, or failure trace.

Simulation packet

Senior Principal / Distinguished ML level overlay

4h 40m including breaks

Use a principal, distinguished engineer, research leader, or equivalent upper IC who can challenge delegated authority and technical depth.

  1. 01

    Enterprise agent architecture 60 min

    Use the enterprise agent-platform case. Probe authority, unknown side effects, memory, regional policy, evaluation, migration, and one deep technical invariant.

  2. 02

    Multi-portfolio technical strategy 55 min

    Balance product, platform, model, security, migration, and retirement investments. Require opportunity cost, checkpoints, and reversal evidence.

  3. 03

    Delegated technical leadership 45 min

    Define principal-owned domains, decision rights, interface review, incident authority, and escalation without centralizing every choice.

  4. 04

    Durability under external change 45 min

    Change a vendor, regulation, region, or organization assumption. Test doctrine, portability, succession, and which decisions reopen.

  5. 05

    Retained domain depth 45 min

    Move from company-level direction to one algorithm, experiment, distributed invariant, security boundary, or performance model.

Simulation packet

Upper-IC reasoning-systems transfer

4h 25m including breaks

Use a research or systems leader who can challenge compute arithmetic, evaluation support, serving constraints, and portfolio claims.

  1. 01

    Reasoning under fixed compute 60 min

    Allocate fixed capacity across training, post-training, evaluation, and serving. Require arithmetic, routing, stopping rules, and graceful degradation.

  2. 02

    Verifier and evaluation strategy 50 min

    Probe verifier support, false acceptance, correlated errors, severe tails, held-out families, and evidence that changes the allocation.

  3. 03

    Serving and recovery 45 min

    Change request mix, KV pressure, or provider availability. Require admission control, fallback, state preservation, and recovery.

  4. 04

    Portfolio and delegated ownership 45 min

    Partition model, verifier, serving, evaluation, and product decisions across technical leaders with explicit review checkpoints.

  5. 05

    Project evidence 35 min

    Test one real capacity, quality, or research decision for ownership, failure, update, and measured outcome.

Simulation packet

Upper-IC ecosystem-ranking transfer

4h 15m including breaks

Use a ranking, marketplace, experimentation, or product leader who can challenge causal claims and participant trade-offs.

  1. 01

    Short-form video ecosystem design 60 min

    Balance viewer value, creator opportunity, quality, revenue, and concentration without hiding policy in one weighted score.

  2. 02

    Causal evidence and feedback loops 50 min

    Probe exposure bias, interference, exploration, cohort maturity, long horizons, and what cannot be learned from one experiment.

  3. 03

    Incident and rollback 45 min

    A watch-time win damages creator retention. Require containment, rollback, diagnosis, compensation, and a safe next experiment.

  4. 04

    Governance and portfolio 45 min

    Assign metric authority and decision rights across ranking, product, integrity, creator, and revenue leaders. Define reversal evidence.

  5. 05

    Retained ranking depth 35 min

    Probe one objective, calibration, counterfactual estimator, exploration policy, or serving constraint below the strategy.

Frontier format overlays

These are not company packets. They model public, current work-sample formats and should be used only when recruiting confirms the corresponding round.

Format-specific packet

Agentic engineering loop

Use only when the confirmed loop includes an authorized coding agent, an existing codebase, or a technical presentation.
  1. Agentic ML codebase60 min

    Run the evaluation-service lab. Score codebase mapping, bounded delegation, diff review, tests, and explanation.

  2. LLM inference design45 min

    Design admission, batching, KV capacity, fairness, overload, and quality-preserving degradation.

  3. Technical presentation45 min

    Present for 30 minutes and reserve 15 for interrupted questions on ownership, evidence, and counterfactuals.

  4. Collaboration and values30 min

    Probe one cross-functional conflict and one decision with a real safety, quality, or user-cost tension.

Format-specific packet

Research work-sample loop

Use only when the confirmed loop includes a notebook, API investigation, research brainstorm, paper discussion, or experiment review.
  1. Black-box investigation90 min

    Produce competing hypotheses, discriminating probes, a boundary condition, and a five-minute readout.

  2. Research depth45 min

    Critique one claim and design an ablation that separates it from the strongest alternative.

  3. Math oral30 min

    Draw two derivations and change one assumption in follow-up.

  4. Technical project defense45 min

    Defend one research-to-system project with failures, evidence, ownership, and transfer.

Format-specific packet

Performance and training-systems loop

Use only for roles that name accelerators, kernels, distributed training, inference performance, or ML infrastructure.
  1. Accelerator work sample120 min

    Use the public challenge and experiment ledger. Require a trace-based hypothesis, measurement, and unchanged correctness tests.

  2. Fault-tolerant training design45 min

    Classify failures, define state consistency, and choose recovery scope.

  3. Frontier training incident45 min

    Diagnose a loss spike from rank-local evidence and preserve experimental validity.

  4. Project deep-dive40 min

    Probe one bottleneck, one failed optimization, and the candidate’s instrumentation choices.

Unfamiliar transfer prompt pool

Use one prompt the candidate has not seen. The goal is transfer of framing, evaluation, and trade-off patterns, not recall of a memorized architecture.

  1. A company has historical project documents and wants to predict which future ML projects will succeed. Scope the problem before proposing a model.
  2. Your model improves the offline metric by 8% but hurts the online primary metric by 2%. Decide what to do next.
  3. Inference quality is fixed; cut annual serving cost by 60% without violating the p95 latency SLO.
  4. A new policy forbids storing raw labels after 24 hours. Redesign training, evaluation, and auditability.
  5. Two teams report opposite experiment results for the same model. Design the investigation.
  6. The strongest paper baseline is missing and the reported gain is within seed variance. Decide what evidence you need.
  7. Eight ML teams use incompatible training and release systems. Choose the first shared contract, migration sequence, ownership model, and stop condition.
  8. Forty teams use agents with incompatible tools and authority. Design the common contracts, delegated ownership, migration, and evidence that can reverse centralization.
  9. A reasoning-model program has a fixed cluster budget. Allocate training, verification, evaluation, and serving capacity, then define evidence that changes the allocation.
  10. A short-form video ranker raises watch time while creator retention falls. Contain the harm, investigate causality, and redesign objectives and decision rights.

Round scorecard

Score 0–2 before giving feedback: 0 = missing or unsafe, 1 = workable but inconsistent, 2 = strong under follow-up.

Go / risk / extend decision

Green

No critical dimension scores 0; every round has a workable structure; follow-ups reveal real depth; remaining misses are bounded.

Amber

One critical round is inconsistent but responds to targeted repair. Keep the loop only if there is time for a spaced retry before it.

Red

More than one critical round scores 0, ML implementation or design cannot reach a baseline in time, or stories lose ownership under follow-up. Extend or move the loop when possible.

After the simulation

  1. Write the failure as an observable behavior, not “study systems more.”
  2. Choose one corrective drill and a due date.
  3. Retry closed-book after spacing.
  4. Graduate only after the repair survives a mixed session.

Open the Workbook →