Before choosing a path
Company titles are inconsistent. Ask for round formats, implementation style, project expectations, and the day-to-day research/modeling/engineering/product mix. If the recruiter cannot provide detail, use the typical rounds in the readiness check and mark uncertainty as risk.
Choose one role
Applied Scientist
Preserve scientific depth, but prove scoping, experimentation, shipping, and influence.
Open 9-step path →Research Scientist
Prove mathematical depth, executable experiments, research taste, and ownership of the central claims.
Open 9-step path →Machine Learning Engineer
Do not substitute ML reading for software execution. Code, debug, and design reliable systems repeatedly.
Open 9-step path →Research Engineer
Balance implementation and systems depth with scientific skepticism. Confirm whether the title is research-heavy or software-heavy.
Open 9-step path →Frontier format overlaysAdd one only when recruiting confirms the format.
Agentic codebase
Use when: The recruiter names an existing repository or authorized coding agent.
Replace one implementation repetition with the agentic evaluation-service lab. Practice mapping, bounded delegation, diff review, and proof.
Open practice →Technical presentation
Use when: The loop includes a project presentation, job talk, or artifact defense.
Prepare 30 minutes around three decisions and reserve 15 minutes for interrupted questions.
Open practice →Research work sample
Use when: The assessment uses a notebook, API, paper, model, or open-ended investigation.
Run the black-box lab and produce a five-minute claim, evidence, uncertainty, and next experiment readout.
Open practice →Values and mission
Use when: The loop includes a dedicated values, safety, mission, or senior-leadership discussion.
Prepare four real tensions and practice independent judgment under principled pushback.
Open practice →Performance trace
Use when: The role centers on kernels, accelerators, training throughput, inference, or systems performance.
Use the released accelerator challenge or inference scheduler and keep an experiment ledger.
Open practice →Domain supplementsAdd only the technical subject used in the actual role.
A domain supplement changes examples and depth. It does not replace implementation, project evidence, behavioral judgment, or confirmed rounds.
LLM and agents
- Reasoning quality under fixed compute
- Coding-product context and evaluation
- Agent authority and runtime safety
- RAG, inference, and tool-use evaluation
Reasoning under budget · AI coding product · Enterprise agent platform · Agent safety control plane · Evaluate an agent
Post-training and environments
- Preference data and reward models
- RL environments and evidence-bearing graders
- Synthetic data and held-out task families
- Verifiable rewards and reward hacking
Post-training design · Preference data · RLVR and GRPO · Synthetic data
Alignment and model behavior
- Scalable oversight and control
- Mechanistic interpretability
- Evaluation validity and contamination
- Adaptive red-team and LLM security
Scalable oversight · Evaluation validity · LLM-as-judge · Red-team design
Multimodal and robotics
- Contrastive alignment and fusion
- Live audio, video, screen, and tool timing
- Grounding and modality-conflict evaluation
- Observation, action, safety, and sim-to-real
Real-time multimodal assistant · Contrastive learning · Multi-task learning · Multimodal models · Robotics policies
Recommendations and search
- Candidate generation and ranking objectives
- Position bias and counterfactual evaluation
- Multi-task learning and feedback loops
- Creator cold start and ecosystem health
Short-form video ecosystem · Ecosystem strategy mock · Ranking losses · Counterfactual ranking · Evaluate a ranker
ML platform and infrastructure
- Foundation-model data provenance and mixtures
- Training and serving workload models
- Distributed communication and memory
- Migration, reliability, adoption, and cost ownership
Foundation-model data platform · Multi-team platform case · ML lineage · Parallelism selection · Distributed profiling
General product ML
- Problem framing and baseline
- Causal identification and experiment design
- Cost-aware thresholds and abstention
- Delayed and selective labels
Causal inference · Delayed labels · Decision thresholds · Choose product metrics · Design an A/B test
Research and modeling
- Moments and distribution choice
- Paired resampling and uncertainty
- Matched baselines and compute
- Reproducible ablations
Statistical moments · Bootstrap · Hypothesis testing · Design an ablation
Specialist-domain boundary
The core library is deepest in LLMs, recommendations, general ML, training, and systems. For specialized CV, speech, or RL loops, use the concept collections below for foundations, but obtain a domain-specific paper/project mock from someone current in that field.
What the core library does not yet cover deeply
This site does not attempt to replace general algorithms or SQL interview preparation. The implementation questions stay here only when the task is fundamentally about ML: attention, training-loop diagnosis, memory-bounded retrieval, or mergeable evaluation metrics. Use external, current material for general data structures, framework internals, streaming-data infrastructure, privacy/compliance implementation, advanced causal inference, or a specialist research frontier.
How to use a path
Attempt each linked question before reading, diagnose with its rubric, and schedule retries. Run the simulation only after every expected round has a baseline.