Training Data & Labels
Models are downstream of data. In production ML work, collecting, labeling, and cleaning data consumes more engineering time than modeling, and interviews mirror that reality: the data discussion is where practical experience is easiest to detect and hardest to fake. This page covers the three buckets of training data, how labels actually get made, and the data pathologies (imbalance, delay, bias, cold start) that interviewers use as probes.