A live 45-minute round on one of the questions reported in NVIDIA interviews, run by an AI interviewer that asks one thing at a time, probes what you say, and weighs what that round is reported to weigh. It ends with a debrief: a score, what was strong, what was missing.
Serve a large language model behind an API: queue the requests, batch them onto scarce GPUs, stream tokens back, and keep every accelerator busy without blowing the latency budget.
You will not see the requirements, the numbers or the diagram before you start — that is the point.
45–60 minutes, often close to accelerators: serving, scheduling or data movement for large models.
Quoting hardware names without the arithmetic. The numbers are the answer here.
From public interview guides and candidate reports, not from NVIDIA. The interviewer is an AI running in that style; it is practice, not a prediction.
Other companies: Anthropic · Google · Meta · Amazon · OpenAI · Microsoft · Uber · Stripe