The staff-level AI system-design round for Qualcomm's Datacenter AI org, worked end to end — the drive script, the napkin math you must produce unprompted, and six full designs (multi-tenant inference service, P/D disaggregation, on-prem RAG appliance, agent control plane, AI gateway, release-gate eval platform) with the pushback drills for each.
At staff level in Qualcomm's Datacenter AI org you are designing on top of the rack, not inside it. Nobody is going to ask you to write an attention kernel. They are going to hand you a rack of AI200-class accelerators and a product requirement, and watch whether you can turn silicon into a service with an SLO, a tenancy model, a cost per million tokens, and a story for what happens at 3am.
Almost every design prompt in this loop collapses to the same three questions:
The one-line bar: IC4 designs the happy path and names the tradeoffs when asked. IC5 proposes the scope unprompted, produces arithmetic out loud, picks two components to go deep on, and names the condition under which their own choice becomes the wrong one.
Read the Qualcomm track first for org context and platform literacy; read the AI system-design framework for the generic eight-stage spine. This page is the staff-level, Qualcomm-shaped version: same spine, harder questions, real numbers.
Then: the coding rounds, solved The 4-week track
The same prompt is graded against a different rubric at each level. Memorize the right-hand column; it is where downlevels happen.
| Dimension | IC4 answer | IC5 (staff) answer | Disqualifier |
|---|---|---|---|
| Scoping | Asks "what scale?" and waits | Proposes the scope: "I'll assume 20 req/s, 4K prompts, 500-token completions, three tenants, 99.5% availability — stop me if that's wrong" | Starts drawing boxes before naming a single requirement |
| Numbers | Estimates when pushed | Estimates unprompted, states assumptions, and identifies the dominant term | "It depends on the hardware" |
| Depth | Even coverage of eight components | Two components at real depth, the rest named and deferred explicitly | Uniform shallow coverage; or one component for 45 minutes |
| Tradeoffs | Names both sides | Names both sides, picks one, and names the observable that would flip the decision | "We could do either" |
| Failure | Mentions retries and monitoring | Walks a specific incident: symptom → metric → root cause → mitigation → the guardrail that prevents recurrence | Happy path only |
| Cost | Mentions it's expensive | Produces $/1M tokens from device-hours and names the two levers that move it most | Ignores cost entirely |
| Org | — | Names what ships in v1 vs v2, what a second team owns, what the migration path is | Designs a system no team could staff |
A useful calibration line from the field: the most common reason a strong candidate gets downleveled from staff is producing a design with senior depth but not staff breadth-and-depth together — one beautiful subsystem, no system.
Forty-five minutes, seven moves. Say the timings out loud; interviewers read it as scope control, not rigidity.
| Min | Move | What you say |
|---|---|---|
| 0–4 | Frame + propose scope | Restate the problem in one sentence. Propose traffic, prompt/completion shape, tenancy, and the availability target. Ask exactly one question: "Is this latency-sensitive interactive traffic, batch, or both?" — the answer changes everything downstream. |
| 4–8 | SLO table | Write TTFT p95, TPOT p95, E2E p95, availability, and goodput (requests/s that meet SLO — not raw throughput). Say: "I'll optimize goodput, not throughput." |
| 8–13 | Napkin math | KV bytes/token → memory per sequence → concurrency per card → cards needed (Little's Law) → $/1M tokens. Out loud. Wrong-but-stated beats right-but-silent. |
| 13–22 | Architecture | Data plane and control plane as separate pictures. Name the request path end to end in one breath, then annotate where each SLO is won or lost. |
| 22–34 | Two deep dives | You choose them: "The two things that decide whether this works are the scheduler and the tenancy model. I'll go deep on both." This is the single highest-signal sentence in the round. |
| 34–40 | Failure modes + ops | Three concrete incidents with detection metric and mitigation. Include one that your own design causes. |
| 40–45 | Cost, cuts, and roadmap | $/1M tokens, the two levers, what ships in v1, what you'd cut if the deadline halved. |
The three sentences that read as staff:
Two for K and V, layers, key/value heads (not query heads — this is the whole point of GQA), head dimension , and bytes per element.
| Model shape | KV/token @ FP16 | 8K ctx | 128K ctx | |||
|---|---|---|---|---|---|---|
| 8B-class (GQA 8) | 32 | 8 | 128 | 128 KiB | 1.0 GiB | 16 GiB |
| 32B-class (GQA 8) | 64 | 8 | 128 | 256 KiB | 2.0 GiB | 32 GiB |
| 70B-class (GQA 8) | 80 | 8 | 128 | 320 KiB | 2.5 GiB | 40 GiB |
| 70B-class, no GQA (MHA 64) | 80 | 64 | 128 | 2.5 MiB | 20 GiB | 320 GiB |
That last row is the one to say out loud: GQA cut the KV cache 8× and that is why long context became affordable at all. Halve it again with FP8 KV; the accuracy cost is usually negligible for KV specifically (it's the activations you have to be careful with).
FP16 ≈ 2 bytes/param. A 70B model is 140 GB in FP16, ~70 GB in FP8, ~53 GB in a 6-bit microscaling format, ~35 GB in INT4. Qualcomm's MXFP6 path compresses weights ~61% and decompresses on the fly to FP16 in the vector engine, overlapping decompression with weight fetch — which matters precisely because decode is bandwidth-bound (§4.4).
Take one AI200-class card at 768 GB, a 70B model at MXFP6 (~55 GB), 8K average context:
usable KV memory = 768 GB − 55 GB (weights) − ~40 GB (activations, fragmentation, headroom)
≈ 670 GB
sequences @ 8K = 670 GB / 2.5 GiB ≈ 250 concurrent sequences on ONE cardNow the punchline, and the most important sentence in this section: that is not your concurrency limit. It's your memory limit. Your real limit is bandwidth (§4.4). Capacity-rich parts are exactly the ones where naive capacity math flatters you.
Every decode step reads all weights once plus the KV of every sequence in the batch:
Two consequences you should state:
Worked, with stated assumptions: 55 GB weights, 2.5 GiB KV/seq, assume 1.2 TB/s effective memory bandwidth.
So 4× the batch bought 2.2× aggregate throughput and cost 44% of per-user speed. That table is the entire TTFT/TPOT-vs-throughput tradeoff, made concrete. Draw it.
Say this: "The AI250's near-memory High Bandwidth Compute architecture targets exactly this term. It's a decode-bottleneck part. On AI200 I'd design for large batches and lean on speculative decoding to buy back per-sequence latency; on AI250 I'd expect to be able to run smaller batches at the same goodput, which relaxes the fairness problem."
70B model, 4K prompt: FLOPs. At an assumed 80 TFLOP/s effective (i.e. a 200 TFLOP/s part at 40% MFU) that's ~7 seconds of TTFT for one request. Nobody's SLO survives that. Which is why the three fixes exist and why you name them in this order:
20 req/s, 40s average session → 800 concurrent sequences → 800 × 2.5 GiB = 2 TB of KV → ~3 cards for KV alone, before you've thought about bandwidth. Then check bandwidth: 800 sequences at your chosen batch size and step time. Whichever number is bigger is your fleet size. Say both numbers and take the max — that's the move.
At \frac{12}{3.96\text{M}} \times 10^6 \approx $3.03$ per million output tokens. Now name the two levers that move it most, in order:
Distant third: quantization. Fourth: model routing (send the easy 60% to an 8B). And the one that actually dominates in agentic products: calls per task, because a 50-step agent turns a cheap per-token price into an expensive per-task bill.
"We have a rack of AI200-class accelerators. Build the service that lets three internal product teams and two external customers hit
/v1/chat/completions."
Scope you propose: 5 tenants, mixed interactive + batch, 8B/32B/70B model catalogue, ~30 req/s peak, prompts 500–32K (bimodal — flag this early), completions ~400 tokens, 99.5% availability, strict tenant isolation on data, soft isolation on capacity.
SLO table (write this, don't say it):
| Class | TTFT p95 | TPOT p95 | E2E p95 | Notes |
|---|---|---|---|---|
| Interactive (chat) | 800 ms | 40 ms | 20 s | The thing users feel |
| Batch (offline) | 60 s | — | 30 min | Throughput-optimized, preemptible |
| Internal eval sweeps | best effort | — | — | Runs on scavenged capacity |
Architecture — data plane:
client
│ OpenAI-compatible HTTP/SSE
▼
[ API gateway ] authn/z · per-tenant token-bucket (RPM + TPM) · request validation
│ idempotency keys · request-id · budget check
▼
[ Router ] model routing · KV-cache-aware worker selection · session affinity
│ admission control (shed, don't queue forever)
▼
[ Engine pool ] vLLM (qaic backend) replicas · continuous batching · paged KV
│ prefix cache (radix/trie over block hashes)
▼
[ KV tier ] on-device HBM/LPDDR → host DRAM → NVMe (offload, cold prefixes)Control plane (separate box, separate on-call): model registry + compiled-artifact store (a model is only "ready" once it has a compiled QPC binary for this hardware — precompilation is not optional on this platform), rollout controller, autoscaler, quota service, eval gate, telemetry pipeline.
Deep dive 1 — the scheduler. This is where the round is won.
The scheduler decides, every step: who gets admitted, who gets preempted, and how much of the token budget goes to prefill vs decode.
max_batched_tokens budget: decode steps cost 1 token each, prefill chunks cost their length. Cap the prefill share so a 200K-token prompt cannot stall 60 decoders. This is the direct answer to the noisy-neighbor question.Deep dive 2 — tenancy. Three levels, and you must be explicit about which you're buying:
| Isolation | Mechanism | Cost |
|---|---|---|
| Data | Tenant ID in every cache key, every index filter, every log. Non-negotiable, always on. | ~free |
| Performance | Fair-share scheduler + per-tenant TPM/RPM quotas + admission control | Some utilization |
| Fault/blast-radius | Separate replica pools (or namespaces) per premium tenant | Big — stranded capacity |
The staff position: buy data isolation always, performance isolation by default, fault isolation only for tenants who pay for it. Then name the leak you just created: the prefix cache is a cross-tenant side channel. If tenant B's request gets a suspiciously fast TTFT because tenant A cached the same prefix, B has learned something about A. Fix: partition the prefix cache by tenant (costs hit rate), or accept sharing only for a whitelisted set of public system prompts. Interviewers love this one because most candidates never see it.
Failure modes to walk:
Retry-After. Say the line: "the queue is where your latency SLO goes to die."What ships first: the gateway with quotas and observability. You cannot debug — or bill — what you can't see, and it's the only component whose value doesn't depend on the rest being finished.
"Should prefill and decode run on the same devices?"
The physics. Prefill is compute-bound: one forward pass over the whole prompt, high arithmetic intensity, wants big compute. Decode is memory-bandwidth-bound: one token at a time, reading the entire weight set per step. Co-locating them means every long prefill injects a latency spike into every active decode stream — you literally cannot tune TTFT and TPOT independently, because one knob (batch composition) controls both.
Disaggregation puts them in separate pools connected by a KV transfer path. Now you tune independently: scale the prefill pool for TTFT, the decode pool for TPOT, and pick the P:D ratio from your prompt:completion ratio. Production systems report large throughput gains from this plus KV-aware routing; the mechanism is well documented and the pattern (SplitWise, DistServe, Mooncake, Dynamo, llm-d) is now standard.
The honest crossover — and this is the staff answer:
| Disaggregate when | Co-locate when |
|---|---|
| Prompts are long relative to completions (RAG, code, doc QA) | Prompts and completions are similar length (chat) |
| You have a fast interconnect for KV transfer | Interconnect is your bottleneck — you'll just move the stall |
| Traffic is high and steady enough to keep both pools busy | Traffic is spiky or low — two pools means two sets of idle capacity |
| You need to hit distinct TTFT and TPOT SLOs | One SLO, and chunked prefill already meets it |
The move: "I'd start co-located with chunked prefill, instrument the TTFT/TPOT correlation, and disaggregate only when I can show that prefill interference is what's breaking TPOT. Disaggregation adds a KV transfer on the critical path and a whole distributed failure mode; I want evidence before I buy that." Then add the condition that flips it: "If the prompt:completion ratio goes above ~10:1 — which it will the moment we ship RAG — I'd expect to disaggregate."
KV-cache-aware routing. Once you have a pool, "least-loaded" routing is wrong: it ignores that one replica already holds 90% of this request's prefix in cache. Route on a score:
score(worker) = w1 · prefix_overlap(request, worker.cache)
− w2 · worker.queue_depth
− w3 · worker.kv_pressurePrefix-aware routing is one of the highest-leverage single changes in a serving stack — reported gains in the 30–60% range in production systems — because it converts a routing decision into skipped prefill FLOPs. The tension to name: cache affinity fights load balancing. Pin too hard and you hotspot; ignore cache and you recompute. Cap it — if a worker's queue exceeds a threshold, fall back to load-based routing regardless of overlap.
KV offload tiering. HBM/on-package → host DRAM → NVMe. Offload turns "evict and recompute" into "fetch," which is a win exactly when transfer is cheaper than recompute — i.e. long prefixes, high reuse. On a 768 GB-per-card part, note that you have far more room before you need this tier than a 192 GB HBM part does. That's a design advantage of the AI200 you should name out loud.
"A bank wants your inference appliance on-prem. 40M documents, per-user ACLs that change hourly, no egress. A leaked chunk is a compliance incident."
This is design a RAG system over 10M docs with three multipliers: air-gap, ACL, and appliance. Don't re-derive chunking from scratch — say "standard hybrid retrieval with reranking, I'll assume 400-token chunks with 15% overlap and semantic boundaries" and spend your time on what's actually different.
Napkin math first: 40M docs × ~8 chunks = 320M chunks. At 1024-dim FP32 that's 320M × 4KB = 1.3 TB of raw vectors — instantly a "this doesn't fit in RAM on one box" conversation. Fixes, in order: int8 scalar quantization (4×, ~1% recall loss), product quantization or binary + rerank (32×, bigger loss), or a disk-backed index (DiskANN-style). Say the number, then pick: "int8 with a full-precision rerank of the top 200. 330 GB fits comfortably."
The three things that are actually different:
(user, doc) and a slower vector index; recompute the user's accessible-set bitmap on permission change, not on query. And re-check at generation time — the retrieved chunk must be authorized again before it enters the prompt, because the index may be seconds stale and "seconds stale" is how a leak happens.Failure mode to walk: a document is deleted from the source system. Its chunks remain in the vector index (orphaned vectors), and it surfaces in an answer three weeks later. This is the single most common real RAG incident. Fix: tombstones + a reconciliation job that diffs source IDs against index IDs on a schedule, plus a hard filter on deleted_at at query time so correctness never depends on the job having run.
"Design the platform that runs long-lived agents for internal teams. Some run for hours."
The reframe that gets you the level: an agent is a distributed workflow with a nondeterministic planner, not a chat request. Everything follows from that.
What must be durable. The agent's state — conversation, plan, tool results, step counter, budget spent — lives in a checkpointed store, not in process memory. A deploy, a crash, or a preemption resumes from the last checkpoint rather than losing an hour of work. This is the mature pattern in 2026: durable-execution runtimes with checkpoint/replay (Temporal-style activities, LangGraph-style checkpointers, Durable Objects), and it's what separates a demo from a platform.
What must be idempotent. Every tool call. If you checkpoint after a tool call and crash during one, replay re-executes it. send_email twice is an incident. Give every tool invocation a caller-generated idempotency key, and make the tool gateway dedupe on it. Say this unprompted — it's the highest-signal detail in the whole design.
The three loops and their guards:
while not done:
step += 1
if step > MAX_STEPS: halt("step budget") # loop guard
if spent > BUDGET: halt("cost budget") # cost guard
if now > DEADLINE: halt("wall-clock guard")
plan = llm(context) # ← nondeterminism lives here
checkpoint(state)
result = tool_gateway.call(plan.tool, plan.args,
idem_key=hash(run_id, step),
policy=tenant_policy) # ← authz lives here
context = compact(context + result) # ← context lives here
checkpoint(state)Three deep-dive candidates; pick two:
Failure mode to walk: an agent enters a two-step loop — call tool, get error, call the same tool identically. Detection: hash (tool, args) per run and alarm on repeats; the step budget is the backstop, but a repeat detector catches it in 3 steps instead of 40.
"Six model endpoints behind one API. Design routing, caching, and failover."
The gateway is the highest-leverage box in any GenAI platform because it's the only place where policy applies to every request. Everything below is a middleware in one chain:
authn → tenant quota (RPM + TPM) → budget check → guardrail(in)
→ cache lookup (exact → prefix → semantic)
→ route (model policy · cost · health · KV affinity)
→ call with timeout/retry/hedge → guardrail(out)
→ cache write → telemetryRouting. Three policies, increasingly ambitious: (1) explicit — the client names the model, you validate and enforce; (2) policy — a tenant-level mapping from alias to model so you can migrate everyone with a config change; (3) dynamic — a classifier sends easy queries to a small model. Ship 1 and 2. Treat 3 as an experiment gated on evals, because a router that misroutes 5% of hard queries to an 8B is a quality regression nobody attributes to the router.
Caching, in three tiers, and know the difference:
| Tier | Key | Hit rate | Risk |
|---|---|---|---|
| Exact | hash(full request) | Low (2–5%) | None |
| Prefix (KV) | block-hash chain of the prompt prefix | High on system prompts / multi-turn | Cross-tenant side channel |
| Semantic | embedding ANN + threshold τ | Medium, workload-dependent | Wrong answers |
Semantic caching is the one to interrogate. It returns a different question's answer because the embeddings were close. The classic killer: negation — "is X safe" and "is X not safe" embed close together. So: (a) calibrate τ against a labeled set and pick the point where false-hit rate ≤ 1%, treating it as an eval problem with a number, not a vibe; (b) key the cache on (embedding, model_version, prompt_template_hash, tool_schema_hash, tenant) so a prompt edit invalidates automatically; (c) never cache tool-calling turns, personalized answers, or anything time-sensitive; (d) ship it disabled and enable per-route.
Resilience. Timeout budget that decreases down the call chain (so an inner retry can't outlive the outer deadline). Retry only idempotent, non-streaming calls, with exponential backoff and jitter — without jitter your retries synchronize and you've built a self-DDoS. Circuit-break per upstream on error rate. Hedge (fire a second request at p95 latency) only for short, cheap calls, and never for streaming — hedging a 500-token generation doubles your bill for a tail you could have fixed with admission control.
Observability. Adopt the OpenTelemetry GenAI semantic conventions rather than inventing span names — they now cover model calls, agent orchestration, and MCP tool calls, and the ecosystem reads them. Note the design principle that matters for a bank or a regulated tenant: prompt and completion content is not captured by default, precisely to avoid PII leakage; content capture is opt-in and needs a redaction story. The four numbers on the dashboard: TTFT p95 by tenant, TPOT p95 by tenant, cache hit rate by tier, and $ per tenant per day.
The cut: if the deadline halves, ship routing and observability, drop semantic caching. Routing is how you migrate models without touching clients; observability is how you keep your job. Semantic caching is a quality risk that needs an eval harness you don't have yet.
"A prompt change and a model-version bump both want to ship today. Design the gate."
The reframe: in a GenAI system, the prompt, the model version, the retrieval index, the tool schemas, and the decoding parameters are all deployable artifacts, and any of them can regress quality without changing a line of code. So the gate isn't a test suite, it's a release process for a five-dimensional artifact.
change (prompt | model | index | tools | params)
→ versioned + hashed, immutable
→ offline eval: golden set (200–1000 labeled cases, stratified by intent)
· task metrics (exact match / rubric score / citation correctness)
· regression set: every case that ever broke in production
· adversarial + safety slice
→ gate: no metric below baseline − ε; zero regressions on the frozen set
→ shadow: mirror 5% of live traffic, compare offline, no user impact
→ canary 1% → 5% → 25% → 100%, auto-rollback on guard metrics
→ online: A/B on the business metric, LLM-judge on sampled live trafficThree staff-level points to make:
Interviewers test staff level by disagreeing. The failure mode is caving; the other failure mode is digging in. The correct shape is: restate their concern, concede the part that's true, hold the part that isn't, and name the evidence that would change your mind.
| They say | You say |
|---|---|
| "Why not just buy more accelerators?" | "That fixes throughput, not tail latency — my p99 problem is queueing and prefill interference, and both get worse with more replicas if routing stays cache-blind. I'd spend on routing and admission control first, and I'd know I was wrong if utilization were already above ~70% at p99 violation." |
| "Isn't disaggregation strictly better?" | "It's better when prompts dominate completions and the interconnect is fast. At our 3:1 prompt:completion ratio and spiky traffic, two pools means two sets of idle capacity plus a KV transfer on the critical path. At 10:1 I'd flip." |
| "Semantic caching would cut costs 40%." | "On hit rate, yes. On correctness, it's the only cache that can return a wrong answer. I'd ship it behind a per-route flag with a calibrated threshold at ≤1% false-hit rate, measured — and I'd start with prefix caching, which gets most of the win with none of the risk." |
| "Why not fine-tune instead of RAG?" | "Fine-tuning teaches form; retrieval supplies facts. Our facts change hourly and need citations, so RAG is the substrate. I'd fine-tune the retriever or a small router model before I'd fine-tune the generator." |
| "This is over-engineered." | "Agreed for v1 — the v1 is the gateway, one engine pool, and prefix caching. Everything else is on the roadmap with a trigger condition attached. Let me mark which boxes are v1." |
| "Our accelerator has less bandwidth than an H100." | "Then it has a different optimum, not a worse one: 4× the memory means I can run much larger batches and much longer contexts before I hit a capacity wall, and batch is exactly what amortizes the bandwidth cost per token. I'd design for high-concurrency, throughput-shaped workloads and use speculative decoding to protect per-user latency." |
Keep one incident from each row ready; interviewers ask "tell me what breaks" and reward specificity.
| Symptom | First metric to check | Usual cause | Fix |
|---|---|---|---|
| p99 TTFT ≫ p50 | queue depth, prefill token share | Long prompts stalling the batch | Chunked prefill; cap prefill share; separate long-prompt lane |
| TPOT degrading over the day | batch size, KV utilization | Batch grew past the roofline knee | Cap batch; admission control; add replicas |
| One tenant slows everyone | per-tenant goodput share | FIFO scheduling, no fair share | Deficit round-robin + per-tenant TPM |
| Cost up 3×, traffic flat | tokens/request, steps/task | Prompt or agent-loop regression | Budget guards; alarm on tokens/request, not just $ |
| Answers confidently wrong | retrieval recall@k, citation rate | Retrieval, not the model | Fix chunking/hybrid/rerank before touching the prompt |
| Stale/deleted docs in answers | index-vs-source ID diff | Orphaned vectors | Tombstones + reconciliation + query-time deleted_at filter |
| Capacity vanishes on deploy | replica-ready count | Simultaneous model reload | Rolling update with surge + cache warm-up |
| Cost with no user | active sequences vs open sockets | No cancellation propagation | Wire client disconnect to engine abort |
| Quality drops with no deploy | model/index/tool version hashes | Something else shipped | Version and gate all five artifact types |
Flashcards
2 · L · H_kv · d_head · b — and H_kv, not H_q, is the GQA win.(W + B·KV) / BW → batch amortizes weights; past the knee, KV dominates.2 · P · T_prompt; fix TTFT with prefix cache → chunked prefill → tensor parallel, in that order.concurrency = arrival_rate × session_time. Take the max of the memory-derived and bandwidth-derived fleet sizes.Qualcomm platform — AI200/AI250 announcement · AI Inference Suite · AI200 Infrastructure Management Suite · vLLM on Cloud AI (qaic) · QEfficient · Speculative decoding + MX formats · Independent Cloud AI 100 Ultra benchmark study
Serving architecture — Prefill/decode disaggregation explained · NVIDIA Dynamo disaggregated serving · Dynamo KV cache offloading · Dynamo + llm-d · Microserving of LLMs · Online routing at scale
Multi-tenancy and SLOs — Cohere: serving fairness · SLOs-Serve · Fairness-aware chunked-prefill scheduling · Prefix-cache admission responsibility · Multi-tenant serving architecture guide
Observability — OpenTelemetry GenAI semantic conventions