← Interview Mastery
IC4IC5IC6

Qualcomm Gen AI (Staff): The System-Design Primer

The staff-level AI system-design round for Qualcomm's Datacenter AI org, worked end to end — the drive script, the napkin math you must produce unprompted, and six full designs (multi-tenant inference service, P/D disaggregation, on-prem RAG appliance, agent control plane, AI gateway, release-gate eval platform) with the pushback drills for each.

45 min read · 22 sections
Prerequisites: /interview/qualcomm-datacenter-ai, /system-design/the-framework

1. Quick anchor

At staff level in Qualcomm's Datacenter AI org you are designing on top of the rack, not inside it. Nobody is going to ask you to write an attention kernel. They are going to hand you a rack of AI200-class accelerators and a product requirement, and watch whether you can turn silicon into a service with an SLO, a tenancy model, a cost per million tokens, and a story for what happens at 3am.

Almost every design prompt in this loop collapses to the same three questions:

  1. Where does the KV cache live, and who is allowed to evict it? Everything about LLM serving — batching, admission, preemption, routing, disaggregation, multi-tenancy — is KV-cache resource management wearing different costumes.
  2. Which SLO are you protecting when two of them conflict? TTFT and TPOT are in direct tension through the batch. You cannot maximize both. Staff-level answers pick, and say why.
  3. What does this accelerator make cheap, and what does it make expensive? The AI200 bet is memory capacity — 768 GB of LPDDR per card, roughly 4× the HBM on a contemporary GPU. Capacity buys you concurrency. Bandwidth buys you tokens per second. A capacity-rich, bandwidth-poorer part moves your bottleneck from "how many sequences fit" to "how fast can I stream weights per decode step" — and a design that doesn't notice that is a design that leaves the machine idle.

The one-line bar: IC4 designs the happy path and names the tradeoffs when asked. IC5 proposes the scope unprompted, produces arithmetic out loud, picks two components to go deep on, and names the condition under which their own choice becomes the wrong one.

Read the Qualcomm track first for org context and platform literacy; read the AI system-design framework for the generic eight-stage spine. This page is the staff-level, Qualcomm-shaped version: same spine, harder questions, real numbers.

Then: the coding rounds, solved The 4-week track

2. What's actually being scored

The same prompt is graded against a different rubric at each level. Memorize the right-hand column; it is where downlevels happen.

Dimension IC4 answer IC5 (staff) answer Disqualifier
Scoping Asks "what scale?" and waits Proposes the scope: "I'll assume 20 req/s, 4K prompts, 500-token completions, three tenants, 99.5% availability — stop me if that's wrong" Starts drawing boxes before naming a single requirement
Numbers Estimates when pushed Estimates unprompted, states assumptions, and identifies the dominant term "It depends on the hardware"
Depth Even coverage of eight components Two components at real depth, the rest named and deferred explicitly Uniform shallow coverage; or one component for 45 minutes
Tradeoffs Names both sides Names both sides, picks one, and names the observable that would flip the decision "We could do either"
Failure Mentions retries and monitoring Walks a specific incident: symptom → metric → root cause → mitigation → the guardrail that prevents recurrence Happy path only
Cost Mentions it's expensive Produces $/1M tokens from device-hours and names the two levers that move it most Ignores cost entirely
Org — Names what ships in v1 vs v2, what a second team owns, what the migration path is Designs a system no team could staff

A useful calibration line from the field: the most common reason a strong candidate gets downleveled from staff is producing a design with senior depth but not staff breadth-and-depth together — one beautiful subsystem, no system.

3. The drive script

Forty-five minutes, seven moves. Say the timings out loud; interviewers read it as scope control, not rigidity.

Min Move What you say
0–4 Frame + propose scope Restate the problem in one sentence. Propose traffic, prompt/completion shape, tenancy, and the availability target. Ask exactly one question: "Is this latency-sensitive interactive traffic, batch, or both?" — the answer changes everything downstream.
4–8 SLO table Write TTFT p95, TPOT p95, E2E p95, availability, and goodput (requests/s that meet SLO — not raw throughput). Say: "I'll optimize goodput, not throughput."
8–13 Napkin math KV bytes/token → memory per sequence → concurrency per card → cards needed (Little's Law) → $/1M tokens. Out loud. Wrong-but-stated beats right-but-silent.
13–22 Architecture Data plane and control plane as separate pictures. Name the request path end to end in one breath, then annotate where each SLO is won or lost.
22–34 Two deep dives You choose them: "The two things that decide whether this works are the scheduler and the tenancy model. I'll go deep on both." This is the single highest-signal sentence in the round.
34–40 Failure modes + ops Three concrete incidents with detection metric and mitigation. Include one that your own design causes.
40–45 Cost, cuts, and roadmap $/1M tokens, the two levers, what ships in v1, what you'd cut if the deadline halved.

The three sentences that read as staff:

  • "Let me put numbers on this before I draw anything."
  • "I'm going to protect TPOT and let TTFT degrade under load — here's why that's the right call for this product."
  • "This design is wrong if the prompt distribution is bimodal; here's the metric I'd watch and what I'd switch to."

4. The napkin math you must produce unprompted

4.1 KV cache per token — the master formula

ƒ
bytes/token=2×L×Hkv×dhead×b\text{bytes/token} = 2 \times L \times H_{kv} \times d_{head} \times b

Two for K and V, LL layers, HkvH_{kv} key/value heads (not query heads — this is the whole point of GQA), head dimension dheadd_{head}, and bb bytes per element.

Model shape LL HkvH_{kv} dheadd_{head} KV/token @ FP16 8K ctx 128K ctx
8B-class (GQA 8) 32 8 128 128 KiB 1.0 GiB 16 GiB
32B-class (GQA 8) 64 8 128 256 KiB 2.0 GiB 32 GiB
70B-class (GQA 8) 80 8 128 320 KiB 2.5 GiB 40 GiB
70B-class, no GQA (MHA 64) 80 64 128 2.5 MiB 20 GiB 320 GiB

That last row is the one to say out loud: GQA cut the KV cache 8× and that is why long context became affordable at all. Halve it again with FP8 KV; the accuracy cost is usually negligible for KV specifically (it's the activations you have to be careful with).

4.2 Weights

FP16 ≈ 2 bytes/param. A 70B model is 140 GB in FP16, ~70 GB in FP8, ~53 GB in a 6-bit microscaling format, ~35 GB in INT4. Qualcomm's MXFP6 path compresses weights ~61% and decompresses on the fly to FP16 in the vector engine, overlapping decompression with weight fetch — which matters precisely because decode is bandwidth-bound (§4.4).

4.3 Concurrency per card — the capacity story

Take one AI200-class card at 768 GB, a 70B model at MXFP6 (~55 GB), 8K average context:

usable KV memory  = 768 GB − 55 GB (weights) − ~40 GB (activations, fragmentation, headroom)
                  ≈ 670 GB
sequences @ 8K    = 670 GB / 2.5 GiB ≈ 250 concurrent sequences on ONE card

Now the punchline, and the most important sentence in this section: that is not your concurrency limit. It's your memory limit. Your real limit is bandwidth (§4.4). Capacity-rich parts are exactly the ones where naive capacity math flatters you.

4.4 Decode roofline — the bandwidth story

Every decode step reads all weights once plus the KV of every sequence in the batch:

ƒ
tstep≈W+B⋅KVseqBWthroughput=Btstept_{step} \approx \frac{W + B \cdot \text{KV}_{seq}}{\text{BW}} \qquad \text{throughput} = \frac{B}{t_{step}}

Two consequences you should state:

  1. Small batch is pure waste. At B=1B{=}1 you pay the full weight read for one token. At B=64B{=}64 you amortize it 64×. This is the argument for continuous batching, in one equation.
  2. There's a batch size beyond which TPOT collapses, because B⋅KVseqB \cdot \text{KV}_{seq} overtakes WW. Past that point every added sequence slows everyone linearly. Find it empirically; it's the knob that trades goodput against tail latency.

Worked, with stated assumptions: 55 GB weights, 2.5 GiB KV/seq, assume 1.2 TB/s effective memory bandwidth.

  • B=32B{=}32: (55+80) GB/1.2 TB/s≈113(55 + 80)\,\text{GB} / 1.2\,\text{TB/s} \approx 113 ms/step → ~9 tok/s per sequence, 283 tok/s aggregate.
  • B=8B{=}8: (55+20)/1200≈63(55 + 20) / 1200 \approx 63 ms/step → ~16 tok/s per sequence, 127 tok/s aggregate.

So 4× the batch bought 2.2× aggregate throughput and cost 44% of per-user speed. That table is the entire TTFT/TPOT-vs-throughput tradeoff, made concrete. Draw it.

Say this: "The AI250's near-memory High Bandwidth Compute architecture targets exactly this term. It's a decode-bottleneck part. On AI200 I'd design for large batches and lean on speculative decoding to buy back per-sequence latency; on AI250 I'd expect to be able to run smaller batches at the same goodput, which relaxes the fairness problem."

4.5 Prefill — the compute story

ƒ
FLOPsprefill≈2⋅P⋅Tprompt,TTFT≈2⋅P⋅TpromptFeff\text{FLOPs}_{prefill} \approx 2 \cdot P \cdot T_{prompt}, \qquad \text{TTFT} \approx \frac{2 \cdot P \cdot T_{prompt}}{F_{eff}}

70B model, 4K prompt: 2×70e9×4096≈5.7e142 \times 70\text{e}9 \times 4096 \approx 5.7\text{e}14 FLOPs. At an assumed 80 TFLOP/s effective (i.e. a 200 TFLOP/s part at 40% MFU) that's ~7 seconds of TTFT for one request. Nobody's SLO survives that. Which is why the three fixes exist and why you name them in this order:

  1. Prefix caching — the cheapest possible fix, because it doesn't reduce the work, it skips it. System prompts, few-shot blocks, and multi-turn history are the same tokens every time.
  2. Chunked prefill — split the 4K prompt into 512-token chunks interleaved with decode steps so one long prompt doesn't stall every active decoder. Trades a bit of TTFT for a lot of TPOT stability.
  3. Tensor parallelism across cards — cuts prefill latency near-linearly, at the cost of collective communication on every layer.

4.6 Little's Law — how many cards

ƒ
Nconcurrent=λ×TsessionN_{concurrent} = \lambda \times T_{session}

20 req/s, 40s average session → 800 concurrent sequences → 800 × 2.5 GiB = 2 TB of KV → ~3 cards for KV alone, before you've thought about bandwidth. Then check bandwidth: 800 sequences at your chosen batch size and step time. Whichever number is bigger is your fleet size. Say both numbers and take the max — that's the move.

4.7 Cost per million tokens

ƒ
$/1M tok=$device-hr×Ndevicestok/s×3600×106\$/1\text{M tok} = \frac{\$_{device\text{-}hr} \times N_{devices}}{\text{tok/s} \times 3600} \times 10^6

At 3/device−hr,4devices,1,100tok/saggregate:3/device-hr, 4 devices, 1,100 tok/s aggregate: \frac{12}{3.96\text{M}} \times 10^6 \approx $3.03$ per million output tokens. Now name the two levers that move it most, in order:

  1. Cache hit rate (prefix + semantic). A 40% prefix-cache hit rate on a workload with long system prompts is a 30%+ cost cut with zero quality change. Nothing else has that profile.
  2. Batch size / goodput, i.e. how close you run to the roofline before violating SLO.

Distant third: quantization. Fourth: model routing (send the easy 60% to an 8B). And the one that actually dominates in agentic products: calls per task, because a 50-step agent turns a cheap per-token price into an expensive per-task bill.

5. Six worked designs

D1 — Multi-tenant, OpenAI-compatible inference service on a rack

"We have a rack of AI200-class accelerators. Build the service that lets three internal product teams and two external customers hit /v1/chat/completions."

Scope you propose: 5 tenants, mixed interactive + batch, 8B/32B/70B model catalogue, ~30 req/s peak, prompts 500–32K (bimodal — flag this early), completions ~400 tokens, 99.5% availability, strict tenant isolation on data, soft isolation on capacity.

SLO table (write this, don't say it):

Class TTFT p95 TPOT p95 E2E p95 Notes
Interactive (chat) 800 ms 40 ms 20 s The thing users feel
Batch (offline) 60 s — 30 min Throughput-optimized, preemptible
Internal eval sweeps best effort — — Runs on scavenged capacity

Architecture — data plane:

client
  │  OpenAI-compatible HTTP/SSE
  ▼
[ API gateway ]  authn/z · per-tenant token-bucket (RPM + TPM) · request validation
  │              idempotency keys · request-id · budget check
  ▼
[ Router ]       model routing · KV-cache-aware worker selection · session affinity
  │              admission control (shed, don't queue forever)
  ▼
[ Engine pool ]  vLLM (qaic backend) replicas · continuous batching · paged KV
  │              prefix cache (radix/trie over block hashes)
  ▼
[ KV tier ]      on-device HBM/LPDDR → host DRAM → NVMe (offload, cold prefixes)

Control plane (separate box, separate on-call): model registry + compiled-artifact store (a model is only "ready" once it has a compiled QPC binary for this hardware — precompilation is not optional on this platform), rollout controller, autoscaler, quota service, eval gate, telemetry pipeline.

Deep dive 1 — the scheduler. This is where the round is won.

The scheduler decides, every step: who gets admitted, who gets preempted, and how much of the token budget goes to prefill vs decode.

  • Admission is fair-share, not FIFO. Deficit round-robin across tenants gives burst-proof fair shares among equal tiers; a priority selector orders within a tenant's share. This is the standard shape and it's worth naming: a rate limiter enforces the quota, then the selector decides who's next.
  • Chunked prefill with a global max_batched_tokens budget: decode steps cost 1 token each, prefill chunks cost their length. Cap the prefill share so a 200K-token prompt cannot stall 60 decoders. This is the direct answer to the noisy-neighbor question.
  • Preemption when KV runs out: evict the most recently scheduled sequence, not the oldest. It has the least sunk cost, and it avoids the starvation cascade where you keep killing the requests closest to finishing. Recompute if the prompt is short, swap KV to host DRAM if it's long — the crossover is roughly when transfer time beats recompute time.
  • Fairness under chunked prefill is a known hard problem, because a tenant with long prompts consumes budget without producing tokens. Say that you'd measure per-tenant goodput share, not request share.

Deep dive 2 — tenancy. Three levels, and you must be explicit about which you're buying:

Isolation Mechanism Cost
Data Tenant ID in every cache key, every index filter, every log. Non-negotiable, always on. ~free
Performance Fair-share scheduler + per-tenant TPM/RPM quotas + admission control Some utilization
Fault/blast-radius Separate replica pools (or namespaces) per premium tenant Big — stranded capacity

The staff position: buy data isolation always, performance isolation by default, fault isolation only for tenants who pay for it. Then name the leak you just created: the prefix cache is a cross-tenant side channel. If tenant B's request gets a suspiciously fast TTFT because tenant A cached the same prefix, B has learned something about A. Fix: partition the prefix cache by tenant (costs hit rate), or accept sharing only for a whitelisted set of public system prompts. Interviewers love this one because most candidates never see it.

Failure modes to walk:

  1. Thundering herd on model rollout — every replica reloads at once, capacity goes to zero. Fix: rolling update with surge capacity, and warm the prefix cache before taking traffic.
  2. Unbounded queue — arrival rate exceeds service rate for 90 seconds, the queue grows, and now every request misses SLO including the ones you could have served. Fix: bounded queue + load shedding with Retry-After. Say the line: "the queue is where your latency SLO goes to die."
  3. Client disconnects mid-stream and nobody cancels the sequence — you burn KV and compute generating tokens into a socket that's gone. Fix: propagate cancellation from the HTTP layer all the way to the engine's running-sequence set. This is a real, common, expensive bug.

What ships first: the gateway with quotas and observability. You cannot debug — or bill — what you can't see, and it's the only component whose value doesn't depend on the rest being finished.


D2 — Prefill/decode disaggregation and KV-aware routing

"Should prefill and decode run on the same devices?"

The physics. Prefill is compute-bound: one forward pass over the whole prompt, high arithmetic intensity, wants big compute. Decode is memory-bandwidth-bound: one token at a time, reading the entire weight set per step. Co-locating them means every long prefill injects a latency spike into every active decode stream — you literally cannot tune TTFT and TPOT independently, because one knob (batch composition) controls both.

Disaggregation puts them in separate pools connected by a KV transfer path. Now you tune independently: scale the prefill pool for TTFT, the decode pool for TPOT, and pick the P:D ratio from your prompt:completion ratio. Production systems report large throughput gains from this plus KV-aware routing; the mechanism is well documented and the pattern (SplitWise, DistServe, Mooncake, Dynamo, llm-d) is now standard.

The honest crossover — and this is the staff answer:

Disaggregate when Co-locate when
Prompts are long relative to completions (RAG, code, doc QA) Prompts and completions are similar length (chat)
You have a fast interconnect for KV transfer Interconnect is your bottleneck — you'll just move the stall
Traffic is high and steady enough to keep both pools busy Traffic is spiky or low — two pools means two sets of idle capacity
You need to hit distinct TTFT and TPOT SLOs One SLO, and chunked prefill already meets it

The move: "I'd start co-located with chunked prefill, instrument the TTFT/TPOT correlation, and disaggregate only when I can show that prefill interference is what's breaking TPOT. Disaggregation adds a KV transfer on the critical path and a whole distributed failure mode; I want evidence before I buy that." Then add the condition that flips it: "If the prompt:completion ratio goes above ~10:1 — which it will the moment we ship RAG — I'd expect to disaggregate."

KV-cache-aware routing. Once you have a pool, "least-loaded" routing is wrong: it ignores that one replica already holds 90% of this request's prefix in cache. Route on a score:

score(worker) = w1 · prefix_overlap(request, worker.cache)
              − w2 · worker.queue_depth
              − w3 · worker.kv_pressure

Prefix-aware routing is one of the highest-leverage single changes in a serving stack — reported gains in the 30–60% range in production systems — because it converts a routing decision into skipped prefill FLOPs. The tension to name: cache affinity fights load balancing. Pin too hard and you hotspot; ignore cache and you recompute. Cap it — if a worker's queue exceeds a threshold, fall back to load-based routing regardless of overlap.

KV offload tiering. HBM/on-package → host DRAM → NVMe. Offload turns "evict and recompute" into "fetch," which is a win exactly when transfer is cheaper than recompute — i.e. long prefixes, high reuse. On a 768 GB-per-card part, note that you have far more room before you need this tier than a 192 GB HBM part does. That's a design advantage of the AI200 you should name out loud.


D3 — On-prem, air-gapped RAG appliance

"A bank wants your inference appliance on-prem. 40M documents, per-user ACLs that change hourly, no egress. A leaked chunk is a compliance incident."

This is design a RAG system over 10M docs with three multipliers: air-gap, ACL, and appliance. Don't re-derive chunking from scratch — say "standard hybrid retrieval with reranking, I'll assume 400-token chunks with 15% overlap and semantic boundaries" and spend your time on what's actually different.

Napkin math first: 40M docs × ~8 chunks = 320M chunks. At 1024-dim FP32 that's 320M × 4KB = 1.3 TB of raw vectors — instantly a "this doesn't fit in RAM on one box" conversation. Fixes, in order: int8 scalar quantization (4×, ~1% recall loss), product quantization or binary + rerank (32×, bigger loss), or a disk-backed index (DiskANN-style). Say the number, then pick: "int8 with a full-precision rerank of the top 200. 330 GB fits comfortably."

The three things that are actually different:

  1. ACL enforcement must be inside retrieval, never after it. Post-filtering is both a correctness bug (you asked for top-50, 45 get filtered, you return 5 bad results) and a timing side channel. Push the permission predicate into the ANN query as a pre-filter, with the ACL set materialized per user. When permissions change hourly: keep a fast-changing ACL store keyed by (user, doc) and a slower vector index; recompute the user's accessible-set bitmap on permission change, not on query. And re-check at generation time — the retrieved chunk must be authorized again before it enters the prompt, because the index may be seconds stale and "seconds stale" is how a leak happens.
  2. Air-gap changes the model supply chain, not the architecture. No hosted APIs, no telemetry egress, no downloading a checkpoint at runtime. Everything — models, compiled artifacts, embedding model, reranker — ships as signed, versioned bundles. Say: "on this platform a model is a compiled binary for this hardware, so the appliance image and the model catalogue are one artifact with one version number. That's actually simpler than the cloud case." Evals run on-prem against a customer-supplied golden set, and results are the only thing that leaves — with the customer's approval.
  3. Appliance means capacity is fixed and known. No autoscaling. So admission control and graceful degradation are not optional — they are the product. Define the degradation ladder explicitly: full pipeline → skip reranker → shrink top-k → route to the smaller model → queue → shed. And write down which quality metric each rung costs you.

Failure mode to walk: a document is deleted from the source system. Its chunks remain in the vector index (orphaned vectors), and it surfaces in an answer three weeks later. This is the single most common real RAG incident. Fix: tombstones + a reconciliation job that diffs source IDs against index IDs on a schedule, plus a hard filter on deleted_at at query time so correctness never depends on the job having run.


D4 — The agent control plane

"Design the platform that runs long-lived agents for internal teams. Some run for hours."

The reframe that gets you the level: an agent is a distributed workflow with a nondeterministic planner, not a chat request. Everything follows from that.

What must be durable. The agent's state — conversation, plan, tool results, step counter, budget spent — lives in a checkpointed store, not in process memory. A deploy, a crash, or a preemption resumes from the last checkpoint rather than losing an hour of work. This is the mature pattern in 2026: durable-execution runtimes with checkpoint/replay (Temporal-style activities, LangGraph-style checkpointers, Durable Objects), and it's what separates a demo from a platform.

What must be idempotent. Every tool call. If you checkpoint after a tool call and crash during one, replay re-executes it. send_email twice is an incident. Give every tool invocation a caller-generated idempotency key, and make the tool gateway dedupe on it. Say this unprompted — it's the highest-signal detail in the whole design.

The three loops and their guards:

while not done:
    step += 1
    if step > MAX_STEPS: halt("step budget")            # loop guard
    if spent > BUDGET:   halt("cost budget")            # cost guard
    if now > DEADLINE:   halt("wall-clock guard")
    plan = llm(context)                                  # ← nondeterminism lives here
    checkpoint(state)
    result = tool_gateway.call(plan.tool, plan.args,
                               idem_key=hash(run_id, step),
                               policy=tenant_policy)     # ← authz lives here
    context = compact(context + result)                  # ← context lives here
    checkpoint(state)

Three deep-dive candidates; pick two:

  • The tool gateway. Every tool call is an authz decision made on behalf of a user by a model that can be talked into things. Tools are declared with schemas, scoped to a tenant, rate-limited per run, and classified read/write/irreversible — with irreversible ones requiring either a human approval hop or a hard allowlist. If MCP is in the picture, note that it standardizes the tool interface and simultaneously widens the supply-chain surface: a tool server can change its description between calls (rug-pull), and tool descriptions are untrusted input that reaches the model's context.
  • Context management. The context window is the real memory hierarchy. Compaction (summarize old turns), externalization (write results to a scratchpad, keep a pointer), and retrieval (pull back on demand). The failure mode: the agent's 40th step has a context full of tool output and no room for the plan.
  • Cost control. Per-run and per-tenant budgets enforced at the gateway, denominated in tokens and dollars, with a hard kill at the ceiling. Agentic workloads are where inference bills go non-linear: 50–200 model calls per task means a 2× regression in step count is a 2× bill, and step count is a function of a prompt you just changed.

Failure mode to walk: an agent enters a two-step loop — call tool, get error, call the same tool identically. Detection: hash (tool, args) per run and alarm on repeats; the step budget is the backstop, but a repeat detector catches it in 3 steps instead of 40.


D5 — The AI gateway

"Six model endpoints behind one API. Design routing, caching, and failover."

The gateway is the highest-leverage box in any GenAI platform because it's the only place where policy applies to every request. Everything below is a middleware in one chain:

authn → tenant quota (RPM + TPM) → budget check → guardrail(in)
      → cache lookup (exact → prefix → semantic)
      → route (model policy · cost · health · KV affinity)
      → call with timeout/retry/hedge → guardrail(out)
      → cache write → telemetry

Routing. Three policies, increasingly ambitious: (1) explicit — the client names the model, you validate and enforce; (2) policy — a tenant-level mapping from alias to model so you can migrate everyone with a config change; (3) dynamic — a classifier sends easy queries to a small model. Ship 1 and 2. Treat 3 as an experiment gated on evals, because a router that misroutes 5% of hard queries to an 8B is a quality regression nobody attributes to the router.

Caching, in three tiers, and know the difference:

Tier Key Hit rate Risk
Exact hash(full request) Low (2–5%) None
Prefix (KV) block-hash chain of the prompt prefix High on system prompts / multi-turn Cross-tenant side channel
Semantic embedding ANN + threshold τ Medium, workload-dependent Wrong answers

Semantic caching is the one to interrogate. It returns a different question's answer because the embeddings were close. The classic killer: negation — "is X safe" and "is X not safe" embed close together. So: (a) calibrate τ against a labeled set and pick the point where false-hit rate ≤ 1%, treating it as an eval problem with a number, not a vibe; (b) key the cache on (embedding, model_version, prompt_template_hash, tool_schema_hash, tenant) so a prompt edit invalidates automatically; (c) never cache tool-calling turns, personalized answers, or anything time-sensitive; (d) ship it disabled and enable per-route.

Resilience. Timeout budget that decreases down the call chain (so an inner retry can't outlive the outer deadline). Retry only idempotent, non-streaming calls, with exponential backoff and jitter — without jitter your retries synchronize and you've built a self-DDoS. Circuit-break per upstream on error rate. Hedge (fire a second request at p95 latency) only for short, cheap calls, and never for streaming — hedging a 500-token generation doubles your bill for a tail you could have fixed with admission control.

Observability. Adopt the OpenTelemetry GenAI semantic conventions rather than inventing span names — they now cover model calls, agent orchestration, and MCP tool calls, and the ecosystem reads them. Note the design principle that matters for a bank or a regulated tenant: prompt and completion content is not captured by default, precisely to avoid PII leakage; content capture is opt-in and needs a redaction story. The four numbers on the dashboard: TTFT p95 by tenant, TPOT p95 by tenant, cache hit rate by tier, and $ per tenant per day.

The cut: if the deadline halves, ship routing and observability, drop semantic caching. Routing is how you migrate models without touching clients; observability is how you keep your job. Semantic caching is a quality risk that needs an eval harness you don't have yet.


D6 — The release gate

"A prompt change and a model-version bump both want to ship today. Design the gate."

The reframe: in a GenAI system, the prompt, the model version, the retrieval index, the tool schemas, and the decoding parameters are all deployable artifacts, and any of them can regress quality without changing a line of code. So the gate isn't a test suite, it's a release process for a five-dimensional artifact.

change (prompt | model | index | tools | params)
  → versioned + hashed, immutable
  → offline eval: golden set (200–1000 labeled cases, stratified by intent)
       · task metrics (exact match / rubric score / citation correctness)
       · regression set: every case that ever broke in production
       · adversarial + safety slice
  → gate: no metric below baseline − ε; zero regressions on the frozen set
  → shadow: mirror 5% of live traffic, compare offline, no user impact
  → canary 1% → 5% → 25% → 100%, auto-rollback on guard metrics
  → online: A/B on the business metric, LLM-judge on sampled live traffic

Three staff-level points to make:

  1. Judges need their own evals. If an LLM-as-judge gates your release, you must measure the judge's agreement with human labels on a held-out set and re-measure it when you change the judge's model. An ungated judge is an unmeasured dependency in your release path.
  2. Choose your error asymmetry and defend it. A gate tuned to catch everything blocks good changes and gets routed around within a month — which is worse than no gate. State it: "I'll take a ~10% false-negative rate on subtle quality regressions to keep the false-positive rate near zero, because a gate people trust is worth more than a gate that's right. I catch the rest in canary, where rollback is 90 seconds."
  3. The regression set is the asset. Golden sets go stale; the set of "things that broke in prod, now frozen as tests" only gets more valuable. Every incident ends with a new case in that set. That's the flywheel, and it's an organizational answer as much as a technical one — which is exactly what the IC6 version of this question is fishing for.

6. Pushback drills

Interviewers test staff level by disagreeing. The failure mode is caving; the other failure mode is digging in. The correct shape is: restate their concern, concede the part that's true, hold the part that isn't, and name the evidence that would change your mind.

They say You say
"Why not just buy more accelerators?" "That fixes throughput, not tail latency — my p99 problem is queueing and prefill interference, and both get worse with more replicas if routing stays cache-blind. I'd spend on routing and admission control first, and I'd know I was wrong if utilization were already above ~70% at p99 violation."
"Isn't disaggregation strictly better?" "It's better when prompts dominate completions and the interconnect is fast. At our 3:1 prompt:completion ratio and spiky traffic, two pools means two sets of idle capacity plus a KV transfer on the critical path. At 10:1 I'd flip."
"Semantic caching would cut costs 40%." "On hit rate, yes. On correctness, it's the only cache that can return a wrong answer. I'd ship it behind a per-route flag with a calibrated threshold at ≤1% false-hit rate, measured — and I'd start with prefix caching, which gets most of the win with none of the risk."
"Why not fine-tune instead of RAG?" "Fine-tuning teaches form; retrieval supplies facts. Our facts change hourly and need citations, so RAG is the substrate. I'd fine-tune the retriever or a small router model before I'd fine-tune the generator."
"This is over-engineered." "Agreed for v1 — the v1 is the gateway, one engine pool, and prefix caching. Everything else is on the roadmap with a trigger condition attached. Let me mark which boxes are v1."
"Our accelerator has less bandwidth than an H100." "Then it has a different optimum, not a worse one: 4× the memory means I can run much larger batches and much longer contexts before I hit a capacity wall, and batch is exactly what amortizes the bandwidth cost per token. I'd design for high-concurrency, throughput-shaped workloads and use speculative decoding to protect per-user latency."

7. Failure-mode catalogue

Keep one incident from each row ready; interviewers ask "tell me what breaks" and reward specificity.

Symptom First metric to check Usual cause Fix
p99 TTFT ≫ p50 queue depth, prefill token share Long prompts stalling the batch Chunked prefill; cap prefill share; separate long-prompt lane
TPOT degrading over the day batch size, KV utilization Batch grew past the roofline knee Cap batch; admission control; add replicas
One tenant slows everyone per-tenant goodput share FIFO scheduling, no fair share Deficit round-robin + per-tenant TPM
Cost up 3×, traffic flat tokens/request, steps/task Prompt or agent-loop regression Budget guards; alarm on tokens/request, not just $
Answers confidently wrong retrieval recall@k, citation rate Retrieval, not the model Fix chunking/hybrid/rerank before touching the prompt
Stale/deleted docs in answers index-vs-source ID diff Orphaned vectors Tombstones + reconciliation + query-time deleted_at filter
Capacity vanishes on deploy replica-ready count Simultaneous model reload Rolling update with surge + cache warm-up
Cost with no user active sequences vs open sockets No cancellation propagation Wire client disconnect to engine abort
Quality drops with no deploy model/index/tool version hashes Something else shipped Version and gate all five artifact types

8. Pitfalls

  • Drawing before counting. The arithmetic is the signal. Start with numbers.
  • Optimizing throughput instead of goodput. Requests that miss SLO have negative value — they cost money and produce a complaint.
  • Treating tenancy as an auth problem. It's a scheduling problem and a cache-key problem too.
  • Forgetting the control plane. Model registry, compiled-artifact versioning, rollout, and quota are half the system on this platform.
  • Ignoring cancellation. The cheapest cost saving in most serving stacks.
  • Uniform coverage. Eight shallow boxes reads as senior. Two deep dives plus explicit deferral reads as staff.
  • Not naming your own design's weakness. Every design has one. Saying it first is the strongest move available to you.

Flashcards

  • KV bytes/token = 2 · L · H_kv · d_head · b — and H_kv, not H_q, is the GQA win.
  • 70B @ FP16 KV, GQA-8: 320 KiB/token, 2.5 GiB per 8K sequence.
  • Decode step time ≈ (W + B·KV) / BW → batch amortizes weights; past the knee, KV dominates.
  • Prefill FLOPs ≈ 2 · P · T_prompt; fix TTFT with prefix cache → chunked prefill → tensor parallel, in that order.
  • Little's Law: concurrency = arrival_rate × session_time. Take the max of the memory-derived and bandwidth-derived fleet sizes.
  • Goodput = requests/s meeting SLO. Optimize this.
  • Preempt the newest sequence, not the oldest.
  • Prefix cache = shared state = cross-tenant side channel.
  • Every tool call needs an idempotency key.
  • Prompt, model, index, tool schema, decoding params — five deployable artifacts, all version-gated.

9. Further reading

Qualcomm platform — AI200/AI250 announcement · AI Inference Suite · AI200 Infrastructure Management Suite · vLLM on Cloud AI (qaic) · QEfficient · Speculative decoding + MX formats · Independent Cloud AI 100 Ultra benchmark study

Serving architecture — Prefill/decode disaggregation explained · NVIDIA Dynamo disaggregated serving · Dynamo KV cache offloading · Dynamo + llm-d · Microserving of LLMs · Online routing at scale

Multi-tenancy and SLOs — Cohere: serving fairness · SLOs-Serve · Fairness-aware chunked-prefill scheduling · Prefix-cache admission responsibility · Multi-tenant serving architecture guide

Observability — OpenTelemetry GenAI semantic conventions

Primary sources
← More in Interview Mastery