Qualcomm is spending 2026 turning itself into a datacenter inference company, and the bet got much bigger this summer. The AI200 rack ships commercially this year; the first anchor customer (HUMAIN, Saudi Arabia) is bringing ~200 MW of racks live in Riyadh and Dammam; and at the June 2026 investor day Qualcomm unveiled the Dragonfly data-center portfolio — a 250+-core C1000 server CPU that Meta has committed to deploy, plus an accelerator roadmap (AI200 → AI250 → AI300) on an annual cadence. Racks and CPUs don't sell themselves — they need a cloud software layer: the AI Inference Suite (OpenAI-compatible endpoints, Python SDK, one-click Hugging Face deployment), serving orchestration, and the applied GenAI/agentic systems that prove the platform. That layer is where an Applied AI Engineer, Datacenter AI in Bengaluru sits: cloud applied AI + app development + AI infra development, on top of the accelerators — not inside them. No compiler work, no NPU microarchitecture, no kernel authoring.
The single most important sentence in this track: the loop will filter on DSA, differentiate on applied LLM depth (agents, RAG, serving, evals), and level you on AI system design — so spend your month exactly in that ratio. One page of platform literacy (Section 3) is cheap and signals genuine intent; a month of quantization-kernel study is expensive and signals you didn't read the room.
Drill this loop in the practice room
The DSA track
This track targets a director-described opening: work on cloud, applied AI, application development, and AI-infra development in the Datacenter AI org — explicitly not low-level compiler/NPU/GPU work. There's no public JD yet, so calibrate against the closest public signals from the same org:
- Qualcomm AI Inference Suite — the product shape of the job: Python SDK and OpenAI-compatible APIs for chat, image generation, multimodal, and RAG; offered as cloud playgrounds, inference-as-a-service, and on-prem appliances. Building and extending this kind of layer is applied AI + app dev + AI infra.
- Datacentre AI Engineering (HUMAIN AI Engineering Center, opened Dec 2025) postings describe "design, operation, and continuous improvement of large-scale AI inference systems… GenAI and agentic AI systems." The Bengaluru software teams feed the same mission.
- AI Platform Engineer reqs in India (2026) ask for Python/Java/C++, distributed backends, Docker/K8s, cloud (AWS/Azure/GCP), and GenAI/RAG experience — the exact applied-infra profile.
- Core AI Software reqs ask for strong software engineering (clean code, testing, Linux, git) plus "familiarity with generative AI model architectures such as LLMs" — engineering rigor first, ML familiarity second.
Context that makes this role bigger than it looked six months ago: Qualcomm now claims two global-scale hyperscaler customers driving ≥$1B revenue within FY2026, the Meta CPU deal is multi-generation, and Bengaluru is Qualcomm's largest engineering workforce outside the US (it taped out a 2nm design in 2026). Datacenter is the growth story, and India is a first-class engineering site for it.
Read the role as three overlapping jobs, and prepare for all three:
You consume the accelerator through vLLM-style serving endpoints and the platform SDKs; you do not program it. Know what's under the floorboards (next section) but budget your time above them.
You need about two pages of fluency here — enough to place your work in the org's story, quote a number when it helps, and ask sharp questions back.
Hardware (the floor below you).
Software (your layer), top to bottom. AI Inference Suite (endpoints, deployment, session management — the applied/cloud product) and the AI200 Infrastructure Management Suite (March 2026, rack-scale fleet management) → vLLM with the qaic backend (continuous batching via --max-num-seq/--block-size, prefix caching, OpenAI-compatible serving; models are pre-compiled offline) and Triton → Efficient Transformers / "QEfficient" (github.com/quic/efficient-transformers, 80+ model families: Llama incl. 4-Scout, Mistral/Mixtral, Qwen3-MoE, Gemma, Phi, Granite, VLMs, FLUX.1, wav2vec2) with a four-step onboarding pipeline: from_pretrained() → export() (ONNX) → compile() (QPC binary, with mxfp6=True-style flags) → run() → Cloud AI SDK / QAIRT / QNN (the compile-and-run layer — know it exists, don't study it). One distinction worth having cold: QNN targets Snapdragon edge devices; QAIC targets Cloud AI datacenter cards — same ONNX input, different worlds.
Interview-grade facts — each one earns more than a paragraph of adjectives:
- Speculative decoding + MX formats ≈ 4× decode throughput on Cloud AI 100 (SpD alone 1.5–2×, MXFP6 alone ~2×). Draft-model tokens are scored in parallel by the target model with modified rejection sampling — zero accuracy loss.
- MXFP6 compresses weights ~61%, decompressed on the fly to FP16 by the vector engine, overlapping decompression with weight fetch. It let OPT-13B fit a single card that previously couldn't hold it.
- The independent arXiv study (2507.00418) measured a 70B model on Cloud AI 100 Ultra at ~148 W vs ~2,983 W for an 8×A100 setup — the perf-per-watt pitch in one number (with honest fine print: results are mixed at the largest models, strongest on small/mid ones).
- SwiftKV / on-device KV cache (April 2025): the cache lives in card DRAM between decode steps with fixed-shape tensors, so ahead-of-time compilation survives dynamic sequence lengths.
- The analyst caveat you should volunteer before they do: AI200 targets cost-per-token and tokens-per-watt, not peak throughput — estimates suggest 2–6× more racks than GPU equivalents for the same raw throughput. The pitch is inference economics (cheap LPDDR capacity vs. HBM), not FLOPs.
Business context (one sentence each). HUMAIN: 200 MW of AI200/AI250 going live from 2026, joint AI Engineering Center in Riyadh, ~$1B deal by market reports, ambitions toward ~6 GW over a decade. Meta: multi-generation Dragonfly C1000 CPU commitment. Qualcomm's datacenter pitch: lowest TCO per token for decode-heavy inference — capacity-rich cards mean bigger batches and longer contexts per node before you shard (the economics lesson gives you that vocabulary).
Across 12+ sourced India experiences (2024–2026), the pattern is online assessment → 2–4 technical rounds → HR/team-match, moderate difficulty (~3.1/5), typically 2–3 weeks end to end (director-referred pipelines compress toward 3–4 weeks total including offer; referrals accelerate but do not skip rounds).
- Online assessment (sometimes waived for referrals and senior profiles). LeetCode medium-to-hard plus CS-fundamentals MCQs (C/C++, OS, aptitude — some with negative marking). Reported problems cluster hard on linked lists (8 of 12 experiences), LRU-style design, trees, bit manipulation: LRU Cache, Merge k Sorted Lists, Number of Islands, Asteroid Collision, Reverse Nodes in k-Group, Binary Tree Maximum Path Sum, Single Number, string compression ("AABCCC" → "2A1B3C"), count-pairs-divisible-by-k, power-of-two.
- Coding round(s). Same DSA territory live, in Python or C++. Two consistent reports: interviewers give hints and grade how you use them, and they grade code quality — names, edge cases, modularity — not just a passing result. Rounds are usually elimination-style.
- Systems fundamentals probing (inside coding or its own round). Qualcomm is a hardware company and its interviewers carry that culture even in applied loops: print odd/even with two threads, mutex vs semaphore, deadlock + the four conditions, malloc/calloc, dangling pointers, endianness, memcpy-with-overlap all recur across experiences. A half-day refresher (scheduled in Week 3) is cheap insurance.
- Applied ML/LLM depth. Reported ML asks: implement causal masking in PyTorch, explain GQA, MoE, encoder vs decoder, CNN backprop, 1×1 convolutions. For this cloud/apps role, layer on what 2025-26 applied loops industry-wide actually probe: design a RAG system end to end, debug a hallucinating RAG service, agent-vs-chain judgment, MCP, fine-tune-vs-RAG, cost-at-scale, hallucination detection, LLM-as-judge evals. Plus a project deep-dive: your resume walked end to end (data → build → deploy → impact) with pointed follow-ups — quantified impact is explicitly what they listen for.
- AI system design. Design a RAG service over a large corpus, a multi-tenant OpenAI-compatible gateway, an eval platform, an agentic workflow system. This is the leveling round — the framework lesson is the drill.
- HR / team-match (~30–60 min). Reported questions: the 20-days-of-work-in-2-days deadline scenario, handling a task nobody can help you with, an urgent customer issue with a 1-hour SLA, five-year plan, relocation/RTO willingness (Qualcomm went 5 days in-office from Sept 2025 — have your answer ready). They listen for process-thinking and quantified impact, not heroics.
Know where you're being slotted before you walk in — the design round is graded against the level, and the offer math is band-driven.
The ladder (India): Associate Engineer (0 YOE, B.Tech) → Engineer (0–2 YOE, M.Tech entry) → Senior Engineer (2–5) → Staff (6–10) → Senior Staff (10–15) → Principal. Specialization can shift this ±1 level; these are inferred from salary data, not published bands.
Bengaluru compensation (levels.fyi + reported offers, 2024–26; treat as calibration, not gospel):
Reported offer anatomy: base + joining bonus (₹5–10L appears repeatedly, incl. a fresher ML offer at ₹19.5L base + ₹9.6L JB ≈ ₹40L CTC) + RSUs (~$16K over 3–4 years at Engineer level) + 8–30% performance bonus + ESPP (15% of base). Negotiation headroom is real but bounded: ±3–5% on base at mid-levels, joining bonus is the flexible component, RSUs move only against competing offers. Timeline: offer 1–2 weeks after the final round; ML/AI roles carry a ~5–10% premium. Appraisals are twice-yearly and rating-driven.
Assumes ~3 focused hours on weekdays, ~5 on weekends (~25 h/week). Every week keeps a daily DSA block — it's the filter round, and consistency beats bingeing. The practice room runs the drills and timed mocks below in-app; the DSA track is the expanded pattern-by-pattern version of the daily block.
The applied-AI core. By Friday you should have working code, not notes.
- DSA (90 min/day): trees and linked lists this week — the Qualcomm-reported cluster: LRU Cache, Merge k Sorted Lists, Reverse Nodes in k-Group, Binary Tree Max Path Sum, plus string compression and Single Number as warm-ups. Narrate edge cases aloud as you code — that's what they grade, and they'll feed you hints: practice using a hint gracefully.
- Agents: the ReAct loop (build it by hand — the most common applied interview topic), tool use & MCP (MCP is now a standard interview probe — know what it standardizes and its security risks), planner-executor.
- RAG: foundations, chunking & embeddings, hybrid search + rerankers. Skim advanced retrieval for vocabulary (ColBERT, GraphRAG, HyDE, contextual retrieval).
- Context engineering: the paradigm — one read, it upgrades every answer you give ("prompt engineering is dead; context engineering is the 2026 framing" is now literally an interview talking point).
- Deliverable: a small FastAPI service exposing (a) a tool-using agent loop and (b) a RAG endpoint over a real document set, with hybrid retrieval. This becomes a live artifact for the project deep-dive.
Learn inference as the operator of a serving stack, not the author of kernels.
- DSA (60–75 min/day): graphs week — Number of Islands, Rotting Oranges, Course Schedule, plus bit-manipulation quickies (power-of-two, single number, bit reversal — all reported).
- Serving concepts: inference economics — prefill vs decode, TTFT vs TPOT (the single highest-leverage lesson for this role), KV cache & paged attention, quantization at decision level, speculative decoding (Qualcomm pairs it with MX formats for ~4× decode — be able to explain why the gains multiply), parallelism & distributed serving for the sharding vocabulary.
- Hands-on infra: run vLLM locally against a small model; hit it with the OpenAI client; watch continuous batching under concurrent load; stream via SSE. Containerize your Week-1 service; deploy both to a local Kubernetes (kind/minikube) with an HPA. Add observability: request logs with TTFT/TPOT, a queue-depth metric.
- Deliverable: a one-page architecture note for "my agent/RAG service behind an OpenAI-compatible gateway on Kubernetes" with honest numbers (TTFT, tokens/sec under 1/10/50 concurrent requests).
The leveling week.
- DSA (60 min/day): mixed timed sets now — one medium in 25 minutes, then review.
- System design: the framework, then one case study per day as a spoken 35-minute mock: RAG over 10M docs, an AI coding agent, an eval platform. Practice out loud — design rounds are a speaking skill.
- Evals: why evals & the loop, LLM-as-judge, the CI eval harness. "How do you know it works?" now appears in applied loops as often as architecture questions — be the candidate with a real answer.
- Platform literacy (one evening): re-read Section 3 until you can reproduce the hardware table, the software layer-cake, and three interview-grade facts from memory. Skim the QEfficient README and the vLLM-qaic page; skim the arXiv study (2507.00418) so you hold one evidence-based opinion — strengths (perf/watt, capacity) and the honest caveat (throughput-per-rack).
- Systems-fundamentals refresher (one evening, insurance): threads vs processes, mutex vs semaphore, print-odd-even-with-two-threads (write it!), the four deadlock conditions, malloc/calloc, endianness, what causes a segfault. Every one of these is reported from real Qualcomm India rounds.
- Deliverable: three recorded system-design run-throughs + a 10-flashcard platform deck.
Convert knowledge into interview performance.
- Mocks: two full timed DSA sets (2 problems / 70 min each) mid-week; two mock design interviews using Week-3 case studies; one mock applied-depth round where you defend your Week-1/2 deliverables against hostile follow-ups. The practice room runs all of these in-app — timed DSA mock sets and AI mock rounds with self-grading.
- Project narrative: script the end-to-end story of your best real project — problem → data/design → build → deployment → quantified impact — in 4 minutes, with depth behind every sentence. Map each resume bullet to the role's three jobs (applied AI / app dev / AI infra). Interviewers explicitly listen for numbers ("cut retrieval failures 15%→4% on a 300-case golden set").
- Behavioral + HR: prep the reported scenarios (impossible deadline, no-help task, 1-hour customer SLA, five-year plan, relocation/5-day RTO) in STAR shape. Prepare questions that show you understand the org in July 2026: "How does the Inference Suite roadmap split between HUMAIN-specific and general cloud offerings?", "Where does this team sit relative to the Dragonfly agentic-AI story and the AI250's disaggregated serving?", "What does the app-layer on-call look like?"
- The offer week: re-read Section 5. Know your level ask, know that joining bonus is the flexible lever, and get competing signals in writing if you have them.
- Rest the day before. Seriously.
[IC4] When would you deliberately NOT build an agent? Walk me through the agent-vs-workflow call for an internal support assistant.
Start from the definition: an agent is a model directing its own control flow in a loop with tools; a workflow is developer-defined steps with LLM calls inside. Choose by predictability of the path: if support tickets follow a known taxonomy — classify, retrieve policy, draft reply, escalate on low confidence — that's a workflow: cheaper, debuggable, evaluable step by step, and its failure modes are bounded. Reach for an agent only where the path genuinely can't be enumerated (open-ended troubleshooting across many tools). The senior answer names the cost asymmetry — agents multiply token spend and compound per-step error — and proposes starting as a workflow with one agentic escape hatch, plus an eval harness before widening autonomy. Drilled in planner-executor and the ReAct loop.
[IC4] Your RAG service answers confidently but wrongly on ~15% of queries. Diagnose it: where in the pipeline do you look, and what do you change first?
Decompose before touching anything: is retrieval failing (right answer not in context) or generation failing (right answer in context, model ignores or garbles it)? Sample the failures and label which. If retrieval: check chunking (are answers split mid-thought?), then run hybrid search — BM25 + dense — because pure embeddings miss exact identifiers and rare terms, then add a reranker for precision at small k. If generation: tighten grounding instructions, cite-then-answer formats, and shrink context to what's relevant (lost-in-the-middle is real). The differentiating move is measurement: build a small golden set and track retrieval hit-rate and faithfulness separately, so "improved" means numbers, not vibes. Drilled in hybrid search & rerank and RAG evaluation.
[IC4] You're building an OpenAI-compatible endpoint on top of an inference rack. Why do continuous batching and streaming matter, and which latency metrics does the app layer actually own?
Decode is memory-bandwidth-bound and one request can't saturate an accelerator, so the server interleaves many requests' decode steps — continuous batching admits and retires requests token-by-token instead of waiting for the slowest in a static batch; it's the throughput lever, and it lives in the serving layer (vLLM-class schedulers), which is exactly the layer this team builds. Streaming exists because total generation time is irreducibly long; sending tokens as they're born converts a 20-second wait into a sub-second perceived response. The app layer owns TTFT-as-experienced (queueing, admission, routing, prompt size, cache hits) and everything around the model: connection handling, backpressure, retry policy. It does not own per-token decode speed — that's the accelerator and the batch size you chose. Quote your SLOs as TTFT p95 and TPOT p95, not a single "latency." Drilled in inference economics.
[IC5] Design a multi-tenant, OpenAI-compatible inference service on a rack of AI200-class accelerators.
Skeleton: stateless API gateway (authN, quotas, token-bucket rate limits per tenant) → router (model + version resolution, tenant isolation policy) → per-model serving pools running continuous batching → shared services: KV/prefix cache, response streaming, usage metering. The design tensions to surface unprompted: (1) fairness vs utilization — per-tenant queues with weighted admission so one bulk tenant can't starve interactive ones; (2) prefill/decode interference — priority scheduling now, disaggregation as scale grows (the AI250's explicit design target); (3) autoscaling on queue depth and TTFT p95, never GPU-util alone, with model cold-start (minutes to load weights) making pre-warmed pools part of the design; (4) failure isolation — health-check wedged replicas out, idempotent retries across replicas, circuit-break pathological tenants; (5) observability — per-tenant TTFT/TPOT/queue-time and tokens-per-second-per-card, because on this platform the pitch is cost efficiency. Capacity-rich cards (768 GB) shift the math toward fewer, bigger replicas with longer contexts before sharding. Drilled via the framework and parallelism & distributed serving.
[IC4] A prompt or model-version change is about to ship in your AI product. What does the CI gate look like that decides whether it can?
A golden dataset of real, versioned input cases; deterministic assertions where possible (structure, required facts, refusal behavior) and LLM-as-judge with a rubric where not — with the judge itself validated against human labels so it's not vibes-judging-vibes; pass/fail thresholds per capability, not one blended score, so a regression in one skill can't hide behind gains in another; run on every prompt/model/retrieval change like unit tests. Ship behind a canary with online quality telemetry because offline sets under-cover real traffic, and log every failure into the next golden-set revision. The senior signal is treating prompts and model versions as code: diffed, reviewed, gated, rollback-able. Drilled in the CI eval harness and LLM-as-judge.
[IC3] Your LLM backend times out on 2% of requests. Design the timeout/retry/fallback policy for the gateway in front of it — and tell me where retries make things worse.
Timeouts must be phase-aware: a TTFT timeout (is it generating yet?) separate from an idle-stream timeout (tokens stopped), never one flat total-duration clock — long generations are legitimate. Retries: only on connection errors and TTFT timeouts, with capped exponential backoff + jitter, an idempotency key so double-submission can't double-charge or double-act, and a retry budget (e.g. ≤1 retry, only below a load threshold). Where retries hurt: mid-stream failures (you'd re-generate and the user re-reads a different answer), and overload — timeouts caused by queue saturation turn retries into a self-inflicted DDoS; that's what circuit breakers and load-shedding are for. Fallback ladder: smaller/faster model → cached or canned response → honest error, chosen by product criticality. This is the bread-and-butter reliability question for the AI-infra half of the role.
[IC4] Fine-tuning vs RAG — your product needs domain knowledge the base model lacks. How do you choose, and what hybrid do you actually ship?
Split the need into knowledge vs behavior. Facts that change, need citations, or must be permission-scoped → RAG: updates are re-indexing, provenance is free, and per-tenant access control is enforceable at retrieval time. Stable style, format discipline, domain vocabulary, tool-use patterns → fine-tuning (LoRA-class): it changes how the model acts, cheaply, but bakes knowledge in — stale the day training ends, unauditable, and un-scopable per tenant. The trap answer is picking one; the shipped answer is almost always RAG for knowledge + a light adapter for behavior, with two more moves before any training: serious prompting/context engineering (often closes the gap alone) and an eval set that defines "better" — you can't justify a fine-tune you can't measure. Cost framing seals it: RAG is an ops cost that scales with corpus churn; fine-tuning is a capex spike plus a re-training tail every time the base model or the domain moves. Drilled across RAG foundations and the fine-tuning pillar.
[IC5] Your AI product serves a million queries a day and the token bill is unsustainable. Cut cost meaningfully without visibly hurting quality.
Work the ladder from free to invasive, measuring at each rung. (1) Caching: exact-match and semantic response caches, and prompt/prefix caching for the shared system-prompt + few-shot header — at 1M queries/day the head of the distribution is fat, and prefix caching alone can cut prefill cost dramatically. (2) Context diet: trim the prompt (instructions, retrieved chunks, history compaction) — tokens are the bill; most products carry 2–3× context bloat. (3) Model routing: classify queries by difficulty and send the easy majority to a small model, escalating on low confidence or user retry — the biggest single lever, often 5–10× on the routed share. (4) Output discipline: max-token caps, structured outputs instead of prose, stop sequences. (5) Serving-side: batching, quantized serving, right-sized hardware — on capacity-rich accelerators, bigger batches per card is exactly the economics this platform sells. The senior frame: state cost-per-query as the metric, gate every rung behind the eval harness so "no visible quality loss" is a measured claim, and name the observability you need first (per-feature token attribution) because you can't cut what you can't see. Drilled in inference economics and prompt caching.
[IC4] What problem does MCP actually solve for agents, and what new security risks does it introduce?
The problem: N agents × M tools previously meant N×M bespoke integrations — every assistant hand-wired its own connectors. MCP standardizes the tool side: a server declares tools/resources/prompts in a uniform schema, any MCP client can discover and call them at runtime, so integrations become N+M. That's why it won: it's boring plumbing, donated to a neutral foundation, and it turns "add a capability" into "point at a server." The risks are the flip side of dynamic capability: prompt injection via tool results (a fetched document instructs the model), tool poisoning / rug-pulls (a server's tool description changes after approval), confused-deputy and excessive-agency failures (the agent wields your credentials across servers), and supply-chain trust in third-party servers. Mitigations to name: least-privilege scoping per server, human confirmation on state-changing tools, treating all tool output as untrusted input, allow-listed servers with pinned versions, and sandboxed execution. Drilled in tool use & MCP.
[IC4] Walk me through taking a Hugging Face checkpoint to a running OpenAI-compatible endpoint on a Qualcomm Cloud AI accelerator.
Four steps, all in the Efficient Transformers library: QEFFAutoModelForCausalLM.from_pretrained() pulls the HF checkpoint (pre-quantized AWQ/GPTQ variants load directly); .export() produces an ONNX graph with the KV cache re-expressed as fixed-shape tensors (that's the trick that lets a dynamic-length decoder compile ahead-of-time); .compile() invokes the qaic compiler to emit a QPC (Qualcomm Program Container) binary for a given core count, batch size, and precision (e.g. MXFP6 weight compression); then either the Python Session API runs it directly, or — for serving — a vLLM with the qaic backend consumes the pre-compiled QPC and exposes the standard OpenAI-compatible API with continuous batching and prefix caching. Two details that signal real understanding: compilation is offline and can take a while, but QPCs are portable artifacts you version and ship like build outputs; and the KV cache lives in the card's own LPDDR between decode steps, so the host never shuttles it. Sources: the QEfficient README and vLLM-qaic docs.
[IC5] Speculative decoding: why does it speed up decode, when does it hurt, and why does Qualcomm pair it with MX formats?
Decode is memory-bandwidth-bound: each step reads all the weights to emit one token, leaving compute idle. Speculative decoding spends that idle compute — a small draft model proposes K tokens, the target model scores all K+1 positions in one forward pass (one weight read instead of K), and modified rejection sampling keeps the output distribution exactly the target's — it's a latency trick, not an approximation. It hurts when acceptance drops: hard or out-of-domain text (draft disagrees with target), or high-temperature sampling — you then pay the draft's cost plus wasted verification for ~1 accepted token. It also competes with large-batch serving: batching already fills the idle compute, so SpD's headroom shrinks; it shines at low-batch, latency-sensitive decode. The Qualcomm pairing is multiplicative because the two attack the same bottleneck differently: MXFP6 shrinks the bytes-per-weight-read (~2×), SpD amortizes reads across tokens (~1.5–2×) — together ~4× reported on Cloud AI 100. One implementation gotcha worth volunteering: rejected draft tokens must be purged from the KV cache or it silently corrupts. Drilled in speculative decoding.
[IC3] Print odd and even numbers alternately using two threads. What synchronization primitive do you reach for, and what breaks without it?
This is Qualcomm's favorite concurrency screen. The clean answer: shared counter + a mutex with a condition variable — each thread locks, waits on the condition until it's its turn (counter % 2 matches its parity), prints, increments, signals the other, unlocks. Two semaphores (each thread releasing the other's) is the equally clean alternative; in Python, two threading.Events or a Condition work. Without synchronization you get interleaving corruption: both threads read the same counter value, print out of order or double-print — a race on check-then-act. The follow-ups they actually probe: why a plain flag + busy-wait is wrong (burns CPU, and without memory barriers the write may never be seen), why wait() must sit in a while loop, not an if (spurious wakeups re-check the predicate), and what deadlocks if you signal before wait. Say "check-then-act must be atomic with the wait" and you've passed.
- Don't prep the wrong altitude. This role is on top of the accelerators. A week on QNN compiler internals is a week not spent on the system-design round that levels you.
- Don't let DSA slide because the role is "applied." Qualcomm India's loop filters on LeetCode-medium regardless of team — linked lists showed up in 8 of 12 sourced experiences. Daily, all four weeks.
- They grade code quality and hint-usage. Names, edge cases, modularity — narrate them. And when the interviewer offers a hint, take it visibly and build on it; hint-driven interviewing is their reported style.
- Don't skip the concurrency staples. Print-odd-even-with-two-threads, mutex vs semaphore, deadlock's four conditions — reported again and again, even for applied roles.
- "How do you know it works?" is the applied-AI trap question. If your answer to any build question doesn't end in an eval, it's incomplete.
- Speak in TTFT/TPOT and tokens/s/W. One latency number marks you as API-consumer; two latency numbers plus a cost metric marks you as the person who runs the service.
- Cite the 2026 story, not the 2025 one. Dragonfly, the Meta CPU deal, AI300, HUMAIN going live — a candidate quoting last October's press release sounds prepped; one quoting June's investor day sounds interested.
- Volunteer the honest tradeoff. "Capacity-first cards win on cost-per-token, not raw throughput-per-rack" — pre-empting the obvious skeptical question converts it from a gotcha into your credibility.
- A referral accelerates; it doesn't exempt. Sourced reports: priority queue and 1–2 weeks faster, but the OA and full loop still happen. Prepare for all of it.
- The project story needs numbers. Interviewers explicitly listen for quantified impact. "Cut retrieval failures 15% → 4% with hybrid + rerank on a 300-case golden set" beats any adjective.
- Know your level before you negotiate. Bands are real; base moves ±3–5% at mid-levels, joining bonus is the flexible lever, RSUs move only against competing offers in writing.
Flashcard. The Qualcomm Datacenter AI applied role = build the cloud layer that sells the racks. Loop: DSA filter (linked lists!) → systems staples (threads/deadlock) → applied LLM depth (agents/RAG/evals/serving-as-consumer) → AI system design levels you → HR (deadline scenarios, 5-day RTO). Month = W1 agents+RAG (build it), W2 serving+infra (run it), W3 design+evals+platform (defend it), W4 mocks+story+offer (sell it). Platform in one breath: AI200 (2026, 768 GB/card, capacity play) → AI250 (2027, HBC near-memory, bandwidth play) → AI300 (2028), Dragonfly C1000 CPU for Meta (H2 2028), HUMAIN 200 MW live, Inference Suite + vLLM-qaic + QEfficient on top — and the pitch is TCO per token, not FLOPs.
Start with the June 2026 investor-day coverage (Dragonfly portfolio) and the AI Inference Suite page — they define the org you're joining. Then the QEfficient README and vLLM-qaic docs (the pipeline you'll be asked to narrate), the two Qualcomm engineering blogs (microscaling, speculative decoding) for quotable numbers, and the arXiv serving study for the independent view. The India interview-experience collections (GeeksforGeeks, CodingKaro, Glassdoor) are what Sections 4–5 are calibrated against — skim two or three raw reports yourself to absorb the texture of how these rounds actually feel.