Interview prep
Notes for backend and GenAI interview loops
50 questions across three sections, plus a DSA pattern reference with a Python template and three classic problems for each. The first two of each are free to read; the rest need a Pro pass (from Rs 199). Click a question to expand it, tick "Reviewed" to track your own progress - saved only in your browser.
Practising problems too? The DSA sheet is free, with hints and tested solutions for a curated set of CSES problems.
GenAI & LLM engineering
RAG, chunking, embeddings, evals, hallucination mitigation, prompt injection, structured output, agents and tools, cost and latency, and determinism.
What is RAG, and why not just fine-tune the model on your data instead? Reviewed
RAG (Retrieval-Augmented Generation) retrieves relevant text at query time and puts it into the prompt, instead of baking facts into the model's weights. Fine-tuning changes how a model behaves or writes, not reliably what it "knows" - it can still hallucinate, and it goes stale the moment your data changes, so you retrain. RAG updates instantly (re-index the new document) and gives you citations back to the actual source, which matters for trust and debugging. The two aren't exclusive: fine-tuning is good for style, format, or narrow task behaviour, while RAG is the better fit for facts that change over time or that need to be traceable to a source.
How do you choose a chunk size and overlap for a retrieval index? Reviewed
Match chunk size to query granularity: short, self-contained answers (FAQ entries, glossary terms) want small chunks; broad conceptual questions want bigger ones so context isn't split across boundaries. A common starting point for prose is 500-1500 characters with 10-15% overlap, then tune against real queries rather than guessing. Overlap exists specifically so a fact sitting right at a chunk boundary doesn't get fragmented in both neighbors. Split on natural boundaries (headings, paragraphs) before falling back to a hard character limit, so a chunk doesn't cut a sentence in half.
Backend & system design
Idempotency, rate limiting, queues and retries, caching, DB indexing, consistency, and two whiteboard-level system designs.
What makes an API endpoint idempotent, and why does it matter for retries? Reviewed
An idempotent endpoint produces the same end state (and ideally the same response) no matter how many times the identical request is applied - calling it twice is safe. It matters because clients and networks retry: a request can time out after the server actually processed it, and the client has no way to know that, so it retries. Without idempotency, that retry can double-charge a payment or create a duplicate record. The standard mechanism is an idempotency key the client generates once per logical operation and sends on every retry; the server stores which keys it has already processed and returns the original result instead of redoing the work.
How would you design a rate limiter? Reviewed
Token bucket is the usual default: each client has a bucket that refills at a fixed rate up to a cap, and a request consumes a token, which naturally allows short bursts while still enforcing an average rate. Sliding window (or sliding window counter, a cheaper approximation) avoids the "burst right at the boundary of two fixed windows" problem a naive fixed-window counter has. For a distributed system, back it with a shared store (Redis, with an atomic increment/expire or a Lua script for the bucket logic) so the limit is enforced correctly across multiple API instances, not per-instance.
Cloud & Kubernetes / SRE basics
Lambda cold starts, Kubernetes probes and CrashLoopBackOff, HPA vs KEDA, Terraform state/locking/drift/import/modules, plan vs apply in CI, and blue-green vs canary.
Why does an AWS Lambda cold start happen, and how do you reduce its impact? Reviewed
A cold start happens when Lambda has to provision a fresh execution environment (download the code, start the runtime, run top-level init code) because no warm one is available - the first invocation after a deploy, after scaling up, or after enough idle time. Mitigations: keep the deployment package small and dependencies lean so init is fast; do expensive setup (DB connections, SDK clients) at module scope so it's reused across warm invocations instead of per-request; use provisioned concurrency for latency-sensitive endpoints where a cold start is unacceptable; and pick a runtime/language with a fast cold-start profile for latency-critical functions.
What are Kubernetes liveness, readiness, and startup probes? Reviewed
A liveness probe asks "is this container still functioning" - failing it gets the container restarted, which is the right response to a genuinely deadlocked process. A readiness probe asks "should traffic be sent to this container right now" - failing it just removes the pod from the Service's load-balancing endpoints without restarting it, which is correct for a pod that's temporarily overloaded or still warming a cache. A startup probe covers slow-starting containers: it holds off liveness/readiness checks until the startup probe itself succeeds, so a slow but healthy boot doesn't get killed by an impatient liveness probe.
DSA patterns
When-to-use signals, a Python template, and three well-known classic problems per pattern (names and links only - the point is recognizing the pattern, not a copied problem statement).
Two pointers Reviewed
A sorted array or string, and a question about a pair/triplet meeting a condition (sum, comparison). Two indices moving toward or away from each other in one pass instead of nested loops.
def two_pointers(arr, target):
left, right = 0, len(arr) - 1
while left < right:
total = arr[left] + arr[right]
if total == target:
return [left, right]
elif total < target:
left += 1
else:
right -= 1
return [] Classic problems
Sliding window Reviewed
A contiguous subarray/substring question asking for the longest/shortest/count meeting a condition. A window that expands and contracts over one pass, avoiding recomputation from scratch each time.
def sliding_window(s):
left = 0
window = {}
best = 0
for right, ch in enumerate(s):
window[ch] = window.get(ch, 0) + 1
while not is_valid(window):
window[s[left]] -= 1
left += 1
best = max(best, right - left + 1)
return best Classic problems