GenAI interview questions with model answers: what the interviewer is actually probing
I did a run of AI engineer interviews this month: recruiter screens, technical screens, a client round. Same shape every time. Nobody asked me to define RAG. Every panel asked some version of "what breaks, how do you know, what do you do about it."
That is the whole game. The definition questions are a warm-up; the real questions are operational.
The questions that actually come up
A sample from the 46 I collected and answered:
- Your RAG answers are fluent and wrong. Walk me through finding out whether retrieval or generation is at fault.
- What is a similarity score threshold, where does it live, and how did you pick the number?
- You have BM25 and vector search. How do you merge two ranked lists, and what does reciprocal rank fusion actually compute?
- The judge model says 4.9 out of 5 and the customer says the answers are useless. Who is right?
- An agent calls the same tool three times with the same arguments. Where is the bug?
- Rate limits: a 429 at 2 a.m. with no human awake. What did you build for that before it happened?
- When does LoRA beat prompting, and how would you prove it with numbers?
Notice the shape: every one of them has a symptom, a way to detect it, a cause, a fix and a way to prevent it. That order is how I wrote every answer, and it is the order a senior engineer thinks in. Interviewers can hear the difference between a lived story and a borrowed definition.
What "model answer" means here
Not a paragraph you memorize. Each answer comes with the mechanism drawn out, and the key drawings are interactive: the retrieval scores, the prompt assembly, the judge rubric, the token budget. You rehearse by saying the answer out loud while the picture moves, then rewrite it around your own projects. An interviewer can tell a borrowed story from a lived one, so the kit tells you to swap in yours.
What is in the kit
Five self-contained interactive pages, no login, works offline:
- GenAI, explained simply. Transformers, RAG (chunking, embeddings, top-k, top-p, temperature with real softmax numbers), agents and orchestration (LangGraph, Microsoft Agent Framework, Bedrock Agents), LLMOps, the tuning ladder. 46 questions with full answers, 45 troubleshooting cards.
- GenAI with the code. A full RAG program in LangChain line by line, agents with tools and memory, an LLMOps program that runs for real (tokens, 429 backoff, caching, tracing, a deploy gate), a LoRA program. 42 code blocks.
- Two hello worlds walked through. A RAG program and an agent program, every line beside a drawing of what it does to real data. 30 frames with the values at each step.
- The refresher.
- Visual answers. Swimlanes, pipeline diagrams, troubleshooting flows.
GenAI Interview Kit — $39 on Whop · both kits together for $69 on the kits page.
One free answer, since you read this far
"How did you pick the threshold?" The honest answer is: I guessed from the embedding model's score bands (unrelated text lands near 0.5, related near 0.7), measured five real questions and five junk questions with the floor off, read the lowest real score and the highest junk score, and set the floor just above the junk. Then I wrote the reason next to the constant in the code, and I re-measure when the model or the notes change. Say that in an interview and watch the panel sit up.
Related notes
- Where are AI-103 practice questions with explanations, not just answers?
Every AI-103 site I found is a question bank. What the exam actually tests, why a dump fails on mechanisms, and the interactive deck I built instead. - I scored 21 RAG explainers and courses with a decision model. None teach evaluation from zero.
21 RAG explainers, courses and videos scored on evaluation mechanics, retrieval mechanics, from-zero and real numbers. Method, table and raw files. - The five rungs of agent engineering
Prompt, context, loop, graph, ownership. Most teams stall on rung one and wonder why their AI never leaves the demo. A ladder for what to fix next.