Free sample: page 05 of the GenAI Interview Kit (visual answers). The full kit — 46 questions with model answers, 77 troubleshooting cards, 42 code blocks, 30 walkthrough frames — is $39 at niktechai.com/kits.
Visual answers — process diagrams, troubleshooting flows, interview questions
Every item is a drawing first: a TL;DR line, the diagram (animates on its own, slowly), three beats to say out loud. No paragraphs. Read any section on its own.
1Agent orchestration — the process, drawn, and where enterprises use it
Every orchestration pattern below is the same six-lane loop — an orchestrator routes work to specialist agents and tools, a critic checks the answer, and a trace records every step — reused with different agents plugged in.
The orchestration process, as a swimlane sequence
An orchestrator routes each step to one agent or tool, a critic checks the draft, and every arrow writes to a trace.
Orchestrator — decides WHICH agent runs next, not the model itself
Critic gate — sends the draft back on missing citations, capped by a guard
Trace — every numbered arrow above also logs step id + latency + tokens
Say the orchestrator is just a router with a checklist — the critic is the part most demos skip.
The same loop, run for real: pump P-7 hello world
Four agents in a fixed order turn a raw log number into a plain-English service alert with the vendor to call.
hours agent — reads the control-room log, returns 4120
policy agent — pulls the 4000-hour threshold from D1 via RAG
eval — groundedness / relevance scoring on real traffic, like a production A/B tests
trace + human gate — every step logged, and a person still owns the risky exits
Say a demo is the circle in the middle — production is everything drawn around it.
The same drawing, mapped onto three frameworks
Every framework fills the same five boxes — orchestrator, agent, tool, memory, safety — with its own names.
Agent Framework — Workflow + Executors + Threads — what I used for the pump/claims examples
LangGraph — nodes/edges are the graph, checkpointer is the memory
Bedrock Agents — action groups are the tools, guardrails are the safety layer
Say learn the five boxes once — every framework question becomes "which box is that?"
2RAG pipeline — the process, drawn
Ingest path, query path, and where AWS sits — the same shape he built on a production RAG build and the Standard, drawn stage by stage.
The two paths in one drawing
Ingest fills the index offline; query hits it live — they only ever meet at the index.
ingest — batch job fills the index, offline
query — live request reads the same index
shared index — the only handoff point between the two
Say "Ingest and query are two separate pipelines that only touch at the index."
Ingest path — the knobs on each stage
Every ingest stage has a setting that changes recall later, and each one gets versioned.
chunk_size/overlap — 1000/150 was a production A/B-tested pick
embedding version — pinned, or old and new vectors mismatch
index version — bumped on every re-embed, never overwritten in place
Say "On a production RAG build we A/B-tested chunking and retrieval settings against groundedness scores before locking them."
Query path — knobs, latency, and cost
Retrieval and generation each have their own knobs, and each stage adds latency or cost.
k / top_n — retrieve wide (k=8), rerank narrow (top_n=4)
temperature — 0 for factual, grounded answers
latency budget — model call is almost always the biggest slice
Say "Reranking narrows what actually goes in the prompt, which controls both cost and hallucination risk."
Hybrid retrieval, drawn
Keyword and vector search run in parallel, then get fused with RRF before reranking.
BM25 — catches exact terms vectors sometimes miss
RRF — fuses two rank lists without needing comparable scores
rerank — a cross-encoder pass to narrow to top_n
Say "Hybrid retrieval covers the case where vector search misses an exact keyword like a part number."
GraphRAG variant, drawn
Text becomes triples in a graph so a question can hop across relationships, not just similarity.
triples — subject-relation-object pulled from configs/RFCs
hops — traverse relationships a flat chunk match can't reach
when — multi-entity questions, not single-fact lookups
Say "On a network-operations build we ran GraphRAG over device configs and RFCs in Neo4j so troubleshooting could hop from a device to the routing policy that touches it."
AWS placement, drawn
Mapped, not claimed in production — ten years of AWS platform work, Bedrock/SageMaker not yet shipped.
S3 → Bedrock KB — managed ingestion, or Lambda for custom parsing
OpenSearch Serverless — the vector store, scales without cluster ops
Guardrails + IAM — content filtering and least-privilege per hop
Say "I map this from ten years of AWS platform work — Bedrock and SageMaker specifically I haven't shipped in production yet."
The Azure twin the author built
Same shape on Azure, shipped in production on a production RAG build with managed identity end to end.
integrated vectorization — AI Search does chunk+embed in the indexer
managed identity — no stored keys between app and OpenAI
App Insights — groundedness/relevance tracked per A/B run on a production RAG build
Say "This is the RAG stack I actually shipped on a production RAG build — Blob through AI Search into Azure OpenAI, with A/B tests on chunking and prompts."
Multi-tenant RAG, drawn
A tenant filter goes into every retrieval call so one query can never see another tenant's chunks.
tenant filter — applied at retrieval, never trusted from the prompt
index-per-tenant — hardest isolation, more ops overhead
metadata ACL — cheaper, filter enforcement is the whole security model
Say "The tenant filter has to live in the retrieval call itself, not in the prompt, or one bad query leaks another tenant's data."
The freshness loop
A change event triggers re-embedding, a version bump, and a golden-set rerun before the new index goes live.
trigger — a source doc changes, not a fixed schedule
version bump — new index never overwrites the live one mid-query
golden set — same idea as a production A/B groundedness checks, gating promotion
Say "A version bump plus a golden-set rerun is what stops a bad re-embed from silently going live."
RAG at enterprise scale — the 120 GB design
120 GB of Azure AI Search knowledge on a science-discovery platform, designed for partitioned throughput and locked-down access.
partitions — split the index so ingestion throughput scales
private network + RBAC — VNet-bound, access locked to role
Foundry IQ — the knowledge layer feeding a six-agent orchestration engine
Say "At a science-discovery platform I designed the 120 GB Azure AI Search knowledge layer that feeds the six-agent architecture."
3Operational troubleshooting — decision flows you can follow under pressure
Ten flows: symptom → the exact check → the concrete fix, so you never freeze on an ops question.
"The answer is wrong" — retrieval vs generation split
Print the retrieved chunk ids first — that one check splits every "wrong answer" into two different fix paths.
print chunk ids — cheapest diagnostic, one debug line before any prompt work
right chunk missing — it's a retrieval bug, tuning the prompt won't help
right chunk present — it's a generation bug, fix the prompt/citation rule
Say "first thing I check is whether the right chunk was even retrieved — that tells me which half of the pipeline is broken."
"The answer says I don't know but the doc exists"
Four separate causes look identical from the outside — check each in order: parsing, stale index, embedding mismatch, filter.
parsing — a scanned page or table can silently drop to empty text
stale index — doc added to the folder but ingest job hasn't run since
filter too strict — a tenant/date metadata filter is excluding it
Say "false negative is usually parsing, staleness, embedder mismatch, or an over-tight filter — I check those four in order."
"Latency is high" — where in the pipeline?
Split the clock into retrieval time and model time before touching any setting — they have different fixes.
retrieval-ms — index size, k, and reranker are the levers
model-ms — max_tokens and context length are the levers
neither moved it — check network hops and provider throttling next
Say "I trace retrieval-ms separately from model-ms before I touch a single setting — they don't share a fix."
"Cost spiked" — trace the token bill
Cost is tokens-per-call times calls; walk context growth, retries, and cache misses before blaming the model.
context growth — chunks or history silently getting longer over time
cache miss rate — repeated prompts not hitting the provider's prompt cache
model choice — routing everything to the biggest model by default
Say "cost is tokens per call times calls per day — I break the bill down into those two before I touch the model."
"429 / throttling"
RPM and TPM are two separate limits — find which one tripped, then back off with jitter and cap concurrency.
RPM vs TPM — the error body/header names which limit tripped
backoff with jitter — the universal retry fix, avoids a retry stampede
PTU / provisioned throughput — the real fix once traffic is sustained
Say "429 first tells me whether I hit requests-per-minute or tokens-per-minute — they need different fixes."
"The agent loops / never finishes"
The trace shows the same tool call repeating — that's the tell, not a guess, so read the trace before changing anything.
same tool call repeats — cap max_steps and feed the tool error back as text
routing keeps flipping — set temperature 0 on the routing decision specifically
schema validation — reject bad args before they ever hit a real tool call
Say "if the trace shows the same tool call over and over I cap max_steps and route the error back in as text, not a stack trace."
"Eval score dropped after a deploy"
Diff exactly what changed — prompt, model, or index — then roll back that one version, not the whole release.
diff first — never guess which of prompt/model/index moved
golden set — the same fixed question set, rerun per candidate version
roll back one thing — not the whole release, just the culprit version
Say "eval drop after a deploy means diff what changed and rerun the golden set on each piece — roll back just that one."
"Users report a leak across tenants"
Treat it as a P0 security incident: verify the filter clause, per-tenant boundary, then audit the trace for who saw what.
filter clause — a missing tenant_id in the retrieval query is the #1 cause
shared index — biggest structural risk, per-tenant index removes the class of bug
treat as incident — audit-trace and disclose, don't quietly patch
Say "a cross-tenant leak is an incident first — I check the filter clause, the index boundary, then pull the audit trace."
"Model output is not valid JSON"
Enforce structured output at the API level first — a parser retry loop is the fallback, not the fix.
retry with the error — feed the exact parser exception back, one retry
lenient parser — strip markdown fences and trailing text as a last resort
Say "structured output mode at the API level first, then one retry with the parser error attached, lenient parsing as the fallback."
The on-call card — five numbers, first
One dashboard glance: p95 latency, error rate, tokens/call, cost/day, and "no source found" rate.
five numbers — p95 latency, error rate, tokens/call, cost/day, "no source found" rate
one red number — routes straight into the matching flow above, no guessing
check this first — before opening logs or asking anyone what happened
Say "on-call, I check five numbers first — whichever one's red tells me which flow to run."
Transformers — what actually happens to a token
Every token becomes a vector, self-attention lets it read every other token, then the vector predicts the next token.
embedding — every token starts as a vector of numbers, not text
self-attention — each token's vector gets updated by weighing every other token
next-token prediction — stack N layers of this, then predict what comes next
Say "a transformer turns tokens into vectors, lets every token look at every other one through attention, stacks that N times, then predicts the next token."
4Transformer questions — each answered with a drawing
Nine core transformer concepts, each as one drawing you can point at instead of explaining in words.
Explain attention simply
Every word looks at every other word and votes how much it matters for the one being decided.
query — the word asking "what matters to me?"
weights — softmax scores that sum to 1 across all words
output — weighted blend, not one winner picked
Say attention is every word scoring every other word and blending them by that score.
What is a token and why is pricing per token?
A sentence is chopped into sub-word pieces first, and every piece in and out costs money.
sub-word — "4000" splits into "40"+"00", not one token
rate — output tokens usually cost more than input tokens
context — every prior message re-billed as input each turn
Say the model doesn't read characters, it reads sub-word tokens, and every one in or out is billed.
What is an embedding?
Text becomes a point in number-space where similar meaning sits physically close.
vector — a fixed list of numbers standing in for meaning
cosine — distance between two vectors = topical closeness
same embedder — query and chunks must use the identical model
Say an embedding turns text into a point in space, so "close meaning" becomes "close numbers."
Why does a long context cost more and get slower?
Every token attends to every prior token, so cost grows as n² while the KV cache keeps growing too.
n² — double the tokens, quadruple the attention math
KV cache — memory grows linearly with tokens kept live
practical fix — trim history, summarize, or use a smaller retrieval window
Say attention cost scales with the square of context length, and the KV cache eats memory the whole turn.
Temperature vs top-p
Temperature reshapes how peaked the probabilities are; top-p cuts the tail off after that.
temperature — divides logits before softmax, flattens or sharpens
top-p — keeps the smallest set of tokens whose probs sum ≥ p
Say temperature reshapes the probability curve, top-p then cuts off its tail.
Encoder vs decoder — why RAG uses two models
RAG pairs a reader that scores meaning with a writer that generates text, because they do opposite jobs.
encoder — bidirectional, good at similarity not generation
decoder — causal, generates fluent text one token forward
why split — a single model doing both is slower and worse at each
Say the embedder is the reader that finds relevant chunks, the chat model is the writer that turns them into an answer.
Why RAG instead of a bigger prompt or fine-tuning?
RAG keeps knowledge current at query time without paying to retrain or stuffing everything into every prompt.
bigger prompt — re-pays the n² attention cost every single call
fine-tuning — changes behavior, not facts; still goes stale
RAG — the doc pool updates, the model doesn't have to
Say RAG is the option where updating a document is cheaper than updating the model.
What happens at the context limit?
Depending on the API, hitting the limit either truncates, errors out, or you summarise history first.
hard error — the call fails outright, must retry shorter
silent truncation — oldest messages quietly dropped, easy to miss
summarise-history — best practice: compress, don't just cut
Say the three failure modes are error, silent truncation, or a summarization step you control.
What is max_tokens?
max_tokens caps only the output length, it never trims what you send in.
input side — governed by context limit, not max_tokens
output side — max_tokens is the only cap on generation length
ops risk — too low silently truncates a valid answer
Say max_tokens only limits what the model writes back, not what you send it.
5The rest of the job description, drawn
Everything else the posting asks for — platform architecture, governance, ops loops, and the lead's actual week — as one drawing each.
What does an enterprise AWS GenAI reference architecture look like?
A request crosses a private VPC through API Gateway → Lambda → Bedrock, hits Knowledge Bases + OpenSearch for retrieval, and every hop is logged and encrypted.
API Gateway → Lambda — thin orchestration layer, no logic in the model call itself
Bedrock + Guardrails — the model call and its input/output filter, one hop
EKS / SageMaker — custom services and classic ML sit beside the LLM, not inside it
Say "I'd put API Gateway and Lambda in front of Bedrock, keep retrieval in OpenSearch Serverless behind Knowledge Bases, and push custom services to EKS with SageMaker for the classic ML side."
Where do security and governance gates sit on that same request path?
Six gates, in order: redact PII in, guardrail in, filter by tenant at retrieval, guardrail out, trace every hop, and require approval before any version ships.
PII redact — strip before the model ever sees it, not after
RBAC + tenant filter — retrieval itself is scoped, not just the answer
Approval gate — a version bump is a change request, not a deploy
Say "Security isn't a wrapper around the model, it's gates on the path — redact, guard in, scope retrieval, guard out, trace, approve."
What does the LLMOps loop actually look like end to end?
Golden set feeds eval, eval gates deploy, deploy is traced, traces become feedback and regression tests, feeding governance and the next golden set.
Golden set — real questions with known-good grounded answers, versioned
Eval gate — groundedness + relevance score decides deploy, not a person's gut
Trace → feedback — production calls become tomorrow's regression tests
Say "This is exactly the A/B loop I ran on a production RAG build — chunking, retrieval and prompt versions all gated by groundedness and relevance before they shipped."
How is MLOps different from LLMOps?
MLOps versions a trained model; LLMOps versions a prompt, an index, and a model choice — three moving parts instead of one.
Same shape — both are build → gate → deploy → monitor → retrain loops
Different unit — MLOps versions weights, LLMOps versions a 3-tuple
Combined risk — a chunking change breaks retrieval even if the model didn't move
Say "MLOps ships a model, LLMOps ships a prompt-index-model combination — that's why I gate all three separately in eval."
What does a scalable data pipeline for AI look like?
Batch and streaming sources both land through clean/validate before chunk/embed/index, with schema and quality gates in between.
Two intakes — batch (Glue/Spark) and streaming (Lambda) converge before cleaning
Two gates — schema gate and quality gate, not one combined check
My depth — platform provisioning + a hand-built Python discovery pipeline, not the Spark jobs themselves
Say "I provision the Databricks/Synapse/ADF layer and I've built the clean-chunk-embed pipeline myself in Python — the Spark job authoring is the part I'd lean on a data engineer for."
How do you integrate a model into an enterprise app safely?
The app never calls the model directly — it goes through a gateway with retries, timeouts, a circuit breaker, and a fallback model behind a feature flag.
Gateway, not direct call — the app depends on an internal contract, not a vendor SDK
Circuit breaker — trips to fallback on sustained errors, not on the first 429
Feature flag — swap models without a redeploy
Say "The app talks to my gateway, never the model directly — that's where retries, the circuit breaker, and the fallback model live."
What's a "reusable accelerator" and do you have one?
A RAG accelerator is IaC + config + an eval kit packaged so another team deploys it without rebuilding the pipeline.
IaC layer — the Bicep pipeline provisions Search + Foundry, not a runbook
Config, not code — chunk size, model, index name are parameters
Eval kit included — the next team gets a groundedness baseline, not a blank slate
Say "My Bicep RAG pipeline is exactly that — another team points it at their own docs and gets the same index-plus-eval setup without rebuilding it."
What does the lead role actually fill your week with?
Mentoring, design review, and standards work sit alongside a 29-person MVP and an offshore team, not heads-down coding.
Mentoring — the offshore team on an insurer's RAG PoC, not a solo build
Standards — the 7-op tool pattern instead of model-written SQL, set once, reused
Scale — a discovery MVP running with 29 people means review beats code time
Say "At that scale the job is design review and standards — the compound-identity tool pattern is one I set so nobody lets the model write raw SQL."
What does responsible AI look like in practice, not policy language?
Bias eval, a human in the loop before impactful actions, visible citations, and pulling only the data the answer actually needs.
Bias/eval — score outputs by cohort, not one blended average
Human in the loop — the P-9 workflow hands off, it doesn't act alone
Citations + minimal data — answers show their source, retrieval scoped to what's needed
Say "My workflow agent routes anything consequential like P-9 to a human handoff — responsible AI means the agent proposes, a person decides."
6More interview questions — visual answers
Ten more likely questions, each answered as one drawing you can point to instead of talk through.
Walk me through how you'd build a RAG feature from scratch here
Six-week plan: discovery → golden set → pipeline v1 → eval gate → hardening → rollout.
golden set — real questions + expected chunk ids, built week 2
eval gate — no rollout until metrics clear a bar
hardening — load, cost, latency, guardrails before ship
Say I'd timebox discovery and the golden set first — everything after is measured against them.
How do you evaluate a RAG system?
Metrics sit on the pipeline stage they measure: retrieval, then generation, then answer.
recall@k / MRR — did the right chunk make the cut, how high
groundedness — answer traceable to retrieved text, not invented
relevance + correctness — judged against a golden Q&A set
Say on a production RAG build we ran A/B on chunking and retrieval settings scored by groundedness and relevance before shipping.
How do you choose between Bedrock, SageMaker, and self-hosted?
Managed models first, SageMaker for custom training, self-hosted only when neither fits.
Bedrock — default: managed models, no infra to run
training + secrets — Azure ML → SageMaker, MI/Key Vault → IAM/Secrets Manager
Say it's a component-for-component swap — the eval gate that proved the Azure build is the same gate that clears the AWS one.
Tell me about a production incident with an AI system
the platform's indexer starved on a missing storage role; found via tracing, fixed via IaC.
detect — tracing surfaced 403s on the storage read path
root cause — managed identity lacked data-plane RBAC, not control-plane
prevent — role assignment moved into IaC so it can't drift again
Say on a science-discovery platform the indexer went quiet, tracing found the 403, root cause was a missing data-plane role on the managed identity, fixed and then locked into IaC.
How do you keep prompts under control across teams?
A versioned prompt registry with review and an eval gate before any prompt ships.
versioned registry — every prompt change is a diff, not an edit in place
peer review — a second set of eyes before it runs against real data
eval gate — scored against the golden set, same bar as a model swap
Say I'd treat prompts like code — versioned, reviewed, and gated by the same eval suite as any other pipeline change.
What is your experience with fine-tuning?
Honest ladder: settings and prompting first, then RAG, LoRA only if those fail.
rungs 1–3 — settings, prompting, RAG: what I've actually shipped
rung 4 — LoRA fine-tune: not in production, named honestly
why the order — cheapest, fastest-to-verify fix tried first every time
Say I climb the ladder — settings, then prompt, then RAG — and only reach for fine-tuning when those three genuinely can't solve it.
How do you mentor engineers on GenAI?
A five-rung curriculum: hello world, RAG, eval, agents, ops — each rung ships something real.
hello world → RAG — smallest working thing first, then retrieval
eval → agents — measurement before autonomy, always in that order
ops — the rung most engineers skip; I put it on the curriculum
Say I run a hands-on ladder — hello world, RAG, eval, agents, ops — every rung ships a real artifact, not a deck.
Why should we hire you for this role?
Fifteen years platform depth carrying real GenAI builds, at the speed of an accelerator.
platform depth — the infra a GenAI system actually runs on, already known
real builds — six-agent build, GraphRAG at a network vendor, production RAG on a production RAG build
accelerator mindset — 30+ home agents, tuned a model from 169s to 0.5s
Say most GenAI candidates learn platform on the job — I bring fifteen years of it plus builds that already shipped.
Built 2026-09-17 by six parallel section-writers (workflow visual-answers).