Free sample: page 05 of the GenAI Interview Kit (visual answers). The full kit — 46 questions with model answers, 77 troubleshooting cards, 42 code blocks, 30 walkthrough frames — is $39 at niktechai.com/kits.

Visual answers — process diagrams, troubleshooting flows, interview questions

Every item is a drawing first: a TL;DR line, the diagram (animates on its own, slowly), three beats to say out loud. No paragraphs. Read any section on its own.

1 · Orchestration processswimlane diagram + enterprise examples2 · RAG pipeline processingest path, query path, AWS placement3 · Troubleshooting flowsdecision trees for what breaks in production4 · Transformer questionseach answered with a drawing5 · JD visualsAWS architecture, security, LLMOps loop, data pipelines6 · More interview questionsvisual answers

1Agent orchestration — the process, drawn, and where enterprises use it

Every orchestration pattern below is the same six-lane loop — an orchestrator routes work to specialist agents and tools, a critic checks the answer, and a trace records every step — reused with different agents plugged in.

The orchestration process, as a swimlane sequence

An orchestrator routes each step to one agent or tool, a critic checks the draft, and every arrow writes to a trace.
userorchestratorspecialist agentstools / datacritictrace1question2task: which agent?3query4rows: D2, D35draft + source ids6draft + citations7✗ no citation → retry8final answer ✓MAX 3guardlogs step id,latency, tokensfor every arrow
Say the orchestrator is just a router with a checklist — the critic is the part most demos skip.

The same loop, run for real: pump P-7 hello world

Four agents in a fixed order turn a raw log number into a plain-English service alert with the vendor to call.
hours agentpolicy agentcalculatorwriter"P-7" →logs: 4120 hrsD1 threshold:4000 hrs4120 ≥ 4000→ overdue 120"P-7 is due —contact Acme Pumps"orchestrator dispatched each stage in order; the same loop as item 1, one agent per box
Say this is the orchestration diagram with real numbers — nothing hidden.

Enterprise pattern: insurance claims triage

Intake extracts fields, RAG pulls the policy terms, a rules tool decides, and only low-confidence claims stop at a human.
intake agentextracts fieldspolicy-lookupagent (RAG)rules /decisioning toolfieldspolicy termsedge case?(low confidence)decisionadjuster(human gate)yesletter writernorulingtied to: an insurer's RAG PoC — underwriters query policy docs, low-confidence claims never auto-decide
Say clear claims never touch a human — only the ones the rules tool can't score confidently.

Enterprise pattern: IT / network ops troubleshooting

Multiple hypothesis agents run in parallel, a graph tool validates each against real topology, and only a confirmed fix executes.
alert(BGP flap)hyp. agent Alink down?hyp. agent BACL change?topology graphtool (Neo4j)validate vs BGP/OSPF/ACLvalidatorconfirmed by graph?runbook executor(human approval)yesno — new hypothesistied to: GraphRAG NetOps — hypotheses validated against the routing graph, not guessed
Say the graph tool is the part that stops the agent from hallucinating a fix — that's what I built at a network vendor.

Enterprise pattern: R&D knowledge assistant

An orchestrator fans a scientist's question out to search, graph, SQL, and model agents in parallel, then merges cited results.
scientistquestionorchestrator(orchestration engine+ Tasks)AI SearchagentGraphRAGagentSQL agentmodel-as-tool agentanswer +citationstied to: a science-discovery platform — six-agent architecture (AI Search, GraphRAG, ontology, SQL, model-as-tool) over 120GB Azure AI Search + Foundry IQ
Say this is the six-agent shape from a science-discovery platform — I contributed the GraphRAG and model-as-tool pieces.

Enterprise pattern: customer support copilot

A ticket is classified, grounded in the knowledge base and order system, and drafted for a human to approve and send.
ticketintentclassifierKB RAGagentorder-systemtooldraft replyagent-assist(human sends)"where's my order?"intent: order_statusKB: return policy chunkorder #4821: shippedreply text generatedhuman never writes the reply from scratch — only approves or edits before it sends
Say it's the RAG-plus-tool loop again — the only new part is the send stays with a person.

Single agent vs multi-agent vs plain pipeline

Use the plain pipeline until steps need to branch, the single agent until roles need separate context, then multi-agent.
plain pipelinestep 1step 2step 3fixed order, no decisions, no retriessingle agent + toolsone agent, looptool Atool Bone model picks which tool, in a loopmulti-agentorch.agent Aagent Bseparate roles, own context, coordinatedrule: pipeline if the steps never change · single agent if one skill set loops over toolsmulti-agent only when roles genuinely need separate context (added latency + cost otherwise)
Say I default to the pipeline and only add agents where a real decision or parallel role shows up.

What production adds around the loop

Guards, trace, eval, and a human gate wrap the same orchestration loop — they don't change what's inside it.
orchestrationloop (item 1)guardstraceevalhuman gateproduction wraps the same loop — it doesn't replace it — with limits, logging, scoring, and a stop point
Say a demo is the circle in the middle — production is everything drawn around it.

The same drawing, mapped onto three frameworks

Every framework fills the same five boxes — orchestrator, agent, tool, memory, safety — with its own names.
generic conceptorchestratoragenttool callmemory / statesafetyMicrosoft AgentFrameworkWorkflowExecutor /ChatAgenttyped functionThreadguard codeLangGraphgraph (nodes/edges)nodetool nodecheckpointerconditional edgeBedrock Agentsagent runtimeaction groupaction groupfnsession memoryguardrails
Say learn the five boxes once — every framework question becomes "which box is that?"

2RAG pipeline — the process, drawn

Ingest path, query path, and where AWS sits — the same shape he built on a production RAG build and the Standard, drawn stage by stage.

The two paths in one drawing

Ingest fills the index offline; query hits it live — they only ever meet at the index.
INGEST (batch / scheduled) source parse doc chunk text embed chunks vectors+meta INDEX QUERY (live) question embed top-k+ids rerank model prompt answer + [D2-0]
Say "Ingest and query are two separate pipelines that only touch at the index."

Ingest path — the knobs on each stage

Every ingest stage has a setting that changes recall later, and each one gets versioned.
chunksize=1000 overlap=150 embedmodel + version pinned tag metadatasource, doc_id, date index v12 knobs that change recall later chunk_size/overlap: too big → diluted match; too small → lost context embedding model + version: swap it → old vectors go stale, must re-embed index version: bump on every re-embed so old/new never mix mid-query
Say "On a production RAG build we A/B-tested chunking and retrieval settings against groundedness scores before locking them."

Query path — knobs, latency, and cost

Retrieval and generation each have their own knobs, and each stage adds latency or cost.
retrievek=8, filters~50-150ms rerankhybrid weight, top_n=4~80-200ms prompttop_n chunks + q modeltemp=0, max_tok~1-3s, $/tok where it accrues latency: reranker and model call dominate — retrieve+rerank ~100-350ms, model ~1-3s cost: per-token on the model call — max_tokens and chunk count both drive it up temperature=0 for grounded answers, higher only for creative rewrites
Say "Reranking narrows what actually goes in the prompt, which controls both cost and hallucination risk."

Hybrid retrieval, drawn

Keyword and vector search run in parallel, then get fused with RRF before reranking.
question keyword (BM25) vector search RRF fuserank blend reranktop_n top-k
Say "Hybrid retrieval covers the case where vector search misses an exact keyword like a part number."

GraphRAG variant, drawn

Text becomes triples in a graph so a question can hop across relationships, not just similarity.
docs (RFCs) extract triples graph (Neo4j) R1 R2 ACL 2-hop query use when the answer needs a relationship chain, not a single passage — device to route to ACL
Say "On a network-operations build we ran GraphRAG over device configs and RFCs in Neo4j so troubleshooting could hop from a device to the routing policy that touches it."

AWS placement, drawn

Mapped, not claimed in production — ten years of AWS platform work, Bedrock/SageMaker not yet shipped.
S3 Bedrock KBor Lambda ingest OpenSearch Serverless Claude (Bedrock) API GW /Lambda / EKS CloudWatch + Guardrails + IAM roles wrap every hop logging/metrics via CloudWatch, content filtering via Bedrock Guardrails, least-privilege IAM per service ■ mapped from ten years of AWS platform work (Control Tower/EKS/Lambda/IAM) — not shipped Bedrock/SageMaker in prod
Say "I map this from ten years of AWS platform work — Bedrock and SageMaker specifically I haven't shipped in production yet."

The Azure twin the author built

Same shape on Azure, shipped in production on a production RAG build with managed identity end to end.
Blob AI Search indexerintegrated vectorization Azure OpenAI Appmanaged identity App Insights
Say "This is the RAG stack I actually shipped on a production RAG build — Blob through AI Search into Azure OpenAI, with A/B tests on chunking and prompts."

Multi-tenant RAG, drawn

A tenant filter goes into every retrieval call so one query can never see another tenant's chunks.
query + tenant_id index-per or ACL? per-tenant index index-per shared idx + ACL filter metadata ACL index-per-tenant: hard isolation, more infra to manage metadata ACL on shared index: cheaper, filter must be enforced on every call
Say "The tenant filter has to live in the retrieval call itself, not in the prompt, or one bad query leaks another tenant's data."

The freshness loop

A change event triggers re-embedding, a version bump, and a golden-set rerun before the new index goes live.
change event re-embed index v+1 golden-set rerun pass → promote v+1 live · fail → hold, alert
Say "A version bump plus a golden-set rerun is what stops a bad re-embed from silently going live."

RAG at enterprise scale — the 120 GB design

120 GB of Azure AI Search knowledge on a science-discovery platform, designed for partitioned throughput and locked-down access.
120 GB index partition 1 partition 2 partition N throughput-tuned private networkVNet-bound RBAC per role Foundry IQ Science platform: 120 GB Azure AI Search + Foundry IQ knowledge design, one input to the six-agent architecture
Say "At a science-discovery platform I designed the 120 GB Azure AI Search knowledge layer that feeds the six-agent architecture."

3Operational troubleshooting — decision flows you can follow under pressure

Ten flows: symptom → the exact check → the concrete fix, so you never freeze on an ops question.

"The answer is wrong" — retrieval vs generation split

Print the retrieved chunk ids first — that one check splits every "wrong answer" into two different fix paths.
wrong / hallucinated answer wrong / hallucinated answer print chunk ids right chunk (e.g. D1) in the printed ids? no yes retrieval: chunking fix: 120/20 → try 300/40 fix: embedder swap fix: k 2 → 5, add hybrid generation: prompt fix: temperature → 0 fix: "cite [id] or say n/a"
Say "first thing I check is whether the right chunk was even retrieved — that tells me which half of the pipeline is broken."

"The answer says I don't know but the doc exists"

Four separate causes look identical from the outside — check each in order: parsing, stale index, embedding mismatch, filter.
"I don't know" — doc exists check: did the PDF parser lose the text? check: last ingest run date vs doc date check: same embedder at index and query time? check: metadata filter excluding the doc? fix: OCR / re-extract fix: re-run the ingest job fix: re-embed with match fix: loosen tenant/date filter
Say "false negative is usually parsing, staleness, embedder mismatch, or an over-tight filter — I check those four in order."

"Latency is high" — where in the pipeline?

Split the clock into retrieval time and model time before touching any setting — they have different fixes.
p95 latency is high split retrieval-ms vs model-ms in the trace retrieval-bound model-bound fix: shrink index / lower k fix: drop/cache reranker fix: ANN index (HNSW) fix: cap max_tokens fix: trim context / chunks fix: smaller / faster model neither? check network / throttling queue
Say "I trace retrieval-ms separately from model-ms before I touch a single setting — they don't share a fix."

"Cost spiked" — trace the token bill

Cost is tokens-per-call times calls; walk context growth, retries, and cache misses before blaming the model.
daily spend jumped context growing retries per call cache miss rate model choice fix: fewer/shorter chunks fix: fix the root error fix: cache warm prompts fix: route easy calls smaller track: tokens/call × calls/day = $/day
Say "cost is tokens per call times calls per day — I break the bill down into those two before I touch the model."

"429 / throttling"

RPM and TPM are two separate limits — find which one tripped, then back off with jitter and cap concurrency.
429 responses error header: RPM or TPM limit? RPM TPM fix: cap concurrency fix: request-level queue fix: trim tokens/call fix: buy provisioned (PTU) always: exponential backoff + jitter
Say "429 first tells me whether I hit requests-per-minute or tokens-per-minute — they need different fixes."

"The agent loops / never finishes"

The trace shows the same tool call repeating — that's the tell, not a guess, so read the trace before changing anything.
agent never finishes trace: same tool call repeating? yes no — routing flips fix: max_steps cap (e.g. 8) fix: feed tool error as text fix: validate args w/ schema fix: temperature → 0 on the routing step only
Say "if the trace shows the same tool call over and over I cap max_steps and route the error back in as text, not a stack trace."

"Eval score dropped after a deploy"

Diff exactly what changed — prompt, model, or index — then roll back that one version, not the whole release.
groundedness ↓ 0.91→0.74 diff: prompt / model / index rerun golden set on each culprit: prompt v14 fix: roll back prompt only model and index stay on latest — one variable at a time
Say "eval drop after a deploy means diff what changed and rerun the golden set on each piece — roll back just that one."

"Users report a leak across tenants"

Treat it as a P0 security incident: verify the filter clause, per-tenant boundary, then audit the trace for who saw what.
⚠ cross-tenant leak report 1 · retrieval filter present? 2 · per-tenant index or shared? 3 · audit trace: who saw it? fix: mandatory tenant_id filter fix: split to per-tenant index fix: log id+tenant every call ✗ never patch silently — this is an incident, not a bug ticket
Say "a cross-tenant leak is an incident first — I check the filter clause, the index boundary, then pull the audit trace."

"Model output is not valid JSON"

Enforce structured output at the API level first — a parser retry loop is the fallback, not the fix.
JSON.parse throws structured mode on? no fix: enable JSON schema mode yes — still broke fix: retry once, feed the parser error back fix: lenient parser (strip fences/trailing text)
Say "structured output mode at the API level first, then one retry with the parser error attached, lenient parsing as the fallback."

The on-call card — five numbers, first

One dashboard glance: p95 latency, error rate, tokens/call, cost/day, and "no source found" rate.
p95 latency 1.8s error rate 0.4% tokens / call 2,140 cost / day $186 "no source" rate 9.1% ▲ ▲ one number red → jump straight to that flow above "no source" 9.1% here → flow 2 (parsing / stale index / filter)
Say "on-call, I check five numbers first — whichever one's red tells me which flow to run."

Transformers — what actually happens to a token

Every token becomes a vector, self-attention lets it read every other token, then the vector predicts the next token.
"Pump P-7 trips" Pump P-7 trips step 1 · each token → its own vector (embedding) self-attention step 2 · "trips" attends to "P-7" — links subject to verb × N transformer layers predict next token: "reset"
Say "a transformer turns tokens into vectors, lets every token look at every other one through attention, stacks that N times, then predicts the next token."

4Transformer questions — each answered with a drawing

Nine core transformer concepts, each as one drawing you can point at instead of explaining in words.

Explain attention simply

Every word looks at every other word and votes how much it matters for the one being decided.
query word: "trips" trips pump weight 0.15 pressure weight 0.70 90 psi weight 0.15 answer for "trips" ≈ 0.70×pressure + 0.15×pump + 0.15×psi
Say attention is every word scoring every other word and blending them by that score.

What is a token and why is pricing per token?

A sentence is chopped into sub-word pieces first, and every piece in and out costs money.
"Pump P-7 must be serviced every 4000 hours" Pump P -7 must serviced every 40 00 hours 9 tokens bill = (input tok × in-rate) + (output tok × out-rate)
Say the model doesn't read characters, it reads sub-word tokens, and every one in or out is billed.

What is an embedding?

Text becomes a point in number-space where similar meaning sits physically close.
meaning space (10 numbers per chunk, drawn in 2D) D1: 4000-hr service Q: "when is P-7 serviced?" close · 0.38 D3: vendor is Acme far · 0.0
Say an embedding turns text into a point in space, so "close meaning" becomes "close numbers."

Why does a long context cost more and get slower?

Every token attends to every prior token, so cost grows as n² while the KV cache keeps growing too.
1000 tokens → ~1,000,000 attention pairs (n²) short: 10² = 100 pairs long: 100² = 10,000 pairs (100× the work) KV cache: every past token's key+value stays in GPU memory until the turn ends
Say attention cost scales with the square of context length, and the KV cache eats memory the whole turn.

Temperature vs top-p

Temperature reshapes how peaked the probabilities are; top-p cuts the tail off after that.
logits [2.0, 1.2, 0.3, 0.1] → softmax probability per setting T = 1.0 56% 25% 10% 8% varied, creative T = 0.3 93% 6.5% focused, safe T → 0 (greedy) 100% always same top pick top-p 0.9 at T=1.0: keep tokens until cumulative ≥ 90% → keeps top 3 (56+25+10 = 91.6%), drops the 8% tail temperature reshapes the whole curve; top-p trims which tokens are even eligible
Say temperature reshapes the probability curve, top-p then cuts off its tail.

Encoder vs decoder — why RAG uses two models

RAG pairs a reader that scores meaning with a writer that generates text, because they do opposite jobs.
ENCODER (reader) reads whole chunk at once outputs one fixed vector used for: embeddings, scoring e.g. BedrockEmbeddings DECODER (writer) generates one token at a time only sees tokens written so far used for: the final answer text e.g. ChatBedrock top chunks RAG needs both: encoder finds the facts, decoder writes the sentence
Say the embedder is the reader that finds relevant chunks, the chat model is the writer that turns them into an answer.

Why RAG instead of a bigger prompt or fine-tuning?

RAG keeps knowledge current at query time without paying to retrain or stuffing everything into every prompt.
bigger prompt ✓ simple to try ✗ cost scales n² ✗ hits context limit freshness: instant RAG ✓ small marginal cost ✓ swap docs, no retrain ✓ answers cite sources freshness: instant, per-doc fine-tuning ✓ bakes in style/format ✗ retrain per data change ✗ hours-days, $$$ freshness: stale until retrained
Say RAG is the option where updating a document is cheaper than updating the model.

What happens at the context limit?

Depending on the API, hitting the limit either truncates, errors out, or you summarise history first.
◆ tokens ≥ context limit? yes, hard API ✗ 400 context_length_exceededfix: trim before sending yes, silent ■ oldest turns truncatedfix: pin system + recent turns managed proactively ✓ summarize old historyLangGraph checkpoint + summary node ops pick: watch token count per turn, summarise or drop before you hit any of these three
Say the three failure modes are error, silent truncation, or a summarization step you control.

What is max_tokens?

max_tokens caps only the output length, it never trims what you send in.
input window (prompt + history + retrieved chunks) sized by the model's context limit, e.g. 128k tokens max_tokens: output cap e.g. 500 — cuts the answer mid-sentence if hit too low → answer truncates before finishing; too high → no cost saving, model still stops naturally set it to the longest answer you actually expect, as a safety cap
Say max_tokens only limits what the model writes back, not what you send it.

5The rest of the job description, drawn

Everything else the posting asks for — platform architecture, governance, ops loops, and the lead's actual week — as one drawing each.

What does an enterprise AWS GenAI reference architecture look like?

A request crosses a private VPC through API Gateway → Lambda → Bedrock, hits Knowledge Bases + OpenSearch for retrieval, and every hop is logged and encrypted.
VPC — private networking (no public model endpoints) 1 API Gateway 2 Lambda 3 Bedrock model+ Guardrails 4 Knowledge Bases→ OpenSearch Serverless 5 EKS services SageMakerclassic ML / fine-tune S3 (docs, artifacts) Secrets Manager + KMS IAM roles (least priv.) CloudWatch + X-Ray every box logs to CloudWatch/X-Ray; every secret in KMS-encrypted Secrets Manager, never in code
Say "I'd put API Gateway and Lambda in front of Bedrock, keep retrieval in OpenSearch Serverless behind Knowledge Bases, and push custom services to EKS with SageMaker for the classic ML side."

Where do security and governance gates sit on that same request path?

Six gates, in order: redact PII in, guardrail in, filter by tenant at retrieval, guardrail out, trace every hop, and require approval before any version ships.
request → PII redactbefore call✗ SSN,card Guardrailin — topic /jailbreak RBAC +tenant filterat retrieval Guardrailout — toxic /grounding Audit tracereq id, user,chunks used ■ every prompt/index/model version bump needs a human approval before promote KMS at rest S3, OpenSearch, logs · private networking end to end: no gate is reachable from outside the VPC
Say "Security isn't a wrapper around the model, it's gates on the path — redact, guard in, scope retrieval, guard out, trace, approve."

What does the LLMOps loop actually look like end to end?

Golden set feeds eval, eval gates deploy, deploy is traced, traces become feedback and regression tests, feeding governance and the next golden set.
Golden set Eval pass? score ≥threshold? Deploy Hold, fix Trace Feedback +regression Governanceprompt / index / model ver. yes no
Say "This is exactly the A/B loop I ran on a production RAG build — chunking, retrieval and prompt versions all gated by groundedness and relevance before they shipped."

How is MLOps different from LLMOps?

MLOps versions a trained model; LLMOps versions a prompt, an index, and a model choice — three moving parts instead of one.
MLOps LLMOps Train Register Deploy Monitor drift Prompt ver. Index ver. Model ver. Eval per combo one artifact to version: the weights three artifacts, all must line up: prompt × index × model ✓ accuracy / F1 vs. holdout ✓ groundedness / relevance vs. golden set
Say "MLOps ships a model, LLMOps ships a prompt-index-model combination — that's why I gate all three separately in eval."

What does a scalable data pipeline for AI look like?

Batch and streaming sources both land through clean/validate before chunk/embed/index, with schema and quality gates in between.
Batch sources Stream sources Glue / Spark Lambda Clean / validate Chunk / embed Index schema gate ⚠ quality gate ⚠ Honest note: my depth is the platform side — provisioning Databricks / Synapse / ADF and IAM/network around this — plus a Python discovery pipeline that does clean/chunk/embed myself
Say "I provision the Databricks/Synapse/ADF layer and I've built the clean-chunk-embed pipeline myself in Python — the Spark job authoring is the part I'd lean on a data engineer for."

How do you integrate a model into an enterprise app safely?

The app never calls the model directly — it goes through a gateway with retries, timeouts, a circuit breaker, and a fallback model behind a feature flag.
App AI gateway / API Primary model Fallback model 1 try, retry×2, timeout circuit open flag: use_fallback gateway guards ■ retries (bounded, backoff) ■ per-call timeout ■ circuit breaker on error rate ■ feature flag routes model ■ fallback model, degraded ok → app never sees a hard 500
Say "The app talks to my gateway, never the model directly — that's where retries, the circuit breaker, and the fallback model live."

What's a "reusable accelerator" and do you have one?

A RAG accelerator is IaC + config + an eval kit packaged so another team deploys it without rebuilding the pipeline.
Bicep RAGpipeline(the author's own) IaC (Bicep) Config (chunk/model) Eval kit next team deploys it ✓ own docs, own eval baseline
Say "My Bicep RAG pipeline is exactly that — another team points it at their own docs and gets the same index-plus-eval setup without rebuilding it."

What does the lead role actually fill your week with?

Mentoring, design review, and standards work sit alongside a 29-person MVP and an offshore team, not heads-down coding.
Lead's week Mentoringoffshore team, an insurer Design reviews6-agent build arch. Standardsagent-tool patterns, no raw SQL Evaluating techFoundry IQ, GraphRAG 29-person a client MVP Stakeholders
Say "At that scale the job is design review and standards — the compound-identity tool pattern is one I set so nobody lets the model write raw SQL."

What does responsible AI look like in practice, not policy language?

Bias eval, a human in the loop before impactful actions, visible citations, and pulling only the data the answer actually needs.
Bias / evalscore by cohort Human in the loopbefore P-9 handoff Transparencycitations [D2-0] shown Data minimisationtenant-scoped retrieval ✓ P-9 due-service alerts route to HumanHandoff, never auto-act ✓ NetOps hypotheses at a network vendor validated against the routing graph before acting
Say "My workflow agent routes anything consequential like P-9 to a human handoff — responsible AI means the agent proposes, a person decides."

6More interview questions — visual answers

Ten more likely questions, each answered as one drawing you can point to instead of talk through.

Walk me through how you'd build a RAG feature from scratch here

Six-week plan: discovery → golden set → pipeline v1 → eval gate → hardening → rollout.
Wk1discovery Wk2golden set Wk3pipeline v1 Wk4eval gate ✓ Wk5hardening Wk6rollout gate: recall@k + groundedness must clear threshold before hardening starts
Say I'd timebox discovery and the golden set first — everything after is measured against them.

How do you evaluate a RAG system?

Metrics sit on the pipeline stage they measure: retrieval, then generation, then answer.
retrieve top-k generate answer final answer recall@kright chunk in top-k?+ MRR (rank) groundednessdoes it cite onlywhat's retrieved? relevance +correctnessvs golden set Production: A/B on chunking, retrieval, prompt, model version — same three metrics
Say on a production RAG build we ran A/B on chunking and retrieval settings scored by groundedness and relevance before shipping.

How do you choose between Bedrock, SageMaker, and self-hosted?

Managed models first, SageMaker for custom training, self-hosted only when neither fits.
need customtraining/fine-tune? no Bedrock (managed) yes data/latency needfull infra control? no SageMaker yes self-hosted — rare, justify it
Say I default to Bedrock, ten years on AWS platform work but honest that Bedrock/SageMaker prod use is what I'd bring next, not what I've shipped yet.

How do you handle hallucinations?

Four defenses stacked: ground every answer, cite it, temperature zero, judge-scored eval.
1 · groundingonly answer fromretrieved chunks 2 · citations[D2-0] tags +human spot-check 3 · temp 0deterministic,no creative drift 4 · judge evalLLM scoresgroundedness D2-0 answer above cited [D2-0] not [D1] — a wrong citation fails step 4 and blocks ship
Say home recruiter agent runs temperature 0 — dropped a 169-second stall to half a second and raised accuracy to 0.94.

How would you migrate an Azure GenAI build to AWS?

Six services swap one-for-one: same shape, different cloud, same eval gate before cutover.
AZURE AWS Azure OpenAI Bedrock AI Search OpenSearch / KB for Bedrock Foundry Agents Bedrock Agents Content Safety Guardrails for Bedrock Azure ML SageMaker MI / Key Vault IAM / Secrets Manager
Say it's a component-for-component swap — the eval gate that proved the Azure build is the same gate that clears the AWS one.

Tell me about a production incident with an AI system

the platform's indexer starved on a missing storage role; found via tracing, fixed via IaC.
1indexer starving,KB stops updating 2detect: tracingshows 403 on read 3root cause: MI missingstorage data-plane RBAC 4fix: grant roleassignment 5prevent: role inIaC, never manual
Say on a science-discovery platform the indexer went quiet, tracing found the 403, root cause was a missing data-plane role on the managed identity, fixed and then locked into IaC.

How do you keep prompts under control across teams?

A versioned prompt registry with review and an eval gate before any prompt ships.
edit prompt v3 peer review eval gatevs golden set ✓ a failing prompt never merges — same gate a chunking or model-version change goes through
Say I'd treat prompts like code — versioned, reviewed, and gated by the same eval suite as any other pipeline change.

What is your experience with fine-tuning?

Honest ladder: settings and prompting first, then RAG, LoRA only if those fail.
1 · settings (temp 0) 2 · prompt engineering 3 · RAG (shipped, a production RAG team) 4 · LoRA fine-tune shipped rungs 1–3 in production; rung 4 not in production — here's how I'd approach it
Say I climb the ladder — settings, then prompt, then RAG — and only reach for fine-tuning when those three genuinely can't solve it.

How do you mentor engineers on GenAI?

A five-rung curriculum: hello world, RAG, eval, agents, ops — each rung ships something real.
1 hello world 2 RAG 3 eval 4 agents 5 ops same ladder as pcep-lessons: each rung is a working artifact, not a slide
Say I run a hands-on ladder — hello world, RAG, eval, agents, ops — every rung ships a real artifact, not a deck.

Why should we hire you for this role?

Fifteen years platform depth carrying real GenAI builds, at the speed of an accelerator.
platform depth10yr AWS, 15yrAzure/Terraform/K8sEKS · Lambda · IAM real GenAI builds6-agent build,a GraphRAG build,a production RAG team prod RAG accelerator mindset30+ home agents,0.5s tuned inference,ships fast, measures
Say most GenAI candidates learn platform on the job — I bring fifteen years of it plus builds that already shipped.

Built 2026-09-17 by six parallel section-writers (workflow visual-answers).