I scored 21 RAG explainers and courses with a decision model. None teach evaluation from zero.
Before building anything I wanted to know whether a simple, visual, from-zero explanation of RAG evaluation already existed. Golden sets, retrieval hit versus citation hit, the judge, the gate, the similarity-score floor and how a person picks it. If it existed, I would consume it rather than build it.
Method
I cataloged 21 resources across five families: interactive explainers, YouTube, free evaluation courses, paid courses and books, and AI-103 exam-prep sites. For each one I wrote a factual description from the page itself (link checked with curl; two demos were dead or shells) and asked the same seven questions of a decision model, Jev (jev-latest, direct API). Jev returns a calibrated probability per question rather than prose. The questions:
- Evaluation mechanics: does it teach a golden set, retrieval hit versus citation hit, an LLM judge with a rubric, or a champion-versus-candidate gate? Mentioning that evaluation exists does not count.
- Retrieval mechanics: a similarity score threshold and how to calibrate it, BM25 fused with cosine (RRF), top-k versus threshold, with numbers?
- From zero: usable with no prior AI vocabulary?
- Real numbers: measured from a real run rather than placeholders?
- Format: article, step-through tool, animated film, code-along video, cohort course, or dead.
Then I scored nine of my own artifacts with the identical questions, from descriptions extracted from the files (titles, counts, spoken lines), not from sales copy. Two caveats you should hold against me: the probabilities are Jev's reading of my descriptions, not of the resources themselves, and I wrote the questions around the gap I suspected. The raw inputs and outputs are linked at the bottom so you can rerun or rewrite them.
Result
No external resource scored above 0.50 on both mechanics columns. The two that teach evaluation (Evidently's free code course at 0.74, Hamel Husain and Shreya Shankar's cohort at 0.70) are for practitioners who already know the words (from-zero 0.05 and 0.07). The only from-zero material (IBM, KodeKloud, ChunkViz) teaches concepts, not numbers. The free RAG visualizers all stop where retrieval ends.
| Resource | Evaluation mechanics | Retrieval mechanics | From zero | Real numbers | Format |
|---|---|---|---|---|---|
| Ours — Hello-world RAG with evaluation (films, unreleased) | 0.96 | 0.83 | 0.71 | 0.88 | animated film |
| Ours — RAG factory 3D (unreleased) | 0.92 | 0.82 | 0.77 | 0.89 | animated film |
| Ours — RAG lab explainer (unreleased) | 0.92 | 0.71 | 0.21 | 0.68 | code along video |
| Hamel Husain & Shreya Shankar — AI Evals (Maven, $2,100) | 0.70 | 0.40 | 0.07 | 0.16 | live or cohort course |
| Boot.dev — Learn RAG (paid) | 0.37 | 0.35 | 0.06 | 0.22 | static text or diagrams |
| Evidently — LLM evaluation course (free) | 0.74 | 0.28 | 0.05 | 0.53 | code along video |
| Ours — GenAI Interview Kit ($39) | 0.40 | 0.19 | 0.08 | 0.53 | interactive stepthrough |
| Ours — RAG from the ground up (film, unreleased) | 0.16 | 0.32 | 0.73 | 0.76 | animated film |
| AI Engineer Roadmap 2027 playbook | 0.16 | 0.12 | 0.13 | 0.09 | static text or diagrams |
| Ours — Chunking strategy films (unreleased) | 0.20 | 0.12 | 0.48 | 0.74 | animated film |
| Ours — AI-103 Flight School ($49) | 0.11 | 0.17 | 0.07 | 0.08 | interactive stepthrough |
| Udemy LangChain & RAG courses | 0.10 | 0.11 | 0.14 | 0.13 | code along video |
| ByteByteGo — Generative AI System Design Interview (book) | 0.11 | 0.09 | 0.12 | 0.04 | static text or diagrams |
| Ours — Evaluation metrics films (unreleased) | 0.94 | 0.08 | 0.79 | 0.84 | animated film |
| Ours — The score floor, five films (unreleased) | 0.08 | 0.96 | 0.55 | 0.87 | animated film |
| DeepLearning.AI — Building and Evaluating Advanced RAG (free) | 0.43 | 0.06 | 0.04 | 0.76 | code along video |
| AI-103 practice-test sites (Udemy, mscertquiz, open-exam-prep, MasteryExamPrep, ExamTopics) | 0.05 | 0.06 | 0.18 | 0.05 | static text or diagrams |
| Logical Lenses — RAG evaluation / RAGAS (YouTube) | 0.13 | 0.04 | 0.30 | 0.12 | static text or diagrams |
| zackproser.com — RAG Visualized | 0.03 | 0.04 | 0.21 | 0.21 | interactive stepthrough |
| jay.tools — RAG Explainer (demo offline) | 0.03 | 0.09 | 0.12 | 0.05 | dead or unusable |
| RAG Playground (ragplay.vercel.app) | 0.03 | 0.04 | 0.19 | 0.15 | interactive stepthrough |
| KodeKloud — RAG Explained for Beginners (YouTube) | 0.03 | 0.04 | 0.60 | 0.03 | animated film |
| Aishwarya Reganti — AI Evals for Everyone (free) | 0.09 | 0.03 | 0.28 | 0.04 | static text or diagrams |
| Cognee — A Picture of RAG | 0.02 | 0.03 | 0.26 | 0.04 | static text or diagrams |
| Vectree — RAG concept tree | 0.02 | 0.03 | 0.42 | 0.03 | interactive stepthrough |
| RAG_Visualizer (GitHub, Vercel demo) | 0.02 | 0.03 | 0.25 | 0.05 | interactive stepthrough |
| ChunkViz — Greg Kamradt | 0.02 | 0.02 | 0.53 | 0.04 | interactive stepthrough |
| IBM Technology — What is RAG (YouTube) | 0.02 | 0.02 | 0.64 | 0.03 | code along video |
| TechViz — Reciprocal Rank Fusion (YouTube) | 0.02 | 0.38 | 0.13 | 0.13 | code along video |
| bbycroft.net/llm — 3D GPT visualization | 0.01 | 0.02 | 0.12 | 0.19 | interactive stepthrough |
What I take from it
Consume, do not rebuild: bbycroft.net/llm for transformer internals, ChunkViz to watch splitters cut, TechViz for RRF arithmetic, DeepLearning.AI for the RAG triad, Evidently for evals in code. The gap is narrow and specific: retrieval mechanics plus evaluation, from zero, with measured numbers, drawn. That is what the kits are, and what the film series behind them will be once narration is licensed.
Raw files
- competitors.jsonl — the 21 descriptions and the seven questions
- competitors.answers.jsonl — Jev's answers
- ours.jsonl · ours.answers.jsonl — the same for my artifacts
Names of products, courses and people above are used to identify those resources. Prices are as listed on their sites on 2026-09-20.
Related notes
- GenAI interview questions with model answers: what the interviewer is actually probing
46 GenAI interview questions with model answers and 45 troubleshooting cards: what AI-engineer panels ask about RAG, agents, LLMOps and fine-tuning. - The five rungs of agent engineering
Prompt, context, loop, graph, ownership. Most teams stall on rung one and wonder why their AI never leaves the demo. A ladder for what to fix next. - Make your agent fail closed
The worst agent bug isn't a wrong answer — it's a confident answer built on a tool that quietly returned nothing. A design rule, and where to enforce it.