RAG: build, inspect, improve

153 concept and worked-example cards in 16 topic sets, with recorded English questions and explanations, teaching diagrams, and 12 executed Python examples.

Open RAG in Flashcards Overlay · Select RAG tutorial · AI engineering in the top Deck picker, then choose a Lesson. All 16 lessons are directly available, with listening and recall controls.

Companion to Harish Neel’s RAG tutorial playlist. The original playlist title can change; this release uses captions retrieved September 20, 2026. Dates on a playlist title are not evidence of a new lesson. No video or audio autoplays here.

Where to start

First explain the difference between a model’s learned weights and its current context. Start with architecture and ingestion, then retrieval. For the code examples, learn Python functions, lists and dictionaries first. Study cosine after dot products and norms. Advanced chunking, fusion and reranking follow the basic pipeline.

The AI101 week-three module introduces architecture on October 2. Later engineering courses revisit implementation and evaluation after prerequisites. These are library sets; they are not 153 new cards assigned to one day and studying them does not automatically award a grade.

MC case: answer a course-policy question

Use the fictional documents in the private AI Operations pack. State the question, select accessible and current evidence, identify exact supporting document IDs, and explain when the system should decline to answer. Compare retrieval quality with answer support. Do not connect a toy retriever to real private records.

Open your private coursework and practice pack

02 · RAG architecture and vector embeddings

22 cards · Original tutorial episode (opens only when selected).

What is retrieval-augmented generation (RAG) at a high level?

RAG combines retrieval from external sources with generation conditioned on selected evidence. In a typical text system, the retrieved passages enter the model context with the question. The entire collection does not need to fit in that context, and indexing documents does not itself retrain model weights. Suitable multimodal systems can also use non-text evidence or outputs. Retrieval does not guarantee truth or freshness: check permissions, versions and claim support.

Worked example. A study assistant has a folder of course notes. When a student asks a question, the assistant searches the notes for relevant passages and gives those passages to the language model along with the question.

RAG high-level flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What must fit within a model request budget?

Instructions, conversation history, the question and retrieved evidence consume input capacity. The model also has output limits; some systems share a total input-plus-output budget. Multimodal inputs may consume capacity too. Check the chosen model and API rather than assuming one universal limit.

Worked example. In a hypothetical shared 4,000-token budget, 3,600 input tokens leave at most 400 for output, subject to any separate output limit.

What must fit within a model request budget?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is a token different from a word?

A token is a unit defined by a tokenizer. It may encode part of a word, a word with nearby whitespace, punctuation, bytes or other sequences. Counts depend on the text and tokenizer; word and character counts are not exact token counts.

Worked example. A made-up identifier such as study_note_v2 may break into several tokens. Measure it with the actual tokenizer instead of counting it as one word.

Why is a token different from a word?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why retrieve selected evidence instead of putting the whole corpus in every prompt?

A corpus may exceed the request budget. Even when it fits, sending everything can add cost, latency and irrelevant or conflicting evidence. Retrieval selects a bounded subset for the question; whether it improves the application must be measured.

Worked example. A fictional MC assistant has 10,000 study notes. A question about one exam policy can use the current policy and its exception, rather than all 10,000 notes.

Why retrieve selected evidence instead of putting the whole corpus in every prompt?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are the main steps in the ingestion pipeline of a RAG system?

The ingestion pipeline typically starts with source documents. These are split into smaller chunks. Each chunk is passed through an embedding model to produce a vector. The vectors are stored in a vector database or a database with vector support. This prepares the knowledge base for retrieval. Index updates and deletions require explicit lifecycle handling.

Worked example. A company's PDFs are chunked into 1,000-token pieces, each piece is embedded, and the vectors are saved in a vector database.

Ingestion pipeline

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

When and why do RAG systems split documents into chunks?

Long or mixed-topic documents are often split into attributable units so retrieval can select focused evidence within embedding and generation limits. Short documents can remain whole. Chunking is a design choice, not a requirement that every document be cut.

Worked example. Keep a short FAQ answer intact, but split a long course handbook by sections while retaining headings, document IDs and page references.

When and why do RAG systems split documents into chunks?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the tradeoff of using overlap between chunks?

Overlap can preserve context that would otherwise be split across a boundary, helping retrieval find coherent passages. However, overlap creates redundancy: the same text appears in multiple chunks, increasing storage and embedding costs and potentially returning duplicate information. Overlap is not universally better.

Worked example. If chunks are 500 tokens with 50 tokens of overlap, the last 50 tokens of one chunk also appear at the start of the next chunk.

Chunk overlap

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does an embedding model differ from a large language model?

An embedding model converts text into a vector representation. It specializes in producing numerical vectors that capture semantic meaning. A large language model generates text. They serve different purposes, although both are based on neural networks. Embeddings encode learned associations, not objective truth.

Worked example. An embedding model turns the sentence 'The cat sat' into a list of numbers, while an LLM might continue the sentence with 'on the mat'.

Embedding vs LLM

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a vector embedding?

A vector embedding is a mathematical representation of text, images, or other data as a list of numbers. Each position in the list is a dimension, and the number at that position is the coordinate on that dimension. The embedding captures learned associations so that items used in similar contexts tend to have similar vectors. Embeddings encode learned associations, not objective truth.

Worked example. A three-dimensional embedding for a word might be [0.34, 0.08, 0.75], where each coordinate is a value on a learned axis.

Vector embedding

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How do an embedding dimension and a coordinate differ?

A dimension is an axis of the vector space; a coordinate is the number on that axis. Dense embedding axes are generally learned latent features, not reliably named concepts such as size or usefulness.

Worked example. A three-dimensional toy vector has coordinates 0.2, 0.4 and -0.1. Those three numbers do not reveal three human-readable facts.

How do an embedding dimension and a coordinate differ?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does semantic similarity in an embedding space mean?

The encoder has learned a representation intended to make relevant relationships measurable. Depending on its training and task, related passages often score more similarly. This is an imperfect learned relationship, not a guarantee that every related pair is close or every unrelated pair is far apart.

Worked example. A search model may connect “exam date” with “assessment schedule,” yet still retrieve an outdated schedule.

What does semantic similarity in an embedding space mean?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What changes when embedding dimensionality increases?

At the same numeric precision, storing more coordinates costs more space per vector. Search and encoding costs depend on the model and index. More dimensions do not automatically provide more information or better retrieval; compare measured quality, latency and storage.

Worked example. With 32-bit floats, 768 coordinates use 3,072 raw bytes and 1,536 use 6,144, before metadata and index overhead.

What changes when embedding dimensionality increases?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why must query and document encoders be compatible?

Similarity scores require representations in a compatible space. Use the intended query/document encoder pair, preprocessing and dimensions. They may be one shared model or different encoders trained to work together. Matching vector length alone is insufficient.

Worked example. Two unrelated models can both return 768 numbers while assigning different meanings to their axes.

Why must query and document encoders be compatible?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a vector database, and why is it used in RAG?

A vector database stores vector embeddings and supports efficient similarity search. In RAG, after chunks are embedded, the vectors are stored in a vector database so that a query vector can be compared against them to find the most relevant chunks. Some general databases also support vector storage and search. Index updates and deletions require explicit lifecycle handling.

Worked example. A vector database might store 10,000 vectors, each with 1,536 dimensions, and quickly return the top 5 closest to a query vector.

Vector database

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What happens at query time in a typical dense text RAG pipeline?

Encode the query with a compatible encoder, search permitted current candidates, select the evidence within a budget, and send the question plus source content to the generator. Preserve source IDs for citations and evaluation.

Worked example. For an exam-date question, retrieve the current schedule and a relevant exception; do not pass the whole handbook.

What happens at query time in a typical dense text RAG pipeline?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does the generator receive in a typical text RAG pipeline?

The search vectors help identify candidates. The generator is usually given the selected source text or a controlled representation of it, with source identifiers. Multimodal RAG may also supply images or structured tables in a supported format.

Worked example. A retrieved vector points to a paragraph. The answering prompt contains that paragraph and its document ID, not merely the list of vector coordinates.

What does the generator receive in a typical text RAG pipeline?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Does a high similarity score mean the retrieved passage is true or correct?

No. Similarity measures how close two vectors are in the embedding space, not whether the content is true, current, or authoritative. A high similarity score only indicates that the passage is semantically related to the query according to the embedding model. It does not guarantee correctness or freshness. Similarity is not probability.

Worked example. A query about a recent policy might retrieve an outdated document that is semantically similar but no longer accurate.

Similarity vs truth

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why do index updates and deletions require explicit lifecycle handling?

When source documents change or are removed, the vector index must be updated accordingly. Otherwise, stale or deleted content may still be retrieved. This requires explicit processes for re-embedding updated chunks, removing old vectors, and ensuring consistency between the source and the index. Index updates and deletions require explicit lifecycle handling.

Worked example. If a policy document is updated, the old chunks must be removed and new chunks embedded; otherwise, the system might retrieve outdated policy text.

Index lifecycle

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Where should API keys be kept in a RAG system?

API keys must be kept out of source code and prompts. They should be stored in secure environment variables or secret managers. Retrieved text cannot grant permissions or execute tools, so keys should never be exposed to the model or included in retrieved content. API keys stay out of source and prompts; retrieved text cannot grant permissions or execute tools.

Worked example. An API key should be loaded from an environment variable, not hardcoded in a Python script or included in a prompt.

API key security

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Does RAG guarantee that answers are true or up to date?

No. RAG can improve grounding by supplying retrieved passages, but it does not guarantee truth, completeness, or freshness. The retrieved sources may be outdated, incorrect, or incomplete. Source permissions, versions, and support still matter, and the model can still make mistakes. RAG does not guarantee truth or freshness.

Worked example. A RAG system might retrieve an old policy document and generate an answer that was correct last year but is no longer valid.

RAG limitations

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Do embeddings encode objective truth?

No. Embeddings encode learned associations from training data. They capture patterns and relationships, not objective truth. Similarity in embedding space reflects statistical co-occurrence and context, not factual correctness. Embeddings encode learned associations, not objective truth.

Worked example. Two passages might be embedded close together because they share vocabulary, even if one is false.

Embeddings

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is a character limit not the same as a token limit?

Tokens are units produced by a tokenizer and can vary in length. A character limit counts individual characters, while a token limit counts tokens. The same text can have different token counts depending on the tokenizer. Therefore, setting a chunk size by characters does not guarantee it fits within a token budget. Character chunk limits are not token limits.

Worked example. A 1,000-character chunk might be 200 tokens or 400 tokens depending on the language and tokenizer.

Characters vs tokens

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

03 · Data ingestion

7 cards · Original tutorial episode (opens only when selected).

What are the main stages of the RAG ingestion pipeline?

The ingestion pipeline typically loads source documents, splits them into chunks, embeds each chunk into a vector, and stores those vectors in a vector database. This prepares the knowledge base for later retrieval.

Worked example. A study assistant ingests lecture notes: it loads the notes, splits them into paragraphs, embeds each paragraph, and stores the vectors so that later a student question can retrieve relevant paragraphs.

RAG Ingestion Pipeline

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should a loaded source document retain?

Keep extracted content with an identity and provenance: source, version and location. Extraction alone does not establish that the text is complete or ordered correctly.

Worked example. A PDF loader returns text plus a document ID and page number. Check whether a two-column page was read in the right order.

What should a loaded source document retain?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

Why inspect extracted text before embedding it?

An index can faithfully store bad extraction. Check representative pages, tables, headings and empty outputs before investing in embeddings. Preserve original evidence so errors can be traced.

Worked example. A table loses its column headers during extraction. Repair the extraction before using its numbers as evidence.

Why inspect extracted text before embedding it?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

Why should each chunk keep source metadata?

A chunk needs a stable relationship to its parent source, location, version and access scope. This supports citations, updates, filtering and debugging. Avoid relying on text alone as identity.

Worked example. Two policy versions contain nearly identical paragraphs; version metadata identifies which is current.

Why should each chunk keep source metadata?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

What should an embedding batch preserve?

Keep a reliable mapping from each input chunk to its output vector and record the intended model configuration. Respect provider limits, checkpoint failures and avoid silently skipping or reordering items.

Worked example. If chunk C fails, do not attach chunk D’s vector to C merely because list positions shifted.

What should an embedding batch preserve?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

How can repeated ingestion avoid accidental duplicate records?

Use stable source/chunk identities and an explicit insert-or-update policy. Detect content or configuration changes and rebuild or replace affected derived records deliberately. Re-running a job is not automatically idempotent.

Worked example. Importing the same unchanged handbook twice should not create two search hits for each paragraph.

How can repeated ingestion avoid accidental duplicate records?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

What must be recorded with an index for reproducible retrieval?

Record the corpus/version, extraction and chunking settings, embedding model and preprocessing, vector dimensions, index configuration and relevant filters. An index file alone may not explain how to recreate its behavior.

Worked example. Changing the encoder requires a compatible re-indexing plan; simply reusing an old directory can mix incompatible representations.

What must be recorded with an index for reproducible retrieval?

Recorded English · Question

Recorded English · Explanation and example

Original operating extension to the tutorial; not a claim of caption coverage.

04 · Document retrieval

15 cards · Original tutorial episode (opens only when selected).

What are the main stages of a basic RAG retrieval pipeline?

The pipeline embeds the user query, searches a vector store for the top-k most similar chunks, and passes those chunks plus the original query to the language model.

Worked example. A study assistant receives the question "What year was the fictional company Nova founded?" It embeds the question, retrieves three notes about Nova, and sends them to the model to compose an answer.

Basic retrieval pipeline

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does top-k control in retrieval?

It bounds the number of returned candidates. In an exact fixed ranking, increasing k keeps the earlier candidates and may increase recall, while adding context and cost. Filters, thresholds, a small corpus or approximate-search behavior may yield fewer than k results.

Worked example. If only two permitted passages meet a threshold, requesting five does not create three more passages.

What does top-k control in retrieval?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a score threshold in retrieval, and what is a tradeoff of setting it high?

A score threshold filters out chunks whose similarity score is below a minimum value. Setting it high reduces noise but may return no chunks if no passage meets the threshold.

Worked example. With a threshold of 0.3, a chunk scoring 0.25 is discarded. If the threshold were 0.9, many queries might return nothing.

Score threshold filter

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does it mean to persist a vector store, and why is it useful?

Persistence saves the embeddings and index to disk so they can be reloaded without recomputing. It is useful for reusing the same index across sessions and avoiding repeated embedding costs.

Worked example. A study assistant saves its index to a local directory. On restart, it loads the index and can immediately answer queries without re-embedding all notes.

Vector store persistence

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Does basic vector similarity search need a generative LLM call?

Searching stored vectors does not itself require a generative LLM. Query encoding still uses an embedding model, and more advanced retrieval can use LLM query rewriting or model-based reranking. Separate these stages when reasoning about cost.

Worked example. The simple path is query embedding followed by index search; a later version may add an LLM to rewrite the query.

Does basic vector similarity search need a generative LLM call?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How should you check whether retrieved passages support a question?

Inspect candidate passages against a clear relevance/support rubric and representative queries. A model judge can assist, but calibrate it against reviewed examples and inspect disagreements. Send private text only to an approved destination.

Worked example. For a policy question, mark whether the current rule and its exception were retrieved; a judge saying “yes” is not independent proof.

How should you check whether retrieved passages support a question?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are synthetic questions in the context of testing a RAG system?

They are pre-written questions, often created to cover known facts in the indexed documents, used to evaluate retrieval quality without needing real user queries.

Worked example. A developer writes ten questions about a fictional company's history and checks whether the retriever returns the correct passages for each.

Synthetic question testing

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In a RAG pipeline, which component generates the final answer?

The language model generates the final answer using the retrieved chunks as context. The retriever only selects chunks.

Worked example. The retriever returns five chunks; the LLM reads them and writes a sentence answering the user's question.

Retrieval vs generation

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does it mean to invoke a retriever?

Invoking a retriever runs the similarity search for a given query and returns the top-k chunks. It is a common method name in RAG frameworks.

Worked example. Calling retriever.invoke("What year was Nova founded?") returns a list of relevant chunks.

Invoke retriever

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why evaluate retrieval separately from answer generation?

A wrong answer can result from missing evidence, misleading evidence or misuse of good evidence. Inspecting each stage distinguishes those failures. A model can guess correctly from memory even when retrieval failed, so answer correctness alone does not prove retrieval worked.

Worked example. If the current schedule was absent but the answer happened to name the right date, retrieval still failed the evidence test.

Why evaluate retrieval separately from answer generation?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are some cost considerations when increasing the number of retrieved chunks?

More chunks increase the amount of text sent to the LLM, raising token usage and cost. They also increase the chance of including irrelevant or contradictory information.

Worked example. Retrieving 20 chunks instead of 5 may quadruple the prompt length and cost, while adding noise.

Retrieval cost

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What determines whether retrieved evidence is current?

Freshness depends on source versions, ingestion/index updates, retrieval scope and cache invalidation. RAG is not automatically current. A model may know a later fact from elsewhere, but that does not make stale retrieved evidence adequate.

Worked example. After a fictional deadline changes, update the source and its indexed representation, invalidate affected caches, and test that the old date is no longer used as current.

What determines whether retrieved evidence is current?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why do source permissions matter in RAG?

Retrieval should respect the same access controls as the source documents. If a user does not have permission to view a document, the retriever should not return it.

Worked example. A confidential HR document should not be retrieved for an employee who lacks access.

Permissions

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is versioning important for RAG indexes?

Documents change over time. Without versioning, the index may mix old and new content, leading to inconsistent or outdated answers. Versioning helps track which content is current.

Worked example. A policy document is updated. If the old version remains in the index, the retriever might return outdated rules.

Versioning

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does it mean for a RAG answer to be supported by its sources?

Its factual claims follow from the authorized evidence supplied for the task. A language model is capable of adding unsupported information; source-bounded answering is a requirement to verify, not a mechanical guarantee.

Worked example. A passage names an exam date but says nothing about extra credit. The answer can cite the date; it must not invent an extra-credit policy.

What does it mean for a RAG answer to be supported by its sources?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

05 · Cosine similarity

9 cards · Original tutorial episode (opens only when selected).

What does cosine similarity measure between two vectors?

It measures the angle between the vectors, not their magnitudes. A value near 1 means similar direction, near -1 means opposite direction, and near 0 means little directional relationship. The raw signed cosine is not universally [0, 1].

Worked example. A study assistant compares a query vector with two passage vectors. The first passage points in nearly the same direction, so its cosine is high; the second points elsewhere, so its cosine is lower.

Angle between vectors

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the formula for cosine similarity between vectors a and b?

Cosine similarity is the dot product of a and b divided by the product of their norms: (a · b) / (||a|| * ||b||). If either vector has zero norm, the formula is undefined.

Worked example. For a = [2, 0] and b = [0, 3], the dot product is 0, so cosine similarity is 0 even though the vectors have different lengths.

Cosine formula
import math

def cosine(a, b):
    if not a or len(a) != len(b):
        raise ValueError('vectors must have equal nonzero length')
    na = math.sqrt(sum(x*x for x in a))
    nb = math.sqrt(sum(x*x for x in b))
    if na == 0 or nb == 0:
        raise ValueError('cosine undefined for a zero vector')
    return sum(x*y for x, y in zip(a, b)) / (na*nb)

print(round(cosine([2, 0], [0, 3]), 4))
Expected output
0.0

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How do you calculate the dot product of two vectors?

Multiply matching components and add the results. For a = [0.6, 0.3, 0.2] and b = [0.7, 0.4, 0.1], the dot product is 0.6*0.7 + 0.3*0.4 + 0.2*0.1 = 0.56.

Worked example. A study assistant compares a query vector [1, 2] with a passage vector [3, 4]. The dot product is 1*3 + 2*4 = 11.

Dot product steps
def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

print(round(dot([0.6, 0.3, 0.2], [0.7, 0.4, 0.1]), 2))
Expected output
0.56

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How is the magnitude (norm) of a vector calculated?

Square each component, add the squares, then take the square root. For v = [0.6, 0.3, 0.2], the norm is sqrt(0.6^2 + 0.3^2 + 0.2^2) = sqrt(0.49) = 0.7.

Worked example. For v = [3, 4], the norm is sqrt(9 + 16) = sqrt(25) = 5.

Vector magnitude
import math

def norm(v):
    return math.sqrt(sum(x * x for x in v))

print(norm([0.6, 0.3, 0.2]))
print(norm([3, 4]))
Expected output
0.7
5.0

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why can cosine similarity simplify to the dot product for normalized vectors?

If both vectors have unit length, their norms are 1, so the denominator is 1 and cosine similarity equals the dot product. Many modern embedding pipelines normalize vectors, but this is not universal; verify normalization for your specific model and index.

Worked example. A study assistant uses an embedding model that outputs unit-length vectors. Comparing a query and passage then only requires the dot product, which is faster.

Normalized simplification
import math

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def norm(v):
    return math.sqrt(sum(x * x for x in v))

a = [0.6, 0.8]
b = [0.8, 0.6]
print(round(dot(a, b), 4))
print(round(norm(a), 4), round(norm(b), 4))
Expected output
0.96
1.0 1.0

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the range of signed cosine similarity, and what do the extremes mean?

Signed cosine similarity lies in [-1, 1]. Values near 1 mean similar direction, near -1 mean opposite direction, and near 0 mean little directional relationship. Some systems map or clamp scores, but raw signed cosine is not universally [0, 1].

Worked example. A study assistant compares a query vector [1, 0] with a passage vector [-1, 0]. The cosine is -1, indicating opposite directions.

Cosine range
import math

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def norm(v):
    return math.sqrt(sum(x * x for x in v))

def cosine(a, b):
    return dot(a, b) / (norm(a) * norm(b))

print(cosine([1, 0], [0, 1]))
print(cosine([1, 0], [-1, 0]))
Expected output
0.0
-1.0

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is cosine similarity undefined for a zero vector?

The formula divides by the product of the norms. A zero vector has norm zero, so the denominator becomes zero and division is undefined.

Worked example. A study assistant tries to compare a query vector [0, 0] with a passage vector [1, 2]. The norm of the query is zero, so cosine similarity cannot be computed.

Zero vector problem
import math

def norm(v):
    return math.sqrt(sum(x * x for x in v))

print(norm([0, 0]))
Expected output
0.0

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Does a high cosine similarity score prove that a retrieved passage is true or current?

No. Cosine similarity measures directional similarity between learned vector representations. It does not guarantee truth, freshness, source permissions, or correctness. Those require separate checks.

Worked example. A study assistant retrieves a passage with cosine 0.98, but the passage is outdated. The high score only means the vectors are similar, not that the content is current.

Similarity vs truth

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

If two vectors have the same number of dimensions, are they always comparable with cosine similarity?

No. Same dimensionality is necessary but not sufficient. The query encoder and document encoder must be compatible, meaning they were trained or designed to produce vectors in the same space. Otherwise the comparison is meaningless.

Worked example. A study assistant uses one model to embed queries and a different, incompatible model to embed documents. Both produce 384-dimensional vectors, but the similarity scores are unreliable.

Compatible encoders

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

06 · First RAG application

5 cards · Original tutorial episode (opens only when selected).

What does the answer generation step do in a RAG pipeline?

It takes the user's query and the retrieved document chunks, combines them into a prompt, and sends that prompt to a language model to produce a response grounded in those chunks. The response is not necessarily plain text; it can be structured or multimodal depending on the model and task.

Worked example. A student asks when a fictional study assistant called StudyBuddy was first released. The retriever finds a chunk stating the year, and the generator uses that chunk to answer.

Answer generation flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why might a RAG prompt instruct the model to answer only from the provided documents?

To reduce the chance that the model uses its own training knowledge, which could be outdated or incorrect, and to keep the answer grounded in the retrieved evidence. This reduces hallucination but does not guarantee truth or freshness.

Worked example. If the documents do not mention a fictional study assistant's price, the model should say it lacks information rather than guessing from memory.

Grounding instruction

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should a RAG system do when the retrieved documents do not contain the answer?

It should respond that it does not have enough information to answer based on the provided documents, rather than fabricating an answer.

Worked example. If a student asks about a feature not mentioned in any retrieved chunk, the system should say it lacks enough information.

Fallback path

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How should a RAG request organize the question and retrieved evidence?

Place the question and bounded evidence in the model API’s supported message structure. Clearly identify sources and treat retrieved text as untrusted data. A single formatted string is one implementation; separate content parts and messages are also possible.

Worked example. An MC request labels a source D4 and quotes its policy text separately from the instruction to answer the learner’s question.

How should a RAG request organize the question and retrieved evidence?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should higher-priority instructions establish in a RAG assistant?

They define the task, evidence-use policy, uncertainty behavior and boundaries. They are not a substitute for server-side access checks or tool permissions, and provider message hierarchies differ.

Worked example. Tell the assistant to cite supplied evidence and acknowledge missing support; separately enforce which documents the caller can retrieve.

What should higher-priority instructions establish in a RAG assistant?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

07 · Conversational retrieval

8 cards · Original tutorial episode (opens only when selected).

In the tutorial’s demonstrated conversation pipeline, in basic RAG, how is each user query treated?

In this demonstrated implementation, each query is treated independently: the retriever searches using the exact question, without considering previous turns.

Worked example. A user asks "What is the Zephyr laptop?" and later "What is its battery life?" In basic RAG, the second query is searched as-is, so "its" is not resolved.

Basic RAG: independent queries

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What extra step does history-aware RAG add before retrieval?

Query reformulation: the system uses the conversation history to rewrite the latest question into a clear, standalone, searchable question.

Worked example. History: "Tell me about the Zephyr laptop." New question: "What is its battery life?" Reformulated: "What is the Zephyr laptop's battery life?"

History-aware RAG: reformulation step

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why can follow-up questions retrieve the wrong evidence without conversation context?

Pronouns and omitted names can make the question ambiguous. The retriever may still find passages, but cannot reliably infer the intended referent from missing context. Use relevant history or ask for clarification.

Worked example. After discussing two courses, “When is its exam?” needs the intended course identified before a reliable search.

Why can follow-up questions retrieve the wrong evidence without conversation context?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What conversation state should a history-aware RAG system retain?

The example stores user and assistant messages. A real system needs a bounded retention/context policy, session isolation and privacy controls; it need not send every past turn on each request. Prior assistant claims are not authoritative source evidence.

Worked example. Retain the current course reference to interpret “its exam,” while keeping another user’s conversation out of this session.

What conversation state should a history-aware RAG system retain?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In the tutorial’s demonstrated conversation pipeline, when is query reformulation performed in history-aware RAG?

In this demonstrated implementation, only when chat history is not empty; if history is empty, the original question is used directly.

Worked example. On the first turn, no reformulation happens. On later turns, the system rewrites the question using the history.

Reformulation condition

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does the reformulation prompt ask the LLM to do?

It provides the chat history and asks the LLM to rewrite the new question into a standalone, searchable question, returning only the rewritten question.

Worked example. System prompt: "Given the chat history, rewrite the new question to be standalone and searchable. Return only the rewritten question."

Reformulation prompt

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In the tutorial’s demonstrated conversation pipeline, which query is used for retrieval in history-aware RAG?

In this demonstrated implementation, the reformulated standalone query, not the original user question.

Worked example. If the user asks "What is its battery life?", the retriever searches for "What is the Zephyr laptop's battery life?"

Retrieval with reformulated query

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In the tutorial’s demonstrated conversation pipeline, how is the final answer generated in history-aware RAG?

In this demonstrated implementation, the system combines the retrieved chunks and the original user question, includes the chat history, and asks the LLM to answer based on the provided documents.

Worked example. The prompt includes the chat history, the retrieved chunks about the Zephyr laptop, and the user's question, then the LLM produces an answer.

Answer generation

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

08 · Chunking strategies

7 cards · Original tutorial episode (opens only when selected).

Where does chunking fit in a typical document-indexing pipeline?

After extraction, before embedding/indexing, chunking determines the initial evidence units. Poor boundaries can weaken retrieval. Neighbor expansion or parent-document retrieval can restore surrounding context if it was preserved, so a split is not inherently irreversible.

Worked example. A boundary separates a rule from its exception. Preserve source offsets so retrieval can expand to the containing section.

Where does chunking fit in a typical document-indexing pipeline?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are the main failure modes of fixed character splitting?

Fixed character splitting cuts at a set character count, so it can split mid-sentence, separate related concepts into different chunks, and lose context across the boundary. It also ignores document structure. The result is poor retrieval because the chunk no longer contains a complete, coherent idea.

Worked example. A 500-character splitter cuts a paragraph between 'supply chain' and 'challenges and inflation.' A query about supply-chain challenges may retrieve only the first half, which lacks the explanation.

Fixed cut failure

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are the tradeoffs between chunks that are too small and chunks that are too large?

Too-small chunks lack surrounding context, so retrieval may return fragments that cannot answer the question. Too-large chunks contain too much noise, can exceed embedding model or context window limits, and dilute relevance. Good chunking balances context against noise and respects natural boundaries.

Worked example. A one-sentence chunk about ATP may omit how ATP is used. A ten-page chunk may include many unrelated topics, so the embedding represents a blur and retrieval precision drops.

Chunk size tradeoff

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does recursive splitting use a separator hierarchy?

Try preferred separators first and fall back to finer separators when a piece is too large. The configured list determines the boundaries. Default character splitters do not necessarily detect sentences; punctuation boundaries must be configured or handled by a sentence-aware method.

Worked example. A hierarchy can try blank lines, line breaks, spaces and individual characters. It prefers larger units but checks the actual resulting chunks.

How does recursive splitting use a separator hierarchy?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does document-specific splitting respect, and why does it matter?

Document-specific splitting respects the structure of the source format: pages, sections, headers, rows, or slides. A PDF, Markdown file, CSV, DOCX, or PPT each gets appropriate treatment. This keeps structural units together, which improves retrieval because a section or table row is more likely to be a coherent answer unit.

Worked example. For a Markdown study guide, splitting at heading boundaries keeps each topic's explanation together. For a CSV, splitting by row keeps each record intact instead of cutting across columns.

Structure-aware split

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How should you choose a chunking strategy?

Choose candidates that fit the format and task, then compare them on the same representative queries, evidence coverage, latency and cost. Fixed-size, recursive, structure-aware, semantic and LLM-guided methods are alternatives; there is no universal quality or cost ranking.

Worked example. Start with section-aware chunks for a structured handbook. Keep the simpler approach if a more expensive method gives no useful improvement.

How should you choose a chunking strategy?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why must chunking strategies be evaluated rather than assumed to be best?

Chunking quality depends on the document, the queries, and the retrieval setup, so no strategy is universally best. A strategy that helps one corpus can hurt another by splitting coherent ideas or merging unrelated ones. You should measure retrieval and answer quality on representative queries and compare strategies on your own data.

Worked example. A fictional study assistant might test recursive splitting against semantic splitting on the same set of exam questions. If recursive splitting retrieves the needed passages just as often at lower cost, the extra semantic compute is not justified.

Evaluate chunking

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

09 · Recursive splitting

3 cards · Original tutorial episode (opens only when selected).

How does a separator-based character splitter build chunks?

It splits at the chosen separator, then merges neighboring pieces toward a size target. The configured length function, separator retention and overlap affect actual lengths. Count separators when they are retained.

Worked example. With zero overlap and a retained two-character separator, pieces of lengths 18, 51 and 19 total 92 characters, not 88: 18 + 2 + 51 + 2 + 19.

How does a separator-based character splitter build chunks?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What happens when a single piece exceeds the chunk size and contains no separator?

The character text splitter keeps that piece intact, producing a chunk larger than the configured limit.

Worked example. A 200-character paragraph with no double newline and a chunk size of 100 remains a single 200-character chunk.

Oversized piece remains intact

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does a recursive character text splitter differ from a basic character text splitter?

It accepts an ordered list of separators and recursively tries them in priority order: if a piece is still too large after splitting with the first separator, it applies the next separator to that piece.

Worked example. With separators [double newline, single newline, period, space], a long paragraph without double newlines is split at periods or spaces only if needed.

Recursive splitting with separator list

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

10 · Semantic chunking

8 cards · Original tutorial episode (opens only when selected).

What is semantic chunking?

Semantic chunking splits a document into meaningful pieces by detecting where topics naturally change, using embeddings to compare adjacent sentences and placing boundaries where similarity drops significantly.

Worked example. A study assistant processes a lecture transcript. Instead of cutting every 500 characters, it groups sentences about 'exam scheduling' together and starts a new chunk when the topic shifts to 'grading policy'.

Semantic chunking flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a percentile threshold?

A percentile is a quantile convention for locating a value in an ordered distribution. Finite samples, ties and interpolation mean it is not always literally a value with exactly that percentage strictly below it. State the calculation method.

Worked example. The median of the values 1, 2, 8 and 9 is 5 under the usual midpoint convention, even though 5 is not an observed value.

What is a percentile threshold?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How can a linearly interpolated percentile be computed?

For sorted values indexed from zero, one common convention uses h=(n-1)p. Interpolate between floor(h) and ceil(h), where p is the percentile expressed from 0 to 1. Other conventions, such as nearest rank, differ.

Worked example. For [0.1, 0.2, 0.8, 0.9] at p=0.75, h=2.25. Interpolate one quarter from 0.8 to 0.9 to get 0.825.

How can a linearly interpolated percentile be computed?
values = [0.1, 0.2, 0.8, 0.9]
p = 0.75
h = (len(values) - 1) * p
lo = int(h)
hi = min(lo + 1, len(values) - 1)
threshold = values[lo] + (h - lo) * (values[hi] - values[lo])
print(round(threshold, 3))
Expected output
0.825

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How can a distance percentile define semantic chunk boundaries?

Compute distances between adjacent sentence representations or sentence windows, then split where distance exceeds a chosen percentile threshold. Equivalently, a low similarity can signal a boundary, but the inequality and percentile direction must match the chosen score.

Worked example. For distances [0.1, 0.2, 0.8, 0.9], a linearly interpolated 75th-percentile threshold is 0.825. With a strict greater-than rule, only the 0.9 boundary splits.

How can a distance percentile define semantic chunk boundaries?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why can a percentile threshold help, and what can it fail to detect?

It adapts a threshold to the observed score distribution instead of using one absolute cutoff for every document. A relative tail can still contain no meaningful topic boundary, and ties or noise can distort the result. Validate boundaries and apply size constraints.

Worked example. A uniform document may have only tiny distance differences. Selecting its largest distances does not prove real topic changes.

Why can a percentile threshold help, and what can it fail to detect?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does changing a percentile affect semantic split frequency?

For a fixed set of distances and a split-above rule, a higher percentile raises the threshold and produces no more splits. For a fixed set of similarities and a split-below rule, a higher percentile raises the threshold and produces no fewer splits. Ties and minimum/maximum size rules also matter.

Worked example. With distances [0.1, 0.2, 0.8, 0.9], a higher distance threshold is more conservative. Do not copy that conclusion to a similarity-below rule.

How does changing a percentile affect semantic split frequency?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a major practical drawback of semantic chunking?

Semantic chunking usually adds embedding/comparison work, for individual sentences or sentence windows depending on the method. Evaluate total ingestion cost, latency and evidence quality rather than assuming it is always too expensive.

Worked example. A handbook collection may work just as well with a simpler section-based split, making extra semantic processing unnecessary for that task.

Cost drawback

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the first step of semantic chunking before embeddings are computed?

The document is split into individual sentences. Each sentence becomes a unit that will later be embedded and compared with its neighbors.

Worked example. A study assistant takes a paragraph of five sentences and separates them into five individual sentences before embedding each one.

Sentence splitting

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

11 · Agent-based chunking

7 cards · Original tutorial episode (opens only when selected).

What is agent-based chunking?

Agent-based chunking is a strategy where a language model is prompted to decide where to split a long text into logical chunks. The model returns the text with explicit split markers, and a program then cuts at those markers to produce the final chunks.

Worked example. A study assistant has a long note mixing photosynthesis and cellular respiration. It asks a model to insert a marker at topic boundaries, then splits on the marker to get two topic-focused chunks.

Agent-based chunking flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are the main steps to perform agent-based chunking?

First, write a prompt that tells the model to split the text into logical chunks, gives size or boundary rules, and asks it to insert a unique split marker. Second, send the text and prompt to the model. Third, programmatically split the returned text at the marker, strip whitespace, and collect the chunks.

Worked example. A study assistant prompts: 'Split this note into chunks of at most 120 characters at topic boundaries. Insert <<SPLIT>> between chunks.' It then splits the response on <<SPLIT>> and strips each piece.

Procedure steps
text = "Photosynthesis makes sugar.<<SPLIT>>Respiration breaks sugar."
chunks = [c.strip() for c in text.split("<<SPLIT>>")]
print(chunks)
Expected output
['Photosynthesis makes sugar.', 'Respiration breaks sugar.']

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What can an LLM contribute to chunk-boundary selection?

It can propose boundaries using contextual cues that a simple separator rule misses. That is a hypothesis to test, not proof of optimal chunks. Preserve original ordering, content and source locations unless a separately designed transformation explicitly changes them.

Worked example. A model may suggest a boundary between an exam policy and a worked example; validate that both extracted spans exactly match the original source.

What can an LLM contribute to chunk-boundary selection?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What additional costs and failure modes can LLM-guided chunking introduce?

Model calls add latency, expense and output variability. A model may omit, alter or duplicate source text, miss length limits, or follow injected instructions. It is not universally the slowest method, but needs measured justification and deterministic output validation.

Worked example. If returned chunks silently drop an exception, a fluent summary is not an acceptable lossless split.

What additional costs and failure modes can LLM-guided chunking introduce?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Is agent-based chunking always the best chunking strategy?

No. It has its place, but it is not universally better. For simple FAQ documents, simpler methods like recursive character splitting may be sufficient. For complex PDFs with images, tables, and layouts, even agent-based chunking may struggle, and specialized extraction tools may be needed.

Worked example. A study assistant for simple flashcards might use fixed-size chunks, while a system for scanned textbooks with tables might need layout-aware extraction before chunking.

Fit by document type

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why might complex PDFs require specialized extraction tools beyond agent-based chunking?

Complex PDFs often contain images, tables, and multi-column layouts where text boundaries are not obvious even to humans. Agent-based chunking works on text, so it may miss or misread layout evidence. Specialized tools use OCR, table transformers, and layout detection to extract structured data and preserve provenance.

Worked example. A PDF with a two-column article and a table may be read in the wrong order by a text-only model, mixing columns. A layout-aware extractor can keep columns separate and preserve table structure.

Complex PDF challenge

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does ETL do when preparing sources for retrieval?

Extract, Transform, Load: obtain source content, convert it into the intended representation, then load it into the target store. Inputs may be structured, semi-structured or unstructured; CSV is commonly tabular structured data, not automatically unstructured.

Worked example. Extract a PDF table, preserve its headers and page reference in records, then load those records into a retrieval index.

What does ETL do when preparing sources for retrieval?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

12 · Multimodal RAG

19 cards · Original tutorial episode (opens only when selected).

What is multimodal RAG?

Multimodal RAG is retrieval-augmented generation over sources that contain more than plain text, such as tables, images, formulas, captions, and layout. It retrieves evidence from those sources and uses a language model to answer with that evidence. It is one form of RAG, not the definition of all RAG.

Worked example. A study assistant indexes a PDF that has paragraphs, a table of layer counts, and a diagram. A user asks about the layer counts. The system retrieves the relevant chunk and sends the raw table to the answering model.

Multimodal RAG scope

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In document parsing, what is an atomic element?

An atomic element is the smallest useful unit a parser recognizes in a document. It has a type, such as title, narrative text, table, image, formula, footer, or figure caption, and it usually keeps metadata such as page number or coordinates.

Worked example. A parser reads a page and returns separate elements: a title element, three narrative-text elements, one table element, and one image element. The paragraph is not split into sentences; it stays one narrative-text element.

Atomic elements

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does title-based chunking group atomic elements?

Title-based chunking starts at a title and collects following atomic elements until it reaches the next title. The result is a chunk that usually corresponds to a section. It may contain only text, or it may mix text with a table or an image.

Worked example. A document has title `2 Background`, then three paragraphs, then title `3 Model Architecture`. The chunker groups the title and three paragraphs into one chunk, then starts a new chunk at `3 Model Architecture`.

Title-based chunking

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is a mixed chunk, and why does its representation matter?

It contains multiple evidence types, such as prose plus a table or image. A text-only embedding model cannot directly consume raw image pixels. A pipeline can use a checked textual surrogate or a compatible multimodal representation.

Worked example. A handbook section pairs a diagram with explanatory prose. Choose a representation that retains the information users will ask about.

What is a mixed chunk, and why does its representation matter?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why might a text-based retriever index a summary of mixed evidence?

A searchable text description can represent evidence that a text-only encoder cannot read directly. It is lossy and may be wrong. Validate it and retain a link to the original text, table or image; other designs use multimodal encoders directly.

Worked example. A diagram summary says which component sends data to another. Retrieval uses that description; the answering model can inspect the preserved diagram when needed.

Why might a text-based retriever index a summary of mixed evidence?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why keep raw text, raw tables, and raw images in metadata after summarizing a chunk?

The summary is used for retrieval, but it is lossy. Final answer generation should use the original evidence so the model can see exact numbers, table structure, and visual relationships. Metadata is where that raw evidence can travel with the chunk. This assumes the raw evidence was preserved during parsing and chunking.

Worked example. A summary says the model uses several layers. The raw table says the base model uses six layers. If only the summary is sent, the exact number may be lost. Sending the raw table lets the model answer precisely.

Summary vs raw evidence

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What two attributes does a LangChain-style document usually have?

It usually has page content and metadata. Page content holds the main text that will be embedded or passed along. Metadata holds supporting information such as source name, page number, raw text, raw table, or raw image.

Worked example. A document for a mixed chunk has page content equal to the searchable summary. Its metadata contains the original paragraph, the table as HTML, and the image bytes.

Document shape

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In multimodal RAG, what should be sent to the answering model after retrieval?

Send the original raw evidence associated with the retrieved chunks, such as raw text, structured table HTML, and raw image data, along with the user question. Do not rely only on the summary that was embedded for retrieval. This assumes the raw evidence was preserved and is available in the chunk metadata.

Worked example. The retriever returns three chunks. For each chunk, the system extracts the raw paragraph, the table HTML, and any image bytes from metadata. It sends those plus the question to the answering model.

Retrieve then answer

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why convert an extracted table to structured HTML instead of relying on plain OCR text?

Plain OCR text often reads left to right and loses which value belongs to which row and column. Structured HTML preserves row and column relationships, so a language model can reason about the table more accurately. HTML is not a guarantee of perfect extraction; complex or merged cells can still be misread.

Worked example. A table has columns `Layer type` and `Complexity per layer`. Plain OCR might produce a jumbled line. HTML keeps each cell under its column, so the model can answer which complexity belongs to which layer type.

Table structure

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is OCR text from an image often not enough for multimodal RAG?

OCR may extract jumbled or partial text from a diagram and can miss spatial relationships, arrows, and labels. The raw image preserves the visual evidence, so it should be kept and sent when the question depends on the diagram. OCR quality varies with image resolution and layout.

Worked example. A diagram has labels `Add`, `Norm`, and `Feed Forward`. OCR might return them in a confusing order. The raw image shows their positions and connections, so the model can describe the architecture correctly.

Image evidence

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What do maximum characters, new-after characters, and combine-under characters control in chunking?

Maximum characters is a hard upper limit for a chunk. New-after characters is a soft target for starting a new chunk. Combine-under characters merges very small chunks with neighbors so the index does not fill with tiny fragments. These are heuristics; the right values depend on the document and retrieval task and should be evaluated.

Worked example. A chunker is set to maximum 3000 characters, new after 2400 characters, and combine under 500 characters. A 200-character fragment is merged with a neighbor, while a 3200-character section is split before exceeding the maximum.

Chunk size controls

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are the main stages of a multimodal RAG ingestion pipeline?

Upload and queue the source, partition it into atomic elements, chunk the elements (for example by title), summarize chunks that contain tables or images, convert the chunks into documents with page content and metadata, then embed and store them in a vector database.

Worked example. A PDF is uploaded, queued, partitioned into text, table, and image elements, chunked by title, and the mixed chunks are summarized. The summaries are embedded and stored, while raw text, tables, and images are kept in metadata.

Multimodal ingestion stages

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should you check when configuring a document parser?

Check extraction strategy, layout handling, table-structure extraction, image extraction and whether original elements or payloads are retained. Parameter names and combinations are library/version-specific; verify the installed API and inspect actual output.

Worked example. Selecting image extraction alone may not include an inline base64 payload. Confirm that the output has the intended image reference or bytes.

What should you check when configuring a document parser?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does base64 do when an image is sent through an API?

It encodes binary bytes as text for a supported transport format. It does not explain the image to a text-only model. The endpoint must accept that image representation and the chosen model must support the image task.

Worked example. A supported multimodal API may accept an image content part containing a data URL. Pasting the same base64 text into an ordinary text prompt is not equivalent.

What does base64 do when an image is sent through an API?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should a summarization prompt for a mixed chunk try to produce?

It should produce a searchable description that covers key facts, numbers, and data points from text and tables, describes charts or diagrams in images, and includes alternative search terms users might use. The goal is findability, not brevity.

Worked example. A prompt asks the model to summarize a paragraph and a table, list possible questions the content could answer, and include alternative search terms. The resulting summary is embedded for retrieval.

Summarization prompt goals

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should a final answer prompt instruct the model to do when using retrieved raw evidence?

Instruct the model to answer from the supplied evidence and acknowledge when evidence is insufficient. Provide text, tables and images through supported input types. Instructions alone do not guarantee grounding; evaluate claim support.

Worked example. A supported multimodal request supplies the question, selected notes, table structure and an image input. Ask for supporting document references or an explicit explanation of missing evidence.

Final answer prompt

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why can document parsers need dependencies outside a Python package?

A parser may call PDF utilities, OCR engines or file-type libraries. Required tools depend on parser mode and version. Install them in the intended runtime, which may be a container or managed environment; they need not be installed globally on a personal computer.

Worked example. A remote parsing container can include an OCR engine while the desktop only submits approved files.

Why can document parsers need dependencies outside a Python package?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How can you inspect the atomic elements inside a composite chunk?

A composite chunk usually exposes its child atomic elements through metadata, often under a key like original elements. Iterating that list lets you see each child's type, text, and other attributes.

Worked example. A chunk's metadata contains an original elements list with a title, narrative text, footer, image, and figure caption. Iterating the list reveals each child element.

Composite chunk internals

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does an orchestrator method do in a multimodal RAG pipeline?

An orchestrator method runs the whole ingestion flow in order: partition the document, chunk it, summarize mixed chunks, create the vector store, and return a retriever or database that can be used for chatting. It packages the steps so they can be rerun with a new source or database location.

Worked example. A method takes a PDF path and a database location, runs partitioning, chunking, summarization, and vector store creation, then returns a retriever. Running it again with a new location creates a second database.

Orchestrator flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

13 · Advanced retrieval

6 cards · Original tutorial episode (opens only when selected).

How do similarity search, score thresholds and MMR differ?

Similarity search ranks candidates for the query. A threshold excludes candidates that fail a score rule. Maximal Marginal Relevance selects from a candidate pool by balancing query relevance and redundancy with already-selected items. Diversity is not a reason to return unrelated topics.

Worked example. For an exam-policy query, MMR may choose the main rule and its exception instead of three near-duplicate copies of the main rule.

How do similarity search, score thresholds and MMR differ?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why can top-k search return irrelevant passages?

If no relevance threshold or abstention rule is applied, the highest-ranked available passages can still be poor matches. The system returns up to k eligible candidates, even when all are weak. Empty or filtered corpora can return fewer or none.

Worked example. A query about astronomy searched against only gardening notes can return the least bad gardening matches.

Why can top-k search return irrelevant passages?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What problem does MMR solve that plain similarity search does not?

Plain similarity search can return multiple chunks that say nearly the same thing, wasting context. MMR selects chunks that are relevant but also diverse from each other, giving the LLM broader context.

Worked example. A study assistant queried with 'tell me about Python' returns three chunks all saying 'Python is a programming language.' MMR instead returns one on syntax, one on applications, and one on libraries.

Redundancy vs diversity

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does MMR select a diverse set of candidates?

Retrieve an initial candidate pool, then select items iteratively using a balance between relevance to the query and similarity to items already selected. In a common convention, lambda near one emphasizes query relevance and near zero emphasizes novelty. The metric and implementation must be stated.

Worked example. After selecting the main policy, a distinct exception passage can add more useful evidence than a near-duplicate copy.

How does MMR select a diverse set of candidates?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

In MMR, what do k, fetch_k, and lambda control?

k is the final number of chunks returned. fetch_k is the initial candidate pool size. lambda balances relevance (near 1) and diversity (near 0); 0.5 is a middle setting.

Worked example. A study assistant sets k=3, fetch_k=10, lambda=0.5: it retrieves 10 candidates, then picks 3 balancing relevance and diversity.

MMR parameters

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

When should you use MMR, and when should you avoid it?

Consider MMR for redundant candidates and questions benefiting from multiple aspects. Its relevance/diversity tradeoff and extra work may not help a narrow lookup. Compare task-specific quality, recall and latency rather than assuming either method is always better.

Worked example. A study assistant researching a broad topic uses MMR to get diverse perspectives. A lookup for a specific fact uses plain similarity for speed and precision.

MMR use cases

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

14 · Multi-query retrieval

5 cards · Original tutorial episode (opens only when selected).

Why generate multiple query variations for a single user question in retrieval?

Different phrasings can surface relevant documents that the original query might miss because they use different vocabulary or angles. This increases recall, but also increases noise and cost, and does not guarantee truth or freshness.

Worked example. A student asks, 'How does spaced repetition improve memory?' A variation like 'What are the benefits of spaced repetition for learning?' might retrieve a chunk about cognitive benefits that the original query missed.

Multi-query retrieval flow

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How can an LLM be used to generate query variations, and what output format is useful?

The LLM is prompted to produce several alternative queries that rephrase the original question from different angles. A structured output such as a JSON object containing a list of strings makes it easy to extract the variations programmatically.

Worked example. Prompt: 'Generate three different variations of this query that would help retrieve relevant documents. Original query: How does photosynthesis work? Return three alternative queries.' The LLM might output: {'queries': ['What are the steps of photosynthesis?', 'How do plants convert light into energy?', 'What is the process of photosynthesis?']}.

LLM generates query variations

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the retrieval step in multi-query retrieval?

For each generated query variation, the system runs a separate retrieval, producing a ranked list of chunks. These lists are stored for later fusion. The number of chunks per list is controlled by a parameter such as k.

Worked example. With three variations and k=5, you get three lists, each containing five chunks. The lists may overlap; the same chunk can appear in multiple lists at different ranks.

Retrieve for each variation

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is it problematic to simply concatenate the ranked lists from multiple queries and take the top results?

Concatenation makes list order dominate, may repeat documents and does not combine evidence of rank across queries. Fusion can combine rank contributions, but repeated retrieval is not proof of truth.

Worked example. If document A is rank 1 in list 1 but absent elsewhere, and document B is rank 2 in all three lists, concatenation might put A first even though B is more consistently relevant.

Naive concatenation problem

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What must be preserved when rewriting queries for multi-query retrieval?

The rewritten queries must retain the original intent and any access permissions. They must not be used to bypass security controls or retrieve documents the user is not authorized to see. Retrieved text cannot grant permissions or execute tools.

Worked example. If the original query is restricted to a certain document set, all variations must be restricted to the same set. A variation that asks for confidential information should not be allowed if the user lacks access.

Preserve intent and permissions

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

15 · Reciprocal rank fusion

10 cards · Original tutorial episode (opens only when selected).

What is the formula for the RRF score of an item?

The RRF score is the sum over all ranked lists where the item appears of 1/(k + rank), where rank is the 1-based position of the item in that list and k is a constant (commonly 60). If the item does not appear in a list, it contributes 0 from that list.

Worked example. At ranks 1 and 3 with k=60, the score is 1/61 + 1/63 = 0.032266458496..., which rounds to 0.03227 at five decimal places. Round the final sum rather than adding rounded intermediate values.

RRF score calculation
def rrf_score(ranks, k=60):
    return sum(1.0 / (k + r) for r in ranks)

# Example: item appears at rank 1 and rank 3
print(round(rrf_score([1, 3]), 5))
Expected output
0.03227

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the role of the constant k in RRF, and why is k=60 commonly used?

The constant k in RRF controls how much lower ranks are penalized. With k=0, scores drop sharply (1, 1/2, 1/3, ...), overemphasizing top positions. A larger k, such as 60, makes the score differences between adjacent ranks smaller, so lower-ranked items can still contribute meaningfully if they appear in multiple lists. k=60 is a widely used default that balances top-rank importance with consensus effects.

Worked example. With k=0, rank 1 scores 1.0 and rank 2 scores 0.5 (a 50% drop). With k=60, rank 1 scores 1/61 ≈ 0.0164 and rank 2 scores 1/62 ≈ 0.0161 (a small drop).

Effect of k on rank scores
def score(rank, k):
    return 1.0 / (k + rank)

print(round(score(1, 0), 2), round(score(2, 0), 2))
print(round(score(1, 60), 4), round(score(2, 60), 4))
Expected output
1.0 0.5
0.0164 0.0161

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the consensus effect in RRF?

The consensus effect means that items appearing in multiple ranked lists receive a boost because their RRF score is the sum of contributions from each list. This rewards items that are consistently retrieved across different query variations or retrieval methods, even if they are not always at the top of every list.

Worked example. If a study note appears at rank 2 in two different query result lists, its RRF score is 1/(60+2) + 1/(60+2) ≈ 0.0323, which may outrank a note that appears at rank 1 in only one list (score ≈ 0.0164).

Consensus effect boosts multi-list items
def rrf_score(ranks, k=60):
    return sum(1.0 / (k + r) for r in ranks)

print(round(rrf_score([2, 2]), 4))
print(round(rrf_score([1]), 4))
Expected output
0.0323
0.0164

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is deduplication important when combining ranked lists with RRF?

Deduplication ensures that each unique item is counted once per list. If the same item appears multiple times in a single ranked list, it should be treated as one entry with its best (or first) rank, otherwise its score would be artificially inflated. RRF assumes each list contains distinct items.

Worked example. If a study note appears twice in List A (at ranks 1 and 3), you should keep only the rank 1 occurrence for scoring, not sum both.

Deduplicate within each list

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is an advantage of RRF over directly combining similarity scores from different retrieval methods?

RRF uses only rank positions, not raw similarity scores. This avoids the problem of mixing scores from different scales or distributions (e.g., one method's scores range 0-1, another's 0-100). RRF is simple and often performs well without score normalization. However, it does not guarantee correctness and cannot recover items absent from all lists.

Worked example. One retriever returns signed cosine scores from minus one to one; another returns BM25 scores on a different scale. Rank fusion avoids directly adding these raw scores.

RRF avoids score scale issues

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What are some limitations of RRF?

RRF does not guarantee that the fused results are correct, truthful, or fresh. It only combines rankings. It cannot recover an item that is absent from all input lists. It also does not consider the actual similarity scores, so it may not reflect fine-grained relevance differences. Additionally, RRF assumes each list is deduplicated and that ranks are meaningful.

Worked example. If a relevant document is not retrieved by any query variation, RRF cannot include it. Also, if two documents have very similar relevance but different ranks, RRF may over-penalize the lower-ranked one depending on k.

RRF limitations

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

When implementing RRF, what are the key steps?

Key steps: (1) Collect multiple ranked lists (e.g., from different query variations). (2) For each list, ensure items are deduplicated. (3) For each item, compute its RRF score by summing 1/(k + rank) across lists where it appears, using a chosen k (often 60). (4) Sort items by descending RRF score. (5) Select top items for downstream use. Implementation requires programming skills and understanding of data structures; it is a later step after learning Python, vectors, and data handling.

Worked example. In a study assistant, you might generate three query variations, retrieve top 5 notes for each, then apply RRF to merge them into a single ranked list of unique notes.

RRF implementation steps

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Given two ranked lists with k=60, how do you compute the RRF score for an item?

For each list where the item appears, compute 1/(60 + rank). Sum these values. For example, if an item is at rank 1 in List A and rank 2 in List B, its score is 1/(60+1) + 1/(60+2) = 1/61 + 1/62 ≈ 0.01639 + 0.01613 = 0.03252.

Worked example. List A: [X, Y, Z]; List B: [Y, X, W]. X is rank 1 in A and rank 2 in B: score = 1/61 + 1/62 ≈ 0.03252. Y is rank 2 in A and rank 1 in B: same score. Z is rank 3 in A only: 1/63 ≈ 0.01587. W is rank 3 in B only: 1/63 ≈ 0.01587.

RRF score for an item in two lists
def rrf_score(ranks, k=60):
    return sum(1.0 / (k + r) for r in ranks)

print(round(rrf_score([1, 2]), 5))
print(round(rrf_score([3]), 5))
Expected output
0.03252
0.01587

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How are ties handled in RRF?

RRF can produce ties when items have the same total score. The method itself does not specify a tie-breaking rule; you can break ties arbitrarily (e.g., by original order, by best rank, or by another criterion). In practice, ties are often acceptable, and the final selection may depend on downstream needs.

Worked example. X and Y have the same reciprocal-rank-fusion score and the same best rank. A documented stable document-ID order resolves this tie reproducibly.

Tie handling in RRF

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does RRF help when there is a limit on how many chunks can be sent to a language model?

RRF ranks the deduplicated union of candidate lists. A separate selection step must enforce the actual context/token budget; the formula does not shrink unique content by itself.

Worked example. Three lists of five entries can contain seven unique chunks. Select a subset after fusion and count actual tokens, including instructions and output capacity.

RRF reduces chunks for context limit

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

16 · Hybrid retrieval

6 cards · Original tutorial episode (opens only when selected).

Why combine lexical and dense retrieval?

They can make complementary errors: lexical methods are useful for identifiers and exact terminology, while dense methods can connect paraphrases. Either can miss evidence. Combine candidates and evaluate whether the added complexity improves the actual task.

Worked example. The exact identifier CS101-Q2 may be easy for lexical search; “second programming quiz” may be easier for semantic retrieval.

Why combine lexical and dense retrieval?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Which factors matter in BM25 scoring?

BM25 variants use query-term matching, term rarity, saturating within-document term frequency and document-length normalization. Tokenization, parameters and the corpus affect the score. Repeating a term is not rewarded linearly forever.

Worked example. A short note containing the rare course ID can be useful, but its rank cannot be calculated from term rarity alone without the other inputs.

Which factors matter in BM25 scoring?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

When can lexical retrieval miss a paraphrase?

A basic token-matching setup may fail when query and document share no matching terms. Real lexical search can add normalization, stemming, synonyms and query expansion, so “keyword search only matches exact strings” is too broad.

Worked example. A plain setup may miss “car” for a query containing “automobile”; an explicit synonym configuration could bridge that gap.

When can lexical retrieval miss a paraphrase?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How should candidate counts be chosen for a production RAG system?

Tune them against measured recall, permission filtering, latency, reranking cost and the final context budget. Production does not always need more candidates than a demo. Fusion deduplicates the union; it does not necessarily produce the sum of list lengths.

Worked example. Two retrievers returning 50 candidates each produce between 50 and 100 unique candidates if both lists contain 50 unique items, depending on overlap.

How should candidate counts be chosen for a production RAG system?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How do weights affect RRF in hybrid retrieval?

Weights scale each list’s contribution in weighted RRF. They need not sum to one: multiplying all weights by the same positive constant preserves the ranking. Tune relative weights on representative evaluation.

Worked example. With vector weight 0.7 and lexical weight 0.3, a rank-one vector hit contributes 0.7/61 before adding any lexical contribution.

Weighted RRF

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is multi-query hybrid search?

It generates multiple variations of the user query, runs hybrid retrieval for each, and then fuses the results (e.g., with RRF). This can increase recall but also increases cost and noise. Each variation must preserve the original intent and permissions.

Worked example. A user query is rewritten into 5 variations. Each runs hybrid search, producing 5 ranked lists. RRF merges them into one list.

Multi-query hybrid

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

17 · Reranking

16 cards · Original tutorial episode (opens only when selected).

What does a reranking stage do?

It scores or reorders an existing candidate set for the task, often before selecting evidence for generation. A trained model is one option; rules and other scoring functions can also rerank. It cannot introduce an absent candidate unless a separate retrieval step adds one.

Worked example. Retrieve twenty permitted passages, then apply a query-aware score and select a smaller supported evidence set.

What does a reranking stage do?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why is reranking described as a two-stage strategy?

Stage one is fast and broad: embeddings or hybrid search cast a wide net and return many candidates. Stage two is precise and focused: a reranker reorders those candidates. Each stage is optimized for a different purpose, like a screening interview followed by a detailed technical interview.

Worked example. A study assistant first retrieves 50 chunks cheaply, then reranks only 20 of them carefully.

Two-stage retrieval

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why might embedding retrieval benefit from reranking?

A passage can be topically similar without answering the question. Joint query-passage scoring can sometimes distinguish finer relationships. Whether it improves accuracy enough to justify cost is an empirical question; embeddings are not universally insufficient.

Worked example. An exam query needs an exception rule, while the top vector match merely describes the course.

Why might embedding retrieval benefit from reranking?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does a bi-encoder process a query and a document?

It encodes the query and the document separately into vectors, then compares those vectors, often with cosine similarity. The model never sees the query and document together. Document encoding often happens ahead of time, and the query is encoded later when the user asks.

Worked example. A study assistant embeds all notes ahead of time, then embeds a new question and compares vectors.

Bi-encoder

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How does a cross-encoder process a query and a document?

It processes the query and document jointly, often as one combined input with a separator token, so the model can attend to their relationship and produce a relevance score.

Worked example. A study assistant sends the question and one note together to a reranker, which returns a relevance score for that pair.

Cross-encoder

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What tradeoff motivates a retrieve-then-rerank design?

Precomputed document vectors support broad candidate search. A more expensive query-dependent scorer can then examine a smaller set. Actual speed, cost and quality depend on the models, corpus and hardware; neither stage is guaranteed superior at every task.

Worked example. Compare direct top-five retrieval with retrieving thirty candidates and reranking them to five, using the same test questions.

What tradeoff motivates a retrieve-then-rerank design?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

When might a reranker be unnecessary?

When the candidate set is already very small, such as five or ten chunks, and the initial retrieval already puts the best chunks near the top. Sending them directly to the LLM may be enough.

Worked example. A study assistant retrieves only five notes and all five are clearly relevant, so reranking adds little.

Small candidate set

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Can a reranker recover a relevant chunk that was never retrieved in stage one?

No. A reranker can only reorder candidates it receives. If a relevant chunk is absent from the candidate set, reranking cannot bring it back.

Worked example. A study assistant's retriever misses the note that answers the question, so the reranker never sees it and cannot rank it.

Missing candidate

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Are all rerankers cross-encoders?

No. A cross-encoder is a common type of reranker that jointly scores query-document pairs, but reranking can also be done by other model types or scoring methods.

Worked example. A study assistant uses a reranking service that scores pairs, but the underlying model is not necessarily a cross-encoder.

Reranker types

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What does the top-n parameter control in a reranking step?

It controls how many of the reranked documents are returned. Without it, the reranker may return the full reordered list; with it, only the top n are kept.

Worked example. A study assistant reranks 18 chunks and asks for the top 5 to send to the LLM.

Top-n

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Why might a reranker move a metric-heavy chunk above a generic announcement chunk?

The query asks about financial performance, so chunks containing specific metrics are more relevant than generic statements about plans. The reranker scores each chunk against the query and reorders accordingly.

Worked example. For a query about a company's financial performance, a chunk with revenue figures moves above a chunk about expanding a factory.

Reordering

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is the cost-benefit tradeoff of adding a reranker?

A reranker adds latency and computational cost, but it can improve the precision of the top chunks. It is most useful when the candidate set is large and accuracy matters.

Worked example. A study assistant adds a reranker for a large exam-prep corpus but skips it for a small FAQ with five entries.

Cost vs benefit

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How do reranking and answer generation differ as tasks?

Reranking selects or orders evidence; generation composes an answer. The same type of model can sometimes perform either task, including an LLM used as a reranker. Separate the role from the model family and compare cost for the actual setup.

Worked example. One call scores candidate policy passages; another writes an answer using selected passages and citations.

How do reranking and answer generation differ as tasks?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What should guide a reranker model or service choice?

Use representative evaluation, language/document support, latency, deployment resources, privacy requirements and actual pricing. Check current model documentation. Hosted and self-hosted options each have tradeoffs; a brand or free tier is not proof of fit.

Worked example. Evaluate two approved candidates on the same permitted MC-style questions before choosing either.

What should guide a reranker model or service choice?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

What is needed to use a cloud reranker API like Cohere's?

An API key, typically stored as an environment variable, is required. The key should be kept out of source code and notebooks that might be shared.

Worked example. A study assistant stores the reranker API key in an environment variable before running the notebook.

API key setup

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

How do you test whether reranking helped?

Compare both rankings against labeled relevance or evidence-support judgments, not merely whether positions changed. Measure the downstream answer, latency and cost on held-out questions too. An unchanged ranking can be correct; a changed ranking can be worse.

Worked example. If the only supporting exception moves from eighth to second, that may improve top-five recall. Confirm the label and compare outcomes.

How do you test whether reranking helped?

Recorded English · Question

Recorded English · Explanation and example

Original teaching companion to auto-generated captions; corrected with source review. On-screen tutorial code has not been independently transcribed.

Corrections and further references

Sentence Transformers retrieval and reranking · Original RAG paper · RAG evaluation · RAG security

Source timestamps · Scope and correction ledger · Release verification

Original study explanations and diagrams. Auto captions can contain errors; on-screen source code was not independently transcribed. Recorded speech is checked with independent transcription and waveform analysis. Subjective listening review is not claimed.