{"article":{"slug":"spring-ai-modular-rag-and-typesafe-jev-retrieve-more-keep-only-what-answers","title":"Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers","subtitle":null,"summary":"Christian Tzolov walks through Spring AI modular RAG plus TypeSafe Jev: rewrite/expand the query before retrieval, then keep only chunks that actually answer—contrasting naive QuestionAnswerAdvisor RAG with a live demo pipeline.","content_type":"blog_post","language":"en","canonical_url":"https://spring.io/blog/2026/10/02/spring-ai-modular-rag-typesafe-jev/","author":{"name":"Christian Tzolov","url":"https://spring.io/blog/2026/10/02/spring-ai-modular-rag-typesafe-jev/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Spring","url":"https://automatictransmission.khoury.northeastern.edu/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"},{"name":"Java","slug":"java","url":"https://listedarticles.com/topics/java"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1487,"reading_minutes":6,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-02T12:10:39.212Z","updated_at":"2026-10-02T12:10:39.212Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/spring-ai-modular-rag-and-typesafe-jev-retrieve-more-keep-only-what-answers","markdown_url":"https://listedarticles.com/articles/spring-ai-modular-rag-and-typesafe-jev-retrieve-more-keep-only-what-answers.md","example":false,"citation":"Christian Tzolov, Spring. \"Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers.\" 2 Oct 2026. https://spring.io/blog/2026/10/02/spring-ai-modular-rag-typesafe-jev/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://spring.io/blog/2026/10/02/spring-ai-modular-rag-typesafe-jev/"},"body_markdown":"# Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers\n\nThis article puts the two together. Two LLM calls clean up and multiply the user's question before retrieval. Jev then judges every retrieved chunk before it reaches the prompt. The whole pipeline is one advisor builder.\n\n\n💡 Demo : The complete example is the [`05-1-modular-rag`](https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-1-modular-rag) module. It sits next to [`05-rag`](https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-rag), the naive version, so you can diff the two. Every output quoted below is from a live run.\n\n\n\n## Why Naive RAG Falls Short\n\nThe classic Spring AI RAG demo is a `QuestionAnswerAdvisor` over a vector store. It embeds the user's text, fetches similar chunks and stuffs them into the prompt. That works on stage and gets shaky in real life, for three reasons:\n\n\n- Users don't type search queries. They type \"I'm a Florida resident and heard a lot about storms last fall. Did Milton actually make landfall in my state, and where?\" Embedding that verbatim drags all the noise into the search.\n\n- Similar is not the same as useful. A vector store returns chunks that are about the topic. A reference list that mentions \"Hurricane Milton\" five times scores high and answers nothing.\n\n- Retrieved text is untrusted input. Whatever sits in your documents goes straight into the prompt, including text written to hijack the model.\n\n\nWe will answer exactly that question over a Wikipedia PDF about Hurricane Milton, and fix each problem with one pluggable component.\n\n## Modular RAG in Spring AI\n\nThe `RetrievalAugmentationAdvisor`, from the `spring-ai-rag` module, implements the architecture described in [Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks](https://arxiv.org/abs/2407.21059). From the outside it is just another advisor: it sits between your prompt and the model, runs the retrieval phases, and appends the result to the user message. The system prompt and the question pass through untouched.\n\n\nInside, each phase is a small interface you can swap:\n\n\n\n\nPhase\nSpring AI interface\nJob\nUsed in this demo\n\n\n\n\nPre-retrieval\n`QueryTransformer`\nrewrite one query into a better one\n`RewriteQueryTransformer`\n\n\nPre-retrieval\n`QueryExpander`\nturn one query into several\n`MultiQueryExpander`\n\n\nRetrieval\n`DocumentRetriever`\nfetch documents for one query\n`VectorStoreDocumentRetriever`\n\n\nRetrieval\n`DocumentJoiner`\nmerge the results of all queries\n`ConcatenationDocumentJoiner` (default)\n\n\nPost-retrieval\n`DocumentPostProcessor`\nfilter, rerank or compress documents\n`JevDocumentFilter`, `JevDocumentReranker`\n\n\nGeneration\n`QueryAugmenter`\nput context and question into the prompt\n`ContextualQueryAugmenter`\n\n\n\nHere is the full flow for one call:\n\n\nThe retrieval branches run in parallel, one per query, and are joined before post-processing. Note one subtle detail: the post-processors and the augmenter see the user's original question, not the rewritten one. The rewrite only exists to improve search.\n\n## Getting Started\n\nAdd the Spring AI RAG module and the two TypeSafe artifacts from the previous article. Spring AI's BOM manages the first one:\n\n```\n\norg.springframework.ai\nspring-ai-rag\n\n\n\norg.springaicommunity\nspring-ai-starter-typesafe\n0.3.0\n\n\n\norg.springaicommunity\ntypesafe-spring-ai\n0.3.0\n\n```\n\nThe starter auto-configures a `TypeSafeClient` bean from your API key:\n\n```\nspring.ai.typesafe.api-key=${TYPESAFE_API_KEY}\n```\n\nThe demo also uses `spring-ai-pdf-document-reader` for ingestion and `spring-ai-starter-model-transformers` for local embeddings.\n\n## The Pipeline\n\nIngestion is the same as in any Spring AI RAG application: read the PDF, split it into chunks, store them.\n\n```\nvectorStore.add(\nTokenTextSplitter.builder().build().split(\nnew PagePdfDocumentReader(hurricaneDocs).read()));\n```\n\nThe interesting part is a single builder. Each numbered comment is one brick from the diagram:\n\n```\n// Separate builder for the LLM-backed stages, so they don't inherit the main client's advisors\nvar ragClientBuilder = chatClientBuilder.clone();\n\nvar modularRag = RetrievalAugmentationAdvisor.builder()\n// 1. Pre-retrieval: rewrite the chatty question into a search-friendly query\n.queryTransformers(RewriteQueryTransformer.builder()\n.chatClientBuilder(ragClientBuilder)\n.build())\n// 2. Pre-retrieval: expand it into several diverse queries (original included)\n.queryExpander(MultiQueryExpander.builder()\n.chatClientBuilder(ragClientBuilder)\n.numberOfQueries(3)\n.build())\n// 3. Retrieval: similarity search per query; the default joiner dedups the union\n.documentRetriever(VectorStoreDocumentRetriever.builder()\n.vectorStore(vectorStore)\n.similarityThreshold(0.5)\n.topK(4)\n.build())\n// 4. Post-retrieval: Jev drops bad passages, then reranks and keeps the top 3\n.documentPostProcessors(\nJevDocumentFilter.builder(typeSafeClient).build(),\nJevDocumentReranker.builder(typeSafeClient).topK(3).build())\n// 5. Generation: stuff the surviving context into the prompt\n.queryAugmenter(ContextualQueryAugmenter.builder()\n.allowEmptyContext(true)\n.build())\n.taskExecutor(taskExecutor) // see \"Things to know\" below\n.build();\n```\n\nUsing it looks like any other advisor:\n\n```\nString answer = chatClient.prompt()\n.advisors(modularRag)\n.user(\"I'm a Florida resident and heard a lot about storms last fall. \"\n+ \"Did Milton actually make landfall in my state, and where?\")\n.call()\n.content();\n```\n\n`DocumentRetriever` and `DocumentPostProcessor` are both single-method interfaces, so a lambda is enough to log what flows through the pipeline. The demo wraps the retriever and adds two printing post-processors around the Jev stages; that is where the output below comes from.\n\n## Jev in the Post-Retrieval Phase\n\nBoth components send one Jev call per passage, with the query and the passage as the state , and turn the answers into a decision.\n\n[JevDocumentFilter](https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentFilter/) asks four `Noul` questions about every passage and applies them in order:\n\n\n\n\nQuestion\nDefault threshold\nOutcome\n\n\n\n\n`contains_prompt_injection`\nabove `0.70`\nexcluded\n\n\n`contradicts_query_premise`\nabove `0.70`\nkept, tagged `CONFLICTING`\n\n\n`is_relevant`\nbelow `0.45`\nexcluded\n\n\n`contains_answer_evidence`\nabove `0.55`\nkept, otherwise excluded\n\n\n\nThis is the [atomic questions](https://spring.io/blog/2026/09/21/spring-ai-typesafe-structured-judgment) idea from the previous article at work: four narrow questions, four thresholds, one call. Contradicting passages survive on purpose. If the user assumes something false, the model should see the evidence that says so. The classification is stored in the document metadata under `jev.classification`, and the thresholds are a `Policy` record you can override.\n\n[JevDocumentReranker](https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentReranker/) asks a single question, could this passage answer the query? , sorts by the score and keeps the top K. The score lands in the metadata under `jev.rerank.score`. Its focus is deliberate: whether the passage states information that answers the query, not merely whether it covers the same subject . That is precisely the gap between similarity and usefulness.\n\nOrder matters. Filter first, so the reranker only pays for the survivors.\n\n## A Live Run\n\nThe rewrite and expansion turned one chatty sentence into four focused searches:\n\n```\n[retrieve] Hurricane Milton 2024 Florida landfall location\n[retrieve] Hurricane Milton track path across Florida from Gulf Coast to Atlantic and National Hurricane Center landfall report\n[retrieve] Siesta Key Sarasota County Milton landfall timeline, storm surge, and areas impacted\n[retrieve] Where did Hurricane Milton make landfall in Florida in October 2024 and at what intensity\n```\n\nThe joiner merged their results into six unique chunks, all above the `0.5` similarity threshold:\n\n```\n[retrieved] 0.84 Hurricane Milton ...\n[retrieved] 0.72 6. \"Hurricane Milton Makes Landfall On Florida's West Coast ...\n[retrieved] 0.70 Hurricane Milton's landfall\" (https://www.wesh.com/...\n[retrieved] 0.69 Archived (http ...\n[retrieved] 0.59 /news/tropical-storm-milton-forms-gulf-of-mexico ...\n[retrieved] 0.54 of Key West. [128] Across the state, about 125 homes were d...\n```\n\nFour of them are reference-list entries and archive links. They mention Milton a lot and answer nothing. `JevDocumentFilter` excluded them for lacking answer evidence, and `JevDocumentReranker` scored the two survivors:\n\n```\n[jev-reranked] 0.97 Hurricane Milton ...\n[jev-reranked] 0.88 6. \"Hurricane Milton Makes Landfall On Florida's West Coast ...\n```\n\nSo `topK(3)` returned only two chunks, and that is the point. The model got less context, all of it relevant:\n\n```\nYes. Hurricane Milton made landfall in Florida, near Siesta Key, on the evening of\nOctober 9, 2024. It had weakened to a Category 3 hurricane by then.\n```\n\nVector similarity answers \"what is this text about?\" . Jev answers \"does this text answer the question?\" . A RAG pipeline needs both.\n\n\n\n### ⚠️ Things to know\n\nThe default executor keeps a command-line app alive. `RetrievalAugmentationAdvisor` runs per-query retrieval on its own thread pool, and those threads are non-daemon. The demo printed its answer and then hung. Passing Spring Boot's auto-configured `TaskExecutor` through `.taskExecutor(...)` fixes it, and with `spring.threads.virtual.enabled=true` you get virtual threads for free.\n\nNewer Claude models reject `temperature`. The Spring AI docs recommend temperature 0 for query transformers. `claude-sonnet-5-5` answers that with HTTP 400, so the demo clones the builder only to keep the main client's advisors out of the rewrite and expand calls.\n\nBoth Jev components fail open. If the Jev API is unreachable, passages pass through unscreened or unscored and a warning is logged. Your pipeline degrades to plain modular RAG instead of breaking.\n\nEvery brick costs latency. This pipeline adds two LLM calls before retrieval and one Jev call per retrieved chunk after it. Jev calls run four at a time by default (`batchOptions(...)` changes that), but measure before you ship.\n\n\n\n## Conclusion\n\nModular RAG turns retrieval from one opaque advisor into a pipeline you can read, log and swap piece by piece. Here are the important insights:\n\n\n- Fix the question before you search. A rewrite plus three variants found better chunks than the raw sentence would.\n\n- Fix the results before you prompt. Similarity found six chunks about Milton; Jev kept the two that answer, and screened every passage for prompt injection on the way.\n\n- Filter, then rerank. Both cost one call per passage, so let the filter shrink the list first.\n\n\n\n## Resources\n\nDemo\n\n\n- [`05-1-modular-rag`](https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-1-modular-rag) — the code in this article\n\n- [`05-rag`](https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-rag) — the naive `QuestionAnswerAdvisor` version, for comparison\n\n\nSpring AI\n\n\n- [Retrieval Augmented Generation](https://docs.spring.io/spring-ai/reference/api/retrieval-augmented-generation.html) — advisors and Modular RAG components\n\n- [Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks](https://arxiv.org/abs/2407.21059) — the paper behind the architecture\n\n\nSpring AI TypeSafe\n\n\n- [JevDocumentFilter](https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentFilter/) and [JevDocumentReranker](https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentReranker/)\n\n- [Reference documentation](https://spring-ai-community.github.io/spring-ai-typesafe/latest/) and [GitHub repository](https://github.com/spring-ai-community/spring-ai-typesafe)\n\n- [TypeSafe cookbooks](https://docs.typesafe.ai/cookbooks) — reranking and RAG classification recipes\n\n\nRelated Spring AI articles\n\n\n- [Spring AI and TypeSafe Jev: Fast, Cheap, Structured Decisions](https://spring.io/blog/2026/09/21/spring-ai-typesafe-structured-judgment)","body_html":"<h1 id=\"spring-ai-modular-rag-and-typesafe-jev-retrieve-more-keep-only-w\">Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers</h1>\n<p>This article puts the two together. Two LLM calls clean up and multiply the user&#39;s question before retrieval. Jev then judges every retrieved chunk before it reaches the prompt. The whole pipeline is one advisor builder.</p>\n<p>💡 Demo : The complete example is the <a href=\"https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-1-modular-rag\" rel=\"nofollow ugc noopener\"><code>05-1-modular-rag</code></a> module. It sits next to <a href=\"https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-rag\" rel=\"nofollow ugc noopener\"><code>05-rag</code></a>, the naive version, so you can diff the two. Every output quoted below is from a live run.</p>\n<h2 id=\"why-naive-rag-falls-short\">Why Naive RAG Falls Short</h2>\n<p>The classic Spring AI RAG demo is a <code>QuestionAnswerAdvisor</code> over a vector store. It embeds the user&#39;s text, fetches similar chunks and stuffs them into the prompt. That works on stage and gets shaky in real life, for three reasons:</p>\n<ul><li>Users don&#39;t type search queries. They type &quot;I&#39;m a Florida resident and heard a lot about storms last fall. Did Milton actually make landfall in my state, and where?&quot; Embedding that verbatim drags all the noise into the search.</li><li>Similar is not the same as useful. A vector store returns chunks that are about the topic. A reference list that mentions &quot;Hurricane Milton&quot; five times scores high and answers nothing.</li><li>Retrieved text is untrusted input. Whatever sits in your documents goes straight into the prompt, including text written to hijack the model.</li></ul>\n<p>We will answer exactly that question over a Wikipedia PDF about Hurricane Milton, and fix each problem with one pluggable component.</p>\n<h2 id=\"modular-rag-in-spring-ai\">Modular RAG in Spring AI</h2>\n<p>The <code>RetrievalAugmentationAdvisor</code>, from the <code>spring-ai-rag</code> module, implements the architecture described in <a href=\"https://arxiv.org/abs/2407.21059\" rel=\"nofollow ugc noopener\">Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks</a>. From the outside it is just another advisor: it sits between your prompt and the model, runs the retrieval phases, and appends the result to the user message. The system prompt and the question pass through untouched.</p>\n<p>Inside, each phase is a small interface you can swap:</p>\n<p>Phase\nSpring AI interface\nJob\nUsed in this demo</p>\n<p>Pre-retrieval\n<code>QueryTransformer</code>\nrewrite one query into a better one\n<code>RewriteQueryTransformer</code></p>\n<p>Pre-retrieval\n<code>QueryExpander</code>\nturn one query into several\n<code>MultiQueryExpander</code></p>\n<p>Retrieval\n<code>DocumentRetriever</code>\nfetch documents for one query\n<code>VectorStoreDocumentRetriever</code></p>\n<p>Retrieval\n<code>DocumentJoiner</code>\nmerge the results of all queries\n<code>ConcatenationDocumentJoiner</code> (default)</p>\n<p>Post-retrieval\n<code>DocumentPostProcessor</code>\nfilter, rerank or compress documents\n<code>JevDocumentFilter</code>, <code>JevDocumentReranker</code></p>\n<p>Generation\n<code>QueryAugmenter</code>\nput context and question into the prompt\n<code>ContextualQueryAugmenter</code></p>\n<p>Here is the full flow for one call:</p>\n<p>The retrieval branches run in parallel, one per query, and are joined before post-processing. Note one subtle detail: the post-processors and the augmenter see the user&#39;s original question, not the rewritten one. The rewrite only exists to improve search.</p>\n<h2 id=\"getting-started\">Getting Started</h2>\n<p>Add the Spring AI RAG module and the two TypeSafe artifacts from the previous article. Spring AI&#39;s BOM manages the first one:</p>\n<pre><code>\norg.springframework.ai\nspring-ai-rag\n\n\n\norg.springaicommunity\nspring-ai-starter-typesafe\n0.3.0\n\n\n\norg.springaicommunity\ntypesafe-spring-ai\n0.3.0\n</code></pre>\n<p>The starter auto-configures a <code>TypeSafeClient</code> bean from your API key:</p>\n<pre><code>spring.ai.typesafe.api-key=${TYPESAFE_API_KEY}</code></pre>\n<p>The demo also uses <code>spring-ai-pdf-document-reader</code> for ingestion and <code>spring-ai-starter-model-transformers</code> for local embeddings.</p>\n<h2 id=\"the-pipeline\">The Pipeline</h2>\n<p>Ingestion is the same as in any Spring AI RAG application: read the PDF, split it into chunks, store them.</p>\n<pre><code>vectorStore.add(\nTokenTextSplitter.builder().build().split(\nnew PagePdfDocumentReader(hurricaneDocs).read()));</code></pre>\n<p>The interesting part is a single builder. Each numbered comment is one brick from the diagram:</p>\n<pre><code>// Separate builder for the LLM-backed stages, so they don&#39;t inherit the main client&#39;s advisors\nvar ragClientBuilder = chatClientBuilder.clone();\n\nvar modularRag = RetrievalAugmentationAdvisor.builder()\n// 1. Pre-retrieval: rewrite the chatty question into a search-friendly query\n.queryTransformers(RewriteQueryTransformer.builder()\n.chatClientBuilder(ragClientBuilder)\n.build())\n// 2. Pre-retrieval: expand it into several diverse queries (original included)\n.queryExpander(MultiQueryExpander.builder()\n.chatClientBuilder(ragClientBuilder)\n.numberOfQueries(3)\n.build())\n// 3. Retrieval: similarity search per query; the default joiner dedups the union\n.documentRetriever(VectorStoreDocumentRetriever.builder()\n.vectorStore(vectorStore)\n.similarityThreshold(0.5)\n.topK(4)\n.build())\n// 4. Post-retrieval: Jev drops bad passages, then reranks and keeps the top 3\n.documentPostProcessors(\nJevDocumentFilter.builder(typeSafeClient).build(),\nJevDocumentReranker.builder(typeSafeClient).topK(3).build())\n// 5. Generation: stuff the surviving context into the prompt\n.queryAugmenter(ContextualQueryAugmenter.builder()\n.allowEmptyContext(true)\n.build())\n.taskExecutor(taskExecutor) // see &quot;Things to know&quot; below\n.build();</code></pre>\n<p>Using it looks like any other advisor:</p>\n<pre><code>String answer = chatClient.prompt()\n.advisors(modularRag)\n.user(&quot;I&#39;m a Florida resident and heard a lot about storms last fall. &quot;\n+ &quot;Did Milton actually make landfall in my state, and where?&quot;)\n.call()\n.content();</code></pre>\n<p><code>DocumentRetriever</code> and <code>DocumentPostProcessor</code> are both single-method interfaces, so a lambda is enough to log what flows through the pipeline. The demo wraps the retriever and adds two printing post-processors around the Jev stages; that is where the output below comes from.</p>\n<h2 id=\"jev-in-the-post-retrieval-phase\">Jev in the Post-Retrieval Phase</h2>\n<p>Both components send one Jev call per passage, with the query and the passage as the state , and turn the answers into a decision.</p>\n<p><a href=\"https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentFilter/\" rel=\"nofollow ugc noopener\">JevDocumentFilter</a> asks four <code>Noul</code> questions about every passage and applies them in order:</p>\n<p>Question\nDefault threshold\nOutcome</p>\n<p><code>contains_prompt_injection</code>\nabove <code>0.70</code>\nexcluded</p>\n<p><code>contradicts_query_premise</code>\nabove <code>0.70</code>\nkept, tagged <code>CONFLICTING</code></p>\n<p><code>is_relevant</code>\nbelow <code>0.45</code>\nexcluded</p>\n<p><code>contains_answer_evidence</code>\nabove <code>0.55</code>\nkept, otherwise excluded</p>\n<p>This is the <a href=\"https://spring.io/blog/2026/09/21/spring-ai-typesafe-structured-judgment\" rel=\"nofollow ugc noopener\">atomic questions</a> idea from the previous article at work: four narrow questions, four thresholds, one call. Contradicting passages survive on purpose. If the user assumes something false, the model should see the evidence that says so. The classification is stored in the document metadata under <code>jev.classification</code>, and the thresholds are a <code>Policy</code> record you can override.</p>\n<p><a href=\"https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentReranker/\" rel=\"nofollow ugc noopener\">JevDocumentReranker</a> asks a single question, could this passage answer the query? , sorts by the score and keeps the top K. The score lands in the metadata under <code>jev.rerank.score</code>. Its focus is deliberate: whether the passage states information that answers the query, not merely whether it covers the same subject . That is precisely the gap between similarity and usefulness.</p>\n<p>Order matters. Filter first, so the reranker only pays for the survivors.</p>\n<h2 id=\"a-live-run\">A Live Run</h2>\n<p>The rewrite and expansion turned one chatty sentence into four focused searches:</p>\n<pre><code>[retrieve] Hurricane Milton 2024 Florida landfall location\n[retrieve] Hurricane Milton track path across Florida from Gulf Coast to Atlantic and National Hurricane Center landfall report\n[retrieve] Siesta Key Sarasota County Milton landfall timeline, storm surge, and areas impacted\n[retrieve] Where did Hurricane Milton make landfall in Florida in October 2024 and at what intensity</code></pre>\n<p>The joiner merged their results into six unique chunks, all above the <code>0.5</code> similarity threshold:</p>\n<pre><code>[retrieved] 0.84 Hurricane Milton ...\n[retrieved] 0.72 6. &quot;Hurricane Milton Makes Landfall On Florida&#39;s West Coast ...\n[retrieved] 0.70 Hurricane Milton&#39;s landfall&quot; (https://www.wesh.com/...\n[retrieved] 0.69 Archived (http ...\n[retrieved] 0.59 /news/tropical-storm-milton-forms-gulf-of-mexico ...\n[retrieved] 0.54 of Key West. [128] Across the state, about 125 homes were d...</code></pre>\n<p>Four of them are reference-list entries and archive links. They mention Milton a lot and answer nothing. <code>JevDocumentFilter</code> excluded them for lacking answer evidence, and <code>JevDocumentReranker</code> scored the two survivors:</p>\n<pre><code>[jev-reranked] 0.97 Hurricane Milton ...\n[jev-reranked] 0.88 6. &quot;Hurricane Milton Makes Landfall On Florida&#39;s West Coast ...</code></pre>\n<p>So <code>topK(3)</code> returned only two chunks, and that is the point. The model got less context, all of it relevant:</p>\n<pre><code>Yes. Hurricane Milton made landfall in Florida, near Siesta Key, on the evening of\nOctober 9, 2024. It had weakened to a Category 3 hurricane by then.</code></pre>\n<p>Vector similarity answers &quot;what is this text about?&quot; . Jev answers &quot;does this text answer the question?&quot; . A RAG pipeline needs both.</p>\n<h3 id=\"things-to-know\">⚠️ Things to know</h3>\n<p>The default executor keeps a command-line app alive. <code>RetrievalAugmentationAdvisor</code> runs per-query retrieval on its own thread pool, and those threads are non-daemon. The demo printed its answer and then hung. Passing Spring Boot&#39;s auto-configured <code>TaskExecutor</code> through <code>.taskExecutor(...)</code> fixes it, and with <code>spring.threads.virtual.enabled=true</code> you get virtual threads for free.</p>\n<p>Newer Claude models reject <code>temperature</code>. The Spring AI docs recommend temperature 0 for query transformers. <code>claude-sonnet-5-5</code> answers that with HTTP 400, so the demo clones the builder only to keep the main client&#39;s advisors out of the rewrite and expand calls.</p>\n<p>Both Jev components fail open. If the Jev API is unreachable, passages pass through unscreened or unscored and a warning is logged. Your pipeline degrades to plain modular RAG instead of breaking.</p>\n<p>Every brick costs latency. This pipeline adds two LLM calls before retrieval and one Jev call per retrieved chunk after it. Jev calls run four at a time by default (<code>batchOptions(...)</code> changes that), but measure before you ship.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Modular RAG turns retrieval from one opaque advisor into a pipeline you can read, log and swap piece by piece. Here are the important insights:</p>\n<ul><li>Fix the question before you search. A rewrite plus three variants found better chunks than the raw sentence would.</li><li>Fix the results before you prompt. Similarity found six chunks about Milton; Jev kept the two that answer, and screened every passage for prompt injection on the way.</li><li>Filter, then rerank. Both cost one call per passage, so let the filter shrink the list first.</li></ul>\n<h2 id=\"resources\">Resources</h2>\n<p>Demo</p>\n<ul><li><a href=\"https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-1-modular-rag\" rel=\"nofollow ugc noopener\"><code>05-1-modular-rag</code></a> — the code in this article</li><li><a href=\"https://github.com/tzolov/voxxeddays2026-demo/tree/main/05-rag\" rel=\"nofollow ugc noopener\"><code>05-rag</code></a> — the naive <code>QuestionAnswerAdvisor</code> version, for comparison</li></ul>\n<p>Spring AI</p>\n<ul><li><a href=\"https://docs.spring.io/spring-ai/reference/api/retrieval-augmented-generation.html\" rel=\"nofollow ugc noopener\">Retrieval Augmented Generation</a> — advisors and Modular RAG components</li><li><a href=\"https://arxiv.org/abs/2407.21059\" rel=\"nofollow ugc noopener\">Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks</a> — the paper behind the architecture</li></ul>\n<p>Spring AI TypeSafe</p>\n<ul><li><a href=\"https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentFilter/\" rel=\"nofollow ugc noopener\">JevDocumentFilter</a> and <a href=\"https://spring-ai-community.github.io/spring-ai-typesafe/latest/rag/JevDocumentReranker/\" rel=\"nofollow ugc noopener\">JevDocumentReranker</a></li><li><a href=\"https://spring-ai-community.github.io/spring-ai-typesafe/latest/\" rel=\"nofollow ugc noopener\">Reference documentation</a> and <a href=\"https://github.com/spring-ai-community/spring-ai-typesafe\" rel=\"nofollow ugc noopener\">GitHub repository</a></li><li><a href=\"https://docs.typesafe.ai/cookbooks\" rel=\"nofollow ugc noopener\">TypeSafe cookbooks</a> — reranking and RAG classification recipes</li></ul>\n<p>Related Spring AI articles</p>\n<ul><li><a href=\"https://spring.io/blog/2026/09/21/spring-ai-typesafe-structured-judgment\" rel=\"nofollow ugc noopener\">Spring AI and TypeSafe Jev: Fast, Cheap, Structured Decisions</a></li></ul>","headings":[{"level":1,"text":"Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers","id":"spring-ai-modular-rag-and-typesafe-jev-retrieve-more-keep-only-w"},{"level":2,"text":"Why Naive RAG Falls Short","id":"why-naive-rag-falls-short"},{"level":2,"text":"Modular RAG in Spring AI","id":"modular-rag-in-spring-ai"},{"level":2,"text":"Getting Started","id":"getting-started"},{"level":2,"text":"The Pipeline","id":"the-pipeline"},{"level":2,"text":"Jev in the Post-Retrieval Phase","id":"jev-in-the-post-retrieval-phase"},{"level":2,"text":"A Live Run","id":"a-live-run"},{"level":3,"text":"⚠️ Things to know","id":"things-to-know"},{"level":2,"text":"Conclusion","id":"conclusion"},{"level":2,"text":"Resources","id":"resources"}]}}