{"article":{"slug":"building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows","title":"Building AI Agents with Spring AI — Tool Calling, Memory, and Autonomous Workflows","subtitle":null,"summary":"A practical Java guide to Spring AI agents: tool calling, memory, and autonomous workflows that move beyond one-shot LLM calls into real multi-step agent loops.","content_type":"tutorial","language":"en","canonical_url":"https://dev.to/ayshriv/building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows-2kgj","author":{"name":"Ayush Shrivastava","url":"https://dev.to/ayshriv","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"DEV Community","url":"https://dev.to/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2810,"reading_minutes":12,"published_at":"2026-09-20T15:12:00.427Z","added_at":"2026-09-20T15:12:00.427Z","updated_at":"2026-09-20T15:12:00.427Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows","markdown_url":"https://listedarticles.com/articles/building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows.md","example":false,"citation":"Ayush Shrivastava, DEV Community. \"Building AI Agents with Spring AI — Tool Calling, Memory, and Autonomous Workflows.\" 20 Sept 2026. https://dev.to/ayshriv/building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows-2kgj (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://dev.to/ayshriv/building-ai-agents-with-spring-ai-tool-calling-memory-and-autonomous-workflows-2kgj"},"body_markdown":"Large Language Models are excellent at generating text.\n\nBut generation alone isn't enough to build truly useful AI applications.\n\nImagine asking an AI assistant:\n\n\n```\nWhat's the status of my order?\n```\nA normal LLM can explain how order tracking works.\n\nBut it cannot magically access your order database.\n\nOr suppose you ask:\n\n\n```\nCancel my order #ORD-10291.\n```\nThe model can tell you how to cancel an order.\n\nBut it cannot actually cancel anything unless your application gives it the ability to perform that action.\n\nThis is where **tool calling** and **AI agents** come in.\n\nInstead of simply generating an answer, an AI application can:\n\n\n```\nUnderstand the request\n        ↓\nDecide what action is required\n        ↓\nSelect a tool\n        ↓\nExecute the tool\n        ↓\nObserve the result\n        ↓\nContinue reasoning\n        ↓\nGenerate the final response\n```\nIn this article, we'll explore how to build this architecture using **Spring AI**.\n\nAn AI agent is an application where an LLM can decide what actions need to be performed and use available tools to accomplish a goal.\n\nA traditional LLM application looks like:\n\n\n```\nUser\n ↓\nPrompt\n ↓\nLLM\n ↓\nResponse\n```\nAn agent-based application looks more like:\n\n\n```\nUser\n ↓\nLLM\n ↓\nDecision\n ↓\nTool\n ↓\nResult\n ↓\nLLM\n ↓\nDecision\n ↓\nAnother Tool\n ↓\nResult\n ↓\nFinal Answer\n```\nThe important difference is:\n\n**The LLM is no longer limited to generating text.**\n\n\nIt can interact with the application through controlled capabilities.\n\nTool calling allows an LLM to request the execution of a function exposed by your application.\n\nFor example, imagine our application provides:\n\n\n```\ngetOrderStatus()\ncancelOrder()\ngetCustomer()\ncreateSupportTicket()\n```\nThe user asks:\n\n\n```\nWhere is my order?\n```\nThe model might determine that it needs:\n\n\n```\ngetOrderStatus()\n```\nThe application executes the function and returns:\n\n\n```\nOrder #10291\nStatus: Shipped\nExpected delivery: September 10\n```\nThe LLM can then generate:\n\n\n```\nYour order has been shipped and is expected to arrive on September 10.\n```\nThe LLM didn't directly access the database.\n\nInstead:\n\n\n```\nLLM\n ↓\nTool Request\n ↓\nApplication\n ↓\nDatabase\n ↓\nTool Result\n ↓\nLLM\n ↓\nAnswer\n```\nThis distinction is extremely important for enterprise applications.\n\nWithout tools:\n\n\n```\nLLM\n ↓\nText\n```\nWith tools:\n\n\n```\nLLM\n ↓\nTools\n ├── Database\n ├── REST APIs\n ├── Search\n ├── Payment systems\n ├── CRM\n ├── Internal services\n └── Business workflows\n```\nThis turns the LLM from a text-generation component into an interface for interacting with your application.\n\nFor example, an AI sales assistant could have:\n\n\n```\ngetCustomer()\ngetCustomerOrders()\ncreateLead()\nupdateLead()\nsendEmail()\nscheduleMeeting()\n```\nA support agent could have:\n\n\n```\nsearchKnowledgeBase()\ngetCustomerAccount()\ngetOrder()\ncreateTicket()\nupdateTicket()\n```\nAn internal developer assistant could have:\n\n\n```\nsearchDocumentation()\nsearchGitRepository()\ngetBuildStatus()\ncreateIssue()\n```\nThe possibilities are much broader than simple question answering.\n\nAt this point, it is useful to distinguish RAG from tool calling.\n\nRAG is primarily about **retrieving information**.\n\nTool calling is about **performing actions or retrieving live data through application capabilities**.\n\nFor example:\n\n\n```\nRAG\n ↓\nRetrieve company documentation\n ↓\nAnswer question\n```\nTool calling:\n\n\n```\nLLM\n ↓\nCall order API\n ↓\nGet live order status\n ↓\nAnswer\n```\nThey can also be combined.\n\nFor example:\n\n\n```\nUser\n ↓\nAI Agent\n ├── RAG → Search company policies\n │\n ├── Tool → Get customer account\n │\n └── Tool → Check order status\n          ↓\n       LLM\n          ↓\n       Answer\n```\nThis combination is extremely powerful.\n\nSpring AI provides abstractions that make it easier to expose application capabilities to language models.\n\nA simplified architecture looks like:\n\n\n```\nSpring Boot\n     │\n     ├── ChatModel\n     │\n     ├── Tools\n     │\n     ├── Advisors\n     │\n     ├── Chat Memory\n     │\n     └── Vector Store\n```\nThe application controls which tools are available.\n\nThe model decides whether a tool is needed.\n\nThis separation is important.\n\nThe model should not have unrestricted access to your application.\n\nInstead, the application exposes specific capabilities.\n\nImagine we have an order service.\n\n\n```\n@Service\npublic class OrderService {\n    public OrderStatus getOrderStatus(String orderId) {\n        // Fetch order from database\n        return orderRepository.findStatus(orderId);\n    }\n}\n```\nWe can expose a controlled method as an AI tool.\n\nConceptually:\n\n\n```\n@Tool(\n    description = \"Get the current status of an order\"\n)\npublic OrderStatus getOrderStatus(String orderId) {\n    return orderService.getOrderStatus(orderId);\n}\n```\nThe description is important.\n\nThe model uses the tool description to understand:\n\n\n```\nWhat does this tool do?\nWhen should I use it?\nWhat parameters does it require?\n```\nFor example:\n\n\n```\nTool:\ngetOrderStatus\nDescription:\nReturns the current shipping and delivery status\nfor a customer order.\nInput:\norderId\n```\nThe model can then determine whether this tool is appropriate.\n\nA tool can be thought of as:\n\n\n```\nTool Name\n     +\nDescription\n     +\nInput Schema\n     +\nExecution Logic\n```\nFor example:\n\n\n```\ngetOrderStatus(\n    orderId: String\n)\n```\nThe model might produce a tool request conceptually like:\n\n\n```\n{\n  \"name\": \"getOrderStatus\",\n  \"arguments\": {\n    \"orderId\": \"ORD-10291\"\n  }\n}\n```\nThe application receives the request and executes the corresponding Java method.\n\nA typical interaction looks like this:\n\n\n```\nUser\n │\n │ \"What's the status of ORD-10291?\"\n ↓\nLLM\n │\n │ Tool Request\n ↓\ngetOrderStatus(\"ORD-10291\")\n │\n ↓\nOrder Service\n │\n ↓\nDatabase\n │\n ↓\nTool Result\n │\n ↓\nLLM\n │\n ↓\nFinal Answer\n```\nNotice something important.\n\nThe LLM doesn't execute Java code itself.\n\nThe application remains responsible for execution.\n\nThe model only requests the action.\n\nSpring AI's `ChatClient` provides a convenient API for interacting with chat models.\n\nConceptually:\n\n\n```\nChatClient chatClient;\nString response = chatClient.prompt()\n        .user(\"What's the status of order ORD-10291?\")\n        .tools(orderTools)\n        .call()\n        .content();\n```\nThe exact APIs may vary depending on the Spring AI version you're using, but the architecture remains the same:\n\n\n```\nChatClient\n   ↓\nChatModel\n   ↓\nTool Selection\n   ↓\nTool Execution\n   ↓\nTool Result\n   ↓\nFinal Response\n```\nA real agent usually has more than one tool.\n\nFor example:\n\n\n```\nCustomerAgentTools\n├── getCustomer()\n├── getCustomerOrders()\n├── getOrderStatus()\n├── createSupportTicket()\n└── updateCustomer()\n```\nNow consider this question:\n\n\n```\nMy order is late. Please check the status\nand create a support ticket if necessary.\n```\nThe model might determine:\n\n\n```\n1. getOrderStatus()\n2. Analyze result\n3. createSupportTicket()\n4. Return final response\n```\nThe application executes each requested operation.\n\nThis is where the concept of an agent starts becoming much more interesting.\n\nA simple agent loop can be represented as:\n\n\n```\n              ┌───────────────┐\n              │     User      │\n              └───────┬───────┘\n                      ↓\n                ┌───────────┐\n                │    LLM    │\n                └─────┬─────┘\n                      ↓\n                Need a Tool?\n                /          \\\n              No            Yes\n              ↓              ↓\n          Final Answer    Tool Call\n                             ↓\n                        Tool Execution\n                             ↓\n                         Tool Result\n                             ↓\n                            LLM\n                             ↓\n                      Need Another Tool?\n```\nThe model can repeatedly interact with tools until it has enough information to produce the final response.\n\nThis distinction is important.\n\nPeople often hear:\n\nAI Agent\n\n\nand immediately think:\n\n\n```\nGive AI access to everything\n        ↓\nLet AI do whatever it wants\n```\nThat is not how production systems should be designed.\n\nA production agent should operate inside clear boundaries.\n\nFor example:\n\n\n```\nAllowed Tools\n     ↓\nAuthorization\n     ↓\nValidation\n     ↓\nExecution\n     ↓\nAudit\n```\nThe application remains in control.\n\nThe model should not be trusted with unrestricted capabilities.\n\nImagine we expose:\n\n\n```\n@Tool\npublic void deleteCustomer(String customerId) {\n    ...\n}\n```\nThis is potentially dangerous.\n\nAn LLM should not automatically receive unrestricted permission to perform destructive operations.\n\nInstead, sensitive tools should have additional controls.\n\nFor example:\n\n\n```\nUser\n ↓\nAuthentication\n ↓\nAuthorization\n ↓\nAgent\n ↓\nTool Request\n ↓\nPermission Check\n ↓\nConfirmation\n ↓\nExecution\n```\nFor destructive operations, you may require explicit user confirmation.\n\nExample:\n\n\n```\nAI:\nI found customer account C-19291.\nDeleting this account is irreversible.\nDo you want me to continue?\nUser:\nYes.\nAI:\nExecuting deletion...\n```\nThe AI should assist with the decision process, not bypass your security model.\n\nA useful production architecture is to classify tools.\n\n\n```\nREAD\ngetCustomer()\ngetOrder()\nsearchDocuments()\ngetInvoice()\n```\nThen:\n\n\n```\nWRITE\ncreateTicket()\nupdateCustomer()\ncreateLead()\n```\nAnd:\n\n\n```\nDESTRUCTIVE\ndeleteCustomer()\ncancelSubscription()\nrefundPayment()\n```\nDifferent permission levels can then be applied.\n\nFor example:\n\n\n```\nREAD\n→ Automatically allowed\nWRITE\n→ Role-based authorization\nDESTRUCTIVE\n→ Authorization + confirmation\n```\nThis makes agent behavior much safer.\n\nTool calling solves one problem.\n\nBut another problem appears quickly:\n\nWhat does the agent remember?\n\n\nConsider this conversation:\n\n\n```\nUser:\nMy order is late.\nAI:\nWhat's your order number?\nUser:\nORD-10291.\nAI:\nLet me check it.\n```\nNow the next message is:\n\n\n```\nCan you create a support ticket for it?\n```\nThe AI needs to understand that:\n\n\n```\n\"it\"\n```\nrefers to:\n\n\n```\nORD-10291\n```\nThis requires conversational context.\n\nThat's where **chat memory** becomes important.\n\nA simple conversation can be represented as:\n\n\n```\nUser:\nMy order is late.\nAssistant:\nWhat's your order number?\nUser:\nORD-10291.\nAssistant:\nLet me check that order.\n```\nThe application maintains the conversation history.\n\nConceptually:\n\n\n```\nConversation ID\n       ↓\nChat Memory\n       ↓\nPrevious Messages\n       ↓\nCurrent Prompt\n       ↓\nLLM\n```\nSpring AI provides abstractions for managing chat memory.\n\nIt is useful to distinguish two concepts.\n\nConversation context.\n\n\n```\nUser:\nMy order is late.\nUser:\nIt's order 10291.\nUser:\nCan you check it?\n```\nThe system remembers the current conversation.\n\nPersistent information about the user.\n\nFor example:\n\n\n```\nCustomer:\nAyush\nPreferences:\nPreferred language = English\nPreferred notification = Email\n```\nLong-term memory usually requires persistence in a database or another storage system.\n\nA production architecture might look like:\n\n\n```\nConversation\n     ↓\nChat Memory Store\n     ↓\nPostgreSQL / Redis\n```\nThe exact storage mechanism depends on the application.\n\nNow we can combine:\n\n\n```\nUser\n ↓\nAgent\n ↓\nMemory\n ↓\nLLM\n ↓\nTools\n ↓\nTool Results\n ↓\nMemory\n ↓\nLLM\n ↓\nAnswer\n```\nThis enables more natural multi-turn interactions.\n\nAnother important Spring AI concept is the **Advisor**.\n\nAdvisors can intercept and influence the interaction between the application and the model.\n\nThey can be used for concerns such as:\n\n\n```\nConversation memory\nRAG\nLogging\nSecurity\nPrompt modification\nContext injection\nObservability\n```\nConceptually:\n\n\n```\nUser\n ↓\nChatClient\n ↓\nAdvisor\n ↓\nChatModel\n ↓\nAdvisor\n ↓\nResponse\n```\nThis allows cross-cutting AI behavior to be separated from business logic.\n\nNow things become much more powerful.\n\nImagine an enterprise support agent.\n\nIt has:\n\n\n```\nRAG\n ↓\nCompany documentation\n```\nTools:\n\n\n```\ngetCustomer()\ngetOrder()\ncreateTicket()\n```\nMemory:\n\n\n```\nConversation history\n```\nThe architecture becomes:\n\n\n```\n                    User\n                     ↓\n                 AI Agent\n                     ↓\n             ┌───────┼────────┐\n             ↓       ↓        ↓\n            RAG    Tools    Memory\n             ↓       ↓        ↓\n        Knowledge   APIs   Conversation\n             │       │        │\n             └───────┼────────┘\n                     ↓\n                    LLM\n                     ↓\n                  Response\n```\nThis is much closer to a production AI application.\n\nConsider the request:\n\n\n```\nMy payment failed for order ORD-19291.\nCan you check what happened and tell me\nwhat I should do?\n```\nThe agent could perform:\n\n\n```\n1. getOrder(\"ORD-19291\")\n2. getPaymentStatus(\"ORD-19291\")\n3. searchKnowledgeBase(\"payment failure\")\n4. Generate explanation\n```\nThe final answer could be:\n\n\n```\nYour payment attempt failed because the transaction\nwas declined by the payment provider.\nAccording to the payment policy, you can retry the\npayment using another payment method.\nWould you like me to create a support ticket?\n```\nThe model combined:\n\n\n```\nLive application data\n+\nKnowledge base\n+\nConversation context\n```\nThis is significantly more useful than a standalone chatbot.\n\nAgents can also perform multi-step workflows.\n\nFor example:\n\n\n```\nUser:\nFind my overdue invoices and send reminders.\n```\nThe agent could reason through:\n\n\n```\ngetCustomer()\n      ↓\ngetInvoices()\n      ↓\nFilter overdue invoices\n      ↓\nsendReminder()\n      ↓\nReturn summary\n```\nThe workflow becomes:\n\n\n```\nGoal\n ↓\nPlan\n ↓\nTool\n ↓\nObserve\n ↓\nNext Decision\n ↓\nTool\n ↓\nObserve\n ↓\nFinal Result\n```\nThis pattern is often called an **agent loop**.\n\nA simplified conceptual implementation looks like:\n\n\n```\nwhile (!completed) {\n    AgentDecision decision =\n            llm.decide(context);\n    if (decision.requiresTool()) {\n        ToolResult result =\n                toolExecutor.execute(\n                        decision.toolCall()\n                );\n        context.add(result);\n    } else {\n        return decision.finalAnswer();\n    }\n}\n```\nIn real applications, frameworks handle much of this interaction.\n\nBut understanding the underlying loop is important.\n\nAn important engineering lesson:\n\n**Not every AI feature needs an agent.**\n\n\nIf your workflow is deterministic:\n\n\n```\nValidate request\n ↓\nCall API\n ↓\nSave result\n ↓\nReturn response\n```\nyou probably don't need an autonomous agent.\n\nA normal service workflow may be better.\n\nAgents become more useful when:\n\n\n```\nThe next step depends on the current result.\n```\nFor example:\n\n\n```\nCheck order\n ↓\nIf delayed\n ↓\nCheck refund policy\n ↓\nIf eligible\n ↓\nAsk for confirmation\n ↓\nCreate refund request\n```\nThe dynamic decision-making is where agents become valuable.\n\n```\nA → B → C → D\n```\nEverything is predetermined.\n\n```\nA\n ↓\nLLM decides\n ├── B\n ├── C\n └── D\n      ↓\n   Observe result\n      ↓\n   Decide again\n```\nAgents provide flexibility.\n\nTraditional workflows provide predictability.\n\nProduction systems often use both.\n\nA practical Spring Boot architecture might look like:\n\n\n```\n                    ┌───────────────┐\n                    │   Frontend    │\n                    └───────┬───────┘\n                            ↓\n                    ┌───────────────┐\n                    │ Spring Boot   │\n                    │     API       │\n                    └───────┬───────┘\n                            ↓\n                     ┌────────────┐\n                     │ ChatClient │\n                     └─────┬──────┘\n                           ↓\n                    ┌──────────────┐\n                    │    Agent     │\n                    └──────┬───────┘\n                           ↓\n              ┌────────────┼────────────┐\n              ↓            ↓            ↓\n           Memory         RAG         Tools\n              ↓            ↓            ↓\n          PostgreSQL    pgvector      APIs\n                                         ↓\n                                    Microservices\n```\nThis architecture fits naturally into existing Spring Boot applications.\n\nAgent systems can become difficult to debug.\n\nImagine an agent performs:\n\n\n```\nTool 1\nTool 2\nTool 3\nTool 4\n```\nand the final response is incorrect.\n\nYou need to know:\n\n\n```\nWhat did the model decide?\nWhich tools were selected?\nWhat arguments were sent?\nHow long did each tool take?\nWhat did each tool return?\nHow many model calls happened?\nHow many tokens were consumed?\n```\nTherefore, observability is critical.\n\nTrack:\n\n\n```\nLLM latency\nTool latency\nRetrieval latency\nToken usage\nTool calls\nTool failures\nModel responses\nAgent iterations\nErrors\n```\nAn agent can potentially continue calling tools indefinitely.\n\nFor example:\n\n\n```\nLLM\n ↓\nTool\n ↓\nLLM\n ↓\nTool\n ↓\nLLM\n ↓\nTool\n ↓\n...\n```\nProduction systems should enforce limits.\n\nFor example:\n\n\n```\nMaximum iterations = 10\nMaximum tool calls = 20\nMaximum execution time = 30 seconds\n```\nYou should also define clear failure behavior.\n\n\n```\nAgent limit reached\n        ↓\nStop execution\n        ↓\nReturn safe response\n        ↓\nLog failure\n```\nNever blindly trust model-generated tool arguments.\n\nSuppose the model requests:\n\n\n```\n{\n  \"orderId\": \"ORD-999999999\"\n}\n```\nYour application should still validate:\n\n\n```\nDoes the order exist?\nDoes the user own the order?\nIs the user authorized?\nIs the order accessible to this tenant?\n```\nThe architecture should be:\n\n\n```\nLLM\n ↓\nTool Request\n ↓\nSchema Validation\n ↓\nAuthorization\n ↓\nBusiness Validation\n ↓\nTool Execution\n```\nThe LLM is not your security boundary.\n\nYour application is.\n\nThis becomes especially important in SaaS applications.\n\nImagine:\n\n\n```\nTenant A\n ├── Customers\n ├── Orders\n └── Documents\nTenant B\n ├── Customers\n ├── Orders\n └── Documents\n```\nAn AI agent must never retrieve Tenant B's information while processing a Tenant A request.\n\nEvery tool and retrieval operation should carry tenant context.\n\nFor example:\n\n\n```\ntenant_id\nuser_id\nroles\npermissions\n```\nThen:\n\n\n```\nUser\n ↓\nAuthentication\n ↓\nTenant Context\n ↓\nAgent\n ↓\nTool\n ↓\nAuthorization\n ↓\nTenant-scoped Data\n```\nThe same principle applies to RAG.\n\n\n```\nVector Search\n +\ntenant_id filter\n```\nshould ensure that retrieved documents belong to the correct tenant.\n\nProduction agents should have explicit guardrails.\n\nExamples:\n\n\n```\nInput validation\nOutput validation\nTool authorization\nRate limiting\nToken limits\nIteration limits\nPII protection\nAudit logging\nHuman approval\n```\nFor high-risk actions:\n\n\n```\nAgent\n ↓\nTool Request\n ↓\nRisk Evaluation\n ↓\nHuman Approval\n ↓\nExecution\n```\nThis creates a human-in-the-loop architecture.\n\nNot every decision should be fully automated.\n\nFor example:\n\n\n```\nRefund amount < $50\n    ↓\nAutomatic\nRefund amount > $50\n    ↓\nHuman approval\n```\nOr:\n\n\n```\nCreate support ticket\n    ↓\nAutomatic\nDelete account\n    ↓\nConfirmation required\n```\nThis gives us a practical balance:\n\n\n```\nAI Automation\n+\nBusiness Rules\n+\nHuman Oversight\n```\nAt this point, we can combine everything we've discussed.\n\n\n```\n                         User\n                          ↓\n                     Spring Boot\n                          ↓\n                      ChatClient\n                          ↓\n                       AI Agent\n                          ↓\n              ┌───────────┼───────────┐\n              ↓           ↓           ↓\n            Memory       RAG         Tools\n              ↓           ↓           ↓\n          PostgreSQL   pgvector    REST APIs\n                                      ↓\n                               Business Services\n                                      ↓\n                                   Database\n```\nThis is a strong foundation for enterprise AI applications.\n\nImagine a sales assistant.\n\nThe user asks:\n\n\n```\nShow me the latest opportunities for Acme\nand tell me which ones are likely to close this month.\n```\nThe agent could:\n\n\n```\n1. getCustomer(\"Acme\")\n2. getOpportunities(\"Acme\")\n3. retrieve sales documentation\n4. analyze opportunity information\n5. generate summary\n```\nNow the user says:\n\n\n```\nCreate a follow-up task for the highest priority opportunity.\n```\nThe agent can:\n\n\n```\n1. Identify opportunity\n2. createFollowUpTask()\n3. Return task details\n```\nThis is where AI starts becoming an application interface rather than simply a chatbot.\n\nThink about the responsibilities this way:\n\n\n```\nLLM\n=\nReasoning + Language\nRAG\n=\nKnowledge Retrieval\nTools\n=\nActions + Live Data\nMemory\n=\nConversation Context\nSpring Boot\n=\nApplication + Security + Business Logic\n```\nTogether:\n\n\n```\nLLM\n +\nRAG\n +\nTools\n +\nMemory\n +\nBusiness Logic\n =\nAI Application\n```\nSpring AI provides abstractions that allow Java developers to work with AI capabilities using familiar Spring patterns.\n\nImportant building blocks include:\n\n\n```\nChatClient\nChatModel\nEmbeddingModel\nVectorStore\nDocument\nAdvisors\nChat Memory\nTools\n```\nThis means an enterprise Java team can integrate AI into an existing Spring Boot architecture instead of creating an entirely separate AI stack.\n\nFor example:\n\n\n```\nExisting Spring Boot Application\n              ↓\n        Spring AI Layer\n              ↓\n      Model + RAG + Tools\n              ↓\n     Existing Microservices\n```\nThis makes AI integration much more practical for Java teams.\n\nA more complete production system might eventually look like:\n\n\n```\n                         ┌───────────────┐\n                         │     User      │\n                         └───────┬───────┘\n                                 ↓\n                         API Gateway\n                                 ↓\n                         Authentication\n                                 ↓\n                         Spring Boot API\n                                 ↓\n                            AI Agent\n                                 ↓\n          ┌──────────────────────┼──────────────────────┐\n          ↓                      ↓                      ↓\n       Memory                  RAG                    Tools\n          ↓                      ↓                      ↓\n     PostgreSQL              pgvector              Microservices\n                                                        ↓\n                                                Business Database\n                                 ↓\n                           LLM Provider\n                                 ↓\n                              Response\n```\nAnd around the entire system:\n\n\n```\nSecurity\nObservability\nRate Limiting\nAudit Logging\nGuardrails\nEvaluation\n```\nThese are not optional concerns in serious enterprise deployments.\n\nIt is useful to understand the difference.\n\n```\nUser\n ↓\nLLM\n ↓\nAnswer\n```\n```\nUser\n ↓\nRetrieve Knowledge\n ↓\nLLM\n ↓\nAnswer\n```\n```\nUser\n ↓\nLLM\n ↓\nTool\n ↓\nResult\n ↓\nAnswer\n```\n```\nUser\n ↓\nAgent\n ↓\nReason\n ↓\nTool\n ↓\nObserve\n ↓\nReason\n ↓\nTool\n ↓\nObserve\n ↓\nFinal Answer\n```\nThe complexity increases at every stage.\n\nWe can now think about the evolution of an AI application:\n\n\n```\nLevel 1\nLLM\n ↓\nText Generation\nLevel 2\nLLM + RAG\n ↓\nKnowledge Retrieval\nLevel 3\nLLM + Tools\n ↓\nActions\nLevel 4\nLLM + Tools + Memory\n ↓\nContextual Assistant\nLevel 5\nLLM + RAG + Tools + Memory\n ↓\nAgent\nLevel 6\nMultiple Agents + Workflows\n ↓\nAgentic System\n```\nThis progression is useful when deciding how much complexity your application actually needs.\n\nThe evolution from traditional AI applications to agentic applications can be summarized as:\n\n\n```\nLLM\n ↓\nGenerate Text\nRAG\n ↓\nRetrieve Knowledge\nTool Calling\n ↓\nTake Actions\nMemory\n ↓\nRemember Context\nAgents\n ↓\nMake Decisions\nWorkflows\n ↓\nCoordinate Multiple Steps\n```\nSpring AI provides Java developers with abstractions for building many of these capabilities inside the Spring ecosystem.\n\nThe most important engineering principle is:\n\n**Let the model decide, but let your application control.**\n\n\nThe LLM can decide which tool may be useful.\n\nYour application should decide whether that tool is actually allowed to execute.\n\nThat separation gives us a much safer architecture for enterprise AI.\n\nWe've now covered three major capabilities:\n\n\n```\nLLM\n ↓\nGenerate\nRAG\n ↓\nRetrieve\nTools\n ↓\nAct\n```\nBut there is another challenge.\n\nWhat happens when a system has:\n\n\n```\nMultiple agents\n        ↓\nMultiple tools\n        ↓\nMultiple services\n        ↓\nMultiple AI models\n```\nHow do these agents communicate?\n\nHow do we standardize tool discovery?\n\nHow can an AI agent securely interact with external tools and services?\n\nThis leads us toward another important concept in modern AI engineering:\n\n**Model Context Protocol — MCP.**\n\nIn the next article, we'll explore:\n\n**Building MCP Clients and Tool-Based AI Applications with Spring AI.**\n\n- Tool calling allows LLMs to interact with application capabilities.\n- The LLM requests tools; your application executes them.\n- Tools can expose APIs, databases, business operations, and services.\n- Chat memory provides conversational context.\n- RAG provides external knowledge.\n- Agents can combine RAG, memory, and tools.\n- Not every workflow requires an agent.\n- Deterministic workflows are often better for predictable business processes.\n- Tool permissions and authorization are critical.\n- Destructive operations should require stronger controls.\n- Multi-tenant applications must enforce tenant isolation at the tool and retrieval layers.\n- Agent loops should have iteration, timeout, and tool-call limits.\n- Observability is essential for debugging agent behavior.\n- Human-in-the-loop approval is useful for high-risk actions.\n- Spring AI provides abstractions that make these patterns accessible to Java and Spring Boot developers.\n\nThe future of enterprise AI isn't just:\n\n\n```\nLLM → Answer\n```\nIt's increasingly:\n\n\n```\nLLM\n ↓\nReason\n ↓\nRetrieve\n ↓\nCall Tools\n ↓\nObserve\n ↓\nAct\n ↓\nRemember\n ↓\nComplete the Goal\n```\nAnd that's where **AI agents with Spring AI** become truly interesting.","body_html":"<p>Large Language Models are excellent at generating text.</p>\n<p>But generation alone isn&#39;t enough to build truly useful AI applications.</p>\n<p>Imagine asking an AI assistant:</p>\n<pre><code>What&#39;s the status of my order?</code></pre>\n<p>A normal LLM can explain how order tracking works.</p>\n<p>But it cannot magically access your order database.</p>\n<p>Or suppose you ask:</p>\n<pre><code>Cancel my order #ORD-10291.</code></pre>\n<p>The model can tell you how to cancel an order.</p>\n<p>But it cannot actually cancel anything unless your application gives it the ability to perform that action.</p>\n<p>This is where <strong>tool calling</strong> and <strong>AI agents</strong> come in.</p>\n<p>Instead of simply generating an answer, an AI application can:</p>\n<pre><code>Understand the request\n        ↓\nDecide what action is required\n        ↓\nSelect a tool\n        ↓\nExecute the tool\n        ↓\nObserve the result\n        ↓\nContinue reasoning\n        ↓\nGenerate the final response</code></pre>\n<p>In this article, we&#39;ll explore how to build this architecture using <strong>Spring AI</strong>.</p>\n<p>An AI agent is an application where an LLM can decide what actions need to be performed and use available tools to accomplish a goal.</p>\n<p>A traditional LLM application looks like:</p>\n<pre><code>User\n ↓\nPrompt\n ↓\nLLM\n ↓\nResponse</code></pre>\n<p>An agent-based application looks more like:</p>\n<pre><code>User\n ↓\nLLM\n ↓\nDecision\n ↓\nTool\n ↓\nResult\n ↓\nLLM\n ↓\nDecision\n ↓\nAnother Tool\n ↓\nResult\n ↓\nFinal Answer</code></pre>\n<p>The important difference is:</p>\n<p><strong>The LLM is no longer limited to generating text.</strong></p>\n<p>It can interact with the application through controlled capabilities.</p>\n<p>Tool calling allows an LLM to request the execution of a function exposed by your application.</p>\n<p>For example, imagine our application provides:</p>\n<pre><code>getOrderStatus()\ncancelOrder()\ngetCustomer()\ncreateSupportTicket()</code></pre>\n<p>The user asks:</p>\n<pre><code>Where is my order?</code></pre>\n<p>The model might determine that it needs:</p>\n<pre><code>getOrderStatus()</code></pre>\n<p>The application executes the function and returns:</p>\n<pre><code>Order #10291\nStatus: Shipped\nExpected delivery: September 10</code></pre>\n<p>The LLM can then generate:</p>\n<pre><code>Your order has been shipped and is expected to arrive on September 10.</code></pre>\n<p>The LLM didn&#39;t directly access the database.</p>\n<p>Instead:</p>\n<pre><code>LLM\n ↓\nTool Request\n ↓\nApplication\n ↓\nDatabase\n ↓\nTool Result\n ↓\nLLM\n ↓\nAnswer</code></pre>\n<p>This distinction is extremely important for enterprise applications.</p>\n<p>Without tools:</p>\n<pre><code>LLM\n ↓\nText</code></pre>\n<p>With tools:</p>\n<pre><code>LLM\n ↓\nTools\n ├── Database\n ├── REST APIs\n ├── Search\n ├── Payment systems\n ├── CRM\n ├── Internal services\n └── Business workflows</code></pre>\n<p>This turns the LLM from a text-generation component into an interface for interacting with your application.</p>\n<p>For example, an AI sales assistant could have:</p>\n<pre><code>getCustomer()\ngetCustomerOrders()\ncreateLead()\nupdateLead()\nsendEmail()\nscheduleMeeting()</code></pre>\n<p>A support agent could have:</p>\n<pre><code>searchKnowledgeBase()\ngetCustomerAccount()\ngetOrder()\ncreateTicket()\nupdateTicket()</code></pre>\n<p>An internal developer assistant could have:</p>\n<pre><code>searchDocumentation()\nsearchGitRepository()\ngetBuildStatus()\ncreateIssue()</code></pre>\n<p>The possibilities are much broader than simple question answering.</p>\n<p>At this point, it is useful to distinguish RAG from tool calling.</p>\n<p>RAG is primarily about <strong>retrieving information</strong>.</p>\n<p>Tool calling is about <strong>performing actions or retrieving live data through application capabilities</strong>.</p>\n<p>For example:</p>\n<pre><code>RAG\n ↓\nRetrieve company documentation\n ↓\nAnswer question</code></pre>\n<p>Tool calling:</p>\n<pre><code>LLM\n ↓\nCall order API\n ↓\nGet live order status\n ↓\nAnswer</code></pre>\n<p>They can also be combined.</p>\n<p>For example:</p>\n<pre><code>User\n ↓\nAI Agent\n ├── RAG → Search company policies\n │\n ├── Tool → Get customer account\n │\n └── Tool → Check order status\n          ↓\n       LLM\n          ↓\n       Answer</code></pre>\n<p>This combination is extremely powerful.</p>\n<p>Spring AI provides abstractions that make it easier to expose application capabilities to language models.</p>\n<p>A simplified architecture looks like:</p>\n<pre><code>Spring Boot\n     │\n     ├── ChatModel\n     │\n     ├── Tools\n     │\n     ├── Advisors\n     │\n     ├── Chat Memory\n     │\n     └── Vector Store</code></pre>\n<p>The application controls which tools are available.</p>\n<p>The model decides whether a tool is needed.</p>\n<p>This separation is important.</p>\n<p>The model should not have unrestricted access to your application.</p>\n<p>Instead, the application exposes specific capabilities.</p>\n<p>Imagine we have an order service.</p>\n<pre><code>@Service\npublic class OrderService {\n    public OrderStatus getOrderStatus(String orderId) {\n        // Fetch order from database\n        return orderRepository.findStatus(orderId);\n    }\n}</code></pre>\n<p>We can expose a controlled method as an AI tool.</p>\n<p>Conceptually:</p>\n<pre><code>@Tool(\n    description = &quot;Get the current status of an order&quot;\n)\npublic OrderStatus getOrderStatus(String orderId) {\n    return orderService.getOrderStatus(orderId);\n}</code></pre>\n<p>The description is important.</p>\n<p>The model uses the tool description to understand:</p>\n<pre><code>What does this tool do?\nWhen should I use it?\nWhat parameters does it require?</code></pre>\n<p>For example:</p>\n<pre><code>Tool:\ngetOrderStatus\nDescription:\nReturns the current shipping and delivery status\nfor a customer order.\nInput:\norderId</code></pre>\n<p>The model can then determine whether this tool is appropriate.</p>\n<p>A tool can be thought of as:</p>\n<pre><code>Tool Name\n     +\nDescription\n     +\nInput Schema\n     +\nExecution Logic</code></pre>\n<p>For example:</p>\n<pre><code>getOrderStatus(\n    orderId: String\n)</code></pre>\n<p>The model might produce a tool request conceptually like:</p>\n<pre><code>{\n  &quot;name&quot;: &quot;getOrderStatus&quot;,\n  &quot;arguments&quot;: {\n    &quot;orderId&quot;: &quot;ORD-10291&quot;\n  }\n}</code></pre>\n<p>The application receives the request and executes the corresponding Java method.</p>\n<p>A typical interaction looks like this:</p>\n<pre><code>User\n │\n │ &quot;What&#39;s the status of ORD-10291?&quot;\n ↓\nLLM\n │\n │ Tool Request\n ↓\ngetOrderStatus(&quot;ORD-10291&quot;)\n │\n ↓\nOrder Service\n │\n ↓\nDatabase\n │\n ↓\nTool Result\n │\n ↓\nLLM\n │\n ↓\nFinal Answer</code></pre>\n<p>Notice something important.</p>\n<p>The LLM doesn&#39;t execute Java code itself.</p>\n<p>The application remains responsible for execution.</p>\n<p>The model only requests the action.</p>\n<p>Spring AI&#39;s <code>ChatClient</code> provides a convenient API for interacting with chat models.</p>\n<p>Conceptually:</p>\n<pre><code>ChatClient chatClient;\nString response = chatClient.prompt()\n        .user(&quot;What&#39;s the status of order ORD-10291?&quot;)\n        .tools(orderTools)\n        .call()\n        .content();</code></pre>\n<p>The exact APIs may vary depending on the Spring AI version you&#39;re using, but the architecture remains the same:</p>\n<pre><code>ChatClient\n   ↓\nChatModel\n   ↓\nTool Selection\n   ↓\nTool Execution\n   ↓\nTool Result\n   ↓\nFinal Response</code></pre>\n<p>A real agent usually has more than one tool.</p>\n<p>For example:</p>\n<pre><code>CustomerAgentTools\n├── getCustomer()\n├── getCustomerOrders()\n├── getOrderStatus()\n├── createSupportTicket()\n└── updateCustomer()</code></pre>\n<p>Now consider this question:</p>\n<pre><code>My order is late. Please check the status\nand create a support ticket if necessary.</code></pre>\n<p>The model might determine:</p>\n<pre><code>1. getOrderStatus()\n2. Analyze result\n3. createSupportTicket()\n4. Return final response</code></pre>\n<p>The application executes each requested operation.</p>\n<p>This is where the concept of an agent starts becoming much more interesting.</p>\n<p>A simple agent loop can be represented as:</p>\n<pre><code>              ┌───────────────┐\n              │     User      │\n              └───────┬───────┘\n                      ↓\n                ┌───────────┐\n                │    LLM    │\n                └─────┬─────┘\n                      ↓\n                Need a Tool?\n                /          \\\n              No            Yes\n              ↓              ↓\n          Final Answer    Tool Call\n                             ↓\n                        Tool Execution\n                             ↓\n                         Tool Result\n                             ↓\n                            LLM\n                             ↓\n                      Need Another Tool?</code></pre>\n<p>The model can repeatedly interact with tools until it has enough information to produce the final response.</p>\n<p>This distinction is important.</p>\n<p>People often hear:</p>\n<p>AI Agent</p>\n<p>and immediately think:</p>\n<pre><code>Give AI access to everything\n        ↓\nLet AI do whatever it wants</code></pre>\n<p>That is not how production systems should be designed.</p>\n<p>A production agent should operate inside clear boundaries.</p>\n<p>For example:</p>\n<pre><code>Allowed Tools\n     ↓\nAuthorization\n     ↓\nValidation\n     ↓\nExecution\n     ↓\nAudit</code></pre>\n<p>The application remains in control.</p>\n<p>The model should not be trusted with unrestricted capabilities.</p>\n<p>Imagine we expose:</p>\n<pre><code>@Tool\npublic void deleteCustomer(String customerId) {\n    ...\n}</code></pre>\n<p>This is potentially dangerous.</p>\n<p>An LLM should not automatically receive unrestricted permission to perform destructive operations.</p>\n<p>Instead, sensitive tools should have additional controls.</p>\n<p>For example:</p>\n<pre><code>User\n ↓\nAuthentication\n ↓\nAuthorization\n ↓\nAgent\n ↓\nTool Request\n ↓\nPermission Check\n ↓\nConfirmation\n ↓\nExecution</code></pre>\n<p>For destructive operations, you may require explicit user confirmation.</p>\n<p>Example:</p>\n<pre><code>AI:\nI found customer account C-19291.\nDeleting this account is irreversible.\nDo you want me to continue?\nUser:\nYes.\nAI:\nExecuting deletion...</code></pre>\n<p>The AI should assist with the decision process, not bypass your security model.</p>\n<p>A useful production architecture is to classify tools.</p>\n<pre><code>READ\ngetCustomer()\ngetOrder()\nsearchDocuments()\ngetInvoice()</code></pre>\n<p>Then:</p>\n<pre><code>WRITE\ncreateTicket()\nupdateCustomer()\ncreateLead()</code></pre>\n<p>And:</p>\n<pre><code>DESTRUCTIVE\ndeleteCustomer()\ncancelSubscription()\nrefundPayment()</code></pre>\n<p>Different permission levels can then be applied.</p>\n<p>For example:</p>\n<pre><code>READ\n→ Automatically allowed\nWRITE\n→ Role-based authorization\nDESTRUCTIVE\n→ Authorization + confirmation</code></pre>\n<p>This makes agent behavior much safer.</p>\n<p>Tool calling solves one problem.</p>\n<p>But another problem appears quickly:</p>\n<p>What does the agent remember?</p>\n<p>Consider this conversation:</p>\n<pre><code>User:\nMy order is late.\nAI:\nWhat&#39;s your order number?\nUser:\nORD-10291.\nAI:\nLet me check it.</code></pre>\n<p>Now the next message is:</p>\n<pre><code>Can you create a support ticket for it?</code></pre>\n<p>The AI needs to understand that:</p>\n<pre><code>&quot;it&quot;</code></pre>\n<p>refers to:</p>\n<pre><code>ORD-10291</code></pre>\n<p>This requires conversational context.</p>\n<p>That&#39;s where <strong>chat memory</strong> becomes important.</p>\n<p>A simple conversation can be represented as:</p>\n<pre><code>User:\nMy order is late.\nAssistant:\nWhat&#39;s your order number?\nUser:\nORD-10291.\nAssistant:\nLet me check that order.</code></pre>\n<p>The application maintains the conversation history.</p>\n<p>Conceptually:</p>\n<pre><code>Conversation ID\n       ↓\nChat Memory\n       ↓\nPrevious Messages\n       ↓\nCurrent Prompt\n       ↓\nLLM</code></pre>\n<p>Spring AI provides abstractions for managing chat memory.</p>\n<p>It is useful to distinguish two concepts.</p>\n<p>Conversation context.</p>\n<pre><code>User:\nMy order is late.\nUser:\nIt&#39;s order 10291.\nUser:\nCan you check it?</code></pre>\n<p>The system remembers the current conversation.</p>\n<p>Persistent information about the user.</p>\n<p>For example:</p>\n<pre><code>Customer:\nAyush\nPreferences:\nPreferred language = English\nPreferred notification = Email</code></pre>\n<p>Long-term memory usually requires persistence in a database or another storage system.</p>\n<p>A production architecture might look like:</p>\n<pre><code>Conversation\n     ↓\nChat Memory Store\n     ↓\nPostgreSQL / Redis</code></pre>\n<p>The exact storage mechanism depends on the application.</p>\n<p>Now we can combine:</p>\n<pre><code>User\n ↓\nAgent\n ↓\nMemory\n ↓\nLLM\n ↓\nTools\n ↓\nTool Results\n ↓\nMemory\n ↓\nLLM\n ↓\nAnswer</code></pre>\n<p>This enables more natural multi-turn interactions.</p>\n<p>Another important Spring AI concept is the <strong>Advisor</strong>.</p>\n<p>Advisors can intercept and influence the interaction between the application and the model.</p>\n<p>They can be used for concerns such as:</p>\n<pre><code>Conversation memory\nRAG\nLogging\nSecurity\nPrompt modification\nContext injection\nObservability</code></pre>\n<p>Conceptually:</p>\n<pre><code>User\n ↓\nChatClient\n ↓\nAdvisor\n ↓\nChatModel\n ↓\nAdvisor\n ↓\nResponse</code></pre>\n<p>This allows cross-cutting AI behavior to be separated from business logic.</p>\n<p>Now things become much more powerful.</p>\n<p>Imagine an enterprise support agent.</p>\n<p>It has:</p>\n<pre><code>RAG\n ↓\nCompany documentation</code></pre>\n<p>Tools:</p>\n<pre><code>getCustomer()\ngetOrder()\ncreateTicket()</code></pre>\n<p>Memory:</p>\n<pre><code>Conversation history</code></pre>\n<p>The architecture becomes:</p>\n<pre><code>                    User\n                     ↓\n                 AI Agent\n                     ↓\n             ┌───────┼────────┐\n             ↓       ↓        ↓\n            RAG    Tools    Memory\n             ↓       ↓        ↓\n        Knowledge   APIs   Conversation\n             │       │        │\n             └───────┼────────┘\n                     ↓\n                    LLM\n                     ↓\n                  Response</code></pre>\n<p>This is much closer to a production AI application.</p>\n<p>Consider the request:</p>\n<pre><code>My payment failed for order ORD-19291.\nCan you check what happened and tell me\nwhat I should do?</code></pre>\n<p>The agent could perform:</p>\n<pre><code>1. getOrder(&quot;ORD-19291&quot;)\n2. getPaymentStatus(&quot;ORD-19291&quot;)\n3. searchKnowledgeBase(&quot;payment failure&quot;)\n4. Generate explanation</code></pre>\n<p>The final answer could be:</p>\n<pre><code>Your payment attempt failed because the transaction\nwas declined by the payment provider.\nAccording to the payment policy, you can retry the\npayment using another payment method.\nWould you like me to create a support ticket?</code></pre>\n<p>The model combined:</p>\n<pre><code>Live application data\n+\nKnowledge base\n+\nConversation context</code></pre>\n<p>This is significantly more useful than a standalone chatbot.</p>\n<p>Agents can also perform multi-step workflows.</p>\n<p>For example:</p>\n<pre><code>User:\nFind my overdue invoices and send reminders.</code></pre>\n<p>The agent could reason through:</p>\n<pre><code>getCustomer()\n      ↓\ngetInvoices()\n      ↓\nFilter overdue invoices\n      ↓\nsendReminder()\n      ↓\nReturn summary</code></pre>\n<p>The workflow becomes:</p>\n<pre><code>Goal\n ↓\nPlan\n ↓\nTool\n ↓\nObserve\n ↓\nNext Decision\n ↓\nTool\n ↓\nObserve\n ↓\nFinal Result</code></pre>\n<p>This pattern is often called an <strong>agent loop</strong>.</p>\n<p>A simplified conceptual implementation looks like:</p>\n<pre><code>while (!completed) {\n    AgentDecision decision =\n            llm.decide(context);\n    if (decision.requiresTool()) {\n        ToolResult result =\n                toolExecutor.execute(\n                        decision.toolCall()\n                );\n        context.add(result);\n    } else {\n        return decision.finalAnswer();\n    }\n}</code></pre>\n<p>In real applications, frameworks handle much of this interaction.</p>\n<p>But understanding the underlying loop is important.</p>\n<p>An important engineering lesson:</p>\n<p><strong>Not every AI feature needs an agent.</strong></p>\n<p>If your workflow is deterministic:</p>\n<pre><code>Validate request\n ↓\nCall API\n ↓\nSave result\n ↓\nReturn response</code></pre>\n<p>you probably don&#39;t need an autonomous agent.</p>\n<p>A normal service workflow may be better.</p>\n<p>Agents become more useful when:</p>\n<pre><code>The next step depends on the current result.</code></pre>\n<p>For example:</p>\n<pre><code>Check order\n ↓\nIf delayed\n ↓\nCheck refund policy\n ↓\nIf eligible\n ↓\nAsk for confirmation\n ↓\nCreate refund request</code></pre>\n<p>The dynamic decision-making is where agents become valuable.</p>\n<pre><code>A → B → C → D</code></pre>\n<p>Everything is predetermined.</p>\n<pre><code>A\n ↓\nLLM decides\n ├── B\n ├── C\n └── D\n      ↓\n   Observe result\n      ↓\n   Decide again</code></pre>\n<p>Agents provide flexibility.</p>\n<p>Traditional workflows provide predictability.</p>\n<p>Production systems often use both.</p>\n<p>A practical Spring Boot architecture might look like:</p>\n<pre><code>                    ┌───────────────┐\n                    │   Frontend    │\n                    └───────┬───────┘\n                            ↓\n                    ┌───────────────┐\n                    │ Spring Boot   │\n                    │     API       │\n                    └───────┬───────┘\n                            ↓\n                     ┌────────────┐\n                     │ ChatClient │\n                     └─────┬──────┘\n                           ↓\n                    ┌──────────────┐\n                    │    Agent     │\n                    └──────┬───────┘\n                           ↓\n              ┌────────────┼────────────┐\n              ↓            ↓            ↓\n           Memory         RAG         Tools\n              ↓            ↓            ↓\n          PostgreSQL    pgvector      APIs\n                                         ↓\n                                    Microservices</code></pre>\n<p>This architecture fits naturally into existing Spring Boot applications.</p>\n<p>Agent systems can become difficult to debug.</p>\n<p>Imagine an agent performs:</p>\n<pre><code>Tool 1\nTool 2\nTool 3\nTool 4</code></pre>\n<p>and the final response is incorrect.</p>\n<p>You need to know:</p>\n<pre><code>What did the model decide?\nWhich tools were selected?\nWhat arguments were sent?\nHow long did each tool take?\nWhat did each tool return?\nHow many model calls happened?\nHow many tokens were consumed?</code></pre>\n<p>Therefore, observability is critical.</p>\n<p>Track:</p>\n<pre><code>LLM latency\nTool latency\nRetrieval latency\nToken usage\nTool calls\nTool failures\nModel responses\nAgent iterations\nErrors</code></pre>\n<p>An agent can potentially continue calling tools indefinitely.</p>\n<p>For example:</p>\n<pre><code>LLM\n ↓\nTool\n ↓\nLLM\n ↓\nTool\n ↓\nLLM\n ↓\nTool\n ↓\n...</code></pre>\n<p>Production systems should enforce limits.</p>\n<p>For example:</p>\n<pre><code>Maximum iterations = 10\nMaximum tool calls = 20\nMaximum execution time = 30 seconds</code></pre>\n<p>You should also define clear failure behavior.</p>\n<pre><code>Agent limit reached\n        ↓\nStop execution\n        ↓\nReturn safe response\n        ↓\nLog failure</code></pre>\n<p>Never blindly trust model-generated tool arguments.</p>\n<p>Suppose the model requests:</p>\n<pre><code>{\n  &quot;orderId&quot;: &quot;ORD-999999999&quot;\n}</code></pre>\n<p>Your application should still validate:</p>\n<pre><code>Does the order exist?\nDoes the user own the order?\nIs the user authorized?\nIs the order accessible to this tenant?</code></pre>\n<p>The architecture should be:</p>\n<pre><code>LLM\n ↓\nTool Request\n ↓\nSchema Validation\n ↓\nAuthorization\n ↓\nBusiness Validation\n ↓\nTool Execution</code></pre>\n<p>The LLM is not your security boundary.</p>\n<p>Your application is.</p>\n<p>This becomes especially important in SaaS applications.</p>\n<p>Imagine:</p>\n<pre><code>Tenant A\n ├── Customers\n ├── Orders\n └── Documents\nTenant B\n ├── Customers\n ├── Orders\n └── Documents</code></pre>\n<p>An AI agent must never retrieve Tenant B&#39;s information while processing a Tenant A request.</p>\n<p>Every tool and retrieval operation should carry tenant context.</p>\n<p>For example:</p>\n<pre><code>tenant_id\nuser_id\nroles\npermissions</code></pre>\n<p>Then:</p>\n<pre><code>User\n ↓\nAuthentication\n ↓\nTenant Context\n ↓\nAgent\n ↓\nTool\n ↓\nAuthorization\n ↓\nTenant-scoped Data</code></pre>\n<p>The same principle applies to RAG.</p>\n<pre><code>Vector Search\n +\ntenant_id filter</code></pre>\n<p>should ensure that retrieved documents belong to the correct tenant.</p>\n<p>Production agents should have explicit guardrails.</p>\n<p>Examples:</p>\n<pre><code>Input validation\nOutput validation\nTool authorization\nRate limiting\nToken limits\nIteration limits\nPII protection\nAudit logging\nHuman approval</code></pre>\n<p>For high-risk actions:</p>\n<pre><code>Agent\n ↓\nTool Request\n ↓\nRisk Evaluation\n ↓\nHuman Approval\n ↓\nExecution</code></pre>\n<p>This creates a human-in-the-loop architecture.</p>\n<p>Not every decision should be fully automated.</p>\n<p>For example:</p>\n<pre><code>Refund amount &lt; $50\n    ↓\nAutomatic\nRefund amount &gt; $50\n    ↓\nHuman approval</code></pre>\n<p>Or:</p>\n<pre><code>Create support ticket\n    ↓\nAutomatic\nDelete account\n    ↓\nConfirmation required</code></pre>\n<p>This gives us a practical balance:</p>\n<pre><code>AI Automation\n+\nBusiness Rules\n+\nHuman Oversight</code></pre>\n<p>At this point, we can combine everything we&#39;ve discussed.</p>\n<pre><code>                         User\n                          ↓\n                     Spring Boot\n                          ↓\n                      ChatClient\n                          ↓\n                       AI Agent\n                          ↓\n              ┌───────────┼───────────┐\n              ↓           ↓           ↓\n            Memory       RAG         Tools\n              ↓           ↓           ↓\n          PostgreSQL   pgvector    REST APIs\n                                      ↓\n                               Business Services\n                                      ↓\n                                   Database</code></pre>\n<p>This is a strong foundation for enterprise AI applications.</p>\n<p>Imagine a sales assistant.</p>\n<p>The user asks:</p>\n<pre><code>Show me the latest opportunities for Acme\nand tell me which ones are likely to close this month.</code></pre>\n<p>The agent could:</p>\n<pre><code>1. getCustomer(&quot;Acme&quot;)\n2. getOpportunities(&quot;Acme&quot;)\n3. retrieve sales documentation\n4. analyze opportunity information\n5. generate summary</code></pre>\n<p>Now the user says:</p>\n<pre><code>Create a follow-up task for the highest priority opportunity.</code></pre>\n<p>The agent can:</p>\n<pre><code>1. Identify opportunity\n2. createFollowUpTask()\n3. Return task details</code></pre>\n<p>This is where AI starts becoming an application interface rather than simply a chatbot.</p>\n<p>Think about the responsibilities this way:</p>\n<pre><code>LLM\n=\nReasoning + Language\nRAG\n=\nKnowledge Retrieval\nTools\n=\nActions + Live Data\nMemory\n=\nConversation Context\nSpring Boot\n=\nApplication + Security + Business Logic</code></pre>\n<p>Together:</p>\n<pre><code>LLM\n +\nRAG\n +\nTools\n +\nMemory\n +\nBusiness Logic\n =\nAI Application</code></pre>\n<p>Spring AI provides abstractions that allow Java developers to work with AI capabilities using familiar Spring patterns.</p>\n<p>Important building blocks include:</p>\n<pre><code>ChatClient\nChatModel\nEmbeddingModel\nVectorStore\nDocument\nAdvisors\nChat Memory\nTools</code></pre>\n<p>This means an enterprise Java team can integrate AI into an existing Spring Boot architecture instead of creating an entirely separate AI stack.</p>\n<p>For example:</p>\n<pre><code>Existing Spring Boot Application\n              ↓\n        Spring AI Layer\n              ↓\n      Model + RAG + Tools\n              ↓\n     Existing Microservices</code></pre>\n<p>This makes AI integration much more practical for Java teams.</p>\n<p>A more complete production system might eventually look like:</p>\n<pre><code>                         ┌───────────────┐\n                         │     User      │\n                         └───────┬───────┘\n                                 ↓\n                         API Gateway\n                                 ↓\n                         Authentication\n                                 ↓\n                         Spring Boot API\n                                 ↓\n                            AI Agent\n                                 ↓\n          ┌──────────────────────┼──────────────────────┐\n          ↓                      ↓                      ↓\n       Memory                  RAG                    Tools\n          ↓                      ↓                      ↓\n     PostgreSQL              pgvector              Microservices\n                                                        ↓\n                                                Business Database\n                                 ↓\n                           LLM Provider\n                                 ↓\n                              Response</code></pre>\n<p>And around the entire system:</p>\n<pre><code>Security\nObservability\nRate Limiting\nAudit Logging\nGuardrails\nEvaluation</code></pre>\n<p>These are not optional concerns in serious enterprise deployments.</p>\n<p>It is useful to understand the difference.</p>\n<pre><code>User\n ↓\nLLM\n ↓\nAnswer</code></pre>\n<pre><code>User\n ↓\nRetrieve Knowledge\n ↓\nLLM\n ↓\nAnswer</code></pre>\n<pre><code>User\n ↓\nLLM\n ↓\nTool\n ↓\nResult\n ↓\nAnswer</code></pre>\n<pre><code>User\n ↓\nAgent\n ↓\nReason\n ↓\nTool\n ↓\nObserve\n ↓\nReason\n ↓\nTool\n ↓\nObserve\n ↓\nFinal Answer</code></pre>\n<p>The complexity increases at every stage.</p>\n<p>We can now think about the evolution of an AI application:</p>\n<pre><code>Level 1\nLLM\n ↓\nText Generation\nLevel 2\nLLM + RAG\n ↓\nKnowledge Retrieval\nLevel 3\nLLM + Tools\n ↓\nActions\nLevel 4\nLLM + Tools + Memory\n ↓\nContextual Assistant\nLevel 5\nLLM + RAG + Tools + Memory\n ↓\nAgent\nLevel 6\nMultiple Agents + Workflows\n ↓\nAgentic System</code></pre>\n<p>This progression is useful when deciding how much complexity your application actually needs.</p>\n<p>The evolution from traditional AI applications to agentic applications can be summarized as:</p>\n<pre><code>LLM\n ↓\nGenerate Text\nRAG\n ↓\nRetrieve Knowledge\nTool Calling\n ↓\nTake Actions\nMemory\n ↓\nRemember Context\nAgents\n ↓\nMake Decisions\nWorkflows\n ↓\nCoordinate Multiple Steps</code></pre>\n<p>Spring AI provides Java developers with abstractions for building many of these capabilities inside the Spring ecosystem.</p>\n<p>The most important engineering principle is:</p>\n<p><strong>Let the model decide, but let your application control.</strong></p>\n<p>The LLM can decide which tool may be useful.</p>\n<p>Your application should decide whether that tool is actually allowed to execute.</p>\n<p>That separation gives us a much safer architecture for enterprise AI.</p>\n<p>We&#39;ve now covered three major capabilities:</p>\n<pre><code>LLM\n ↓\nGenerate\nRAG\n ↓\nRetrieve\nTools\n ↓\nAct</code></pre>\n<p>But there is another challenge.</p>\n<p>What happens when a system has:</p>\n<pre><code>Multiple agents\n        ↓\nMultiple tools\n        ↓\nMultiple services\n        ↓\nMultiple AI models</code></pre>\n<p>How do these agents communicate?</p>\n<p>How do we standardize tool discovery?</p>\n<p>How can an AI agent securely interact with external tools and services?</p>\n<p>This leads us toward another important concept in modern AI engineering:</p>\n<p><strong>Model Context Protocol — MCP.</strong></p>\n<p>In the next article, we&#39;ll explore:</p>\n<p><strong>Building MCP Clients and Tool-Based AI Applications with Spring AI.</strong></p>\n<ul><li>Tool calling allows LLMs to interact with application capabilities.</li><li>The LLM requests tools; your application executes them.</li><li>Tools can expose APIs, databases, business operations, and services.</li><li>Chat memory provides conversational context.</li><li>RAG provides external knowledge.</li><li>Agents can combine RAG, memory, and tools.</li><li>Not every workflow requires an agent.</li><li>Deterministic workflows are often better for predictable business processes.</li><li>Tool permissions and authorization are critical.</li><li>Destructive operations should require stronger controls.</li><li>Multi-tenant applications must enforce tenant isolation at the tool and retrieval layers.</li><li>Agent loops should have iteration, timeout, and tool-call limits.</li><li>Observability is essential for debugging agent behavior.</li><li>Human-in-the-loop approval is useful for high-risk actions.</li><li>Spring AI provides abstractions that make these patterns accessible to Java and Spring Boot developers.</li></ul>\n<p>The future of enterprise AI isn&#39;t just:</p>\n<pre><code>LLM → Answer</code></pre>\n<p>It&#39;s increasingly:</p>\n<pre><code>LLM\n ↓\nReason\n ↓\nRetrieve\n ↓\nCall Tools\n ↓\nObserve\n ↓\nAct\n ↓\nRemember\n ↓\nComplete the Goal</code></pre>\n<p>And that&#39;s where <strong>AI agents with Spring AI</strong> become truly interesting.</p>","headings":[]}}