{"article":{"slug":"how-to-build-a-reliable-ai-assistant-with-the-claude-api","title":"How to Build a Reliable AI Assistant with the Claude API","subtitle":null,"summary":"A freeCodeCamp tutorial building ShopHelper with the Claude API: conversation history, tools, multi-block responses, workflow patterns, and evaluating whether prompt changes actually help.","content_type":"tutorial","language":"en","canonical_url":"https://www.freecodecamp.org/news/how-to-build-a-reliable-ai-assistant-with-the-claude-api/","author":{"name":"Chidozie Managwu","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"freeCodeCamp","url":"https://www.freecodecamp.org","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2132,"reading_minutes":9,"published_at":"2026-09-29T00:00:00.000Z","added_at":"2026-09-30T12:14:06.208Z","updated_at":"2026-09-30T12:14:06.208Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/how-to-build-a-reliable-ai-assistant-with-the-claude-api","markdown_url":"https://listedarticles.com/articles/how-to-build-a-reliable-ai-assistant-with-the-claude-api.md","example":false,"citation":"Chidozie Managwu, freeCodeCamp. \"How to Build a Reliable AI Assistant with the Claude API.\" 29 Sept 2026. https://www.freecodecamp.org/news/how-to-build-a-reliable-ai-assistant-with-the-claude-api/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.freecodecamp.org/news/how-to-build-a-reliable-ai-assistant-with-the-claude-api/"},"body_markdown":"Large language models can answer questions, summarise documents, write code, and interact with external systems. But building a reliable AI application requires more than sending a prompt and displaying the response.\n\nA production-ready application must manage conversation history, provide relevant context, use tools safely, handle different response types, and evaluate whether the generated output is useful.\n\nIn this tutorial, we’ll build **ShopHelper**, a customer-support assistant for an imaginary online shop. By the end, ShopHelper will be able to:\n\n- Answer general questions in a consistent tone\n- Remember what a customer said earlier\n- Look up order statuses by calling a function in your code\n- Handle Claude’s multi-block responses safely\n- Process support tickets using workflows\n- Evaluate whether prompt changes improve results\n\nEach section adds one piece, so you can follow along in your own editor.\n\n## Table of Contents\n\n## Prerequisites\n\nYou should have:\n\n- Basic Python knowledge\n- Python 3.9 or later\n- An Anthropic API key\n- Familiarity with functions and JSON\n\n## How to Set Up the Project and Keep Your API Key Secure\n\nCreate a virtual environment and install the Anthropic Python SDK:\n\n```\npython -m venv .venv\nsource .venv/bin/activate\npip install anthropic python-dotenv\n```\nOn Windows:\n\n```\n.venv\\Scripts\\activate\n```\nCreate a `.env` file:\n\n```\nANTHROPIC_API_KEY=your_api_key_here\n```\nAn API key is a secret credential. Never place it in browser JavaScript, mobile-app code, or client-side configuration. Never commit it to a repository:\n\n```\necho \".env\" >> .gitignore\n```\nIf you add a web interface later, keep the key on your backend:\n\n```\nBrowser → Your backend → Claude API\n```\nCreate `app.py`:\n\n```\nimport os\nfrom anthropic import Anthropic\nfrom dotenv import load_dotenv\nload_dotenv()\nMODEL = \"claude-sonnet-5\"\nclient = Anthropic(\n    api_key=os.environ[\"ANTHROPIC_API_KEY\"]\n)\n```\n`load_dotenv()` loads the value from `.env`. The `MODEL` constant means you only need to change the model name in one place. Confirm that the model identifier is available to your account before running the example.\n\n## How to Make Your First Request\n\n```\nresponse = client.messages.create(\n    model=MODEL,\n    max_tokens=500,\n    messages=[\n        {\n            \"role\": \"user\",\n            \"content\": \"Explain what an API is in simple terms.\"\n        }\n    ],\n)\nanswer = \"\".join(\n    block.text\n    for block in response.content\n    if block.type == \"text\"\n)\nprint(answer)\n```\nA request contains three important parts:\n\n- `model` selects the Claude model that handles the request. Models can differ in capability, speed, and cost.\n- `max_tokens` limits the maximum amount of text Claude can generate. A smaller value can reduce latency, but Claude may stop before completing its answer.\n- `messages` contains the conversation. Each message has a`role` and`content` . The role is usually`user` or`assistant` .\n\nFor example, a one-off request contains one user message. A multi-turn conversation contains earlier user and assistant messages.\n\nClaude returns `response.content`, which is a list of typed content blocks. Common blocks include:\n\n| Block type | Meaning | \n|---|---|\n| `text` | Generated text | \n| `tool_use` | A request for your application to call a tool | \n| `thinking` | Reasoning content when enabled | \n\nThe example collects text blocks instead of assuming `response.content[0]` is always text.\n\nYou can inspect usage information for monitoring:\n\n```\nprint(response.usage.input_tokens)\nprint(response.usage.output_tokens)\n```\n## How to Manage Conversation History\n\nClaude doesn't automatically remember separate API requests. Send relevant history with every request:\n\n```\nmessages = [\n    {\n        \"role\": \"user\",\n        \"content\": \"What is your returns policy?\"\n    },\n    {\n        \"role\": \"assistant\",\n        \"content\": \"Items can be returned within 30 days.\"\n    },\n    {\n        \"role\": \"user\",\n        \"content\": \"How long do I have?\"\n    },\n]\nresponse = client.messages.create(\n    model=MODEL,\n    max_tokens=300,\n    messages=messages,\n)\n```\nThe assistant message records Claude’s earlier answer, allowing the final question to be interpreted in context.\n\nA simple chat function can maintain the history:\n\n```\ndef chat(history, user_text):\n    history.append({\n        \"role\": \"user\",\n        \"content\": user_text,\n    })\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=500,\n        messages=history,\n    )\n    reply = \"\".join(\n        block.text\n        for block in response.content\n        if block.type == \"text\"\n    )\n    history.append({\n        \"role\": \"assistant\",\n        \"content\": reply,\n    })\n    return reply\nhistory = []\nprint(chat(history, \"What is your returns policy?\"))\nprint(chat(history, \"How long do I have?\"))\n```\nEach call adds the new user message, sends the complete history, and stores Claude’s response for the next turn. In production, store histories by customer or session ID.\n\n### How to Manage History as it Grows\n\nUnlimited history increases input size and may make it harder for Claude to focus. One option is to retain only recent messages:\n\n```\ndef trim_history(history, max_messages=10):\n    trimmed = history[-max_messages:]\n    while trimmed and trimmed[0][\"role\"] != \"user\":\n        trimmed.pop(0)\n    return trimmed\n```\nAnother option is to summarise older turns while keeping recent messages:\n\n```\ndef summarise_history(history, keep_last=6):\n    old = history[:-keep_last]\n    recent = history[-keep_last:]\n    transcript = \"\\n\".join(\n        f\"{message['role']}: {message['content']}\"\n        for message in old\n    )\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=250,\n        messages=[{\n            \"role\": \"user\",\n            \"content\": (\n                \"Summarise this conversation in under 100 words. \"\n                \"Keep order numbers and unresolved issues.\\n\\n\"\n                f\"<conversation>{transcript}</conversation>\"\n            ),\n        }],\n    )\n    summary = \"\".join(\n        block.text\n        for block in response.content\n        if block.type == \"text\"\n    )\n    return summary, recent\n```\nKeep the summary as separate application state and include it as context in the next request. Don't insert it as an additional user message before `recent`, because that can create invalid consecutive user messages.\n\nSensitive information should also be redacted before storage or transmission:\n\n```\nimport re\ndef redact(text):\n    return re.sub(\n        r\"\\b(?:\\d[ -]?){13,16}\\b\",\n        \"[REDACTED CARD]\",\n        text,\n    )\n```\n## How to Structure Prompts with Clear Boundaries\n\nXML-style tags are ordinary text, not special API commands. They make each part of a prompt explicit:\n\n```\nprompt = \"\"\"\n<customer_reviews>\nThe product is comfortable, but the available colours are limited.\nCustomers also describe it as durable.\n</customer_reviews>\n<sales_data>\nJanuary: 120 units\nFebruary: 150 units\nMarch: 98 units\n</sales_data>\n<task>\nCompare the reviews with the sales data.\nIdentify possible relationships and state uncertainty.\n</task>\n\"\"\"\n```\nHere, `<customer_reviews>` identifies reference material, `<sales_data>` identifies the data, and `<task>` identifies the instruction. Use similar boundaries for policies, user-generated content, examples, and output requirements.\n\n## How to Use a System Prompt\n\nA system prompt defines ShopHelper’s general behaviour:\n\n```\nsystem_prompt = \"\"\"\nYou are ShopHelper, a friendly customer-support assistant.\nKeep answers concise and clear.\nDo not invent prices, policies, or order details.\nIf information is missing, ask for it.\n\"\"\"\n```\nPass it separately from the conversation:\n\n```\nresponse = client.messages.create(\n    model=MODEL,\n    max_tokens=500,\n    system=system_prompt,\n    messages=[\n        {\"role\": \"user\", \"content\": \"Where is my order?\"}\n    ],\n)\n```\nBecause the customer didn't provide an order number, ShopHelper should ask for one instead of guessing.\n\n## How to Add Tools\n\nClaude can't directly access your database. A tool gives it a structured way to request information from your application:\n\n```\ndef get_order_status(order_id):\n    orders = {\n        \"ORD-1001\": \"shipped\",\n        \"ORD-1002\": \"processing\",\n    }\n    return {\n        \"order_id\": order_id,\n        \"status\": orders.get(order_id, \"not_found\"),\n    }\n```\nThe function accepts an order ID, looks it up, and returns predictable data. In production, the dictionary would be replaced by a database query. Claude doesn't execute the function. Your application does.\n\nDescribe the function with a schema:\n\n```\ntools = [{\n    \"name\": \"get_order_status\",\n    \"description\": \"Get the current status of a customer order.\",\n    \"input_schema\": {\n        \"type\": \"object\",\n        \"properties\": {\n            \"order_id\": {\n                \"type\": \"string\",\n                \"description\": \"An order ID such as ORD-1001.\"\n            }\n        },\n        \"required\": [\"order_id\"],\n    },\n}]\n```\nClaude may return a `tool_use` block instead of a final answer:\n\n```\ntype=\"tool_use\"\nid=\"toolu_example\"\nname=\"get_order_status\"\ninput={\"order_id\": \"ORD-1001\"}\n```\nThe `name` identifies the function, `input` contains its arguments, and `id` is needed when returning the result. A `stop_reason` of `\"tool_use\"` means your application should handle the request before asking Claude to continue.\n\n## How to Handle a Tool-Use Response\n\nA tool-use response is a response containing the `tool_use` block described above.\n\nValidate the tool name, arguments, and user permissions before execution:\n\n```\nimport re\nORDER_ID_PATTERN = re.compile(r\"^ORD-\\d{4}$\")\ndef validate_tool_request(name, tool_input, current_user):\n    if name != \"get_order_status\":\n        return False, \"Unknown tool\"\n    order_id = tool_input.get(\"order_id\")\n    if not isinstance(order_id, str):\n        return False, \"order_id must be a string\"\n    if not ORDER_ID_PATTERN.fullmatch(order_id):\n        return False, \"Invalid order ID format\"\n    if order_id not in current_user[\"order_ids\"]:\n        return False, \"The customer cannot access this order\"\n    return True, None\n```\nA complete loop can then validate and execute the request:\n\n```\ndef run_conversation(user_text, current_user):\n    messages = [{\"role\": \"user\", \"content\": user_text}]\n    while True:\n        response = client.messages.create(\n            model=MODEL,\n            max_tokens=500,\n            system=system_prompt,\n            tools=tools,\n            messages=messages,\n        )\n        if response.stop_reason != \"tool_use\":\n            return \"\".join(\n                block.text\n                for block in response.content\n                if block.type == \"text\"\n            )\n        messages.append({\n            \"role\": \"assistant\",\n            \"content\": response.content,\n        })\n        results = []\n        for block in response.content:\n            if block.type != \"tool_use\":\n                continue\n            valid, error = validate_tool_request(\n                block.name,\n                block.input,\n                current_user,\n            )\n            if valid:\n                result = get_order_status(block.input[\"order_id\"])\n                results.append({\n                    \"type\": \"tool_result\",\n                    \"tool_use_id\": block.id,\n                    \"content\": str(result),\n                })\n            else:\n                results.append({\n                    \"type\": \"tool_result\",\n                    \"tool_use_id\": block.id,\n                    \"content\": error,\n                    \"is_error\": True,\n                })\n        messages.append({\n            \"role\": \"user\",\n            \"content\": results,\n        })\n```\nThe `tool_use_id` connects the result to the original request. The application remains responsible for authorisation and execution.\n\n## Claude Responses Can Contain Multiple Blocks\n\nThis assumption is fragile:\n\n```\nanswer = response.content[0].text\n```\nIt assumes that the first block exists and is text. Instead, inspect each block:\n\n```\nfor block in response.content:\n    if block.type == \"text\":\n        print(block.text)\n    elif block.type == \"tool_use\":\n        print(\"Validate and execute:\", block.name)\n    elif block.type == \"thinking\":\n        continue\n    else:\n        print(\"Unhandled block type:\", block.type)\n```\nShopHelper displays text, validates and executes approved tool requests, doesn't display internal thinking, and logs unknown block types.\n\n## Workflows vs Agents\n\nA workflow follows a predefined sequence:\n\n```\nReceive ticket\n↓\nExtract details\n↓\nDraft reply\n↓\nReview reply\n```\n```\ndef ask(prompt, max_tokens=500):\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=max_tokens,\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n    )\n    return \"\".join(\n        block.text\n        for block in response.content\n        if block.type == \"text\"\n    )\ndef handle_ticket_workflow(ticket):\n    details = ask(\n        f\"<ticket>{ticket}</ticket>\\n\"\n        \"<task>Extract the problem and desired outcome.</task>\"\n    )\n    draft = ask(\n        f\"<details>{details}</details>\\n\"\n        \"<task>Draft a concise support reply.</task>\"\n    )\n    review = ask(\n        f\"<draft>{draft}</draft>\\n\"\n        \"<task>List unsupported promises, or say OK.</task>\"\n    )\n    return draft, review\n```\nAn agent is more flexible: Claude decides whether to use a tool and what to do next. Agents still require validation and a maximum step count. The `run_conversation()` function above can be reused inside an agent loop.\n\nUse workflows when the steps are known and repeatability matters. Use agents when the next action depends on the current result.\n\n## Chaining, Parallelisation, Routing, and Evaluator-Optimizer\n\n**Chaining** passes each result to the next stage:\n\n```\ndef chained_reply(ticket, policy):\n    draft = ask(\n        f\"<ticket>{ticket}</ticket>\\n\"\n        \"<task>Draft a support reply.</task>\"\n    )\n    issues = ask(\n        f\"<policy>{policy}</policy>\\n\"\n        f\"<draft>{draft}</draft>\\n\"\n        \"<task>List unsupported claims.</task>\"\n    )\n    return ask(\n        f\"<draft>{draft}</draft>\\n\"\n        f\"<issues>{issues}</issues>\\n\"\n        \"<task>Rewrite the final reply.</task>\"\n    )\n```\n**Parallelisation** runs independent tasks concurrently:\n\n```\nfrom concurrent.futures import ThreadPoolExecutor\ntickets = [\n    \"My headphones arrived broken.\",\n    \"I was charged twice.\",\n    \"How do I change my address?\",\n]\ndef summarise(ticket):\n    return ask(\n        f\"<ticket>{ticket}</ticket>\\n\"\n        \"<task>Summarise in one sentence.</task>\",\n        max_tokens=100,\n    )\nwith ThreadPoolExecutor(max_workers=3) as pool:\n    summaries = list(pool.map(summarise, tickets))\ndigest = ask(\n    \"<summaries>\\n\"\n    + \"\\n\".join(summaries)\n    + \"\\n</summaries>\\n\"\n    \"<task>Summarise today's support themes.</task>\"\n)\n```\n**Routing** classifies a request before selecting a specialised workflow:\n\n```\ndef route(ticket):\n    label = ask(\n        f\"<ticket>{ticket}</ticket>\\n\"\n        \"<task>Return exactly refund, delivery, or general.</task>\",\n        max_tokens=10,\n    ).strip().lower()\n    return label if label in {\"refund\", \"delivery\", \"general\"} else \"general\"\n```\n**Evaluator-optimizer** generates, reviews, and revises an answer:\n\n```\ndef improve_reply(ticket, rounds=2):\n    reply = ask(\n        f\"<ticket>{ticket}</ticket>\\n\"\n        \"<task>Write a support reply.</task>\"\n    )\n    for _ in range(rounds):\n        review = ask(\n            f\"<reply>{reply}</reply>\\n\"\n            \"<task>List accuracy or tone problems, or say PASS.</task>\"\n        )\n        if review.strip().upper() == \"PASS\":\n            break\n        reply = ask(\n            f\"<reply>{reply}</reply>\\n\"\n            f\"<review>{review}</review>\\n\"\n            \"<task>Rewrite the reply.</task>\"\n        )\n    return reply\n```\nUse chaining for dependent stages, parallelisation for independent work, routing for specialised paths, and evaluator-optimizer loops when additional quality justifies extra API calls.\n\n## How to Evaluate Prompt Quality\n\nUse representative test cases:\n\n```\ntest_cases = [\n    {\n        \"ticket\": \"I want a refund for broken headphones.\",\n        \"expected\": \"refund\",\n    },\n    {\n        \"ticket\": \"Where is ORD-1002?\",\n        \"expected\": \"delivery\",\n    },\n    {\n        \"ticket\": \"Do you sell gift cards?\",\n        \"expected\": \"general\",\n    },\n]\n```\nThese cases cover different request types. Run the same cases after changing the system prompt, examples, model, token limit, or routing instructions:\n\n```\ndef evaluate(route_fn, cases):\n    passed = 0\n    for case in cases:\n        result = route_fn(case[\"ticket\"])\n        if result == case[\"expected\"]:\n            passed += 1\n        else:\n            print(\"Failed:\", case[\"ticket\"], result)\n    score = passed / len(cases)\n    print(f\"{passed}/{len(cases)} passed\")\n    return score\n```\nUse code-based graders for labels and JSON. Use human or model-based graders for tone, accuracy, and helpfulness.\n\n## Conclusion\n\nBuilding with the Claude API involves more than writing prompts. A reliable application needs structured context, managed conversation state, validated tool execution, deliberate response handling, suitable workflows, and repeatable evaluation.\n\nThe goal isn't to find one perfect prompt. It's to build a system around Claude that provides the right context, limits unsafe actions, handles uncertainty, and measures whether changes improve the result.","body_html":"<p>Large language models can answer questions, summarise documents, write code, and interact with external systems. But building a reliable AI application requires more than sending a prompt and displaying the response.</p>\n<p>A production-ready application must manage conversation history, provide relevant context, use tools safely, handle different response types, and evaluate whether the generated output is useful.</p>\n<p>In this tutorial, we’ll build <strong>ShopHelper</strong>, a customer-support assistant for an imaginary online shop. By the end, ShopHelper will be able to:</p>\n<ul><li>Answer general questions in a consistent tone</li><li>Remember what a customer said earlier</li><li>Look up order statuses by calling a function in your code</li><li>Handle Claude’s multi-block responses safely</li><li>Process support tickets using workflows</li><li>Evaluate whether prompt changes improve results</li></ul>\n<p>Each section adds one piece, so you can follow along in your own editor.</p>\n<h2 id=\"table-of-contents\">Table of Contents</h2>\n<h2 id=\"prerequisites\">Prerequisites</h2>\n<p>You should have:</p>\n<ul><li>Basic Python knowledge</li><li>Python 3.9 or later</li><li>An Anthropic API key</li><li>Familiarity with functions and JSON</li></ul>\n<h2 id=\"how-to-set-up-the-project-and-keep-your-api-key-secure\">How to Set Up the Project and Keep Your API Key Secure</h2>\n<p>Create a virtual environment and install the Anthropic Python SDK:</p>\n<pre><code>python -m venv .venv\nsource .venv/bin/activate\npip install anthropic python-dotenv</code></pre>\n<p>On Windows:</p>\n<pre><code>.venv\\Scripts\\activate</code></pre>\n<p>Create a <code>.env</code> file:</p>\n<pre><code>ANTHROPIC_API_KEY=your_api_key_here</code></pre>\n<p>An API key is a secret credential. Never place it in browser JavaScript, mobile-app code, or client-side configuration. Never commit it to a repository:</p>\n<pre><code>echo &quot;.env&quot; &gt;&gt; .gitignore</code></pre>\n<p>If you add a web interface later, keep the key on your backend:</p>\n<pre><code>Browser → Your backend → Claude API</code></pre>\n<p>Create <code>app.py</code>:</p>\n<pre><code>import os\nfrom anthropic import Anthropic\nfrom dotenv import load_dotenv\nload_dotenv()\nMODEL = &quot;claude-sonnet-5&quot;\nclient = Anthropic(\n    api_key=os.environ[&quot;ANTHROPIC_API_KEY&quot;]\n)</code></pre>\n<p><code>load_dotenv()</code> loads the value from <code>.env</code>. The <code>MODEL</code> constant means you only need to change the model name in one place. Confirm that the model identifier is available to your account before running the example.</p>\n<h2 id=\"how-to-make-your-first-request\">How to Make Your First Request</h2>\n<pre><code>response = client.messages.create(\n    model=MODEL,\n    max_tokens=500,\n    messages=[\n        {\n            &quot;role&quot;: &quot;user&quot;,\n            &quot;content&quot;: &quot;Explain what an API is in simple terms.&quot;\n        }\n    ],\n)\nanswer = &quot;&quot;.join(\n    block.text\n    for block in response.content\n    if block.type == &quot;text&quot;\n)\nprint(answer)</code></pre>\n<p>A request contains three important parts:</p>\n<ul><li><code>model</code> selects the Claude model that handles the request. Models can differ in capability, speed, and cost.</li><li><code>max_tokens</code> limits the maximum amount of text Claude can generate. A smaller value can reduce latency, but Claude may stop before completing its answer.</li><li><code>messages</code> contains the conversation. Each message has a<code>role</code> and<code>content</code> . The role is usually<code>user</code> or<code>assistant</code> .</li></ul>\n<p>For example, a one-off request contains one user message. A multi-turn conversation contains earlier user and assistant messages.</p>\n<p>Claude returns <code>response.content</code>, which is a list of typed content blocks. Common blocks include:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Block type</th><th>Meaning</th></tr></thead><tbody><tr><td><code>text</code></td><td>Generated text</td></tr><tr><td><code>tool_use</code></td><td>A request for your application to call a tool</td></tr><tr><td><code>thinking</code></td><td>Reasoning content when enabled</td></tr></tbody></table></div>\n<p>The example collects text blocks instead of assuming <code>response.content[0]</code> is always text.</p>\n<p>You can inspect usage information for monitoring:</p>\n<pre><code>print(response.usage.input_tokens)\nprint(response.usage.output_tokens)</code></pre>\n<h2 id=\"how-to-manage-conversation-history\">How to Manage Conversation History</h2>\n<p>Claude doesn&#39;t automatically remember separate API requests. Send relevant history with every request:</p>\n<pre><code>messages = [\n    {\n        &quot;role&quot;: &quot;user&quot;,\n        &quot;content&quot;: &quot;What is your returns policy?&quot;\n    },\n    {\n        &quot;role&quot;: &quot;assistant&quot;,\n        &quot;content&quot;: &quot;Items can be returned within 30 days.&quot;\n    },\n    {\n        &quot;role&quot;: &quot;user&quot;,\n        &quot;content&quot;: &quot;How long do I have?&quot;\n    },\n]\nresponse = client.messages.create(\n    model=MODEL,\n    max_tokens=300,\n    messages=messages,\n)</code></pre>\n<p>The assistant message records Claude’s earlier answer, allowing the final question to be interpreted in context.</p>\n<p>A simple chat function can maintain the history:</p>\n<pre><code>def chat(history, user_text):\n    history.append({\n        &quot;role&quot;: &quot;user&quot;,\n        &quot;content&quot;: user_text,\n    })\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=500,\n        messages=history,\n    )\n    reply = &quot;&quot;.join(\n        block.text\n        for block in response.content\n        if block.type == &quot;text&quot;\n    )\n    history.append({\n        &quot;role&quot;: &quot;assistant&quot;,\n        &quot;content&quot;: reply,\n    })\n    return reply\nhistory = []\nprint(chat(history, &quot;What is your returns policy?&quot;))\nprint(chat(history, &quot;How long do I have?&quot;))</code></pre>\n<p>Each call adds the new user message, sends the complete history, and stores Claude’s response for the next turn. In production, store histories by customer or session ID.</p>\n<h3 id=\"how-to-manage-history-as-it-grows\">How to Manage History as it Grows</h3>\n<p>Unlimited history increases input size and may make it harder for Claude to focus. One option is to retain only recent messages:</p>\n<pre><code>def trim_history(history, max_messages=10):\n    trimmed = history[-max_messages:]\n    while trimmed and trimmed[0][&quot;role&quot;] != &quot;user&quot;:\n        trimmed.pop(0)\n    return trimmed</code></pre>\n<p>Another option is to summarise older turns while keeping recent messages:</p>\n<pre><code>def summarise_history(history, keep_last=6):\n    old = history[:-keep_last]\n    recent = history[-keep_last:]\n    transcript = &quot;\\n&quot;.join(\n        f&quot;{message[&#39;role&#39;]}: {message[&#39;content&#39;]}&quot;\n        for message in old\n    )\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=250,\n        messages=[{\n            &quot;role&quot;: &quot;user&quot;,\n            &quot;content&quot;: (\n                &quot;Summarise this conversation in under 100 words. &quot;\n                &quot;Keep order numbers and unresolved issues.\\n\\n&quot;\n                f&quot;&lt;conversation&gt;{transcript}&lt;/conversation&gt;&quot;\n            ),\n        }],\n    )\n    summary = &quot;&quot;.join(\n        block.text\n        for block in response.content\n        if block.type == &quot;text&quot;\n    )\n    return summary, recent</code></pre>\n<p>Keep the summary as separate application state and include it as context in the next request. Don&#39;t insert it as an additional user message before <code>recent</code>, because that can create invalid consecutive user messages.</p>\n<p>Sensitive information should also be redacted before storage or transmission:</p>\n<pre><code>import re\ndef redact(text):\n    return re.sub(\n        r&quot;\\b(?:\\d[ -]?){13,16}\\b&quot;,\n        &quot;[REDACTED CARD]&quot;,\n        text,\n    )</code></pre>\n<h2 id=\"how-to-structure-prompts-with-clear-boundaries\">How to Structure Prompts with Clear Boundaries</h2>\n<p>XML-style tags are ordinary text, not special API commands. They make each part of a prompt explicit:</p>\n<pre><code>prompt = &quot;&quot;&quot;\n&lt;customer_reviews&gt;\nThe product is comfortable, but the available colours are limited.\nCustomers also describe it as durable.\n&lt;/customer_reviews&gt;\n&lt;sales_data&gt;\nJanuary: 120 units\nFebruary: 150 units\nMarch: 98 units\n&lt;/sales_data&gt;\n&lt;task&gt;\nCompare the reviews with the sales data.\nIdentify possible relationships and state uncertainty.\n&lt;/task&gt;\n&quot;&quot;&quot;</code></pre>\n<p>Here, <code>&lt;customer_reviews&gt;</code> identifies reference material, <code>&lt;sales_data&gt;</code> identifies the data, and <code>&lt;task&gt;</code> identifies the instruction. Use similar boundaries for policies, user-generated content, examples, and output requirements.</p>\n<h2 id=\"how-to-use-a-system-prompt\">How to Use a System Prompt</h2>\n<p>A system prompt defines ShopHelper’s general behaviour:</p>\n<pre><code>system_prompt = &quot;&quot;&quot;\nYou are ShopHelper, a friendly customer-support assistant.\nKeep answers concise and clear.\nDo not invent prices, policies, or order details.\nIf information is missing, ask for it.\n&quot;&quot;&quot;</code></pre>\n<p>Pass it separately from the conversation:</p>\n<pre><code>response = client.messages.create(\n    model=MODEL,\n    max_tokens=500,\n    system=system_prompt,\n    messages=[\n        {&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;Where is my order?&quot;}\n    ],\n)</code></pre>\n<p>Because the customer didn&#39;t provide an order number, ShopHelper should ask for one instead of guessing.</p>\n<h2 id=\"how-to-add-tools\">How to Add Tools</h2>\n<p>Claude can&#39;t directly access your database. A tool gives it a structured way to request information from your application:</p>\n<pre><code>def get_order_status(order_id):\n    orders = {\n        &quot;ORD-1001&quot;: &quot;shipped&quot;,\n        &quot;ORD-1002&quot;: &quot;processing&quot;,\n    }\n    return {\n        &quot;order_id&quot;: order_id,\n        &quot;status&quot;: orders.get(order_id, &quot;not_found&quot;),\n    }</code></pre>\n<p>The function accepts an order ID, looks it up, and returns predictable data. In production, the dictionary would be replaced by a database query. Claude doesn&#39;t execute the function. Your application does.</p>\n<p>Describe the function with a schema:</p>\n<pre><code>tools = [{\n    &quot;name&quot;: &quot;get_order_status&quot;,\n    &quot;description&quot;: &quot;Get the current status of a customer order.&quot;,\n    &quot;input_schema&quot;: {\n        &quot;type&quot;: &quot;object&quot;,\n        &quot;properties&quot;: {\n            &quot;order_id&quot;: {\n                &quot;type&quot;: &quot;string&quot;,\n                &quot;description&quot;: &quot;An order ID such as ORD-1001.&quot;\n            }\n        },\n        &quot;required&quot;: [&quot;order_id&quot;],\n    },\n}]</code></pre>\n<p>Claude may return a <code>tool_use</code> block instead of a final answer:</p>\n<pre><code>type=&quot;tool_use&quot;\nid=&quot;toolu_example&quot;\nname=&quot;get_order_status&quot;\ninput={&quot;order_id&quot;: &quot;ORD-1001&quot;}</code></pre>\n<p>The <code>name</code> identifies the function, <code>input</code> contains its arguments, and <code>id</code> is needed when returning the result. A <code>stop_reason</code> of <code>&quot;tool_use&quot;</code> means your application should handle the request before asking Claude to continue.</p>\n<h2 id=\"how-to-handle-a-tool-use-response\">How to Handle a Tool-Use Response</h2>\n<p>A tool-use response is a response containing the <code>tool_use</code> block described above.</p>\n<p>Validate the tool name, arguments, and user permissions before execution:</p>\n<pre><code>import re\nORDER_ID_PATTERN = re.compile(r&quot;^ORD-\\d{4}$&quot;)\ndef validate_tool_request(name, tool_input, current_user):\n    if name != &quot;get_order_status&quot;:\n        return False, &quot;Unknown tool&quot;\n    order_id = tool_input.get(&quot;order_id&quot;)\n    if not isinstance(order_id, str):\n        return False, &quot;order_id must be a string&quot;\n    if not ORDER_ID_PATTERN.fullmatch(order_id):\n        return False, &quot;Invalid order ID format&quot;\n    if order_id not in current_user[&quot;order_ids&quot;]:\n        return False, &quot;The customer cannot access this order&quot;\n    return True, None</code></pre>\n<p>A complete loop can then validate and execute the request:</p>\n<pre><code>def run_conversation(user_text, current_user):\n    messages = [{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: user_text}]\n    while True:\n        response = client.messages.create(\n            model=MODEL,\n            max_tokens=500,\n            system=system_prompt,\n            tools=tools,\n            messages=messages,\n        )\n        if response.stop_reason != &quot;tool_use&quot;:\n            return &quot;&quot;.join(\n                block.text\n                for block in response.content\n                if block.type == &quot;text&quot;\n            )\n        messages.append({\n            &quot;role&quot;: &quot;assistant&quot;,\n            &quot;content&quot;: response.content,\n        })\n        results = []\n        for block in response.content:\n            if block.type != &quot;tool_use&quot;:\n                continue\n            valid, error = validate_tool_request(\n                block.name,\n                block.input,\n                current_user,\n            )\n            if valid:\n                result = get_order_status(block.input[&quot;order_id&quot;])\n                results.append({\n                    &quot;type&quot;: &quot;tool_result&quot;,\n                    &quot;tool_use_id&quot;: block.id,\n                    &quot;content&quot;: str(result),\n                })\n            else:\n                results.append({\n                    &quot;type&quot;: &quot;tool_result&quot;,\n                    &quot;tool_use_id&quot;: block.id,\n                    &quot;content&quot;: error,\n                    &quot;is_error&quot;: True,\n                })\n        messages.append({\n            &quot;role&quot;: &quot;user&quot;,\n            &quot;content&quot;: results,\n        })</code></pre>\n<p>The <code>tool_use_id</code> connects the result to the original request. The application remains responsible for authorisation and execution.</p>\n<h2 id=\"claude-responses-can-contain-multiple-blocks\">Claude Responses Can Contain Multiple Blocks</h2>\n<p>This assumption is fragile:</p>\n<pre><code>answer = response.content[0].text</code></pre>\n<p>It assumes that the first block exists and is text. Instead, inspect each block:</p>\n<pre><code>for block in response.content:\n    if block.type == &quot;text&quot;:\n        print(block.text)\n    elif block.type == &quot;tool_use&quot;:\n        print(&quot;Validate and execute:&quot;, block.name)\n    elif block.type == &quot;thinking&quot;:\n        continue\n    else:\n        print(&quot;Unhandled block type:&quot;, block.type)</code></pre>\n<p>ShopHelper displays text, validates and executes approved tool requests, doesn&#39;t display internal thinking, and logs unknown block types.</p>\n<h2 id=\"workflows-vs-agents\">Workflows vs Agents</h2>\n<p>A workflow follows a predefined sequence:</p>\n<pre><code>Receive ticket\n↓\nExtract details\n↓\nDraft reply\n↓\nReview reply</code></pre>\n<pre><code>def ask(prompt, max_tokens=500):\n    response = client.messages.create(\n        model=MODEL,\n        max_tokens=max_tokens,\n        messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}],\n    )\n    return &quot;&quot;.join(\n        block.text\n        for block in response.content\n        if block.type == &quot;text&quot;\n    )\ndef handle_ticket_workflow(ticket):\n    details = ask(\n        f&quot;&lt;ticket&gt;{ticket}&lt;/ticket&gt;\\n&quot;\n        &quot;&lt;task&gt;Extract the problem and desired outcome.&lt;/task&gt;&quot;\n    )\n    draft = ask(\n        f&quot;&lt;details&gt;{details}&lt;/details&gt;\\n&quot;\n        &quot;&lt;task&gt;Draft a concise support reply.&lt;/task&gt;&quot;\n    )\n    review = ask(\n        f&quot;&lt;draft&gt;{draft}&lt;/draft&gt;\\n&quot;\n        &quot;&lt;task&gt;List unsupported promises, or say OK.&lt;/task&gt;&quot;\n    )\n    return draft, review</code></pre>\n<p>An agent is more flexible: Claude decides whether to use a tool and what to do next. Agents still require validation and a maximum step count. The <code>run_conversation()</code> function above can be reused inside an agent loop.</p>\n<p>Use workflows when the steps are known and repeatability matters. Use agents when the next action depends on the current result.</p>\n<h2 id=\"chaining-parallelisation-routing-and-evaluator-optimizer\">Chaining, Parallelisation, Routing, and Evaluator-Optimizer</h2>\n<p><strong>Chaining</strong> passes each result to the next stage:</p>\n<pre><code>def chained_reply(ticket, policy):\n    draft = ask(\n        f&quot;&lt;ticket&gt;{ticket}&lt;/ticket&gt;\\n&quot;\n        &quot;&lt;task&gt;Draft a support reply.&lt;/task&gt;&quot;\n    )\n    issues = ask(\n        f&quot;&lt;policy&gt;{policy}&lt;/policy&gt;\\n&quot;\n        f&quot;&lt;draft&gt;{draft}&lt;/draft&gt;\\n&quot;\n        &quot;&lt;task&gt;List unsupported claims.&lt;/task&gt;&quot;\n    )\n    return ask(\n        f&quot;&lt;draft&gt;{draft}&lt;/draft&gt;\\n&quot;\n        f&quot;&lt;issues&gt;{issues}&lt;/issues&gt;\\n&quot;\n        &quot;&lt;task&gt;Rewrite the final reply.&lt;/task&gt;&quot;\n    )</code></pre>\n<p><strong>Parallelisation</strong> runs independent tasks concurrently:</p>\n<pre><code>from concurrent.futures import ThreadPoolExecutor\ntickets = [\n    &quot;My headphones arrived broken.&quot;,\n    &quot;I was charged twice.&quot;,\n    &quot;How do I change my address?&quot;,\n]\ndef summarise(ticket):\n    return ask(\n        f&quot;&lt;ticket&gt;{ticket}&lt;/ticket&gt;\\n&quot;\n        &quot;&lt;task&gt;Summarise in one sentence.&lt;/task&gt;&quot;,\n        max_tokens=100,\n    )\nwith ThreadPoolExecutor(max_workers=3) as pool:\n    summaries = list(pool.map(summarise, tickets))\ndigest = ask(\n    &quot;&lt;summaries&gt;\\n&quot;\n    + &quot;\\n&quot;.join(summaries)\n    + &quot;\\n&lt;/summaries&gt;\\n&quot;\n    &quot;&lt;task&gt;Summarise today&#39;s support themes.&lt;/task&gt;&quot;\n)</code></pre>\n<p><strong>Routing</strong> classifies a request before selecting a specialised workflow:</p>\n<pre><code>def route(ticket):\n    label = ask(\n        f&quot;&lt;ticket&gt;{ticket}&lt;/ticket&gt;\\n&quot;\n        &quot;&lt;task&gt;Return exactly refund, delivery, or general.&lt;/task&gt;&quot;,\n        max_tokens=10,\n    ).strip().lower()\n    return label if label in {&quot;refund&quot;, &quot;delivery&quot;, &quot;general&quot;} else &quot;general&quot;</code></pre>\n<p><strong>Evaluator-optimizer</strong> generates, reviews, and revises an answer:</p>\n<pre><code>def improve_reply(ticket, rounds=2):\n    reply = ask(\n        f&quot;&lt;ticket&gt;{ticket}&lt;/ticket&gt;\\n&quot;\n        &quot;&lt;task&gt;Write a support reply.&lt;/task&gt;&quot;\n    )\n    for _ in range(rounds):\n        review = ask(\n            f&quot;&lt;reply&gt;{reply}&lt;/reply&gt;\\n&quot;\n            &quot;&lt;task&gt;List accuracy or tone problems, or say PASS.&lt;/task&gt;&quot;\n        )\n        if review.strip().upper() == &quot;PASS&quot;:\n            break\n        reply = ask(\n            f&quot;&lt;reply&gt;{reply}&lt;/reply&gt;\\n&quot;\n            f&quot;&lt;review&gt;{review}&lt;/review&gt;\\n&quot;\n            &quot;&lt;task&gt;Rewrite the reply.&lt;/task&gt;&quot;\n        )\n    return reply</code></pre>\n<p>Use chaining for dependent stages, parallelisation for independent work, routing for specialised paths, and evaluator-optimizer loops when additional quality justifies extra API calls.</p>\n<h2 id=\"how-to-evaluate-prompt-quality\">How to Evaluate Prompt Quality</h2>\n<p>Use representative test cases:</p>\n<pre><code>test_cases = [\n    {\n        &quot;ticket&quot;: &quot;I want a refund for broken headphones.&quot;,\n        &quot;expected&quot;: &quot;refund&quot;,\n    },\n    {\n        &quot;ticket&quot;: &quot;Where is ORD-1002?&quot;,\n        &quot;expected&quot;: &quot;delivery&quot;,\n    },\n    {\n        &quot;ticket&quot;: &quot;Do you sell gift cards?&quot;,\n        &quot;expected&quot;: &quot;general&quot;,\n    },\n]</code></pre>\n<p>These cases cover different request types. Run the same cases after changing the system prompt, examples, model, token limit, or routing instructions:</p>\n<pre><code>def evaluate(route_fn, cases):\n    passed = 0\n    for case in cases:\n        result = route_fn(case[&quot;ticket&quot;])\n        if result == case[&quot;expected&quot;]:\n            passed += 1\n        else:\n            print(&quot;Failed:&quot;, case[&quot;ticket&quot;], result)\n    score = passed / len(cases)\n    print(f&quot;{passed}/{len(cases)} passed&quot;)\n    return score</code></pre>\n<p>Use code-based graders for labels and JSON. Use human or model-based graders for tone, accuracy, and helpfulness.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>Building with the Claude API involves more than writing prompts. A reliable application needs structured context, managed conversation state, validated tool execution, deliberate response handling, suitable workflows, and repeatable evaluation.</p>\n<p>The goal isn&#39;t to find one perfect prompt. It&#39;s to build a system around Claude that provides the right context, limits unsafe actions, handles uncertainty, and measures whether changes improve the result.</p>","headings":[{"level":2,"text":"Table of Contents","id":"table-of-contents"},{"level":2,"text":"Prerequisites","id":"prerequisites"},{"level":2,"text":"How to Set Up the Project and Keep Your API Key Secure","id":"how-to-set-up-the-project-and-keep-your-api-key-secure"},{"level":2,"text":"How to Make Your First Request","id":"how-to-make-your-first-request"},{"level":2,"text":"How to Manage Conversation History","id":"how-to-manage-conversation-history"},{"level":3,"text":"How to Manage History as it Grows","id":"how-to-manage-history-as-it-grows"},{"level":2,"text":"How to Structure Prompts with Clear Boundaries","id":"how-to-structure-prompts-with-clear-boundaries"},{"level":2,"text":"How to Use a System Prompt","id":"how-to-use-a-system-prompt"},{"level":2,"text":"How to Add Tools","id":"how-to-add-tools"},{"level":2,"text":"How to Handle a Tool-Use Response","id":"how-to-handle-a-tool-use-response"},{"level":2,"text":"Claude Responses Can Contain Multiple Blocks","id":"claude-responses-can-contain-multiple-blocks"},{"level":2,"text":"Workflows vs Agents","id":"workflows-vs-agents"},{"level":2,"text":"Chaining, Parallelisation, Routing, and Evaluator-Optimizer","id":"chaining-parallelisation-routing-and-evaluator-optimizer"},{"level":2,"text":"How to Evaluate Prompt Quality","id":"how-to-evaluate-prompt-quality"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}