{"article":{"slug":"building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows","title":"Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows","subtitle":null,"summary":"Ambient agents respond to events such as an Amazon S3 upload, a schedule, or an alert instead of waiting for a chat prompt. This post walks through building framework-agnostic ambient agents on Amazon Bedrock AgentCore using Amazon SQS, AWS Lambda, and Amazon DynamoDB, with a single ask_human tool and a Jobs page for human-in-the-loop review.","content_type":"tutorial","language":"en","canonical_url":"https://aws.amazon.com/blogs/machine-learning/building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows/","author":{"name":"Juan Albarran, Andy Widjaja, Kenton Blacutt, Omar Hamden","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Amazon Web Services","url":"https://aws.amazon.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"}],"about_listings":[],"cover_image_url":"https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/25/ML-19851-featured-image.png","license":"all-rights-reserved","word_count":5504,"reading_minutes":24,"published_at":"2026-10-01T16:40:24.000Z","added_at":"2026-10-01T18:10:11.404Z","updated_at":"2026-10-01T18:10:11.404Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows","markdown_url":"https://listedarticles.com/articles/building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows.md","example":false,"citation":"Juan Albarran, Andy Widjaja, Kenton Blacutt, Omar Hamden, Amazon Web Services. \"Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows.\" 1 Oct 2026. https://aws.amazon.com/blogs/machine-learning/building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://aws.amazon.com/blogs/machine-learning/building-ambient-agents-with-amazon-bedrock-agentcore-from-event-driven-signals-to-human-in-the-loop-workflows/"},"body_markdown":"## Artificial Intelligence\n\n# Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows\n\nTeams that process documents at scale know the routine: files land in storage, someone notices, opens each one, decides what it needs, and routes it for review. Monitoring alerts queue up the same way, waiting for a person to act on them. The hours lost to manual triage are the operational problem ambient agents solve. Imagine a document lands in your Amazon Simple Storage Service (Amazon S3) bucket and within seconds a job appears on your Jobs page, ready to run (or already running if you configured it that way). The agent analyzes the file, surfaces the findings, and asks you for approval before taking the next step. The event itself is the prompt. That is an ambient agent: it responds to event streams, pauses for human input through a single `ask_human` tool when it needs to, and resumes from where it left off once the human answers.\n\nMost AI agent experiences today follow a different pattern: a user opens a chat interface, types a prompt, and waits for a response. That works for one-time questions, but it limits the agent to one conversation at a time and requires a human to describe what happened before anything can act on it. For scenarios where agents should react to events happening across your infrastructure (file uploads, database changes, scheduled tasks, system alerts), that chat-only model breaks down.\n\nAmbient agents describe a different paradigm, one that LangChain among others has articulated. Instead of waiting for users to initiate conversations, ambient agents listen to an event stream and act on it, potentially handling many events in parallel. They aren’t solely triggered by human messages, and multiple agents can run simultaneously. Crucially, they aren’t fully autonomous: a production design pays careful attention to when the agent pauses to interact with humans. When a signal fires, the agent executes its workflow and only interrupts a human when clarification, approval, or review is needed. This human-in-the-loop component lowers the stakes for deploying agents to production, builds user trust, and lets agents learn and improve over time through feedback.\n\nOrganizations running on AWS already have the event-driven infrastructure in place: Amazon S3 event notifications, Amazon EventBridge rules, AWS Lambda triggers, and Amazon DynamoDB streams. The missing piece is connecting those event sources to intelligent agents that can reason about what happened, act, and loop in humans when the situation calls for it. Fully automated pipelines like AWS Step Functions can orchestrate workflows but can’t reason through ambiguity or ask clarifying questions. Chat-based agents can reason but require someone to start the conversation. Ambient agents bridge this gap.\n\nAmazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. AgentCore Runtime provides the execution environment that makes this pattern work: container-based agent hosting with support for long-running workloads, built-in session isolation, and integration with Amazon Bedrock foundation models. AgentCore Runtime supports sessions long enough to cover the signal → agent → human-in-the-loop (HITL) flow shown here. The reference implementation caps each agent turn at the Lambda 15-minute timeout, which is more than enough headroom in practice. Combined with AWS Lambda for event processing and Amazon DynamoDB for state management, the result is a fully serverless ambient-agent platform.\n\nIn this post we walk through the pattern end-to-end on Amazon Bedrock AgentCore. You will come away understanding:\n\n- How an Amazon S3 or scheduled event becomes a job that an agent runs on AgentCore Runtime, with or without a human in the loop.\n- How a single `ask_human` tool plus a canonical response envelope is enough to support the full range of human-in-the-loop interactions.\n- What you get from the reference sample, and what you write on top for your own use case.\n\n## Prerequisites\n\nBefore deploying the reference implementation, make sure you have the following in place:\n\n- An AWS account with permissions to create AWS Identity and Access Management (IAM) roles, Lambda functions, DynamoDB tables, S3 buckets, Amazon Simple Queue Service (Amazon SQS) queues, Amazon API Gateway APIs, Amazon CloudFront distributions, Amazon Cognito user pools, Amazon Elastic Container Registry (Amazon ECR) repositories, and Bedrock AgentCore runtimes. Administrator access on a sandbox account is a good starting point.\n- The AWS Command Line Interface (AWS CLI) configured with credentials for that account and a default AWS Region of `us-east-1` (the sample defaults are wired up for that Region).\n- The AWS Cloud Development Kit (AWS CDK) v2 installed and bootstrapped in your account and Region (`cdk bootstrap` ).\n- Docker installed and running locally. The agent container is built and pushed to Amazon ECR as part of deployment.\n- Python 3.11 or later for the backend Lambda functions and the agent build, and `Node.js` 18 or later for the React frontend.\n- Access to the Anthropic Claude Sonnet 4.5 model in Amazon Bedrock in your target Region. If you haven’t used Bedrock before, follow Manage access to Amazon Bedrock foundation models (FMs) to enable the model. Model availability varies by AWS Region. Check the Amazon Bedrock documentation for the current list of models supported in your target Region. Switching models is a one-line config change later.\n\n## Understanding ambient agents\n\nBefore we get into the architecture, it helps to look at what makes an ambient agent different from a typical chatbot and at the building blocks the rest of the post relies on: the event-driven trigger model, the ambient signal abstraction, and the single human-in-the-loop tool that ties them together.\n\n### Event-driven compared to user-initiated agents\n\nUser-initiated agents follow a request-response pattern:\n\nAmbient agents follow an event-driven pattern:\n\nThe key difference is the trigger mechanism. Ambient agents are activated by system events rather than explicit user requests, which makes them a natural fit for document-processing pipelines, monitoring and alerting, scheduled analysis, and multi-step workflows that need approval gates along the way.\n\n### Ambient signals: The trigger mechanism\n\nAn ambient signal is a configuration that maps an event source to an agent. When the event occurs, the platform automatically creates a job for the agent. What happens next depends on one setting on the signal:\n\n- With `autoExecute: false` (the default), the job lands on the Jobs page in`idle` status and waits for a human to review and run it. This is the safe, review-first flow you want when a signal could fire on unknown input or when the agent has high-stakes tools available.\n- With `autoExecute: true` , the signal processor enqueues the job straight onto the worker queue, the agent runs immediately, and a human is only pulled in if the agent itself calls`ask_human` . This is the fully autonomous flow.\n\nThe pattern covers several signal event sources. The reference sample ships the first two. The rest are extension points you add by writing a new handler Lambda function and a corresponding form field on the Signals page:\n\n- **Amazon S3 file uploads** (ships): Trigger when files are uploaded to specific buckets and prefixes.\n- **Scheduled events** (ships): Trigger agents on a cron-like schedule. Driven by jobs carrying`jobType: \"scheduled\"` rather than by a signal on the Signals page.\n- **API webhooks** (extension point): Respond to external system notifications.\n- **Database changes** (extension point): React to Amazon DynamoDB streams or Amazon Relational Database Service (Amazon RDS) events.\n\n### Human-in-the-loop: One tool, one envelope, one view\n\nAmbient agents need structured ways to interact with humans. In this sample the agent surfaces those interactions through a single tool (`ask_human`) and returns a canonical response envelope. In that envelope, `status` is one of `completed`, `interrupted`, or `error`, and the matching field is `result`, `question`, or `error`. The platform additionally threads `session_id` and `job_id` through every response so continuation turns can be correlated. Those are correlation metadata, not part of the core contract your agent must implement. When the agent returns `interrupted`, the platform moves the job into `interrupted` status and sets its `requiresAction` flag to `true`. The reference React frontend surfaces these on the **Interrupted** tab of the Jobs page with a warning indicator on each row, so there is no separate review queue to poll. The same Jobs view shows pending questions, proposed actions awaiting approval, final results, and failed jobs, giving a user one place to see everything their agents are doing instead of monitoring multiple chat windows or email threads.\n\nThe same mechanism supports several prompting patterns that a reader may recognize from the wider agents literature: a Notify turn where the agent simply reports a result, a Question turn where it asks for clarification, a Review turn where it proposes an action and waits for `APPROVE` / `REJECT` / `MODIFY`, and an Error turn where the failure is captured on the job record and the user decides whether to retry. These are conventions for how the agent writes its question, not separate runtime modes. At the platform level there is exactly one code path and exactly one envelope.\n\n## Architecture overview\n\nThe platform is a small set of serverless components stitched together by an event pipeline. This section walks through the end-to-end flow first, then describes each component in turn.\n\nEvents flow through the platform end-to-end as follows. Amazon S3 emits an `s3:ObjectCreated` notification, which a Signal Processor Lambda function receives. The Signal Processor queries a global secondary index (GSI) on the ambient-signals table to find any matching signal for the event bucket, then creates a job record for each match. The API tier (or the scheduler) enqueues the job onto an Amazon SQS queue. The same Job Execution Lambda function that serves the API path also drains the queue through an attached SQS event source, invokes the agent on Amazon Bedrock AgentCore Runtime, and writes results (and any human-input requests) back to Amazon DynamoDB. A React frontend served from Amazon S3 through Amazon CloudFront polls a small Amazon API Gateway and Lambda tier for updates and lets the user respond to pending interactions.\n\nThe major components are:\n\n- **Amazon S3** with event notifications serves as the entry point for signals when files are uploaded. Prefix and suffix filters are pushed down into the bucket’s notification configuration so the signal processor is only invoked for events that could plausibly match a signal.\n- **Amazon SQS** decouples the API Gateway request from the agent call. A job-execution queue holds pending work. A dead-letter queue (DLQ) captures messages the worker can’t process after the configured number of retries.\n- **Three pipeline Lambda functions** carry the event from intake to agent (a separate management tier behind API Gateway is described later in this section):\n  - *Signal Processor* matches incoming events to configured signal definitions and creates jobs.\n  - *Job Execution* has two entry paths in one Lambda function: an API handler that enqueues messages, and an SQS worker that consumes them and invokes AgentCore Runtime with the job context.\n  - *Scheduler* fires on a one-minute cron and enqueues due scheduled jobs onto the same SQS queue.\n- **Amazon Bedrock AgentCore Runtime** runs agent code in isolated containers and supports long-running workloads.\n- **Amazon DynamoDB** stores the agent registry, job records, ambient signal definitions, chat threads, conversation history, and Powertools idempotency records. Conversation messages are appended atomically with`UpdateItem` +`list_append` so concurrent writers do not clobber each other.\n- **Five management-tier Lambda functions behind Amazon API Gateway** expose the REST API the frontend consumes (`agent_management` ,`job_management` ,`signal_management` ,`chat_management` ,`conversation_management` ), plus a`chat_execution` worker Lambda that`chat_management` invokes asynchronously so chat API calls return immediately. See Production deployment for how they are provisioned.\n- **A React frontend** served from Amazon S3 through Amazon CloudFront provides the Agent Management UI where users monitor jobs, chat with agents, respond to questions, and review pending actions.\n\n## Building the event infrastructure\n\nWith the architecture in mind, the next step is wiring up the event source that turns an Amazon S3 upload into an agent job. This section creates the bucket, points its event notifications at the Signal Processor Lambda function, and walks through how the processor matches events against configured signals.\n\n### Setting up the Amazon S3 signal trigger\n\nCreate the Amazon S3 bucket and configure event notifications that drive the Signal Processor:\n\nThe bucket is configured to send `s3:ObjectCreated:*` events directly to the Signal Processor Lambda function. In this sample the notification configuration is installed dynamically by the `signal_management` Lambda function when a signal is created or updated, so adding a new signal for a new prefix doesn’t require a redeploy.\n\n### Signal Processor Lambda function\n\nThe Signal Processor receives Amazon S3 events, finds matching signal definitions in DynamoDB, and creates a job per match. At its heart the handler is the shape shown in the following example. The real handler in `backend/functions/multi_agent/signal_processor.py` also uses AWS Lambda Powertools for structured logging and idempotency, queries the `bucketName-signalId-index` GSI on the signals table, applies the prefix and suffix checks configured on each signal, and writes a `signal_triggered` job row.\n\n### Signal configuration data model\n\nSignals are stored in DynamoDB with this structure:\n\nYou only set `configuration.bucketName` when creating a signal through the API. The top-level `bucketName` shown in the preceding example is populated by the platform. DynamoDB GSI partition keys can’t be nested inside a map attribute, so `signal_management` mirrors `configuration.bucketName` out to a top-level `bucketName` on every write so the `bucketName-signalId-index` GSI can fan out to matching signals on every Amazon S3 event.\n\nThe `autoExecute` flag is the single switch that decides whether the agent fires autonomously or a human reviews the job first. The Signals form in the Agent Management UI exposes it as a checkbox alongside the familiar `enabled` setting, so changing the behavior is a quick edit without touching code or the database.\n\n## Deploying agents on AgentCore Runtime\n\nAmazon Bedrock AgentCore Runtime hosts the agent container and exposes it through an `InvokeAgentRuntime` API that the worker Lambda function calls on every job. This section walks through the agent layout that ships with the sample, the configuration that drives it, and the small orchestrator that turns LangGraph tool calls into the platform’s response envelope.\n\n### Agent packaging and structure\n\nAgentCore Runtime provides a container-based execution environment for agents, so you can bring any Python agent framework. The sample uses a modular design:\n\n### Agent configuration\n\n`config.yaml` defines agent behavior, tools, and the system prompt. The sample defaults to Anthropic Claude Sonnet 4.5 on Amazon Bedrock for this walkthrough, a fit for the multi-step tool calling and long-context reasoning the human-in-the-loop workflow relies on. Switching to Claude Haiku, Amazon Nova, or another tool-calling model available on Amazon Bedrock (availability varies by Region) is a one-line change to `model_id` in the following configuration. `max_iterations: 10` gives the graph enough headroom for around ten model-to-tool round trips before it halts (the orchestrator doubles the value to compute LangGraph’s recursion limit, since each round trip traverses two graph nodes), which covers the multi-step tool use the sample’s S3, calculator, and `ask_human` tools expect. The file also contains an `execution:` block (loop-detector, circuit-breaker, session-cache thresholds) omitted here for brevity. See `agent/config.example.yaml` for the full file.\n\n### Core agent implementation\n\nThe agent is built on `langchain.agents.create_agent`, a tool-calling agent compiled as a LangGraph. LangChain handles the orchestration layer because it brings pre-built tool-calling patterns, a broad open ecosystem of integrations, and APIs many teams already know, while AgentCore Runtime supplies the managed hosting, session isolation, and scaling underneath. The two layers are complementary. The platform wrapper around the graph does three things: it builds the message list for the turn (including any conversation history and, for signal-triggered jobs, the Amazon S3 bucket and key that fired the signal so the model can pick up the file without being told), it invokes the graph, and it scans the tool output for the `ask_human` sentinel so a tool call can be turned into an `interrupted` response.\n\nIn its simplest form (the full version in `agent/core/agent_core.py` also handles per-session history caching, loop detection, and an execution trace), the orchestrator is this:\n\n### Deploying the agent\n\nDeploy the agent to AgentCore Runtime using the provided script:\n\nThe script returns an Agent Runtime Amazon Resource Name (ARN) which the backend stores in the agent registry so the Job Execution Lambda function can invoke it.\n\n## State management with DynamoDB\n\nAmazon DynamoDB is the system of record for everything that needs to outlive a single Lambda invocation: which agents are registered, which jobs are in flight, the conversation history that gives an agent continuity across turns, and the signal definitions and chat threads the UI reads. The next subsections describe the table layout, the session model, and how the Job Execution Lambda function uses both to drive a job to completion.\n\n### Database design\n\nThe state layer is backed by DynamoDB. The core tables exercised in this post are:\n\n**Agent registry table** (one record per registered agent runtime):\n\n**Job registry table** (one record per agent invocation). With the default `autoExecute: false`, a new signal-triggered job lands here in `idle` status, waiting for a user to choose Execute:\n\nWhile a job is running the worker flips `status` to `busy`. If the agent pauses by calling `ask_human`, `status` becomes `interrupted` and `requiresAction` becomes `true`, which surfaces the job on the Jobs page’s Interrupted tab.\n\n**Conversation store table** (full message history per session, with a 30-day time to live (TTL)):\n\nAdditional tables exist for ambient signals, chat threads, and Powertools idempotency records. Scheduled execution reuses the job-registry table through a `jobType` and `nextRun` GSI rather than having its own table.\n\n### Session management\n\nSessions provide conversation continuity across job executions and chat turns. The `conversation_management` Lambda function owns DynamoDB persistence: every turn (human + AI pair) is appended to the session with an atomic `UpdateItem` + `list_append` so two concurrent writers on the same session can’t clobber each other, and a 30-day TTL takes care of cleanup.\n\n### Job execution with conversation continuity\n\nThe Job Execution Lambda function has two entry paths. The API path sends an SQS message (adding Powertools idempotency and a `userId` ownership check) and returns 202 Accepted immediately, so the frontend never waits for the model. The SQS worker path is where the real work happens: it loads conversation history, folds any human response into a continuation prompt, calls AgentCore Runtime, and writes the result back. That worker looks roughly like this:\n\n## Implementing human-in-the-loop patterns\n\nThe agent snippet earlier returned a sentinel-based `interrupted` envelope when the graph produced an `ask_human` tool call. This section zooms in on the `ask_human` tool itself and shows exactly how the orchestrator catches that sentinel without breaking the reasoning loop.\n\n### The human input tool\n\nThe core of human-in-the-loop functionality is the `ask_human` tool. It doesn’t raise an exception: because LangGraph’s tool node captures tool exceptions as error observations and feeds them back to the model, the tool instead returns a sentinel string. The orchestrator detects the sentinel on the `ToolMessage` stream after the graph finishes and converts it into an `interrupted` response. Per-invocation state (metrics, session ID, and so on) isn’t a tool argument. The orchestrator binds it on a `ContextVar` that tools read through `current_execution_state.get()`, which keeps concurrent invocations inside the same container isolated from each other.\n\n### Agent-side usage\n\nThe agent uses the tool naturally as part of its tool-calling loop. For example, when analyzing an invoice with multiple line items the agent may call:\n\nThe tool returns the sentinel, the orchestrator stops the graph, persists the question as the latest AI turn in the conversation store, and surfaces an `interrupted` job to the UI. When the user replies, the Job Execution Lambda function invokes the agent again with the answer folded into a continuation prompt.\n\n## Building the Agent Management UI\n\nThe Agent Management UI is the human side of the platform: where users browse jobs, answer pending questions, and chat with agents directly. It’s a React single-page application that talks to the same REST API the rest of the post has been describing. The next two subsections cover the layout and the user experience flow that ties signals, jobs, and human input together.\n\n### Frontend architecture\n\nThe frontend is a React and Cloudscape Design single-page application served from Amazon S3 behind Amazon CloudFront. It exposes a five-tab navigation:\n\n- **Workflows** : A gallery and CRUD surface for workflow-level definitions that group agents and signals.\n- **Chat** : A standalone chat page (`/chat` ,`/chat/:threadId` ) for user-initiated conversations with a registered agent.\n- **Agents** : Manages registered agent runtimes (name, ARN, capabilities, status).\n- **Jobs** : Lists all jobs with status filters and a detail view that includes an interactive Chat tab, an execution trace, and metadata.\n- **Signals** : Defines and toggles ambient signals. The shipped form covers Amazon S3 prefix and suffix filters and schedules. Webhook, Amazon EventBridge, and DynamoDB stream fields are added when you wire those extension points.\n\nThe same chat component is reused by both the standalone Chat page and the Jobs detail Chat tab. While a job is running (`status === \"busy\"`), the panel polls `/conversations/:sessionId` on a short interval so turns from the agent or from a second browser tab show up within a few seconds. After the job settles, polling stops. Because it’s the same component in both places, a user can freely converse with an agent from the Jobs detail view while a job is active. They are not limited to answering a single pending question.\n\n### User experience flow\n\n- **Signal fires** : A new job appears in the Jobs list with`jobType: \"signal_triggered\"` . If the signal has`autoExecute: false` (the default) the job lands in`idle` and the user runs it manually. With`autoExecute: true` the worker is already drafting a response by the time the list refreshes. User-initiated jobs carry`jobType: \"user_initiated\"` and render with a different badge so you can tell the two apart at a glance.\n- **Human input requested** : Any job that calls`ask_human` moves to the Interrupted tab and its`requiresAction` flag flips to`true` , which renders a warning indicator on the row.\n- **Responding** : The detail view surfaces a Provide Response button so the user can answer the pending question without leaving the job context.\n- **Execution feedback** : Status transitions (`idle` →`busy` →`completed` |`interrupted` |`error` ) are polled and reflected in the list and in the Chat tab header.\n- **Interactive chat** : The Jobs detail Chat tab and the`/chat` page both render full conversation history and accept new user messages, so humans can answer pending questions or nudge the agent with additional context without leaving the UI.\n- **Response submission** : Sending a message calls the Chat Execution or Job Execution Lambda function, which invokes the agent on AgentCore Runtime with the job context and the human response folded into the next turn.\n- **Completion and audit** : The final result is stored on the job record and the full conversation history is preserved on the session, available in the Chat tab for audit and re-use.\n\nAuthentication is handled by Amazon Cognito and an API Gateway `CognitoUserPoolsAuthorizer` (see `backend/infrastructure/multi_agent_stack.py`).\n\nReal-time updates are implemented as lightweight polling of the conversation and job endpoints, so there is no WebSocket infrastructure to run.\n\n## Extending the sample\n\nThe reference implementation is deliberately small so you can see the contract before you start adding to it. This section describes what ships in the box, what you build on top, and where to plug in new event sources or tools.\n\n### What ships in the sample compared to what you build\n\nBefore extending, it helps to know where the seams are. This sample is a reference implementation of the ambient-agent pattern, not a turnkey product. What ships out of the box:\n\n- **The platform** : Signal intake, the SQS-backed job pipeline, the HITL interrupt mechanism, job and conversation state in DynamoDB, the React UI with the Jobs page and Chat surfaces, and AgentCore integration with Cognito-backed auth.\n- **A reference agent** : A containerized LangChain agent with four tools (calculator,`ask_human` ,`list_s3_files` ,`read_s3_file` ) and a generic system prompt. Useful for validating the pipeline end-to-end, not for solving your business problem out of the gate.\n- **One signal source wired through the Signals UI** (Amazon S3 file uploads, matched by bucket, prefix, and suffix), plus a job-level scheduler that runs on a one-minute cron for jobs created with`jobType: \"scheduled\"` . Adding new signal types such as webhooks, Amazon EventBridge events, or DynamoDB streams requires extending the`signalType` enum in both the backend and the Signals page form.\n\nWhat you build on top, and what the platform expects from each:\n\n- **Your agent** : A containerized Python agent tailored to your domain, with its own tools and its own system prompt. It must speak the two conventions the platform relies on: return the three-status envelope (`completed` ,`interrupted` , or`error` with the matching`result` ,`question` , or`error` field) and use the`ask_human` sentinel to request human input. That is roughly 50 lines of adapter code around whatever agent framework you prefer.\n- **New signal types** if Amazon S3 uploads and cron are not enough: A new handler Lambda function that queries the signals table and writes a`signal_triggered` job row (shown in the following template), plus a corresponding form field on the Signals page so users can configure it.\n- **New tools** for your agent: A new module under`agent/tools/` , a registration entry in`agent/core/tool_factory.py` , and an enable flag in`config.yaml` . The orchestrator and HITL flow do not change. Any tool that returns a string participates in the same graph.\n\nIn short: the platform does the plumbing, you bring the brain. The three places you will write code (agent logic, tool implementations, and new signal handlers) all slot into the existing contract without touching the rest of the stack.\n\n### Configuration over code\n\nEnabling a new capability for an existing agent is a config change rather than a code change. The agent’s `config.yaml` toggles tools on or off and supplies their descriptions, so adding a new Amazon S3 prefix handler or a new calculator mode doesn’t require rebuilding the container. Signal definitions live in DynamoDB and are edited from the Signals page in the UI.\n\n### Adding a new signal type\n\nSignals today trigger on Amazon S3 events and on cron schedules. Adding a new type (a webhook, an Amazon EventBridge rule, a DynamoDB stream, a Kafka topic) follows the same three-step contract. The downstream pipeline is signal-agnostic: everything from the SQS worker through the AgentCore call through the UI treats all `signal_triggered` jobs the same, so you only need to write the bridge from your event source into a job row.\n\nYou add the Lambda function and its API Gateway route in the CDK stack, extend `signalType` in the frontend’s TypeScript union and the Signals page form, and the rest (the Jobs page UI, the `ask_human` interrupt handling, the conversation persistence, the `autoExecute` toggle, the contract tests) all apply without change. Like the production Signal Processor, a real webhook handler would add AWS Lambda Powertools idempotency so a retried delivery doesn’t create duplicate jobs.\n\n## Production deployment\n\nAfter the pieces fit together locally, the next step is provisioning them in an account. The platform deploys as a single AWS CDK stack. The following subsections describe what that stack creates, the deploy commands, the metrics that matter when it is running, and the security and scaling defaults the sample ships with.\n\n### Infrastructure as code\n\nThe whole platform is an AWS CDK stack defined in `backend/infrastructure/multi_agent_stack.py`. For storage and messaging, it provisions the DynamoDB tables (agent registry, job registry, conversation store, ambient signals, chat threads, idempotency records) with pay-per-request billing and point-in-time recovery, plus the SQS job-execution queue and its dead-letter queue.\n\nFor compute and delivery, it provisions the three pipeline Lambda functions from the Architecture section (Signal Processor, Job Execution, Scheduler), the five management-tier Lambda functions that back the REST API (`agent_management`, `job_management`, `signal_management`, `chat_management`, `conversation_management`) and the `chat_execution` worker Lambda that `chat_management` invokes asynchronously, a Cognito-backed REST API, and the CloudFront distribution that serves the React frontend.\n\nIAM grants flow directly from the code: DynamoDB `grant_read_write_data` on each table, SQS `grant_send_messages` and `grant_consume_messages` as appropriate, and a scoped `bedrock-agentcore:InvokeAgentRuntime` policy on the worker role.\n\n### Deployment steps\n\n### Monitoring and observability\n\nThe Lambda functions emit custom CloudWatch metrics under the `AmbientAgents` namespace for the events that matter in this pattern: signal matches, jobs enqueued, jobs completed, jobs interrupted, and jobs errored. Combined with the default Lambda and SQS metrics (invocation count, errors, queue depth, DLQ depth), these give you a dashboard where the signal → job → agent → human cycle is visible end-to-end. Set an alarm on DLQ depth as the first line of defense against silent failures.\n\n### Security best practices\n\n- **IAM least privilege** : Grant only the permissions required by each Lambda function (reads on its own table, targeted`bedrock-agentcore:InvokeAgentRuntime` , targeted Amazon S3 prefixes).\n- **Encryption at rest** : All DynamoDB tables use AWS-managed keys. Amazon S3 buckets use S3-managed encryption. CloudFront logs live in a dedicated logging bucket.\n- **Encryption in transit** : All API Gateway endpoints and CloudFront distributions require HTTPS. Amazon S3 bucket policies enforce SSL.\n- **Audit logging** : AWS CloudTrail captures API calls across the stack. DynamoDB streams can be enabled on the job table for richer job-lifecycle audit.\n- **Responsible AI** : Apply Amazon Bedrock Guardrails to agent inputs and outputs so content filters, denied topics, and contextual grounding checks run before findings surface on the Jobs page or reach a reviewer through the`ask_human` prompt. This protects the human-in-the-loop interaction and keeps agent outputs grounded in the source documents.\n\n### Scaling strategies\n\n- **Lambda concurrency** : Reserve concurrency for the Job Execution worker Lambda function to cap downstream Bedrock invocations at a predictable per-account ceiling.\n- **DynamoDB capacity** : Pay-per-request is the default. Switch to provisioned capacity with auto scaling when traffic patterns stabilize.\n- **Cost optimization** : Use DynamoDB TTL to retire old conversations automatically, Amazon S3 lifecycle policies to age out processed documents, and CloudWatch metric filters to track per-agent model invocation cost.\n\n## Clean up\n\nTo avoid ongoing charges after you are finished evaluating the sample, tear the stack back down in the reverse order it was created.\n\nFirst, delete the agent runtime from Amazon Bedrock AgentCore so the container stops being billed:\n\nNext, empty the Amazon S3 buckets the stack provisions (the documents bucket, the frontend hosting bucket, and the CloudFront logging bucket) so CloudFormation can delete them:\n\nThen destroy the AWS CDK stack itself, which removes the Lambda functions, API Gateway, CloudFront distribution, DynamoDB tables, SQS queues, Cognito user pool, and IAM roles:\n\nFinally, delete the Amazon ECR repository created by the agent deploy script if you do not plan to redeploy:\n\nIf you enabled Amazon Bedrock model access only for this walkthrough and no longer need it, you can revoke it from the **Model access** page of the Amazon Bedrock console.\n\n## Summary\n\nAmbient agents shift AI automation from “wait for a user” to “respond to signals”. Time-to-action drops from hours to seconds, and around-the-clock monitoring becomes possible without constant human attention. Each agent operates independently with isolated sessions, so the platform handles many concurrent events in parallel without coordination overhead. Humans stay in the loop only when their input is truly needed. The unified Jobs page removes context switching across multiple tools, and full conversation history means you never lose track of what an agent did or why it paused.\n\n### When to use ambient patterns\n\nThe following use cases describe the ambient-agent pattern in general. The reference sample ships the plumbing for Amazon S3 and the scheduler today. Webhooks, database changes, and external integrations are extension points you wire up on top.\n\nThe pattern is a strong fit whenever an event in your environment should drive work that a model can reason about and a human should occasionally weigh in on. Document-processing pipelines are the canonical example: a file lands, an agent reads it, summarizes it, and asks for approval before acting. Monitoring and alerting flows benefit from the same shape, with the agent triaging an alert and asking a human only when escalation is needed. Scheduled processing, analysis, and reporting fit naturally on the cron path, and any multi-step workflow that needs an approval gate or a human checkpoint along the way maps cleanly onto the `ask_human` interrupt.\n\nThere are also workloads where ambient agents are the wrong tool. Real-time chat applications belong in a traditional chat agent that is purpose-built for low-latency conversational turn-taking. Simple request-response APIs do not need an agent at all. A plain API endpoint is faster, cheaper, and easier to reason about. And for purely deterministic workflows that do not need large language model (LLM) reasoning, AWS Step Functions is still the right choice. Note that ambient agents with `autoExecute: true` already cover the fully autonomous case if you want LLM reasoning without a human gate.\n\n### Next steps\n\nClone the reference implementation and follow the top-level README for deployment and configuration.\n\nFurther reading:\n\nStart small. Pick a single use case (document processing or a monitoring alert) and expand from there. Clone the sample, point it at your own event source, and see how much routine work the Jobs page takes off your plate.","body_html":"<h2 id=\"artificial-intelligence\">Artificial Intelligence</h2>\n<h1 id=\"building-ambient-agents-with-amazon-bedrock-agentcore-from-event\">Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows</h1>\n<p>Teams that process documents at scale know the routine: files land in storage, someone notices, opens each one, decides what it needs, and routes it for review. Monitoring alerts queue up the same way, waiting for a person to act on them. The hours lost to manual triage are the operational problem ambient agents solve. Imagine a document lands in your Amazon Simple Storage Service (Amazon S3) bucket and within seconds a job appears on your Jobs page, ready to run (or already running if you configured it that way). The agent analyzes the file, surfaces the findings, and asks you for approval before taking the next step. The event itself is the prompt. That is an ambient agent: it responds to event streams, pauses for human input through a single <code>ask_human</code> tool when it needs to, and resumes from where it left off once the human answers.</p>\n<p>Most AI agent experiences today follow a different pattern: a user opens a chat interface, types a prompt, and waits for a response. That works for one-time questions, but it limits the agent to one conversation at a time and requires a human to describe what happened before anything can act on it. For scenarios where agents should react to events happening across your infrastructure (file uploads, database changes, scheduled tasks, system alerts), that chat-only model breaks down.</p>\n<p>Ambient agents describe a different paradigm, one that LangChain among others has articulated. Instead of waiting for users to initiate conversations, ambient agents listen to an event stream and act on it, potentially handling many events in parallel. They aren’t solely triggered by human messages, and multiple agents can run simultaneously. Crucially, they aren’t fully autonomous: a production design pays careful attention to when the agent pauses to interact with humans. When a signal fires, the agent executes its workflow and only interrupts a human when clarification, approval, or review is needed. This human-in-the-loop component lowers the stakes for deploying agents to production, builds user trust, and lets agents learn and improve over time through feedback.</p>\n<p>Organizations running on AWS already have the event-driven infrastructure in place: Amazon S3 event notifications, Amazon EventBridge rules, AWS Lambda triggers, and Amazon DynamoDB streams. The missing piece is connecting those event sources to intelligent agents that can reason about what happened, act, and loop in humans when the situation calls for it. Fully automated pipelines like AWS Step Functions can orchestrate workflows but can’t reason through ambiguity or ask clarifying questions. Chat-based agents can reason but require someone to start the conversation. Ambient agents bridge this gap.</p>\n<p>Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. AgentCore Runtime provides the execution environment that makes this pattern work: container-based agent hosting with support for long-running workloads, built-in session isolation, and integration with Amazon Bedrock foundation models. AgentCore Runtime supports sessions long enough to cover the signal → agent → human-in-the-loop (HITL) flow shown here. The reference implementation caps each agent turn at the Lambda 15-minute timeout, which is more than enough headroom in practice. Combined with AWS Lambda for event processing and Amazon DynamoDB for state management, the result is a fully serverless ambient-agent platform.</p>\n<p>In this post we walk through the pattern end-to-end on Amazon Bedrock AgentCore. You will come away understanding:</p>\n<ul><li>How an Amazon S3 or scheduled event becomes a job that an agent runs on AgentCore Runtime, with or without a human in the loop.</li><li>How a single <code>ask_human</code> tool plus a canonical response envelope is enough to support the full range of human-in-the-loop interactions.</li><li>What you get from the reference sample, and what you write on top for your own use case.</li></ul>\n<h2 id=\"prerequisites\">Prerequisites</h2>\n<p>Before deploying the reference implementation, make sure you have the following in place:</p>\n<ul><li>An AWS account with permissions to create AWS Identity and Access Management (IAM) roles, Lambda functions, DynamoDB tables, S3 buckets, Amazon Simple Queue Service (Amazon SQS) queues, Amazon API Gateway APIs, Amazon CloudFront distributions, Amazon Cognito user pools, Amazon Elastic Container Registry (Amazon ECR) repositories, and Bedrock AgentCore runtimes. Administrator access on a sandbox account is a good starting point.</li><li>The AWS Command Line Interface (AWS CLI) configured with credentials for that account and a default AWS Region of <code>us-east-1</code> (the sample defaults are wired up for that Region).</li><li>The AWS Cloud Development Kit (AWS CDK) v2 installed and bootstrapped in your account and Region (<code>cdk bootstrap</code> ).</li><li>Docker installed and running locally. The agent container is built and pushed to Amazon ECR as part of deployment.</li><li>Python 3.11 or later for the backend Lambda functions and the agent build, and <code>Node.js</code> 18 or later for the React frontend.</li><li>Access to the Anthropic Claude Sonnet 4.5 model in Amazon Bedrock in your target Region. If you haven’t used Bedrock before, follow Manage access to Amazon Bedrock foundation models (FMs) to enable the model. Model availability varies by AWS Region. Check the Amazon Bedrock documentation for the current list of models supported in your target Region. Switching models is a one-line config change later.</li></ul>\n<h2 id=\"understanding-ambient-agents\">Understanding ambient agents</h2>\n<p>Before we get into the architecture, it helps to look at what makes an ambient agent different from a typical chatbot and at the building blocks the rest of the post relies on: the event-driven trigger model, the ambient signal abstraction, and the single human-in-the-loop tool that ties them together.</p>\n<h3 id=\"event-driven-compared-to-user-initiated-agents\">Event-driven compared to user-initiated agents</h3>\n<p>User-initiated agents follow a request-response pattern:</p>\n<p>Ambient agents follow an event-driven pattern:</p>\n<p>The key difference is the trigger mechanism. Ambient agents are activated by system events rather than explicit user requests, which makes them a natural fit for document-processing pipelines, monitoring and alerting, scheduled analysis, and multi-step workflows that need approval gates along the way.</p>\n<h3 id=\"ambient-signals-the-trigger-mechanism\">Ambient signals: The trigger mechanism</h3>\n<p>An ambient signal is a configuration that maps an event source to an agent. When the event occurs, the platform automatically creates a job for the agent. What happens next depends on one setting on the signal:</p>\n<ul><li>With <code>autoExecute: false</code> (the default), the job lands on the Jobs page in<code>idle</code> status and waits for a human to review and run it. This is the safe, review-first flow you want when a signal could fire on unknown input or when the agent has high-stakes tools available.</li><li>With <code>autoExecute: true</code> , the signal processor enqueues the job straight onto the worker queue, the agent runs immediately, and a human is only pulled in if the agent itself calls<code>ask_human</code> . This is the fully autonomous flow.</li></ul>\n<p>The pattern covers several signal event sources. The reference sample ships the first two. The rest are extension points you add by writing a new handler Lambda function and a corresponding form field on the Signals page:</p>\n<ul><li><strong>Amazon S3 file uploads</strong> (ships): Trigger when files are uploaded to specific buckets and prefixes.</li><li><strong>Scheduled events</strong> (ships): Trigger agents on a cron-like schedule. Driven by jobs carrying<code>jobType: &quot;scheduled&quot;</code> rather than by a signal on the Signals page.</li><li><strong>API webhooks</strong> (extension point): Respond to external system notifications.</li><li><strong>Database changes</strong> (extension point): React to Amazon DynamoDB streams or Amazon Relational Database Service (Amazon RDS) events.</li></ul>\n<h3 id=\"human-in-the-loop-one-tool-one-envelope-one-view\">Human-in-the-loop: One tool, one envelope, one view</h3>\n<p>Ambient agents need structured ways to interact with humans. In this sample the agent surfaces those interactions through a single tool (<code>ask_human</code>) and returns a canonical response envelope. In that envelope, <code>status</code> is one of <code>completed</code>, <code>interrupted</code>, or <code>error</code>, and the matching field is <code>result</code>, <code>question</code>, or <code>error</code>. The platform additionally threads <code>session_id</code> and <code>job_id</code> through every response so continuation turns can be correlated. Those are correlation metadata, not part of the core contract your agent must implement. When the agent returns <code>interrupted</code>, the platform moves the job into <code>interrupted</code> status and sets its <code>requiresAction</code> flag to <code>true</code>. The reference React frontend surfaces these on the <strong>Interrupted</strong> tab of the Jobs page with a warning indicator on each row, so there is no separate review queue to poll. The same Jobs view shows pending questions, proposed actions awaiting approval, final results, and failed jobs, giving a user one place to see everything their agents are doing instead of monitoring multiple chat windows or email threads.</p>\n<p>The same mechanism supports several prompting patterns that a reader may recognize from the wider agents literature: a Notify turn where the agent simply reports a result, a Question turn where it asks for clarification, a Review turn where it proposes an action and waits for <code>APPROVE</code> / <code>REJECT</code> / <code>MODIFY</code>, and an Error turn where the failure is captured on the job record and the user decides whether to retry. These are conventions for how the agent writes its question, not separate runtime modes. At the platform level there is exactly one code path and exactly one envelope.</p>\n<h2 id=\"architecture-overview\">Architecture overview</h2>\n<p>The platform is a small set of serverless components stitched together by an event pipeline. This section walks through the end-to-end flow first, then describes each component in turn.</p>\n<p>Events flow through the platform end-to-end as follows. Amazon S3 emits an <code>s3:ObjectCreated</code> notification, which a Signal Processor Lambda function receives. The Signal Processor queries a global secondary index (GSI) on the ambient-signals table to find any matching signal for the event bucket, then creates a job record for each match. The API tier (or the scheduler) enqueues the job onto an Amazon SQS queue. The same Job Execution Lambda function that serves the API path also drains the queue through an attached SQS event source, invokes the agent on Amazon Bedrock AgentCore Runtime, and writes results (and any human-input requests) back to Amazon DynamoDB. A React frontend served from Amazon S3 through Amazon CloudFront polls a small Amazon API Gateway and Lambda tier for updates and lets the user respond to pending interactions.</p>\n<p>The major components are:</p>\n<ul><li><strong>Amazon S3</strong> with event notifications serves as the entry point for signals when files are uploaded. Prefix and suffix filters are pushed down into the bucket’s notification configuration so the signal processor is only invoked for events that could plausibly match a signal.</li><li><strong>Amazon SQS</strong> decouples the API Gateway request from the agent call. A job-execution queue holds pending work. A dead-letter queue (DLQ) captures messages the worker can’t process after the configured number of retries.</li><li><strong>Three pipeline Lambda functions</strong> carry the event from intake to agent (a separate management tier behind API Gateway is described later in this section):<ul><li><em>Signal Processor</em> matches incoming events to configured signal definitions and creates jobs.</li><li><em>Job Execution</em> has two entry paths in one Lambda function: an API handler that enqueues messages, and an SQS worker that consumes them and invokes AgentCore Runtime with the job context.</li><li><em>Scheduler</em> fires on a one-minute cron and enqueues due scheduled jobs onto the same SQS queue.</li></ul></li><li><strong>Amazon Bedrock AgentCore Runtime</strong> runs agent code in isolated containers and supports long-running workloads.</li><li><strong>Amazon DynamoDB</strong> stores the agent registry, job records, ambient signal definitions, chat threads, conversation history, and Powertools idempotency records. Conversation messages are appended atomically with<code>UpdateItem</code> +<code>list_append</code> so concurrent writers do not clobber each other.</li><li><strong>Five management-tier Lambda functions behind Amazon API Gateway</strong> expose the REST API the frontend consumes (<code>agent_management</code> ,<code>job_management</code> ,<code>signal_management</code> ,<code>chat_management</code> ,<code>conversation_management</code> ), plus a<code>chat_execution</code> worker Lambda that<code>chat_management</code> invokes asynchronously so chat API calls return immediately. See Production deployment for how they are provisioned.</li><li><strong>A React frontend</strong> served from Amazon S3 through Amazon CloudFront provides the Agent Management UI where users monitor jobs, chat with agents, respond to questions, and review pending actions.</li></ul>\n<h2 id=\"building-the-event-infrastructure\">Building the event infrastructure</h2>\n<p>With the architecture in mind, the next step is wiring up the event source that turns an Amazon S3 upload into an agent job. This section creates the bucket, points its event notifications at the Signal Processor Lambda function, and walks through how the processor matches events against configured signals.</p>\n<h3 id=\"setting-up-the-amazon-s3-signal-trigger\">Setting up the Amazon S3 signal trigger</h3>\n<p>Create the Amazon S3 bucket and configure event notifications that drive the Signal Processor:</p>\n<p>The bucket is configured to send <code>s3:ObjectCreated:*</code> events directly to the Signal Processor Lambda function. In this sample the notification configuration is installed dynamically by the <code>signal_management</code> Lambda function when a signal is created or updated, so adding a new signal for a new prefix doesn’t require a redeploy.</p>\n<h3 id=\"signal-processor-lambda-function\">Signal Processor Lambda function</h3>\n<p>The Signal Processor receives Amazon S3 events, finds matching signal definitions in DynamoDB, and creates a job per match. At its heart the handler is the shape shown in the following example. The real handler in <code>backend/functions/multi_agent/signal_processor.py</code> also uses AWS Lambda Powertools for structured logging and idempotency, queries the <code>bucketName-signalId-index</code> GSI on the signals table, applies the prefix and suffix checks configured on each signal, and writes a <code>signal_triggered</code> job row.</p>\n<h3 id=\"signal-configuration-data-model\">Signal configuration data model</h3>\n<p>Signals are stored in DynamoDB with this structure:</p>\n<p>You only set <code>configuration.bucketName</code> when creating a signal through the API. The top-level <code>bucketName</code> shown in the preceding example is populated by the platform. DynamoDB GSI partition keys can’t be nested inside a map attribute, so <code>signal_management</code> mirrors <code>configuration.bucketName</code> out to a top-level <code>bucketName</code> on every write so the <code>bucketName-signalId-index</code> GSI can fan out to matching signals on every Amazon S3 event.</p>\n<p>The <code>autoExecute</code> flag is the single switch that decides whether the agent fires autonomously or a human reviews the job first. The Signals form in the Agent Management UI exposes it as a checkbox alongside the familiar <code>enabled</code> setting, so changing the behavior is a quick edit without touching code or the database.</p>\n<h2 id=\"deploying-agents-on-agentcore-runtime\">Deploying agents on AgentCore Runtime</h2>\n<p>Amazon Bedrock AgentCore Runtime hosts the agent container and exposes it through an <code>InvokeAgentRuntime</code> API that the worker Lambda function calls on every job. This section walks through the agent layout that ships with the sample, the configuration that drives it, and the small orchestrator that turns LangGraph tool calls into the platform’s response envelope.</p>\n<h3 id=\"agent-packaging-and-structure\">Agent packaging and structure</h3>\n<p>AgentCore Runtime provides a container-based execution environment for agents, so you can bring any Python agent framework. The sample uses a modular design:</p>\n<h3 id=\"agent-configuration\">Agent configuration</h3>\n<p><code>config.yaml</code> defines agent behavior, tools, and the system prompt. The sample defaults to Anthropic Claude Sonnet 4.5 on Amazon Bedrock for this walkthrough, a fit for the multi-step tool calling and long-context reasoning the human-in-the-loop workflow relies on. Switching to Claude Haiku, Amazon Nova, or another tool-calling model available on Amazon Bedrock (availability varies by Region) is a one-line change to <code>model_id</code> in the following configuration. <code>max_iterations: 10</code> gives the graph enough headroom for around ten model-to-tool round trips before it halts (the orchestrator doubles the value to compute LangGraph’s recursion limit, since each round trip traverses two graph nodes), which covers the multi-step tool use the sample’s S3, calculator, and <code>ask_human</code> tools expect. The file also contains an <code>execution:</code> block (loop-detector, circuit-breaker, session-cache thresholds) omitted here for brevity. See <code>agent/config.example.yaml</code> for the full file.</p>\n<h3 id=\"core-agent-implementation\">Core agent implementation</h3>\n<p>The agent is built on <code>langchain.agents.create_agent</code>, a tool-calling agent compiled as a LangGraph. LangChain handles the orchestration layer because it brings pre-built tool-calling patterns, a broad open ecosystem of integrations, and APIs many teams already know, while AgentCore Runtime supplies the managed hosting, session isolation, and scaling underneath. The two layers are complementary. The platform wrapper around the graph does three things: it builds the message list for the turn (including any conversation history and, for signal-triggered jobs, the Amazon S3 bucket and key that fired the signal so the model can pick up the file without being told), it invokes the graph, and it scans the tool output for the <code>ask_human</code> sentinel so a tool call can be turned into an <code>interrupted</code> response.</p>\n<p>In its simplest form (the full version in <code>agent/core/agent_core.py</code> also handles per-session history caching, loop detection, and an execution trace), the orchestrator is this:</p>\n<h3 id=\"deploying-the-agent\">Deploying the agent</h3>\n<p>Deploy the agent to AgentCore Runtime using the provided script:</p>\n<p>The script returns an Agent Runtime Amazon Resource Name (ARN) which the backend stores in the agent registry so the Job Execution Lambda function can invoke it.</p>\n<h2 id=\"state-management-with-dynamodb\">State management with DynamoDB</h2>\n<p>Amazon DynamoDB is the system of record for everything that needs to outlive a single Lambda invocation: which agents are registered, which jobs are in flight, the conversation history that gives an agent continuity across turns, and the signal definitions and chat threads the UI reads. The next subsections describe the table layout, the session model, and how the Job Execution Lambda function uses both to drive a job to completion.</p>\n<h3 id=\"database-design\">Database design</h3>\n<p>The state layer is backed by DynamoDB. The core tables exercised in this post are:</p>\n<p><strong>Agent registry table</strong> (one record per registered agent runtime):</p>\n<p><strong>Job registry table</strong> (one record per agent invocation). With the default <code>autoExecute: false</code>, a new signal-triggered job lands here in <code>idle</code> status, waiting for a user to choose Execute:</p>\n<p>While a job is running the worker flips <code>status</code> to <code>busy</code>. If the agent pauses by calling <code>ask_human</code>, <code>status</code> becomes <code>interrupted</code> and <code>requiresAction</code> becomes <code>true</code>, which surfaces the job on the Jobs page’s Interrupted tab.</p>\n<p><strong>Conversation store table</strong> (full message history per session, with a 30-day time to live (TTL)):</p>\n<p>Additional tables exist for ambient signals, chat threads, and Powertools idempotency records. Scheduled execution reuses the job-registry table through a <code>jobType</code> and <code>nextRun</code> GSI rather than having its own table.</p>\n<h3 id=\"session-management\">Session management</h3>\n<p>Sessions provide conversation continuity across job executions and chat turns. The <code>conversation_management</code> Lambda function owns DynamoDB persistence: every turn (human + AI pair) is appended to the session with an atomic <code>UpdateItem</code> + <code>list_append</code> so two concurrent writers on the same session can’t clobber each other, and a 30-day TTL takes care of cleanup.</p>\n<h3 id=\"job-execution-with-conversation-continuity\">Job execution with conversation continuity</h3>\n<p>The Job Execution Lambda function has two entry paths. The API path sends an SQS message (adding Powertools idempotency and a <code>userId</code> ownership check) and returns 202 Accepted immediately, so the frontend never waits for the model. The SQS worker path is where the real work happens: it loads conversation history, folds any human response into a continuation prompt, calls AgentCore Runtime, and writes the result back. That worker looks roughly like this:</p>\n<h2 id=\"implementing-human-in-the-loop-patterns\">Implementing human-in-the-loop patterns</h2>\n<p>The agent snippet earlier returned a sentinel-based <code>interrupted</code> envelope when the graph produced an <code>ask_human</code> tool call. This section zooms in on the <code>ask_human</code> tool itself and shows exactly how the orchestrator catches that sentinel without breaking the reasoning loop.</p>\n<h3 id=\"the-human-input-tool\">The human input tool</h3>\n<p>The core of human-in-the-loop functionality is the <code>ask_human</code> tool. It doesn’t raise an exception: because LangGraph’s tool node captures tool exceptions as error observations and feeds them back to the model, the tool instead returns a sentinel string. The orchestrator detects the sentinel on the <code>ToolMessage</code> stream after the graph finishes and converts it into an <code>interrupted</code> response. Per-invocation state (metrics, session ID, and so on) isn’t a tool argument. The orchestrator binds it on a <code>ContextVar</code> that tools read through <code>current_execution_state.get()</code>, which keeps concurrent invocations inside the same container isolated from each other.</p>\n<h3 id=\"agent-side-usage\">Agent-side usage</h3>\n<p>The agent uses the tool naturally as part of its tool-calling loop. For example, when analyzing an invoice with multiple line items the agent may call:</p>\n<p>The tool returns the sentinel, the orchestrator stops the graph, persists the question as the latest AI turn in the conversation store, and surfaces an <code>interrupted</code> job to the UI. When the user replies, the Job Execution Lambda function invokes the agent again with the answer folded into a continuation prompt.</p>\n<h2 id=\"building-the-agent-management-ui\">Building the Agent Management UI</h2>\n<p>The Agent Management UI is the human side of the platform: where users browse jobs, answer pending questions, and chat with agents directly. It’s a React single-page application that talks to the same REST API the rest of the post has been describing. The next two subsections cover the layout and the user experience flow that ties signals, jobs, and human input together.</p>\n<h3 id=\"frontend-architecture\">Frontend architecture</h3>\n<p>The frontend is a React and Cloudscape Design single-page application served from Amazon S3 behind Amazon CloudFront. It exposes a five-tab navigation:</p>\n<ul><li><strong>Workflows</strong> : A gallery and CRUD surface for workflow-level definitions that group agents and signals.</li><li><strong>Chat</strong> : A standalone chat page (<code>/chat</code> ,<code>/chat/:threadId</code> ) for user-initiated conversations with a registered agent.</li><li><strong>Agents</strong> : Manages registered agent runtimes (name, ARN, capabilities, status).</li><li><strong>Jobs</strong> : Lists all jobs with status filters and a detail view that includes an interactive Chat tab, an execution trace, and metadata.</li><li><strong>Signals</strong> : Defines and toggles ambient signals. The shipped form covers Amazon S3 prefix and suffix filters and schedules. Webhook, Amazon EventBridge, and DynamoDB stream fields are added when you wire those extension points.</li></ul>\n<p>The same chat component is reused by both the standalone Chat page and the Jobs detail Chat tab. While a job is running (<code>status === &quot;busy&quot;</code>), the panel polls <code>/conversations/:sessionId</code> on a short interval so turns from the agent or from a second browser tab show up within a few seconds. After the job settles, polling stops. Because it’s the same component in both places, a user can freely converse with an agent from the Jobs detail view while a job is active. They are not limited to answering a single pending question.</p>\n<h3 id=\"user-experience-flow\">User experience flow</h3>\n<ul><li><strong>Signal fires</strong> : A new job appears in the Jobs list with<code>jobType: &quot;signal_triggered&quot;</code> . If the signal has<code>autoExecute: false</code> (the default) the job lands in<code>idle</code> and the user runs it manually. With<code>autoExecute: true</code> the worker is already drafting a response by the time the list refreshes. User-initiated jobs carry<code>jobType: &quot;user_initiated&quot;</code> and render with a different badge so you can tell the two apart at a glance.</li><li><strong>Human input requested</strong> : Any job that calls<code>ask_human</code> moves to the Interrupted tab and its<code>requiresAction</code> flag flips to<code>true</code> , which renders a warning indicator on the row.</li><li><strong>Responding</strong> : The detail view surfaces a Provide Response button so the user can answer the pending question without leaving the job context.</li><li><strong>Execution feedback</strong> : Status transitions (<code>idle</code> →<code>busy</code> →<code>completed</code> |<code>interrupted</code> |<code>error</code> ) are polled and reflected in the list and in the Chat tab header.</li><li><strong>Interactive chat</strong> : The Jobs detail Chat tab and the<code>/chat</code> page both render full conversation history and accept new user messages, so humans can answer pending questions or nudge the agent with additional context without leaving the UI.</li><li><strong>Response submission</strong> : Sending a message calls the Chat Execution or Job Execution Lambda function, which invokes the agent on AgentCore Runtime with the job context and the human response folded into the next turn.</li><li><strong>Completion and audit</strong> : The final result is stored on the job record and the full conversation history is preserved on the session, available in the Chat tab for audit and re-use.</li></ul>\n<p>Authentication is handled by Amazon Cognito and an API Gateway <code>CognitoUserPoolsAuthorizer</code> (see <code>backend/infrastructure/multi_agent_stack.py</code>).</p>\n<p>Real-time updates are implemented as lightweight polling of the conversation and job endpoints, so there is no WebSocket infrastructure to run.</p>\n<h2 id=\"extending-the-sample\">Extending the sample</h2>\n<p>The reference implementation is deliberately small so you can see the contract before you start adding to it. This section describes what ships in the box, what you build on top, and where to plug in new event sources or tools.</p>\n<h3 id=\"what-ships-in-the-sample-compared-to-what-you-build\">What ships in the sample compared to what you build</h3>\n<p>Before extending, it helps to know where the seams are. This sample is a reference implementation of the ambient-agent pattern, not a turnkey product. What ships out of the box:</p>\n<ul><li><strong>The platform</strong> : Signal intake, the SQS-backed job pipeline, the HITL interrupt mechanism, job and conversation state in DynamoDB, the React UI with the Jobs page and Chat surfaces, and AgentCore integration with Cognito-backed auth.</li><li><strong>A reference agent</strong> : A containerized LangChain agent with four tools (calculator,<code>ask_human</code> ,<code>list_s3_files</code> ,<code>read_s3_file</code> ) and a generic system prompt. Useful for validating the pipeline end-to-end, not for solving your business problem out of the gate.</li><li><strong>One signal source wired through the Signals UI</strong> (Amazon S3 file uploads, matched by bucket, prefix, and suffix), plus a job-level scheduler that runs on a one-minute cron for jobs created with<code>jobType: &quot;scheduled&quot;</code> . Adding new signal types such as webhooks, Amazon EventBridge events, or DynamoDB streams requires extending the<code>signalType</code> enum in both the backend and the Signals page form.</li></ul>\n<p>What you build on top, and what the platform expects from each:</p>\n<ul><li><strong>Your agent</strong> : A containerized Python agent tailored to your domain, with its own tools and its own system prompt. It must speak the two conventions the platform relies on: return the three-status envelope (<code>completed</code> ,<code>interrupted</code> , or<code>error</code> with the matching<code>result</code> ,<code>question</code> , or<code>error</code> field) and use the<code>ask_human</code> sentinel to request human input. That is roughly 50 lines of adapter code around whatever agent framework you prefer.</li><li><strong>New signal types</strong> if Amazon S3 uploads and cron are not enough: A new handler Lambda function that queries the signals table and writes a<code>signal_triggered</code> job row (shown in the following template), plus a corresponding form field on the Signals page so users can configure it.</li><li><strong>New tools</strong> for your agent: A new module under<code>agent/tools/</code> , a registration entry in<code>agent/core/tool_factory.py</code> , and an enable flag in<code>config.yaml</code> . The orchestrator and HITL flow do not change. Any tool that returns a string participates in the same graph.</li></ul>\n<p>In short: the platform does the plumbing, you bring the brain. The three places you will write code (agent logic, tool implementations, and new signal handlers) all slot into the existing contract without touching the rest of the stack.</p>\n<h3 id=\"configuration-over-code\">Configuration over code</h3>\n<p>Enabling a new capability for an existing agent is a config change rather than a code change. The agent’s <code>config.yaml</code> toggles tools on or off and supplies their descriptions, so adding a new Amazon S3 prefix handler or a new calculator mode doesn’t require rebuilding the container. Signal definitions live in DynamoDB and are edited from the Signals page in the UI.</p>\n<h3 id=\"adding-a-new-signal-type\">Adding a new signal type</h3>\n<p>Signals today trigger on Amazon S3 events and on cron schedules. Adding a new type (a webhook, an Amazon EventBridge rule, a DynamoDB stream, a Kafka topic) follows the same three-step contract. The downstream pipeline is signal-agnostic: everything from the SQS worker through the AgentCore call through the UI treats all <code>signal_triggered</code> jobs the same, so you only need to write the bridge from your event source into a job row.</p>\n<p>You add the Lambda function and its API Gateway route in the CDK stack, extend <code>signalType</code> in the frontend’s TypeScript union and the Signals page form, and the rest (the Jobs page UI, the <code>ask_human</code> interrupt handling, the conversation persistence, the <code>autoExecute</code> toggle, the contract tests) all apply without change. Like the production Signal Processor, a real webhook handler would add AWS Lambda Powertools idempotency so a retried delivery doesn’t create duplicate jobs.</p>\n<h2 id=\"production-deployment\">Production deployment</h2>\n<p>After the pieces fit together locally, the next step is provisioning them in an account. The platform deploys as a single AWS CDK stack. The following subsections describe what that stack creates, the deploy commands, the metrics that matter when it is running, and the security and scaling defaults the sample ships with.</p>\n<h3 id=\"infrastructure-as-code\">Infrastructure as code</h3>\n<p>The whole platform is an AWS CDK stack defined in <code>backend/infrastructure/multi_agent_stack.py</code>. For storage and messaging, it provisions the DynamoDB tables (agent registry, job registry, conversation store, ambient signals, chat threads, idempotency records) with pay-per-request billing and point-in-time recovery, plus the SQS job-execution queue and its dead-letter queue.</p>\n<p>For compute and delivery, it provisions the three pipeline Lambda functions from the Architecture section (Signal Processor, Job Execution, Scheduler), the five management-tier Lambda functions that back the REST API (<code>agent_management</code>, <code>job_management</code>, <code>signal_management</code>, <code>chat_management</code>, <code>conversation_management</code>) and the <code>chat_execution</code> worker Lambda that <code>chat_management</code> invokes asynchronously, a Cognito-backed REST API, and the CloudFront distribution that serves the React frontend.</p>\n<p>IAM grants flow directly from the code: DynamoDB <code>grant_read_write_data</code> on each table, SQS <code>grant_send_messages</code> and <code>grant_consume_messages</code> as appropriate, and a scoped <code>bedrock-agentcore:InvokeAgentRuntime</code> policy on the worker role.</p>\n<h3 id=\"deployment-steps\">Deployment steps</h3>\n<h3 id=\"monitoring-and-observability\">Monitoring and observability</h3>\n<p>The Lambda functions emit custom CloudWatch metrics under the <code>AmbientAgents</code> namespace for the events that matter in this pattern: signal matches, jobs enqueued, jobs completed, jobs interrupted, and jobs errored. Combined with the default Lambda and SQS metrics (invocation count, errors, queue depth, DLQ depth), these give you a dashboard where the signal → job → agent → human cycle is visible end-to-end. Set an alarm on DLQ depth as the first line of defense against silent failures.</p>\n<h3 id=\"security-best-practices\">Security best practices</h3>\n<ul><li><strong>IAM least privilege</strong> : Grant only the permissions required by each Lambda function (reads on its own table, targeted<code>bedrock-agentcore:InvokeAgentRuntime</code> , targeted Amazon S3 prefixes).</li><li><strong>Encryption at rest</strong> : All DynamoDB tables use AWS-managed keys. Amazon S3 buckets use S3-managed encryption. CloudFront logs live in a dedicated logging bucket.</li><li><strong>Encryption in transit</strong> : All API Gateway endpoints and CloudFront distributions require HTTPS. Amazon S3 bucket policies enforce SSL.</li><li><strong>Audit logging</strong> : AWS CloudTrail captures API calls across the stack. DynamoDB streams can be enabled on the job table for richer job-lifecycle audit.</li><li><strong>Responsible AI</strong> : Apply Amazon Bedrock Guardrails to agent inputs and outputs so content filters, denied topics, and contextual grounding checks run before findings surface on the Jobs page or reach a reviewer through the<code>ask_human</code> prompt. This protects the human-in-the-loop interaction and keeps agent outputs grounded in the source documents.</li></ul>\n<h3 id=\"scaling-strategies\">Scaling strategies</h3>\n<ul><li><strong>Lambda concurrency</strong> : Reserve concurrency for the Job Execution worker Lambda function to cap downstream Bedrock invocations at a predictable per-account ceiling.</li><li><strong>DynamoDB capacity</strong> : Pay-per-request is the default. Switch to provisioned capacity with auto scaling when traffic patterns stabilize.</li><li><strong>Cost optimization</strong> : Use DynamoDB TTL to retire old conversations automatically, Amazon S3 lifecycle policies to age out processed documents, and CloudWatch metric filters to track per-agent model invocation cost.</li></ul>\n<h2 id=\"clean-up\">Clean up</h2>\n<p>To avoid ongoing charges after you are finished evaluating the sample, tear the stack back down in the reverse order it was created.</p>\n<p>First, delete the agent runtime from Amazon Bedrock AgentCore so the container stops being billed:</p>\n<p>Next, empty the Amazon S3 buckets the stack provisions (the documents bucket, the frontend hosting bucket, and the CloudFront logging bucket) so CloudFormation can delete them:</p>\n<p>Then destroy the AWS CDK stack itself, which removes the Lambda functions, API Gateway, CloudFront distribution, DynamoDB tables, SQS queues, Cognito user pool, and IAM roles:</p>\n<p>Finally, delete the Amazon ECR repository created by the agent deploy script if you do not plan to redeploy:</p>\n<p>If you enabled Amazon Bedrock model access only for this walkthrough and no longer need it, you can revoke it from the <strong>Model access</strong> page of the Amazon Bedrock console.</p>\n<h2 id=\"summary\">Summary</h2>\n<p>Ambient agents shift AI automation from “wait for a user” to “respond to signals”. Time-to-action drops from hours to seconds, and around-the-clock monitoring becomes possible without constant human attention. Each agent operates independently with isolated sessions, so the platform handles many concurrent events in parallel without coordination overhead. Humans stay in the loop only when their input is truly needed. The unified Jobs page removes context switching across multiple tools, and full conversation history means you never lose track of what an agent did or why it paused.</p>\n<h3 id=\"when-to-use-ambient-patterns\">When to use ambient patterns</h3>\n<p>The following use cases describe the ambient-agent pattern in general. The reference sample ships the plumbing for Amazon S3 and the scheduler today. Webhooks, database changes, and external integrations are extension points you wire up on top.</p>\n<p>The pattern is a strong fit whenever an event in your environment should drive work that a model can reason about and a human should occasionally weigh in on. Document-processing pipelines are the canonical example: a file lands, an agent reads it, summarizes it, and asks for approval before acting. Monitoring and alerting flows benefit from the same shape, with the agent triaging an alert and asking a human only when escalation is needed. Scheduled processing, analysis, and reporting fit naturally on the cron path, and any multi-step workflow that needs an approval gate or a human checkpoint along the way maps cleanly onto the <code>ask_human</code> interrupt.</p>\n<p>There are also workloads where ambient agents are the wrong tool. Real-time chat applications belong in a traditional chat agent that is purpose-built for low-latency conversational turn-taking. Simple request-response APIs do not need an agent at all. A plain API endpoint is faster, cheaper, and easier to reason about. And for purely deterministic workflows that do not need large language model (LLM) reasoning, AWS Step Functions is still the right choice. Note that ambient agents with <code>autoExecute: true</code> already cover the fully autonomous case if you want LLM reasoning without a human gate.</p>\n<h3 id=\"next-steps\">Next steps</h3>\n<p>Clone the reference implementation and follow the top-level README for deployment and configuration.</p>\n<p>Further reading:</p>\n<p>Start small. Pick a single use case (document processing or a monitoring alert) and expand from there. Clone the sample, point it at your own event source, and see how much routine work the Jobs page takes off your plate.</p>","headings":[{"level":2,"text":"Artificial Intelligence","id":"artificial-intelligence"},{"level":1,"text":"Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows","id":"building-ambient-agents-with-amazon-bedrock-agentcore-from-event"},{"level":2,"text":"Prerequisites","id":"prerequisites"},{"level":2,"text":"Understanding ambient agents","id":"understanding-ambient-agents"},{"level":3,"text":"Event-driven compared to user-initiated agents","id":"event-driven-compared-to-user-initiated-agents"},{"level":3,"text":"Ambient signals: The trigger mechanism","id":"ambient-signals-the-trigger-mechanism"},{"level":3,"text":"Human-in-the-loop: One tool, one envelope, one view","id":"human-in-the-loop-one-tool-one-envelope-one-view"},{"level":2,"text":"Architecture overview","id":"architecture-overview"},{"level":2,"text":"Building the event infrastructure","id":"building-the-event-infrastructure"},{"level":3,"text":"Setting up the Amazon S3 signal trigger","id":"setting-up-the-amazon-s3-signal-trigger"},{"level":3,"text":"Signal Processor Lambda function","id":"signal-processor-lambda-function"},{"level":3,"text":"Signal configuration data model","id":"signal-configuration-data-model"},{"level":2,"text":"Deploying agents on AgentCore Runtime","id":"deploying-agents-on-agentcore-runtime"},{"level":3,"text":"Agent packaging and structure","id":"agent-packaging-and-structure"},{"level":3,"text":"Agent configuration","id":"agent-configuration"},{"level":3,"text":"Core agent implementation","id":"core-agent-implementation"},{"level":3,"text":"Deploying the agent","id":"deploying-the-agent"},{"level":2,"text":"State management with DynamoDB","id":"state-management-with-dynamodb"},{"level":3,"text":"Database design","id":"database-design"},{"level":3,"text":"Session management","id":"session-management"},{"level":3,"text":"Job execution with conversation continuity","id":"job-execution-with-conversation-continuity"},{"level":2,"text":"Implementing human-in-the-loop patterns","id":"implementing-human-in-the-loop-patterns"},{"level":3,"text":"The human input tool","id":"the-human-input-tool"},{"level":3,"text":"Agent-side usage","id":"agent-side-usage"},{"level":2,"text":"Building the Agent Management UI","id":"building-the-agent-management-ui"},{"level":3,"text":"Frontend architecture","id":"frontend-architecture"},{"level":3,"text":"User experience flow","id":"user-experience-flow"},{"level":2,"text":"Extending the sample","id":"extending-the-sample"},{"level":3,"text":"What ships in the sample compared to what you build","id":"what-ships-in-the-sample-compared-to-what-you-build"},{"level":3,"text":"Configuration over code","id":"configuration-over-code"},{"level":3,"text":"Adding a new signal type","id":"adding-a-new-signal-type"},{"level":2,"text":"Production deployment","id":"production-deployment"},{"level":3,"text":"Infrastructure as code","id":"infrastructure-as-code"},{"level":3,"text":"Deployment steps","id":"deployment-steps"},{"level":3,"text":"Monitoring and observability","id":"monitoring-and-observability"},{"level":3,"text":"Security best practices","id":"security-best-practices"},{"level":3,"text":"Scaling strategies","id":"scaling-strategies"},{"level":2,"text":"Clean up","id":"clean-up"},{"level":2,"text":"Summary","id":"summary"},{"level":3,"text":"When to use ambient patterns","id":"when-to-use-ambient-patterns"},{"level":3,"text":"Next steps","id":"next-steps"}]}}