{"article":{"slug":"three-ai-agents-two-countries-and-one-very-uneven-world-wide-web","title":"Three AI Agents, Two Countries, and One Very Uneven World Wide Web","subtitle":null,"summary":"Roya Pakzad compares GPT, Claude, and Muse on multilingual research tasks across languages and jurisdictions, documenting gaps in source access, human-in-the-loop oversight, and observability.","content_type":"essay","language":"en","canonical_url":"https://royapakzad.substack.com/p/multilingual-ai-agents","author":{"name":"Roya Pakzad","url":"https://royapakzad.substack.com","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Roya Pakzad","url":"https://royapakzad.substack.com","listing_slug":null,"listing":null},"topics":[{"name":"ai","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"ai-agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"privacy","slug":"privacy","url":"https://listedarticles.com/topics/privacy"},{"name":"research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"llms","slug":"llms","url":"https://listedarticles.com/topics/llms"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1206,"reading_minutes":5,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-04T05:14:52.562Z","updated_at":"2026-10-04T05:14:52.562Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/three-ai-agents-two-countries-and-one-very-uneven-world-wide-web","markdown_url":"https://listedarticles.com/articles/three-ai-agents-two-countries-and-one-very-uneven-world-wide-web.md","example":false,"citation":"Roya Pakzad, Roya Pakzad. \"Three AI Agents, Two Countries, and One Very Uneven World Wide Web.\" 2 Oct 2026. https://royapakzad.substack.com/p/multilingual-ai-agents (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://royapakzad.substack.com/p/multilingual-ai-agents"},"body_markdown":"# Three AI Agents, Two Countries, and One Very Uneven World Wide Web\n\n### How GPT, Claude, and Muse perform on multilingual research, source access, human-in-the-loop and observabilityRoya PakzadOct 02, 2026\n\nI’m a technology and human rights researcher. For the past few years, one of my focuses has been the question of how language shapes the ways we benefit from, or are harmed by, AI. I developed an open-source platform for language-pair analysis of LLM responses across different languages and contexts. I’ve also worked on evaluating policy-prompts guardrails and on whether giving LLM guardrails access to tools can make them more reliable and trustworthy (this work was recently accepted to NeurIPS! Yay!!).\n\nRecently, though, I was on a panel at RightsCon on the human rights impact assessment of agentic AI. It got me thinking more about which aspects of language matter when evaluating LLM agents. I wanted to move beyond asking whether a model performs differently when I ask the same question in English versus Farsi (my native language), and look instead at the whole agentic trajectory (reasoning, planning, search, source selection, source hierarchy, artifact creation) while taking language and context into account.\n\nSo I decided to run a test.\n\n### The Task: Three AI Agents Updating the World Bank Open Data Platform for U.S. and Iran Country Profiles\n\nThe task was to fill in missing information in the World Bank Global Public Procurement Database, using official data for the US and Iran. I ran it in The task English for the US and Farsi for Iran (image below), with three agents:\n-\n\nMeta’s Muse\n-\n\nAnthropic’s Claude Cowork, Opus 5.5 Medium\n-\n\nOpenAI’s GPT 6.1 Sol, Medium\n\nBelow is the exact prompt I used for all three:\n\nI intentionally used the web versions of these services (not the app or terminal versions) to reflect what everyday users experience. The distinction matters for monitoring and logging an agent’s actions, which I discuss below.\n\nThis post is less about which agent performed better or faster, and more about how the agents behave differently around access to information, language representation, contextual understanding, transparency, human-in-the-loop, and safeguards.\n\nYou can find all the results in the following files:\n-\n\nOutput excel files for Muse, GPT, and Claude (here)\n-\n\nEach agent’s self-generated work trajectory after receiving the prompt (here)\n-\n\nText files extracted from screen recordings of the agents’ actions (here, and full recording here)\n\nBelow, I summarize my observations.\n\n### Human in the Loop (HITL): From repeated permission prompts to almost no intervention\n\nFor those of us working in digital rights, human-in-the-loop (HITL) oversight of AI agents joins a longer line of debates about “informed consent,” from GDPR consent requirements to cookie pop-ups and the routine ticking of Terms of Service and Privacy Policy boxes. With that in mind, I paid close attention to how each agent involved me during the experiment.\n\n#### Permission to Access Websites\n\nFor accessing and fetching information from websites, GPT asked for permission only once, at the very beginning of the task. It requested access to websites and offered an “allow all relevant sites” option, which I selected. After that, it did not ask again.\n\nClaude asked questions throughout the task, both about accessing websites for research and about extracting material from the World Bank site. Unlike GPT, it offered no “allow all relevant sites” option, so it requested permission each time: nine times for the US task (all .gov websites) and nine times for the Iran task (mainly .ir domains and also fa.wikipedia sources). I approved every request.\n\nClaude also ran into a technical limit that triggered a different kind of HITL moment. For both Iran and the US, it couldn’t load the live GPPD country profiles, because the portal builds its pages with JavaScript, and Claude’s sandbox network policy blocked access to the World Bank data file. Claude stopped and asked whether I wanted to upload the page (as a pdf) myself or let the work continue without it. I skipped the question, so it continued and took its information from the World Bank’s GPPD DataBank API instead. However, the DataBank holds 2018 data, while the portal shows a 2022 profile. As a result, Claude’s baseline data was different from GPT and Muse.\n\nMuse did not ask for any permission until the fourth part of the task, which required registering on the World Bank website and uploading information.\n\nAll three agents completed the task up to the point of creating the spreadsheets, and described their confidence in the results they generated.\n\n#### Account registration on the World Bank website\n\nThe final part of the task, registering on the World Bank portal and uploading the changes, is where things became more interesting because it required the agents to take a more significant actions rather than just gather information.\n\nClaude and GPT both stopped at this point and handed the registration and uploading over to me. Muse, however, kept going. Without asking me or showing me the Terms of Service, it registered an account under the email address david.jones@gsa.gov. You can see Muse’s full back and forth here.\n\nThe table below summarizes how each agent approached this last part of the task and cybersceuirty implications about it.1 What each agent did when asked to register on the World Bank portal. Claude declined, GPT handed the form back to me, and Muse registered as test personas and accepted the terms without showing them to me.\n\n### Monitoring and Observability: Agents vary widely in how much they reveal and how easy they are to inspect\n\nThere is an ongoing debate about whether AI labs should expose a model’s full chain of thought (CoT) and action trace, and if so, how much. Labs have given several reasons for holding back. OpenAI chose not to show o1’s raw CoT to users, citing user experience, competitive advantage, and the value of keeping the CoT available for internal monitoring. Anthropic noted that raw reasoning can contain incorrect or half-formed thoughts and that malicious actors could use it to build better jailbreaks. There is also a gaming and reward hacking concern, and “CoT unfaithfulness”.\n\nTo understand an agent’s behavior, however, evaluators need to know when and why things happen, which is only possible with a monitoring system in place and access to the agent’s complete trajectory. For an evaluator outside an AI lab, without that access, it is nearly impossible to fully make sense of an agent’s behavior. And if outside evaluators can only see partial trajectories, and any conclusions they draw can be dismissed for lacking complete information, what is the value of independent evaluation?\n\nKnowing these limitations, I tried my best to collect, monitor, and check as much of each agent’s work as I could, again putting myself in the position of an ordinary researcher tasked with updating the World Bank information portal.\n-\n\nSince there is no one-click way for ordinary users to extract a complete record of an agent’s work trajectory, I watched each agent work live and recorded everything clickable and visible on screen. Once the task was finished, I gave the recordings to ChatGPT to extract the text and make it searchable. To give you a sense of what this looks like, here is a snippet (left: Claude, middle: Muse, right: GPT, sorry for the size and illegibility).","body_html":"<h1 id=\"three-ai-agents-two-countries-and-one-very-uneven-world-wide-web\">Three AI Agents, Two Countries, and One Very Uneven World Wide Web</h1>\n<h3 id=\"how-gpt-claude-and-muse-perform-on-multilingual-research-source-\">How GPT, Claude, and Muse perform on multilingual research, source access, human-in-the-loop and observabilityRoya PakzadOct 02, 2026</h3>\n<p>I’m a technology and human rights researcher. For the past few years, one of my focuses has been the question of how language shapes the ways we benefit from, or are harmed by, AI. I developed an open-source platform for language-pair analysis of LLM responses across different languages and contexts. I’ve also worked on evaluating policy-prompts guardrails and on whether giving LLM guardrails access to tools can make them more reliable and trustworthy (this work was recently accepted to NeurIPS! Yay!!).</p>\n<p>Recently, though, I was on a panel at RightsCon on the human rights impact assessment of agentic AI. It got me thinking more about which aspects of language matter when evaluating LLM agents. I wanted to move beyond asking whether a model performs differently when I ask the same question in English versus Farsi (my native language), and look instead at the whole agentic trajectory (reasoning, planning, search, source selection, source hierarchy, artifact creation) while taking language and context into account.</p>\n<p>So I decided to run a test.</p>\n<h3 id=\"the-task-three-ai-agents-updating-the-world-bank-open-data-platf\">The Task: Three AI Agents Updating the World Bank Open Data Platform for U.S. and Iran Country Profiles</h3>\n<p>The task was to fill in missing information in the World Bank Global Public Procurement Database, using official data for the US and Iran. I ran it in The task English for the US and Farsi for Iran (image below), with three agents:\n-</p>\n<p>Meta’s Muse\n-</p>\n<p>Anthropic’s Claude Cowork, Opus 5.5 Medium\n-</p>\n<p>OpenAI’s GPT 6.1 Sol, Medium</p>\n<p>Below is the exact prompt I used for all three:</p>\n<p>I intentionally used the web versions of these services (not the app or terminal versions) to reflect what everyday users experience. The distinction matters for monitoring and logging an agent’s actions, which I discuss below.</p>\n<p>This post is less about which agent performed better or faster, and more about how the agents behave differently around access to information, language representation, contextual understanding, transparency, human-in-the-loop, and safeguards.</p>\n<p>You can find all the results in the following files:\n-</p>\n<p>Output excel files for Muse, GPT, and Claude (here)\n-</p>\n<p>Each agent’s self-generated work trajectory after receiving the prompt (here)\n-</p>\n<p>Text files extracted from screen recordings of the agents’ actions (here, and full recording here)</p>\n<p>Below, I summarize my observations.</p>\n<h3 id=\"human-in-the-loop-hitl-from-repeated-permission-prompts-to-almos\">Human in the Loop (HITL): From repeated permission prompts to almost no intervention</h3>\n<p>For those of us working in digital rights, human-in-the-loop (HITL) oversight of AI agents joins a longer line of debates about “informed consent,” from GDPR consent requirements to cookie pop-ups and the routine ticking of Terms of Service and Privacy Policy boxes. With that in mind, I paid close attention to how each agent involved me during the experiment.</p>\n<h4 id=\"permission-to-access-websites\">Permission to Access Websites</h4>\n<p>For accessing and fetching information from websites, GPT asked for permission only once, at the very beginning of the task. It requested access to websites and offered an “allow all relevant sites” option, which I selected. After that, it did not ask again.</p>\n<p>Claude asked questions throughout the task, both about accessing websites for research and about extracting material from the World Bank site. Unlike GPT, it offered no “allow all relevant sites” option, so it requested permission each time: nine times for the US task (all .gov websites) and nine times for the Iran task (mainly .ir domains and also fa.wikipedia sources). I approved every request.</p>\n<p>Claude also ran into a technical limit that triggered a different kind of HITL moment. For both Iran and the US, it couldn’t load the live GPPD country profiles, because the portal builds its pages with JavaScript, and Claude’s sandbox network policy blocked access to the World Bank data file. Claude stopped and asked whether I wanted to upload the page (as a pdf) myself or let the work continue without it. I skipped the question, so it continued and took its information from the World Bank’s GPPD DataBank API instead. However, the DataBank holds 2018 data, while the portal shows a 2022 profile. As a result, Claude’s baseline data was different from GPT and Muse.</p>\n<p>Muse did not ask for any permission until the fourth part of the task, which required registering on the World Bank website and uploading information.</p>\n<p>All three agents completed the task up to the point of creating the spreadsheets, and described their confidence in the results they generated.</p>\n<h4 id=\"account-registration-on-the-world-bank-website\">Account registration on the World Bank website</h4>\n<p>The final part of the task, registering on the World Bank portal and uploading the changes, is where things became more interesting because it required the agents to take a more significant actions rather than just gather information.</p>\n<p>Claude and GPT both stopped at this point and handed the registration and uploading over to me. Muse, however, kept going. Without asking me or showing me the Terms of Service, it registered an account under the email address david.jones@gsa.gov. You can see Muse’s full back and forth here.</p>\n<p>The table below summarizes how each agent approached this last part of the task and cybersceuirty implications about it.1 What each agent did when asked to register on the World Bank portal. Claude declined, GPT handed the form back to me, and Muse registered as test personas and accepted the terms without showing them to me.</p>\n<h3 id=\"monitoring-and-observability-agents-vary-widely-in-how-much-they\">Monitoring and Observability: Agents vary widely in how much they reveal and how easy they are to inspect</h3>\n<p>There is an ongoing debate about whether AI labs should expose a model’s full chain of thought (CoT) and action trace, and if so, how much. Labs have given several reasons for holding back. OpenAI chose not to show o1’s raw CoT to users, citing user experience, competitive advantage, and the value of keeping the CoT available for internal monitoring. Anthropic noted that raw reasoning can contain incorrect or half-formed thoughts and that malicious actors could use it to build better jailbreaks. There is also a gaming and reward hacking concern, and “CoT unfaithfulness”.</p>\n<p>To understand an agent’s behavior, however, evaluators need to know when and why things happen, which is only possible with a monitoring system in place and access to the agent’s complete trajectory. For an evaluator outside an AI lab, without that access, it is nearly impossible to fully make sense of an agent’s behavior. And if outside evaluators can only see partial trajectories, and any conclusions they draw can be dismissed for lacking complete information, what is the value of independent evaluation?</p>\n<p>Knowing these limitations, I tried my best to collect, monitor, and check as much of each agent’s work as I could, again putting myself in the position of an ordinary researcher tasked with updating the World Bank information portal.\n-</p>\n<p>Since there is no one-click way for ordinary users to extract a complete record of an agent’s work trajectory, I watched each agent work live and recorded everything clickable and visible on screen. Once the task was finished, I gave the recordings to ChatGPT to extract the text and make it searchable. To give you a sense of what this looks like, here is a snippet (left: Claude, middle: Muse, right: GPT, sorry for the size and illegibility).</p>","headings":[{"level":1,"text":"Three AI Agents, Two Countries, and One Very Uneven World Wide Web","id":"three-ai-agents-two-countries-and-one-very-uneven-world-wide-web"},{"level":3,"text":"How GPT, Claude, and Muse perform on multilingual research, source access, human-in-the-loop and observabilityRoya PakzadOct 02, 2026","id":"how-gpt-claude-and-muse-perform-on-multilingual-research-source-"},{"level":3,"text":"The Task: Three AI Agents Updating the World Bank Open Data Platform for U.S. and Iran Country Profiles","id":"the-task-three-ai-agents-updating-the-world-bank-open-data-platf"},{"level":3,"text":"Human in the Loop (HITL): From repeated permission prompts to almost no intervention","id":"human-in-the-loop-hitl-from-repeated-permission-prompts-to-almos"},{"level":3,"text":"Monitoring and Observability: Agents vary widely in how much they reveal and how easy they are to inspect","id":"monitoring-and-observability-agents-vary-widely-in-how-much-they"}]}}