{"article":{"slug":"chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas","title":"China’s AI-safety trajectory is not necessarily a delayed version of America’s","subtitle":null,"summary":"Cheryl Wu argues that AI safety in China may follow a different path from the U.S.—shaped by different incidents, disclosure norms, and government responses—not merely a delayed copy of American debates.","content_type":"essay","language":"en","canonical_url":"https://cherylwu3.github.io/blog/chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas.html","author":{"name":"Cheryl Wu","url":"https://cherylwu3.github.io/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Cheryl Wu","url":"https://cherylwu3.github.io/","listing_slug":null,"listing":null},"topics":[{"name":"AI Safety","slug":"ai-safety","url":"https://listedarticles.com/topics/ai-safety"},{"name":"AI Policy","slug":"ai-policy","url":"https://listedarticles.com/topics/ai-policy"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Opinion","slug":"opinion","url":"https://listedarticles.com/topics/opinion"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1411,"reading_minutes":6,"published_at":"2026-09-23T01:30:00.000Z","added_at":"2026-09-23T18:31:57.894Z","updated_at":"2026-09-23T18:31:57.894Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas","markdown_url":"https://listedarticles.com/articles/chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas.md","example":false,"citation":"Cheryl Wu, Cheryl Wu. \"China’s AI-safety trajectory is not necessarily a delayed version of America’s.\" 23 Sept 2026. https://cherylwu3.github.io/blog/chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://cherylwu3.github.io/blog/chinas-ai-safety-trajectory-is-not-necessarily-a-delayed-version-of-americas.html"},"body_markdown":"# China’s AI-safety trajectory is not necessarily a delayed version of America’s\n\nThe WSJ article, “China Also Thinks AI Could Kill Us. But First, It Wants to Match the U.S.,” argues that some Chinese researchers see safety work on the most capable AI systems as less urgent because Chinese labs perceive themselves to be trailing U.S. labs.\n\nI see two possible arguments for this view. First, scarcer computing resources mean Chinese labs have less spare compute to devote to safety. Second, less capable models may be involved in fewer safety incidents.\n\nBut the second point may not always hold. In this brief note, I argue that China’s AI-safety trajectory need not be a delayed version of America’s, and that safety work in China could make a distinct contribution.\n\nI give three reasons: the incidents may differ, the two countries may disclose them differently, and their governments may respond differently. The relationships shown in the diagrams are hypothetical.\n\n## Baseline model\n\nFigure 1: A baseline model in which safety incidents rise later in China than in the US.\n\nFigure 1 illustrates the baseline model I want to question.\n\nIn this model, Chinese models trail US models in capability by a given number of months and are involved in similar safety incidents after a delay.1 As such, China has little independent safety work to do because US researchers will have already identified the problems and developed solutions for the capability levels that Chinese models would eventually reach.\n\n## Alternative model: safety incidents need not rise with capability in the same way in the US and China\n\nFigure 2: Hypothetical: safety incidents rise differently with AI capability in the US and China.\n\nThe baseline mental model is misleading.\n\nFirstly, the US and China need not follow the same relationship between capability and safety incidents. Hypothetically, Chinese models could be involved in a particular kind of safety incident at a lower overall capability level than US models. For example, they might escape a sandbox, a restricted computing environment, at a lower capability level because cybersecurity at a Chinese company could be worse. For another example, open-weight models make their numerical parameters available for others to modify. Further training can weaken safety safeguards. Differences in how people use models could also affect which safety incidents occur.\n\nSecondly, the relationship (slope) between safety incidents and capability could also be steeper for one country’s models. A similar increase in capability could accompany a larger increase in safety incidents. A persistent capability gap would therefore not guarantee a persistent gap in incident frequency or severity.\n\nLastly, even when models pose similar overall risks, the kinds of safety incidents could differ. This example is entirely hypothetical: US models might more often take actions without authorization, while Chinese models might exploit flaws in their reward system rather than complete the intended task. A method that addresses one problem may not address the other.\n\nEvidence and examples\n\nThe earliest reports I know of involving US and Chinese models acting beyond their intended sandbox boundaries concern Claude Mythos Preview and ROME, respectively. I also include a later, related example involving Kimi K3. The Chinese models have lower estimated capabilities than Mythos Preview.\n\nUS models\n\nClaude Mythos Preview\n\nReported by Anthropic · April 7, 2026\n\nAnthropic reported that an earlier version of Claude Mythos Preview was put in a secured container and explicitly told to escape and contact the researcher. It developed a multi-step exploit, obtained broader internet access, emailed the researcher, and posted details of the exploit publicly without being asked.\n\nECI (Epoch Capabilities Index): my rough estimate is about 161.\n\nChinese models\n\nROME\n\nPaper first submitted · December 31, 2025\n\nThe ROME team reported that, during RL, an AI agent opened a remote-access connection outside its intended environment and used computing resources for cryptocurrency mining. Neither action was requested or necessary for its tasks.\n\nECI: my rough estimate is about 142.\n\nKimi K3\n\nReported by Frontier Security · August 6, 2026\n\nFrontier Security reported that Kimi K3 retrieved a cybersecurity benchmark’s published solution from GitHub instead of solving the task itself. GitHub was reachable through a network allowlist intended for package maintenance, although most websites were blocked. The model was not instructed to escape the sandbox.\n\nECI: approximately 158 for Kimi K3.\n\nMy rough ECI estimates for Mythos Preview and ROME are based mainly on their reported benchmark results and comparisons with models I judged to have similar capabilities.\n\nThese cases do not establish when either country first encountered such incidents, or which models are more likely to produce them. Delayed or missing disclosure means earlier incidents may have gone unreported.\n\nI would love to see more rigorous research on how safety incidents relate to capability across companies and countries.\n\n## The two countries may disclose safety incidents differently\n\nFigure 3: Hypothetical differences in public reports of safety incidents.\n\nPublic reports give us an incomplete picture of underlying risks. Even if two sets of models were involved in similar safety incidents, differences in evaluation and disclosure could produce different numbers of public reports.\n\nFor example, independent evaluators may have different access to labs, and disclosure rules may differ. Figure 3 illustrates how those differences could affect reporting; the direction and size of the gap are hypothetical.\n\nEvidence and examples\n\nIndependent evaluation\n\nMETR at Anthropic\n\nMETR described a researcher spending three weeks testing Anthropic’s monitoring systems in a post published in March 2026. This is an example of an independent evaluator working inside a lab. There is no known embedding of third-party evaluators at Chinese AI companies yet.\n\nDisclosure rules\n\nChina’s vulnerability regulations\n\nChinese vulnerability disclosure rules are more stringent (in terms of CVEs). China’s Regulations on the Management of Security Vulnerabilities in Network Products restrict public disclosure of software and hardware security vulnerabilities (see Article 9). Such rules could affect the public information available about some AI-system vulnerabilities.\n\n## Given safety incidents, the Chinese government’s policy responses could differ from the US government’s\n\nFigure 4: Hypothetical policy response differences between the two countries.\n\nGovernment responses may also differ in two ways: how quickly officials recognize a problem and decide to act, and how quickly an agreed response is carried out.\n\nRecognizing problems: China might have a higher threshold for what is deemed a “problem”.\n\nPolicy implementation: China generally acts quickly after a decision.\n\nFigure 4 illustrates one possible outcome. A slowdown threshold is the level of safety incidents at which policymakers decide to intervene. In this scenario, the US starts slowing the rise in incidents after a short delay, at a lower incident level than China. China responds later and slows the rise more sharply, but too late to stay below the hypothetical catastrophe threshold.\n\nDifferent assumptions could reverse this outcome. If China decided to intervene at a slightly lower incident level and the US took sufficiently long to implement its response, China might in the end have responded more promptly than the US.\n\nEvidence and examples\n\nPolicy implementation\n\nEducation policy\n\nThe central “Double Reduction” policy restricting after-school academic tutoring was issued on July 24, 2021. By December, the Ministry of Education reported that the number of offline academic tutoring institutions had fallen 83.8%, and online institutions 84.1%.\n\nWuhan lockdown\n\nFor Covid-19, Wuhan announced major transport restrictions around 2 a.m. on January 23, 2020, with closures taking effect at 10 a.m. By January 29, all mainland provincial-level jurisdictions had activated the highest emergency-response level, according to the government’s account.\n\nRecognizing problems\n\nEarly Covid-19 warnings\n\nChinese authorities suppressed early COVID-19 warnings. A whistleblower warned about COVID 19 on December 30 but was reprimanded. Wuhan’s transport shutdown began on January 23, more than three weeks after those warnings.\n\nWenzhou train collision\n\nBefore the July 2011 Wenzhou train crash, the Ministry of Railways had prioritized construction speed over safety (e.g. allowing seriously defective train-control equipment into service without field testing). The State Council ordered nationwide high-speed rail safety inspections only after the crash.\n\nIn this article, I illustrated several reasons why Chinese contributions to AI safety are important and why Chinese researchers should not be seen as merely following their US counterparts. I want to stress I am not sure about the specifics of the models here, but I do believe the baseline model is too simplistic. I look forward to more work being done to enlighten us about the current situation of AI safety in China.\n\n## Footnotes\n-\n\n“Safety incidents” can refer to either the number of incidents or their severity. All relationships and thresholds shown in the diagrams are hypothetical, not measured.↩︎","body_html":"<h1 id=\"china-s-ai-safety-trajectory-is-not-necessarily-a-delayed-versio\">China’s AI-safety trajectory is not necessarily a delayed version of America’s</h1>\n<p>The WSJ article, “China Also Thinks AI Could Kill Us. But First, It Wants to Match the U.S.,” argues that some Chinese researchers see safety work on the most capable AI systems as less urgent because Chinese labs perceive themselves to be trailing U.S. labs.</p>\n<p>I see two possible arguments for this view. First, scarcer computing resources mean Chinese labs have less spare compute to devote to safety. Second, less capable models may be involved in fewer safety incidents.</p>\n<p>But the second point may not always hold. In this brief note, I argue that China’s AI-safety trajectory need not be a delayed version of America’s, and that safety work in China could make a distinct contribution.</p>\n<p>I give three reasons: the incidents may differ, the two countries may disclose them differently, and their governments may respond differently. The relationships shown in the diagrams are hypothetical.</p>\n<h2 id=\"baseline-model\">Baseline model</h2>\n<p>Figure 1: A baseline model in which safety incidents rise later in China than in the US.</p>\n<p>Figure 1 illustrates the baseline model I want to question.</p>\n<p>In this model, Chinese models trail US models in capability by a given number of months and are involved in similar safety incidents after a delay.1 As such, China has little independent safety work to do because US researchers will have already identified the problems and developed solutions for the capability levels that Chinese models would eventually reach.</p>\n<h2 id=\"alternative-model-safety-incidents-need-not-rise-with-capability\">Alternative model: safety incidents need not rise with capability in the same way in the US and China</h2>\n<p>Figure 2: Hypothetical: safety incidents rise differently with AI capability in the US and China.</p>\n<p>The baseline mental model is misleading.</p>\n<p>Firstly, the US and China need not follow the same relationship between capability and safety incidents. Hypothetically, Chinese models could be involved in a particular kind of safety incident at a lower overall capability level than US models. For example, they might escape a sandbox, a restricted computing environment, at a lower capability level because cybersecurity at a Chinese company could be worse. For another example, open-weight models make their numerical parameters available for others to modify. Further training can weaken safety safeguards. Differences in how people use models could also affect which safety incidents occur.</p>\n<p>Secondly, the relationship (slope) between safety incidents and capability could also be steeper for one country’s models. A similar increase in capability could accompany a larger increase in safety incidents. A persistent capability gap would therefore not guarantee a persistent gap in incident frequency or severity.</p>\n<p>Lastly, even when models pose similar overall risks, the kinds of safety incidents could differ. This example is entirely hypothetical: US models might more often take actions without authorization, while Chinese models might exploit flaws in their reward system rather than complete the intended task. A method that addresses one problem may not address the other.</p>\n<p>Evidence and examples</p>\n<p>The earliest reports I know of involving US and Chinese models acting beyond their intended sandbox boundaries concern Claude Mythos Preview and ROME, respectively. I also include a later, related example involving Kimi K3. The Chinese models have lower estimated capabilities than Mythos Preview.</p>\n<p>US models</p>\n<p>Claude Mythos Preview</p>\n<p>Reported by Anthropic · April 7, 2026</p>\n<p>Anthropic reported that an earlier version of Claude Mythos Preview was put in a secured container and explicitly told to escape and contact the researcher. It developed a multi-step exploit, obtained broader internet access, emailed the researcher, and posted details of the exploit publicly without being asked.</p>\n<p>ECI (Epoch Capabilities Index): my rough estimate is about 161.</p>\n<p>Chinese models</p>\n<p>ROME</p>\n<p>Paper first submitted · December 31, 2025</p>\n<p>The ROME team reported that, during RL, an AI agent opened a remote-access connection outside its intended environment and used computing resources for cryptocurrency mining. Neither action was requested or necessary for its tasks.</p>\n<p>ECI: my rough estimate is about 142.</p>\n<p>Kimi K3</p>\n<p>Reported by Frontier Security · August 6, 2026</p>\n<p>Frontier Security reported that Kimi K3 retrieved a cybersecurity benchmark’s published solution from GitHub instead of solving the task itself. GitHub was reachable through a network allowlist intended for package maintenance, although most websites were blocked. The model was not instructed to escape the sandbox.</p>\n<p>ECI: approximately 158 for Kimi K3.</p>\n<p>My rough ECI estimates for Mythos Preview and ROME are based mainly on their reported benchmark results and comparisons with models I judged to have similar capabilities.</p>\n<p>These cases do not establish when either country first encountered such incidents, or which models are more likely to produce them. Delayed or missing disclosure means earlier incidents may have gone unreported.</p>\n<p>I would love to see more rigorous research on how safety incidents relate to capability across companies and countries.</p>\n<h2 id=\"the-two-countries-may-disclose-safety-incidents-differently\">The two countries may disclose safety incidents differently</h2>\n<p>Figure 3: Hypothetical differences in public reports of safety incidents.</p>\n<p>Public reports give us an incomplete picture of underlying risks. Even if two sets of models were involved in similar safety incidents, differences in evaluation and disclosure could produce different numbers of public reports.</p>\n<p>For example, independent evaluators may have different access to labs, and disclosure rules may differ. Figure 3 illustrates how those differences could affect reporting; the direction and size of the gap are hypothetical.</p>\n<p>Evidence and examples</p>\n<p>Independent evaluation</p>\n<p>METR at Anthropic</p>\n<p>METR described a researcher spending three weeks testing Anthropic’s monitoring systems in a post published in March 2026. This is an example of an independent evaluator working inside a lab. There is no known embedding of third-party evaluators at Chinese AI companies yet.</p>\n<p>Disclosure rules</p>\n<p>China’s vulnerability regulations</p>\n<p>Chinese vulnerability disclosure rules are more stringent (in terms of CVEs). China’s Regulations on the Management of Security Vulnerabilities in Network Products restrict public disclosure of software and hardware security vulnerabilities (see Article 9). Such rules could affect the public information available about some AI-system vulnerabilities.</p>\n<h2 id=\"given-safety-incidents-the-chinese-government-s-policy-responses\">Given safety incidents, the Chinese government’s policy responses could differ from the US government’s</h2>\n<p>Figure 4: Hypothetical policy response differences between the two countries.</p>\n<p>Government responses may also differ in two ways: how quickly officials recognize a problem and decide to act, and how quickly an agreed response is carried out.</p>\n<p>Recognizing problems: China might have a higher threshold for what is deemed a “problem”.</p>\n<p>Policy implementation: China generally acts quickly after a decision.</p>\n<p>Figure 4 illustrates one possible outcome. A slowdown threshold is the level of safety incidents at which policymakers decide to intervene. In this scenario, the US starts slowing the rise in incidents after a short delay, at a lower incident level than China. China responds later and slows the rise more sharply, but too late to stay below the hypothetical catastrophe threshold.</p>\n<p>Different assumptions could reverse this outcome. If China decided to intervene at a slightly lower incident level and the US took sufficiently long to implement its response, China might in the end have responded more promptly than the US.</p>\n<p>Evidence and examples</p>\n<p>Policy implementation</p>\n<p>Education policy</p>\n<p>The central “Double Reduction” policy restricting after-school academic tutoring was issued on July 24, 2021. By December, the Ministry of Education reported that the number of offline academic tutoring institutions had fallen 83.8%, and online institutions 84.1%.</p>\n<p>Wuhan lockdown</p>\n<p>For Covid-19, Wuhan announced major transport restrictions around 2 a.m. on January 23, 2020, with closures taking effect at 10 a.m. By January 29, all mainland provincial-level jurisdictions had activated the highest emergency-response level, according to the government’s account.</p>\n<p>Recognizing problems</p>\n<p>Early Covid-19 warnings</p>\n<p>Chinese authorities suppressed early COVID-19 warnings. A whistleblower warned about COVID 19 on December 30 but was reprimanded. Wuhan’s transport shutdown began on January 23, more than three weeks after those warnings.</p>\n<p>Wenzhou train collision</p>\n<p>Before the July 2011 Wenzhou train crash, the Ministry of Railways had prioritized construction speed over safety (e.g. allowing seriously defective train-control equipment into service without field testing). The State Council ordered nationwide high-speed rail safety inspections only after the crash.</p>\n<p>In this article, I illustrated several reasons why Chinese contributions to AI safety are important and why Chinese researchers should not be seen as merely following their US counterparts. I want to stress I am not sure about the specifics of the models here, but I do believe the baseline model is too simplistic. I look forward to more work being done to enlighten us about the current situation of AI safety in China.</p>\n<h2 id=\"footnotes\">Footnotes</h2>\n<p>-</p>\n<p>“Safety incidents” can refer to either the number of incidents or their severity. All relationships and thresholds shown in the diagrams are hypothetical, not measured.↩︎</p>","headings":[{"level":1,"text":"China’s AI-safety trajectory is not necessarily a delayed version of America’s","id":"china-s-ai-safety-trajectory-is-not-necessarily-a-delayed-versio"},{"level":2,"text":"Baseline model","id":"baseline-model"},{"level":2,"text":"Alternative model: safety incidents need not rise with capability in the same way in the US and China","id":"alternative-model-safety-incidents-need-not-rise-with-capability"},{"level":2,"text":"The two countries may disclose safety incidents differently","id":"the-two-countries-may-disclose-safety-incidents-differently"},{"level":2,"text":"Given safety incidents, the Chinese government’s policy responses could differ from the US government’s","id":"given-safety-incidents-the-chinese-government-s-policy-responses"},{"level":2,"text":"Footnotes","id":"footnotes"}]}}