The U.S. debate over how to navigate China’s role in developing and governing advanced artificial intelligence is stuck between two extremes. One camp calls for full-blown international treaties even though the two countries cannot yet agree on AI’s most serious risks, let alone how to manage them. The other is deeply skeptical of dialogue, believing progress will only be possible if the Chinese side is desperate.
A grand bargain may be too ambitious, yet focusing only on what can be agreed on today could lock both sides into a policy regime that becomes obsolete as technology evolves. Doing nothing is even worse because regardless of whether it was developed in San Francisco or Hangzhou, an AI model that makes it easier for terrorists to synthesize deadly pathogens or hack into the infrastructure that underpins financial markets threatens everyone. These risks are becoming all too real as AI systems demonstrate cyber- and biological capabilities that match or exceed those of human experts on increasingly complex tasks.
The U.S. debate over how to navigate China’s role in developing and governing advanced artificial intelligence is stuck between two extremes. One camp calls for full-blown international treaties even though the two countries cannot yet agree on AI’s most serious risks, let alone how to manage them. The other is deeply skeptical of dialogue, believing progress will only be possible if the Chinese side is desperate.
A grand bargain may be too ambitious, yet focusing only on what can be agreed on today could lock both sides into a policy regime that becomes obsolete as technology evolves. Doing nothing is even worse because regardless of whether it was developed in San Francisco or Hangzhou, an AI model that makes it easier for terrorists to synthesize deadly pathogens or hack into the infrastructure that underpins financial markets threatens everyone. These risks are becoming all too real as AI systems demonstrate cyber- and biological capabilities that match or exceed those of human experts on increasingly complex tasks.
The recent OpenAI-Hugging Face incident, in which AI agents collaborated to break out of their testing environment and gain access to the internal infrastructure of external organizations, made clear that AI already operates beyond human control in ways that can unwittingly impact third parties. It is only a matter of time before these threats cross borders, and Beijing will certainly want clarity from Washington should one of its companies become the next accidental victim of a U.S. model going rogue—and vice versa.
As Tsinghua University’s Xiao Qian recently wrote in Foreign Policy, “Sustained dialogue is a strategic necessity.” She is right. And there is a middle path forward. When U.S. and Chinese delegations gather this month to discuss AI risks, rather than pursuing a comprehensive agreement, they should start with a series of narrower measures focused on preventing extreme risks, managing crises, and investing in verification tools. Each small agreement helps build the understanding and trust needed for more ambitious deals, allowing the two countries to cooperate where they already agree without requiring consensus on everything else.
Of course, any bilateral agreement must recognize the powerful incentives for governments and companies on both sides to continue competing against each other. But even within this environment, each country can take steps that advance its own interests while building tools that make reciprocal commitments more credible down the line.
One of the biggest obstacles for negotiators is that neither country has developed tools that can confidently evaluate whether AI models are safe, nor safeguards capable of mitigating AI’s most extreme risks. These challenges affect both countries but are particularly acute for China, whose AI safety evaluations are improving but still less mature than the West’s. Given the transnational nature of AI risks, it is in both countries’ interests to have robust safety infrastructure, but governments have been hesitant to participate in technical exchanges, afraid of revealing information that could enhance a rival’s capabilities.
But it is possible to have carefully scoped, working-level technical exchanges that discuss best practices without exposing models’ inner workings, vulnerabilities, or advanced capabilities. In general, the two sides should compare broad approaches but withhold tactical details and refrain from granular, riskier discussions of the specific methods they use to elicit unsafe model behavior. They could, for instance, discuss how to scale evaluations to cover more models, behaviors, and potential threats, as well as how to use AI systems themselves to help identify and respond to threats. They could also discuss important, publicly available lessons on monitoring AI models to ensure they don’t escape their sandboxes—a topic that should be of mutual concern following the OpenAI-Hugging Face incident.
The two sides should also discuss safeguards using the same principles, comparing general approaches to detecting harmful activity, restricting dangerous use, and managing risks once models are deployed without discussing specific techniques for bypassing safeguards or proprietary details about how models are built.
U.S. experts might be able to offer insights from how they determined which actors should have access to Mythos, Anthropic’s most advanced model, explaining the broad criteria they used to evaluate safeguard effectiveness. They should also talk about the limitations to their current safeguards, explaining how they decided when dangerous, dual-use capabilities warranted tighter restrictions on model access. Similarly, U.S. third-party evaluators could use the public report from METR and Redwood Research on the OpenAI-Hugging Face incident to examine why safeguards were insufficient.
If these discussions go well, a next step could be to create technical working groups to establish voluntary best practices, soliciting input from companies and independent third-party organizations.
China would ultimately assess lessons from these exchanges on its own terms, but insights that align with its national security interests could shape decisions that reduce risks for everyone.
Even if the United States and China have a mutual understanding of AI risks and do everything they can to mitigate them, serious crises can still unfold.
Though China has the world’s most extensive AI regulations, it lacks clear rules for catastrophic risk management. The biggest U.S. AI companies are now subject to California’s SB 53 law, which requires them to disclose safety practices, report certain critical safety incidents, and protect whistleblowers. But these rules have not yet been federally adopted, limiting what the government can do when a serious incident occurs.
Thus, the United States and China should discuss how to share key information and communicate effectively in a crisis. On the domestic front, each side should commit to setting up its own reporting systems to track cases where AI caused, or may cause, material harm—for example, if a model escapes a contained sandbox to attack telecommunications infrastructure or gain access to sensitive medical patient data. These incident-reporting systems can vary according to each country’s preferences, but they should collect broadly similar information and have similar categories and thresholds for severity.
In parallel, the countries should explore creating a bilateral information-sharing channel, which would be voluntary at first, with the ability to change requirements later on. U.S. companies such as OpenAI and Anthropic have already voluntarily shared information on certain safety incidents, but having a private channel can help both sides navigate cross-border incidents and proactively warn counterparts if a model is behaving in unintended ways that might be perceived as malicious.
Critically, though, crisis channels must fit how governments actually make decisions. Chinese interlocutors, for example, have explained in Track 2 conversations that they face strong disincentives from answering phone calls on “hotlines” because no individual answering the call has the authority to respond without consulting superiors. Using written communication could circumvent some of those challenges.
Even if the United States and China wanted to reach a bilateral AI agreement, they might not trust that the other side will follow through on its commitments. Tools that enable both countries to verify the other’s compliance can allow both sides to feel confident about the other’s safety practices. This mutual confidence, in turn, could enable further agreements and governance applications.
The first verification prototypes are just now being developed, and it may take years until they are sufficiently mature and trusted by both countries. That’s why the United States and China need to begin talking about the technology now, explicitly supporting technical research to develop verification tools for bilateral AI agreements.
China has been historically hesitant to participate in hardware-based verification regimes. These concerns could be alleviated if the two countries discussed at these early stages what types of verification technologies could be politically acceptable and who should build them to make them credible to both sides. And if specific proposals are brought to the table, the delegations could have a more concrete discussion on a given technology’s strengths and limitations.
The United States and China may never need some of these tools, and the agreements they could enable may ultimately not be desirable. But developing credible verification technologies takes both time and institutional capacity, and to be useful, they must be designed and built long before they are needed.
The upcoming dialogue should not be evaluated on whether the United States and China reach full agreement on all these issues. It will be a success if it lays the foundation for a serious second meeting and eventually a permanent high-level channel supported by both leaders. If talks break down early, the two sides might not get back to the negotiation table for months or even years.
These goals may seem modest, given the urgent need for international cooperation to combat extreme risks. But sequenced strategically, small diplomatic breakthroughs may lead to more ambitious measures, and each minor agreement will open possibilities for the next.
The United States and China may be competing to build the potentially defining technology of our time, but rivalry need not foreclose pragmatism. This September, both countries can take the first step in that direction.


