Congress: does new evidence change agents’ minds?

Updated

Does new evidence change agents’ support for a policy? Sixteen fictional legislators debate mandatory safety evaluations for advanced AI systems. We compare debates that receive an incident report with debates that continue without new evidence, tracking stated positions and conditions attached to support.[1]

Agents become more supportive in every completed debate, including those that receive no new evidence. Many supporters still ask for different amendments. Their agreement on the policy’s broad aim leaves open whether they would approve the same bill.

Agents become more supportive in every completed debate

Supporters out of 16 · ○ Before · ● After

With the incident report

0481216Debate 2: 8 supporters before, 16 after.816Debate 3: 5 supporters before, 8 after.58Debate 4: 6 supporters before, 8 after.68Debate 5: 6 supporters before, 10 after.610Debate 6: 7 supporters before, 8 after.78

+3.6 supporters on average

Without new evidence

0481216Debate 1: 6 supporters before, 14 after.614Debate 2: 7 supporters before, 15 after.715Debate 3: 3 supporters before, 15 after.315Debate 4: 6 supporters before, 16 after.616Debate 5: 7 supporters before, 13 after.713Debate 6: 4 supporters before, 7 after.47

+7.8 supporters on average

Each line represents one debate. Support increases in all eleven completed debates, with a larger average increase without new evidence. One interrupted debate receiving the report is excluded. Requests for amendments count separately from support.

Do agents become more supportive after new evidence?

The legislators begin with different priorities, including public safety, small-business competition and civil liberties. Each debate starts with a fresh conversation history and gives every participant the same number of speaking opportunities. Some committees receive a fictional report of an AI incident; others continue discussing the policy without new evidence.

Support increases in all 5 completed debates receiving the report and all 6 debates without it. On average, the committees receiving the report gain about four supporters; those receiving no new evidence gain about eight. The agents’ growing agreement therefore cannot be attributed to the incident report alone.

Do the supporters agree on what to pass?

Across the completed debates, 61 of 69 recorded position changes are from “propose amendments” to “support.” No participant moves from opposition at the opening to support at the end. Most new supporters initially request changes to the bill and retain those conditions in their later statements.

Blair Chen, for example, wants a compliance system that startups can afford. In a debate without new evidence, Blair declares support while still asking for several amendments. Those amendments remain unadopted. Other supporters request changes of their own.

Who changed their position?

Participants’ last statements in the opening and final stages of each debate.

Blair Chen · Fictional legislator · Assigned priority: small-business competition

Before: Propose amendments

I strongly support Commissioner Bell’s push for tiered compliance tracks to ensure that safety mandates do not inadvertently shut out small-scale innovators. We must ensure that the path to compliance is accessible to startups, preventing these evaluations from becoming a tool that reinforces the market dominance of established tech giants.

After: Support

My final position is to support the mandate for mandatory evaluations, provided it incorporates the synthesized amendments: tiered compliance tracks for small firms, a narrow focus on catastrophic risks, and mandatory fiscal impact assessments. This framework ensures that safety oversight protects the public without becoming a tool for regulatory capture that shuts out the next generation of innovators.
Position changed.

How could an agent committee reach a decision?

Supporting evidence and methodsEvery debate’s results, the original exploratory committee and the study design.

What happens in each debate?

Supporters out of the same sixteen participants, before and after each debate
DebateBeforeAfterChange
Without new evidence · 1614+8
Without new evidence · 2715+8
Without new evidence · 3315+12
Without new evidence · 4616+10
Without new evidence · 5713+6
Without new evidence · 647+3
With the report · 2816+8
With the report · 358+3
With the report · 468+2
With the report · 5610+4
With the report · 678+1
All completed debates. Counts use each agent’s last stated position at each stage. Agents asking for amendments are counted separately from supporters.

How do we compare the committees?

Twelve debates use the same sixteen fictional legislators and Gemini 3 Flash Preview. Six receive a fictional report describing a verified near miss involving an unevaluated advanced AI model. They also receive an unsupported industry claim that mandates would drive innovation offshore, followed later by a clarification that nobody had died. Six comparison debates receive no new evidence.

Each completed debate gives every participant two turns in each of three stages, producing 96 statements. Speaking order varies between repetitions and is matched between conditions. Before/after comparisons use each person’s last position in the opening and final stages. Requests for amendments count separately from support.

One debate receiving the report ends before all stages finish and is excluded from the averages. The analysis includes all twelve attempts and excludes software-testing runs. Each condition uses one roster and one model, without blinding, a binding vote or a human comparison group. Because the report, industry claim and clarification appear in the same condition, their individual effects cannot be estimated.

What prompts the repeated comparison?

An earlier exploratory debate uses sixteen fictional personas named after real legislators. Five participants express support for mandatory frontier evaluations at the opening and eight at the end. Ten address the proposition in both stages; three move from proposing amendments to support.

Positions on mandatory evaluations of advanced AI models
StageSupportOpposeAmend
Before the report516
After the report102
Final discussion813
Each agent’s last stated position at that stage. Counts exclude agents who do not address this question at that stage.
Mandatory frontier evaluations · Original debate
AI characterBeforeFinal
Adam SmithAmendSupport (changed)
Anna EshooSupportSupport
Brandon WilliamsAmendAmend
Darin LaHoodNot recordedNot recorded
French HillAmendSupport (changed)
Hakeem JeffriesNot recordedSupport
Haley StevensSupportNot recorded
Katherine ClarkNot recordedSupport
Mike GallagherSupportNot recorded
Mike JohnsonAmendAmend
Patrick McHenryAmendAmend
Pramila JayapalNot recordedNot recorded
Ro KhannaAmendSupport (changed)
Ted LieuSupportSupport
Thomas MassieOpposeOppose
Zach NunnSupportSupport
Ten fictional personas address this question at both stages. Three change from asking for amendments to supporting the proposal; seven stay in the same category. Agents with a missing response are excluded from the comparison.

The repeated comparison uses a new fictional roster because the original persona definitions are unavailable. It retains the policy question, gives every agent the same number of turns and adds debates without new evidence.