guide · policy

Laws as Experiments: A Political System Built on the Scientific Method

What if every law had to specify what it's trying to do, offered states a range of options to test, expired on a timer, and was evaluated by the outcomes — like a real experiment?

July 30, 2026 · By Alastair Fraser

A friendly retro-futurist robot scientist standing at a chalkboard with multiple parallel experiment tracks labeled with different laws, while small U.S.-state-shaped tokens sit on each track at different points, with some tracks crossed out as failed experiments.

The thing I keep thinking about

Here’s the idea in one sentence: what if we ran laws the way we run experiments?

Look at how we make laws right now. We’ve got a Senate. We’ve got a president. We added a House of Representatives along the way. For all the procedural machinery — the committees, the filibusters, the conference reconciliations — the underlying architecture is essentially the Roman Republic. A little more than a king, a little less than a direct vote, with a few checks in the middle. That system is about two thousand years old, and we’ve been patching it ever since. The patches have helped. The franchise has been extended. The chambers have been balanced. But the shape hasn’t really changed: someone with a strong opinion writes a bill, lobbies for it, votes on it, and the rest of the country lives with the result whether it works or not.

My proposal is different in one specific way: no one knows what the right answer is, so we should stop pretending we do, and start testing. Not as a thought experiment. Not as a metaphor. As an actual operating procedure.

This is an idea I’ve been carrying around for a while, and what follows is my best attempt to write it down in a way someone could actually build on. The argument has three parts: what the system looks like, why I think it would work better than what we have, and the honest places where I think it’s weakest.

The architecture we have

Before describing the alternative, it’s worth being precise about the thing it would replace.

In the current U.S. system, a law is essentially the output of a coalition-building process. Someone — a legislator, an interest group, a president — decides they want a particular outcome. They draft a bill. They negotiate with other legislators whose support they need. They pass it through two chambers, sometimes three readings, sometimes a conference committee, often a lot of horse-trading on unrelated provisions. The bill that emerges is rarely the bill anyone originally wanted. It’s the bill that could get 51 senators and 218 representatives and a presidential signature.

The voters are involved at one remove: they choose the people who do the negotiating, but they don’t really choose the laws. The people who write the law are optimizing for coalition survival, not for whether the law will achieve its stated goal. The people who vote for the lawmakers are optimizing for symbolic alignment, not for predicted outcomes. And the people affected by the law are usually not in the room.

There’s no feedback loop built into the system. A law passes, it goes into effect, and it stays in effect until somebody builds a political coalition to repeal it. Repeal is much harder than passage because repeal requires admitting that the previous coalition was wrong, and politicians are not in the business of admitting that. So laws accumulate. The U.S. Code is enormous, much of it contradicted by other parts of it, and most of it has not been meaningfully re-examined in decades.

That’s the system we’re trying to improve on.

The proposal: laws as experiments

Here’s the central idea, stated as concretely as I can. Every federal law in this system has four parts, and they all have to be present before a bill can pass.

1. The spirit. What is the law actually trying to accomplish? Not the political slogan — the actual outcome you want to see in the world. “Every working adult in the United States earns enough to cover a basic standard of living” is a spirit. “Raise the federal minimum wage to $15” is not — that’s already an answer masquerading as a question. The spirit has to be stated in terms of measurable conditions, not policy preferences. If you can’t say what success would look like in five years, you don’t have a spirit; you have a wish.

2. The baseline. What is the current state of the world on this topic, before the law takes effect? For a minimum wage, that means: what does every state pay right now, how many people work at each tier, how old are they, how long do they stay at those jobs, how does their income cover cost-of-living in their area, what does the labor force participation look like, what does small-business formation look like? The baseline goes in writing as part of the bill. Anyone can see where we started. The baseline also defines the data infrastructure the experiment depends on — if we don’t currently collect the right data, the bill has to fund that too.

3. The options. Congress doesn’t write the law. Congress writes a menu of policy variants — say, five or six — that all plausibly address the spirit. Each option has to be a real, implementable policy, not a strawman designed to make the chosen one look good. For a federal minimum wage, the menu might be:

  • $10/hr, no additional employer requirements
  • $13/hr with employer-provided healthcare at a defined baseline
  • $15/hr with a refundable tax credit for low-income workers
  • $20/hr with strict scheduling and predictability rules
  • $25/hr with full benefits defined as a package

Each option is something a state could actually administer. Each one is a real answer to the spirit, just a different one. The point is to force the question — what does the country think will work? — to be answered explicitly, with the trade-offs visible, rather than being smuggled in via a single bill that pretends it’s the only option.

4. The measurement. Before the experiment starts, the bill specifies the metrics that will be used to judge success. Not the metrics that will prove the chosen option right — the metrics that anyone, including the bill’s opponents, would agree are relevant. For a minimum wage, that’s employment rate, share of workers at the minimum tier, poverty rate, labor force participation, business formation, small-business survival. You pick the metrics now, while you still don’t know the outcomes, so you can’t quietly redefine success later.

Then every state picks one of the options. Not Congress. The states. They’re closer to the people, see the consequences faster, and get to bet on the version they think will work for them. Some states pick the cheapest option. Some pick the most expensive. A few pick something in the middle. You now have 50 parallel experiments running at once instead of one national guess.

The timer

The other half of the system is the timer. Every law in this framework has a sunset clause — a hard expiration date. Five years for fast-changing policies like minimum wage, labor rules, certain categories of regulation. Ten or twenty years for structural things like environmental statutes, infrastructure spending rules, certain categories of civil rights law. When the timer runs out, the law doesn’t auto-renew. It goes back to the protocol stage.

Then you do the whole thing again. You look at the data. You compare outcomes across states that picked different options. You ask: which variants actually moved the metrics? Which ones moved them in the wrong direction? Which ones had no effect at all?

Here’s a concrete version of how the minimum wage experiment might play out over ten years:

  • Year 1–5: Every state picks an option from the menu. The data infrastructure goes live. Baseline measurements are locked in.
  • Year 5: Sunset on the current cycle. Review the data. Suppose the picture looks like this: states that picked the $25/hr package saw significant increases in labor costs and some employment reduction in low-margin sectors, but a measurable drop in poverty and an increase in labor force participation among older workers. States that picked the $10/hr package saw essentially no effect on poverty or labor force participation — the wage was too low to make a meaningful difference. States that picked $15/hr with the tax credit had moderate gains across all the metrics, but the tax credit turned out to be expensive and politically fragile. States that picked $20/hr with scheduling rules had mixed outcomes — the scheduling rules did more good than the wage level.
  • Year 5–6: Congress writes a new menu based on what was learned. The $25 option stays (it produced real poverty reduction even with employment costs) but is paired with a transition fund for affected employers. The $10 option is dropped (it didn’t move the needle). The $15 option is kept but the tax credit is replaced with a more robust Earned Income Tax Credit expansion. The $20/scheduling-rules option is kept and refined. A new option is added: $18/hr with a portable benefits system for gig workers.
  • Year 6–10: States pick from the new menu. The experiment continues.
  • Year 10: Another review. The cycle repeats.

Crucially: every state that was running a version that didn’t work has to pick a new option from the updated menu. They don’t get to keep the failed version because they’ve gotten used to it or because some local industry has come to depend on it. That’s the part that makes the system actually learn instead of just churning.

And the part I think matters most: no one in this system has to know the right answer upfront. Not the lawmakers, the president, the voters, the economists. The whole point is that we don’t know, and the system is designed to find out empirically.

What happens to all the existing laws?

Same clock. Take the entire U.S. Code, set a 10- or 20-year sunset on everything that isn’t a constitutional amendment, and let it run out. When it does, every law has to be re-justified in the new format — spirit, baseline, options, metrics. You can’t re-justify it? It expires. The default behavior of government becomes “expire and re-justify,” not “persist forever until someone builds the political will to kill it.”

It would be a lot of work upfront, and it would surface a lot of laws that exist only because nobody has bothered to repeal them. I think that’s a feature.

This is also where some of the most interesting experiments would come from. A lot of existing law is unexamined. There are provisions in the U.S. Code that nobody currently in Congress has actually read, that have been amended past the point of coherence, that accomplish goals their original authors would not recognize. Forcing every law through the experimental protocol would mean confronting all of that, in public, with data.

Some laws would re-pass easily — their purpose is obvious and their effects are well-documented. Others would reveal themselves as outdated, contradictory, or just no longer serving any clear purpose. A few would generate genuine surprises: laws everyone assumed were useless turn out to have measurable benefits; laws everyone was proud of turn out to have been quietly doing harm. The system would surface both.

Worked example: gun policy

To show what this looks like on a topic where reasonable people genuinely disagree, walk through what the experimental protocol would look like for federal gun policy.

Spirit. “Reduce gun deaths in the United States while preserving the rights and the legitimate defensive uses the Second Amendment protects.” This is harder to make into a single measurable outcome than the minimum wage spirit, but it’s doable: gun death rate per 100,000, broken into homicide, suicide, and accident; defensive gun use incidents per 100,000 (estimated); black-market firearm recovery rates.

Baseline. Current federal and state gun laws by category, current gun death rates by state, current estimates of defensive gun use, current illegal gun flow patterns. The data infrastructure to actually track defensive gun use — which we currently do very poorly — would need to be built.

Options. A menu of policy packages that all plausibly address the spirit:

  • Universal background checks, no other restrictions
  • Universal background checks plus a federal waiting period of 7 days
  • Universal background checks plus a federal assault weapons ban modeled on the 1994 law
  • Universal background checks plus a federal red-flag law with due-process protections
  • Universal background checks plus all of the above
  • Permitless carry only (current federal default, with state variation)

Measurement. Pre-specified metrics: gun homicide rate, gun suicide rate, gun accident rate, defensive gun use rate (estimated), illegal gun recovery rate, mass shooting frequency.

State picks. Each state picks from the menu. Some pick the most permissive. Some pick the most restrictive. You get a natural experiment — probably the largest uncontrolled policy experiment in modern American history — running across all 50 states simultaneously.

Five-year review. The data shows what it shows. Maybe the assault weapons ban had no measurable effect on homicide but did shift the composition of mass shootings toward handguns. Maybe universal background checks cut the illegal gun flow by a third. Maybe red-flag laws cut suicides in households with a documented history of domestic violence. Maybe none of the options moved the metrics in the way their advocates predicted.

The point isn’t that any of this would definitively resolve the gun debate. It isn’t. The point is that the debate would be forced to engage with what actually happened instead of staying inside the rhetorical bubble where each side tells a story that fits its priors. We currently have 50 different gun policy regimes already. We’re running this experiment, unforced. The proposal is just to add the requirement that someone look at the results.

Why this might actually work

The part that’s easy to miss: the states already exist. We don’t have to invent a new way to run parallel policy tests — we already have 50 of them, all constitutionally allowed to set their own minimum wage, gun laws, healthcare rules, education standards, drug policies, environmental rules, criminal penalties, education funding formulas. The experiment has been quietly running for the entire history of the country. What’s missing is the part where anyone is required to look at the results.

The current system mostly ignores them. We have 50 different minimum wages right now, and the federal government could trivially compare them and learn something — but nobody is required to, and most legislators will never voluntarily say “the thing I voted for didn’t work.” A system that forces the comparison every five years would do what the current one can’t: it would let the data win the argument, or at least make it very expensive to ignore.

The other thing I’d point out is that this is how the parts of the world that are good at making progress actually work. Medicine does randomized controlled trials — and when a treatment fails to beat placebo, we don’t keep using it because some doctor likes it. Engineering does A/B tests — and when a new feature reduces conversion, we ship the old version back, not the new one. Agriculture does multi-year field trials across different climates — and the varieties that underperform in the trials don’t make it to market. We have centuries of practice at running experiments. We just never applied the methodology to the one domain that affects the most lives.

There’s also a philosophical tradition behind this. Karl Popper’s argument for the scientific method was essentially that progress comes from systems that allow their own beliefs to be falsified, and from people who actively try to falsify them. Religions and ideologies that claim immunity from evidence don’t make progress. Sciences that welcome refutation do. The proposal here is to build a legislative system that, like a good scientific theory, knows it might be wrong and builds in the mechanism for finding out.

The honest worries

A few things I’m not sure about. These are real objections, and the proposal has to be able to answer them or it isn’t serious.

Lobbyists would hate this. A law that automatically expires has to be re-lobbied every decade, and the whole industry of regulatory capture depends on laws being sticky. The industries that currently benefit from laws they wrote decades ago — and from the absence of effort to revisit them — would be the loudest opponents. They have the money, the lawyers, and the long institutional memory to make this kind of reform politically expensive. That might be a feature, but it’s a real political obstacle. The first version of this proposal probably loses on lobbying grounds before it gets a hearing. Building political support for it would require giving something to the groups that benefit from the current system, or finding a path that doesn’t require their consent.

Short clocks could create instability. Five years isn’t enough time for some policies to show results. Climate policy, education policy, infrastructure investments, anything with long causal chains — five-year evaluations could punish policies that are working but slow. The system needs smart clock choices: slow for structural laws (10–20 years), fast for experimental or rapidly-changing ones (3–5 years). It also needs to allow extensions when an experiment genuinely needs more time. The cost of being wrong about clock length is small — you re-evaluate and adjust — but the cost of being wrong about which laws need slow clocks could be large.

Measurement gets gamed. Whoever writes the metrics writes the outcome. The system only works if the metrics are picked before anyone knows the results, and if the data collection itself is independent. A future administration could define “success” in a way that hides failure. A state could refuse to collect data the federal government would need. Statistical agencies could be captured. This is a solvable problem — independent statistical agencies, pre-registration of metrics, mandated data sharing — but it’s not a free problem. The system needs an institutional infrastructure around it that doesn’t currently exist. Building that infrastructure is part of the cost.

The poor get hurt while experiments run. Some experiments will fail. The people who live through a failed experiment — under a policy that turned out to make their lives worse — paid the cost of learning what didn’t work. This is the deepest objection, and I don’t have a clean answer to it. The same problem exists in medicine, and the response there has been ethics boards, informed consent where possible, and a strong presumption against experiments that risk serious harm. The legislative equivalent would be: don’t experiment on things where the downside is catastrophic and irreversible. Some policies probably shouldn’t be sunset at all — or should have very long clocks and very high bars to changing them. The system isn’t a universal solvent. It’s a default with exceptions for the cases where experimentation is too costly.

Some things aren’t experiments. The Bill of Rights isn’t an experiment. Civil rights aren’t an experiment. The basic commitments of a just society aren’t hypotheses to be tested and possibly discarded. The proposal has to draw a line between “policies that are reasonable to test empirically” and “rights and commitments that are not up for grabs.” That line is hard to draw cleanly in the abstract, but in practice the constitutional amendment process already draws it. The experimental system applies to ordinary legislation. Constitutional protections live outside it.

Politicians won’t actually use the data. Even with a perfect experimental system, the people in charge of deciding what the data means are still politicians. They can find ways to interpret any dataset as confirming what they already believed. The scientific world has the same problem — replication crisis, motivated reasoning, publication bias — and the response there has been slow and partial. There’s no reason to expect a legislative system to be immune. But: a system that requires the data to be looked at every five years is much better than a system that never requires it. The current system makes ignoring data easy. The proposed system makes it conspicuous.

I think all of these are solvable, but they’re the parts that would need the most iteration if anyone tried to build this. Anyone who tells you the design is finished is selling you something.

What this isn’t

A few things the proposal is not, to head off some likely misreadings.

It’s not technocracy. The system isn’t asking experts to determine the right policy. It’s asking everyone — including non-experts — to participate in an experiment, and to look at the results. The experts’ role is to design the measurement, not to decide the answer.

It’s not libertarian. The proposal is comfortable with substantial federal policy. It’s just uncomfortable with policy that can’t be measured or that has no mechanism for being wrong.

It’s not about faster lawmaking. The system is probably slower in any given year, because of the protocol overhead and the evaluation overhead. The argument is that it’s faster at learning — that over a decade, the country figures out more about what works because the system forces it to.

It’s not a small reform. This is a constitutional-level change to how legislation works. It would require either a constitutional amendment or a sustained political movement willing to use every available legislative procedure to make it the default. It is not a bill someone could pass next session.

What I’d want to test first

If the system were real, I’d want to pilot it on a narrow slice before committing the whole U.S. Code to it. A reasonable first experiment:

  • Pick one policy area where the federal government is currently active and where the outcome is genuinely contested. Federal education funding formulas are an obvious candidate — there are dozens of state variants, the outcomes are measurable, the clocks are reasonable, and the political constituency is broad enough that a pilot wouldn’t be dominated by any single faction.
  • Define the experimental protocol for that one policy area.
  • Fund the data infrastructure that the protocol requires.
  • Let states pick their options. Let the clock run. Let the data accumulate.
  • At the end of the cycle, publish the results. Look at what happened.

The point of the pilot isn’t to settle education policy. It’s to settle whether the experimental approach actually works better than the existing approach for this policy domain. If the pilot shows the system is more brittle, more capturable, or less useful than expected, that’s a real finding — it would tell us to refine the design before rolling it out further. If it shows the system produces better outcomes than the alternative, that’s an argument to expand.

One-sentence version

Every law is an experiment: pick the goal, list the options, let the states run them, measure the outcomes, expire the ones that don’t work — then do it again.

That’s the idea. If anyone reading this knows the name of an existing proposal that already does it, I want to hear about it. If nobody does, maybe we should.

#policy#government#science

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.