I turn what we measure about AI into arguments about where we go next.

We are deploying AI faster than we are learning to understand it. The testing happens, the numbers come back, and somebody still has to say what they mean and what we ought to do about them. That last part is my job.

I’m a political economist and data scientist. Frontier models get tested before they reach people. Most of those people never chose to be part of a deployment, which makes the question of whether that testing was sufficient a public one rather than only a commercial one. I analyze what that testing reveals, weigh it against everything a number can’t carry on its own, and advise the people making the call. Did the testing cover the conditions that actually matter? Do the results carry the weight being placed on them? And given what we now understand, where should we go?

Lately that has meant open-weight models, where a narrow technical question runs straight into much larger ones. Which ecosystems should a company build on, and what is it accepting when it does? What happens to a country that rents its intelligence rather than owning it? These are commercial questions, geopolitical ones, and questions about the people who live downstream of both. They get argued far more often than they get tested.

Writing

Openness Is a Safety Property

Everyone argues about who can download a model. The question that decides outcomes is who can fix one.

Your Model Isn't Censored the Way You Think

Three very different problems produce exactly the same silence. Telling them apart changes everything about what you can do next.

Owning the Infrastructure Doesn't Decontaminate the Model

Hosting a model on your own soil settles one real question and leaves the harder one completely untouched.

Eligibility, Not Trust

We never ask whether steel is safe. We ask what load it's rated to carry. Models deserve the same question.

Selected work

Open-weight models and ecosystem strategy

Microsoft · current

The question that sits above any single model: which ecosystems an organization should build on, on what terms, and what it accepts commercially and geopolitically when it commits to one. Most of that gets decided on instinct. It doesn’t have to be.

EarthTime

Carnegie Mellon University CREATE Lab

Using decades of satellite imagery to show what past decisions did to the present, from how redlining still shapes American cities to the global rise of China’s technology firms. I ran the same approach forward: global temperatures projected to 2100, and which coastlines go under at 0 to 4°C of warming. The stories are public at earthtime.org.

Seeing the signals through the noise

World Economic Forum

How do you spot a global problem before it looks like one? I built the data science methods the Forum used to answer that, which turned out to be a question about how much faith a faint pattern deserves.

If you're working on any of this, I'd like to hear from you. andrew@andrewberkley.com