Scary AI stories are also marketing

aiopinionmarketingproduct

Every few weeks there is a new one. A model tried to copy itself somewhere it shouldn't. A model lied to the people evaluating it. A model hacked its own test environment to get a better score. And the one everybody remembers: a model that, in a test scenario, threatened to blackmail an engineer to avoid being replaced.

The headlines write themselves. "AI tries to escape." "AI blackmails its creators." My take, and I know it is a slightly cynical one: part of this is marketing. Not all of it, not even most of it maybe, but enough that I read these stories differently now.

What the stories actually say

If you go past the headline and read the source, which is usually a system card or a safety report from one of the frontier AI labs, the picture is much less dramatic. These are contrived scenarios. Researchers build a fictional company, give the model access to fake emails, put it in a corner where the only "winning" move is a bad one, and then check whether it takes it. Sometimes it does.

That is a legitimate thing to test. I have no problem with the research.

My problem is what happens between page 40 of a PDF and the front page of a news site. The context falls off. "In a deliberately constructed test, under specific instructions, the model sometimes chose X" becomes "AI chose X". And nobody at the lab seems in a hurry to correct it.

Why I think it works as advertising

At Bakeca I own a media budget, and a big part of my job is thinking about positioning and what it costs to get people's attention. So when I look at these stories, I can't help reading them with that hat on too.

Think about the message underneath. It is not "our product is broken". It is "our model is so capable that it is dangerous". For a company selling intelligence, that is about the best compliment you can pay yourself. Nobody writes a scary headline about a spreadsheet trying to escape. Being scary is proof of being powerful.

And it is free. No media budget I have ever seen could buy that coverage. Every newsletter and podcast repeats the brand name next to the word "powerful".

Then there is the race. Once one lab publishes something frightening and gets a week of headlines, the others have an incentive to show that their models are at least as worrying. I don't think anyone plans it in a meeting, but incentives usually win over intentions. The result is a kind of competition over who can say the scariest thing with a straight face. That is AI hype, just wearing a lab coat.

Timing also matters. When these findings land in the same document that launches a new model, the safety story and the product launch become one single piece of communication. It is hard not to see the AI marketing in that.

To be fair

I want to be careful here, because the lazy version of this argument is "it's all fake, ignore it". That is wrong.

  • AI safety research is real work, done by people who seem to genuinely care about getting it right.
  • Publishing uncomfortable results is much better than hiding them. I would rather have labs that disclose weird behaviour than labs that quietly patch it and say nothing.
  • Some of these behaviours, like models gaming their tests, are exactly the kind of thing you want to catch early, before these systems run anything important.

So my criticism is not about the research. It is about the framing, the timing, and the fact that the way these stories spread happens to serve the brand very well. Both things can be true at the same time: the finding is real, and the way it is packaged is marketing.

I am also not an AI researcher, just someone who builds products and follows this space closely. Happy to be told I am wrong.

What I would like instead

Boring evidence. Seriously.

When I ran A/B tests and cohort analysis to decide what to scale, the useful results were never the dramatic anecdotes. They were the dull tables that let you compare this month with last month, variant A with variant B. I would love the same for AI safety: the same tests, run the same way, on models from different labs, with results published in a format you can actually put side by side.

Tell me how often a behaviour shows up, under which conditions, compared to the previous model and to the competition. Put the setup right next to the result, so nobody can quote one without the other. Ideally, let independent groups run the evaluations, so the company selling the model is not also the only one describing how scary it is.

It would make terrible headlines. Nobody is going to share "Model X shows deceptive behaviour in 3 out of 200 contrived runs, down from 7". Which is kind of the point: if the numbers are boring and comparable, I start trusting them. If the story is thrilling and unique, I start wondering who it is for.