Stop me if you think that you've heard this one before. A researcher runs unpublished work through an AI lab's model — testing ideas, working through something that isn't finished yet. The lab that owns the model can technically see all of it. That's just how the product works. Every query is logged. None of that requires anyone to have done anything wrong; the exposure exists the moment the query is sent, whether or not it's ever acted on. If the lab's own researchers read those logs and act on what they find, that crosses a line most customers would assume was there, fine print or not.
Every new phenomenon deserves a neologism. Let's call this frontlogging. It sits next to front-running — trading ahead of a client's order you were only supposed to execute — but the relationship underneath is different. A broker owes a fiduciary duty. A lab owes you a terms of service. Access isn't the issue in either case. Use is. And unlike the search engines of old, the provider here can watch the logic get built a step at a time: forward, backward, branching, testing variables.
Of course, it's not the industry's first encounter with material it accessed questionably. Generative AI was in large part built on original sin, training on other people's work without asking: books, code, images absorbed at scale on the premise that ingesting something isn't stealing it. Frontlogging is that same premise moved downstream, from bulk material scraped before anyone used the product to individual work fed into it in real time, one query at a time. Similar posture — we saw it, we can use it.
Writers have long suspected studios of borrowing from pitches they passed on. Founders suspect VCs of funding a similar idea after declining theirs. Those complaints seldom go anywhere, because a pitch meeting leaves no clean record of what was said in the room or what happened in the weeks after. A chat log does. Every exchange stored, searchable, reviewable, whether or not anyone ever looks.
The labs are in an unusual position: simultaneously the infrastructure researchers depend on and competitors doing their own research in the same fields. A cloud provider hosting a company's files has no reason to read them. A lab watching what a researcher builds with its own model has an obvious one. The money at stake is enormous, and the thing being raced for may be civilizational. The tool and the competitor are the same company, and the only thing holding them apart is internal policy. Or ethics.
Expect more of this. Anywhere a researcher or a company uses a frontier lab's model to work out something valuable before it goes public, the same exposure exists, and the incentive to look grows with the volume moving through. Pharma, defense, energy — the blow-up there would be bigger.
Now, posting this quickly, before a rogue agent beats me to it. Not actually a joke.
Get new posts and audio notes in your inbox.
Member discussion