Martín Morales

On OpenAI plagiarism

 ·  Spanish version

A new controversy about OpenAI and one of the Millennium math problems happened recently. In case you didn’t notice, in short: two mathematicians were investigating the Navier-Stokes problem using OpenAI’s Codex. One of them works for Anthropic. OpenAI, investing a lot of resources, declared they had resolved the problem on their own with an internal model, published the finding, and claimed the discovery as their own.

This case is plagiarism because one of the mathematicians suspects that the model could have been trained on the content of his private Codex sessions, without his authorization. When he asked about it, OpenAI only confirmed that Codex doesn’t access user data, without clarifying whether that data was used in some other way to train the model. Besides, OpenAI proposed to give the finding’s credit only to the mathematician who doesn’t work with the competitor, but this guy refused.

I leave the news with more details here.

However, AI has reached a point of collaboration with people in a fascinating way. It has improved processes and found solutions that benefit us. But the corporate side of this has proven to be insensitive and grasping, looking to get credit at any cost.

This makes me wonder if privacy is really respected when we use these tools. And what would happen if another discovery like this occurred?

I understand that competition pushed OpenAI to reach this milestone, but it seems too selfish to me to leave these people aside just because they’re a “competitor”. In fact, when reviewing the “about” section on OpenAI’s site, it states:

OpenAI is an AI research and deployment company. Our mission is to ensure that artificial general intelligence benefits all of humanity.

The mission they claim to be so proud of conflicts with what they did this time, and this says more about their real priorities than any “about”. Meanwhile, we need to be more careful about what we share on these tools: it is not always clear who gets the credit—or the data.