OpenAI on Tuesday, Oct. 6, published 722 mathematics papers that it says were produced by an unreleased internal AI model. The papers sit in a public collection on GitHub, a site where people share code and files, and are grouped into 372 families of related papers, which can include a main result along with companion arguments, consequences or alternative proofs. OpenAI says it tests its models on open research problems, questions mathematicians have not yet answered, and widened those tests after its models aced its existing math tests.
OpenAI says the model was given about 4,000 problems, and that the collection keeps the results it judged significant enough. On average, each result used about as much computing power as three hours of thinking on ChatGPT Pro, one of its paid plans. OpenAI says the vast majority of results came from the same standard procedure with that model. It names two exceptions, and says people edited the write-up of one of them for readability.
Many of the proofs, though not all, come with versions in Lean, a programming language that lets a computer check a proof step by step. Proofs checked this way are all but certain to be correct, according to a news report. OpenAI says the results are at different stages of checking, that some of the unchecked ones could have problems, and that it will try to fix them quickly. Many of the results are not yet understood even by OpenAI's own mathematicians, a company spokesperson said, according to a news report. It also published summaries of how the model reasoned its way to 10 of the results, and figures on how many problems it tried.
The release comes a week after recommendations from the Advisory Group on Mathematics and Artificial Intelligence, an independent group at the Institute for Advanced Study. On Sept. 29, after a survey of mathematicians that drew more than 600 replies, the group asked AI companies to stop testing hard math problems on models outside mathematicians' reach, to put results in collections no AI company controls, and to publish the model's name, prompts, time and cost for each result. OpenAI says it drew on that advice. Its papers sit on OpenAI's own GitHub account for now, while it looks at community-run alternatives, the model has no public name, and it gave an average cost figure rather than one for each result. OpenAI also says it will pay for workshops and conferences to help mathematicians understand major results produced by AI.
Mathematicians are likely to need months to work through the collection and judge whether the proofs hold new ideas or mostly recombine known methods, according to a news report. OpenAI says it is working to release the model responsibly but has not said when, or on what terms.