HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

Sharing AI progress in mathematics

951 pointsby OfficialTurkey 12 hours ago894 comments

https://github.com/openai/math https://github.com/openai/math/tree/main/preprints

Discussion

Loading discussion
  • senderista · 12 hours ago

    Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.

    • sebzim4500 · 11 hours ago

      This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.

      • senderista · 10 hours ago

        Yeah I'm not sure they met them even halfway.

  • karahime · 12 hours ago

    Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.

    • bravoetch · 12 hours ago

      In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.

      • whimsicalism · 11 hours ago

        Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results. No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.

        • Ancapistani · 7 hours ago

          Can you show where it was discounted? Last I heard OpenAI was declining to deny it, presumably while they thoroughly confirmed.

          • whimsicalism · 6 hours ago

            > “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.” https://openai.com/index/navier-stokes-solution/

            • youoy · 3 hours ago

              Come on, they were working on this for more than 2 months. Dont fall for the corporate half truths.

        • margorczynski · 52 minutes ago

          I really don't get how people are still continuing with this "stolen results" narrative after today. Like NS was kinda insignificant compared to treasure trove they released now, thinking that the LLM needs to "steal" from some human is simply coping.

    • xpct · 12 hours ago

      Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press. There's no gatekeeping here!

    • hgoel · 11 hours ago

      After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time. We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work). If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.

      • make3 · 10 hours ago

        It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic

        • hgoel · 9 hours ago

          I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.

    • kzrdude · 6 hours ago

      They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.

  • ravenical · 12 hours ago

    https://github.com/openai/math

  • binlog · 12 hours ago

    So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.

    • fph · 12 hours ago

      Most mathematical results are shared on Arxiv. Journals add peer review.

    • traes · 12 hours ago

      GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.

    • adverbly · 11 hours ago

      End of an age for journals?

  • rafterydj · 12 hours ago

    I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.

    • osiris970 · 12 hours ago

      You want them to stop doing math research?

  • k2xl · 12 hours ago

    Can someone knowledgeable about the subject outline the most significant portions of the results?

  • sebmellen · 12 hours ago

    It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

    • ndriscoll · 12 hours ago

      > Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability! No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.

    • adverbly · 11 hours ago

      > Look at one of their examples of an initial prompt Interesting that its only an excerpt. I wonder what else they include but didn't share.

      • philipwhiuk · 10 hours ago

        Attempts to edit the problem description on Wikipedia ;) https://wikimediafoundation.org/news/2026/10/05/openai-rogue...

    • cubefox · 5 hours ago

      These are not reasoning traces, these are summaries of excerpts of reasoning traces.

  • ed · 12 hours ago

    Actual results: https://github.com/openai/math/blob/main/overview.pdf

    • ks2048 · 12 hours ago

      HTML version, https://github.com/openai/math/blob/main/CONTENTS.md

  • gizmodo59 · 12 hours ago

    This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans

    • fspeech · 12 hours ago

      Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.

      • gizmodo59 · 12 hours ago

        >So until we can comprehend it there really isn't much progress. Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as

        • fspeech · 12 hours ago

          If it changes how we think then yes it has an effect.

        • le-mark · 10 hours ago

          But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.

      • fspeech · 12 hours ago

        Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.

      • warkdarrior · 12 hours ago

        > Math theorems are tautologies Proven math theorems are tautologies.

        • fspeech · 12 hours ago

          True.

        • fspeech · 12 hours ago

          FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.

      • binlog · 12 hours ago

        What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.

        • fspeech · 11 hours ago

          I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.

      • gpt5 · 12 hours ago

        Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet. We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.

        • fspeech · 11 hours ago

          This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.

      • caaqil · 11 hours ago

        > until we can comprehend it there really isn't much progress Who is "we" here exactly?

        • fspeech · 11 hours ago

          Whoever wants to study the result.

          • caaqil · 11 hours ago

            > Whoever wants to study the result. Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee. Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.

            • fspeech · 11 hours ago

              AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.

      • yieldcrv · 11 hours ago

        Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants

        • fspeech · 11 hours ago

          If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.

      • fspeech · 11 hours ago

        I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/... He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.

        • fspeech · 9 hours ago

          Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve. BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.

      • AIblemblio · 1 hour ago

        It is progress on another / the next evolutionary later: A AI/AGI/ASI system. Which either replaces us in the long term, augments us or makes us better (gentherapy).

    • traes · 12 hours ago

      Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.

      • xpct · 12 hours ago

        I just did a quick search on this and apparently the misspellings are German surnames as well: https://en.wikipedia.org/wiki/Reimann https://en.wikipedia.org/wiki/Reinmann

        • conformist · 11 hours ago

          Yes sure but they are different surnames and pronounced differently.

          • xpct · 11 hours ago

            I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!

        • traes · 11 hours ago

          I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?

          • NewsaHackO · 11 hours ago

            People just don’t spell that seriously buddy, especially when it is so immaterial to the point.

            • traes · 11 hours ago

              My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.

              • xanderlewis · 11 hours ago

                You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.

          • ndriscoll · 11 hours ago

            Maybe they skipped straight to Lebeg integrals.

            • raegis · 4 hours ago

              Thanks for the laugh!

            • thunspa · 1 hour ago

              very good lol

          • jryb · 10 hours ago

            Autocorrect might be doing it

          • paulhebert · 10 hours ago

            My last name is Hebert. There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert. Even in situations where they just read it or I just said it. I’ve had Herbert soccer trophies, health insurance cards, etc. The mind fills in a lot of blanks and doesnt always get them right.

            • jbaber · 10 hours ago

              I sympathize. -- Not Barber

            • Agentlien · 3 hours ago

              My last name is Kvick - the Swedish word for quick. I live in Sweden. I get a lot of people thinking my name is spelled Kvik, Quick, Kwick, ... My father once got a mail addressed to Mr. Kvack (Swedish for quack, like a duck).

          • neutronicus · 8 hours ago

            iPhone would be my guess

      • lanyard-textile · 11 hours ago

        They're mathematicians, not linguists :)

        • traes · 11 hours ago

          The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.

          • pixl97 · 11 hours ago

            Uh oh, no true scottsman....

            • vector_spaces · 10 hours ago

              It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.

              • quacktopia · 6 hours ago

                I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic. Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something. Google existed then and now and we could look them up if needed.

          • gpm · 10 hours ago

            One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too. Mathematicians aren't exactly known for being well rounded.

            • jwilber · 9 hours ago

              Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.

              • senderista · 6 hours ago

                google "Grothendieck prime"

            • scrame · 8 hours ago

              Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.

          • bananaflag · 5 hours ago

            Still, enough misspell Lebesgue as Lebesque.

        • bootsmann · 4 hours ago

          If they’re mathematicians they have written this name down about a 100 different times throughout a standard Real Analysis course. Riemann was foundational in that field.

      • jere · 11 hours ago

        “How many Ns in Riemann?”

        • sdenton4 · 10 hours ago

          Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.

    • zone411 · 11 hours ago

      There was A LOT of drama about this release.

    • cyclopeanutopia · 4 hours ago

      > Point the repo to your agent and ask for the significance! Wow, this comment really shows how low this community fell.

    • robotpepi · 7 minutes ago

      > This is significant progress and released without all the drama. I feel gaslighted.

  • enoether · 12 hours ago

    Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal! [0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

    • impossiblefork · 12 hours ago

      Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.

      • davemp · 9 hours ago

        TCS being theoretical computer science? I have not seen that acronym before.

        • jhanschoo · 9 hours ago

          Yes, TCS is theoretical computer science, I commonly use that acronym too.

    • gregdeon · 11 hours ago

      This was the biggest highlight for me as well. Astounding...

    • inkysigma · 11 hours ago

      I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.

    • amluto · 11 hours ago

      I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem: > A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)). I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it. 1. e is maybe a name of a list. 2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks. 3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k. 4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e , it's for some e, and the goal is to count them. So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple. Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this. If this were my paper, or if I were trying to train a model to write math , I'd want something like: A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third. A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).

      • danbruc · 1 hour ago

        […] a finite edge set E = (V × V) […] E ⊆ V × V

  • Catloafdev · 12 hours ago

    This is a pretty hilarious thing to read juxtaposed with AGMAI's requests. Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

  • open592 · 12 hours ago

    Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do? Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

    • goalieca · 12 hours ago

      Don’t paste your research into these AI because they will train on it and then scoop you.

      • esafak · 12 hours ago

        I think that happened after word of the project reached OpenAI and they allocated resources to it.

    • binlog · 12 hours ago

      Use whatever is published as the new base for your research. Use AI tools to help you going forward.

      • xpct · 12 hours ago

        In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD. It has to feel awful to be in this position.

        • torben-friis · 12 hours ago

          Could be worse, imagine having years of experience in a profession these things can now handle by themselves. :)

          • jltsiren · 10 hours ago

            It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.

    • caaqil · 12 hours ago

      > what do I do? Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

    • aaraujo002 · 12 hours ago

      This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.

      • CaptainNegative · 9 hours ago

        Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...). It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago. There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

    • dcl · 12 hours ago

      This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.

    • dekhn · 12 hours ago

      Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.

      • thimotedupuch · 11 hours ago

        Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?

        • dekhn · 11 hours ago

          No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen ( https://research.google/blog/groundbreaking-simulations-by-g... ). My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

          • vasco · 10 hours ago

            So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.

            • globular-toast · 12 minutes ago

              That, and also it's just a completely different approach which might later on turn out to be useful. People should remember that artificial neural networks were developed decades before they were useful. People were doing all kinds of other approaches to ML like support vector machines before advances in hardware made deep neural nets feasible and therefore interesting again. ANNs were never obsoleted by SVMs.

      • boznz · 9 hours ago

        For every door that shuts another one opens - great if you're not a cabinet-maker.

    • vinyl7 · 12 hours ago

      Look forward to being obsolete I guess

    • bobmarleybiceps · 12 hours ago

      I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/

    • moralestapia · 11 hours ago

      That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).

    • hgoel · 11 hours ago

      It could still be interesting if your approach to the problem was different to theirs.

    • claaams · 11 hours ago

      Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.

    • yieldcrv · 11 hours ago

      Yes, and?

    • glitchc · 11 hours ago

      Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.

    • ex-aws-dude · 11 hours ago

      That’s always been a thing, it’s called “getting scooped”

      • vouaobrasil · 10 hours ago

        Killing with knives has always been a thing. Now, we have the machine gun.

        • ex-aws-dude · 8 hours ago

          The scoop gun

    • bamboozled · 10 hours ago

      Ask OpenAI for money when you don't have a job or future? I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

    • vouaobrasil · 10 hours ago

      > Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do? I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.

      • s3graham · 8 hours ago

        You might enjoy https://asteriskmag.com/issues/15/so-you-think-you-could-be-... if you didn't see it recently.

    • pratikdeoghare · 10 hours ago

      > what do I do? Very hard question. Your work makes you one of the very few people who really understands the problem and solution and its significance.

    • netsec_burn · 9 hours ago

      Verification is equally important, if not more so.

    • porcoda · 6 hours ago

      As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time. What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results. I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.

      • rubikscube09 · 4 hours ago

        math will just be black boxed away. no one will "need" to understand it.

    • katatue · 5 hours ago

      At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).

    • DCKP · 2 hours ago

      I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication. A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.

    • rfgplk · 2 hours ago

      Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.

  • ks2048 · 12 hours ago

    I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)

    • xpct · 12 hours ago

      Presumably they don't because they're training the audience (us) to trust the machine , not its verifiers, even if they were included.

      • alexgoodhart · 11 hours ago

        I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.

    • agnosticmantis · 11 hours ago

      Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}. 1: Author 2: Verifier /s

    • chiwilliams · 11 hours ago

      There are competitive reasons that they don't want to share all the people on the team.

    • chrisjj · 11 hours ago

      > I think they should put human names on the papers as someone who has reviewed the result Assume the empty list you see is complete. :)

    • make3 · 10 hours ago

      I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"

      • ks2048 · 9 hours ago

        Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe". With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).

    • procedurecall · 9 hours ago

      Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.

      • kzrdude · 6 hours ago

        And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.

        • rubikscube09 · 4 hours ago

          the reference group can recommend all they want, no one will review 700 plus papers.

  • aaraujo002 · 12 hours ago

    The Advisory Group states in its recommendations [1]: "We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too? [1] https://agmai.org/general-sep29/

    • osiris970 · 12 hours ago

      Comical ask

    • medler · 12 hours ago

      The rest of that document makes a pretty compelling case for why this is a bad practice

      • esafak · 12 hours ago

        I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.

        • pixl97 · 7 hours ago

          Then make 2 AI's and force them to challenge each other.

    • warkdarrior · 12 hours ago

      The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians. > "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]" https://mathstodon.xyz/@tao/117395269325940185

      • binlog · 7 hours ago

        There are a dozen+ AIs available to you that can do that right now.

    • jhrmnn · 12 hours ago

      It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.

      • andriy_koval · 6 hours ago

        without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.

    • mattr03 · 12 hours ago

      I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.

      • bravoetch · 11 hours ago

        > Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. It's been a while since I was reminded of this xkcd: https://xkcd.com/435/

      • zeroonetwothree · 10 hours ago

        Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.

    • fph · 12 hours ago

      ...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)

      • Yamata · 6 hours ago

        It harms lives? How so?

    • tchalla · 12 hours ago

      Why did you leave out the entire quote? > At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models. To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.

      • aaraujo002 · 12 hours ago

        Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?

        • adrian_m · 11 hours ago

          The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.

          • agnosticmantis · 11 hours ago

            Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access? These models are too expensive for broad access unfortunately.

            • Jtarii · 11 hours ago

              ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.

              • Jweb_Guru · 10 hours ago

                Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.

          • strange_quark · 10 hours ago

            I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results. They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.

            • blurbleblurble · 7 hours ago

              It's honestly shit marketing that only cultivates spite and erodes their whole brand. This is a total ego trip.

        • TeeWEE · 9 hours ago

          The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish

    • perching_aix · 12 hours ago

      The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick. How this maps back to math, idk.

    • bmitc · 11 hours ago

      Advocating purely for progress and not humanitarian value is how we'll all get enslaved.

  • pavitheran · 12 hours ago

    From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

    • password54321 · 12 hours ago

      Oh cool, we will all now have a math genius on our computer.

      • an0malous · 11 hours ago

        Well, on their computers. But you can rent them for a price.

        • binlog · 7 hours ago

          An open source model will reproduce it 6 months later

      • jrflo · 11 hours ago

        It was using their internal math model, so not yet for us

        • password54321 · 11 hours ago

          I used future tense. It was implied this will be available.

    • scrlk · 12 hours ago

      Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?

      • inferencecoder · 11 hours ago

        It doesn't imply that, it's just measuring the amount of compute.

        • bigmadshoe · 10 hours ago

          But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?

          • inferencecoder · 7 hours ago

            Not necessarily, could be agent swarm with low N

            • bigmadshoe · 6 hours ago

              At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.

    • orlp · 11 hours ago

      I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost. Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

      • timjver · 11 hours ago

        >OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...] That doesn't sound right

        • orlp · 11 hours ago

          Oops, edited.

      • machomaster · 11 hours ago

        They did say that. "3 hours of ChatGPT Pro thinking compute"

        • orlp · 11 hours ago

          Yes, what does that mean?

          • mh- · 5 hours ago

            It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours. If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.

      • pixl97 · 11 hours ago

        Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.

        • orlp · 11 hours ago

          I'm not denying that, but I'd still like to know what that cost.

    • Jtarii · 11 hours ago

      That estimate is obviously going to conveniently ignore all the failed runs.

  • mathisfun123 · 12 hours ago

    With so many results in so many different areas no way they even remotely spot checked well enough. Prediction: one of these is wrong and this (publicity stunt) will backfire. Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

    • bravoetch · 12 hours ago

      What does a backfire look like? It's ok to be wrong in the science/math world.

      • mathisfun123 · 12 hours ago

        of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.

        • bravoetch · 11 hours ago

          Do they claim that's the case? I don't think they do.

          • mathisfun123 · 11 hours ago

            does company A making product B claim that the product is robust and consistent? is this a serious question?

            • zamadatix · 11 hours ago

              If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space. The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.

    • jojva · 11 hours ago

      You have not read their readme: > Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.

      • mathisfun123 · 11 hours ago

        i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.

        • stevenhuang · 10 hours ago

          I don't think anyone would particularly care if only one of them is wrong, if most are correct. If they are all wrong, that's when it would backfire.

    • orlp · 10 hours ago

      It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...

  • dekhn · 12 hours ago

    I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions ( https://en.wikipedia.org/wiki/Compressed_sensing ). I am curious if any of the results have immediate applications in any kind of engineering or science. It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

    • brandonpelfrey · 10 hours ago

      Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.

      • OutOfHere · 10 hours ago

        Please share your findings.

    • qnleigh · 4 hours ago

      I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment. From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.