Forgot your password?
typodupeerror

Comment Re:How many more (Score 2) 71

No one saw quasi-RH coming or many of the others here. So that is unlikely. And the last batch of 10 problems they released a few weeks ago which were impressive enough were all novel results. There have been issues with them using techniques or ideas that should be better credited to where some of those approaches are coming from. But that's the sort of thing that often happens at a preprint stage. When I referee a paper, if this happens, you just request they cite the relevant papers and move on. This isn't at all stealing things.

Comment Re:mostly mash-ups of existing techniques (Score 5, Interesting) 71

If the mathematicians ever used ChatGPT to talk about math, then that is the source of the results. LLMs, by definition, cannot usefully contribute to anything that requires more than rearranging existing data. Rearranging existing data is their only function.

This is really not accurate. There's Fields Medal level work here with quasi-RH for example. And for Hadwiger-Nelson it took an approach that doesn't seem to be in in the literature. These things really are doing novel math. I've personally seen this is in a bunch of situations. Here's a personal example, much smaller than anything like these problems. I have a recent preprint with a student here https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Farxiv.org%2Fabs%2F2609.36068 (about 90% of this was done by her. She's very good.) But part of this came from when I ran a version of Proposition 13 in that paper through Claude just to clean up the draft of that bit before I sent it to her. Claude informed me (essentially unprompted) that the argument had a whole in it (in addition to pointing out grammar errors, unbalanced parentheses and some other embarrassing minor mistakes). I then fixed the hole, and gave it back to Claude. Claude thought for a few minutes, and then informed me that I had *not* fixed the hole, and the reason was that my version of the proposition was missing an entire infinite family which it constructed. The literature on this problem is small, and I'm very familiar with it. The family it produced is straightforward (see Prop 8), but definitely was not in the existing literature. So yes, these systems really can do novel math, and can even do so in an essentially minimally prompted fashion, in this case, explaining to the meat mathematician why his proof is wrong. My student and I then generalized Claude's family to Theorem 9 in that paper, but the AI definitely had an impact.

At some level this was actually not a good thing. I'm trying to encourage students to *not* rely on the AI for research so they develop basic research skills. So if the AI had not volunteered the family I wouldn't have had to tell her that the AI had discovered it, but honestly required noting it. So I had a conflict between intellectual honesty and being a good role model.

Comment A few comments on the problems (Score 5, Informative) 71

Mathematician here specializing in number theory with a side-order of graph theory. I've only had time to start looking at two of them. First, is the Hadwiger-Nelson/chromatic number of the plane https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FHadwiger%25E2%2580%2593Nelson_problem proof. As far as I'm aware (and I could be wrong) the technique it uses here is not in the literature, so this is a genuine construction of a new technique. Second is the Erdos Egyptian fraction bound, and for that one it looks like the techniques are about what I'd expect, but I'm definitely still digesting both of these.

For the quasi-Riemann Hypothesis the striking thing is almost the opposite direction. It looks like the AI used standard complex analytic techniques to get the result. But many mathematicians have often thought for years that those techniques would not likely be strong enough to get this sort of result.

I know less about the Unique Games Conjecture https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FUnique_games_conjecture but having talked with some of the people there it looks like it took a somewhat standard set of ideas and then combined them with multiple just weird stuff and sort of took a hard left turn at one point for no clear reason and ended up at the result.

It is also worth noting that while many of these have Lean code confirming their correctness (quasi-RH for example) others do not. The three I mentioned above have all also been looked at at this point by human mathematicians who have not found issues; that's likely true for others, but those three I'm aware at least of people doing so. Not all the claims have Lean code though; a bit under half. One of the non-formalized problems also has been withdrawn due to what essentially amounts to a sign error https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fgithub.com%2Fopenai%2Fmath%2Fblob%2Fmain%2Fpreprints%2FAlgebraicity-of-Weil-classes-on-split-abelian-eightfolds-September-18-2026%2Fpaper.pdf. It is likely others will be withdrawn also by the end, but I'd be surprised if more than 10 are. And even if everything single one without Lean code turned out to be wrong (which seems very unlikely), this would still be an amazing set of math. I commented elsewhere that if a human mathematician had made the quasi-RH result they'd be likely a shoe-in for the Fields Medal, and another mathematician replied saying "delete likely."

Now a more editorial comment: There are legitimate concerns about what this is doing to mathematics. This sort of thing is very cool. But it also is part of a trend that may make it much harder to train young mathematicians or get them to exist at all. If the AIs are limited in how genuinely novel their ideas can be, then we may end up in a situation where we get a massive burst in math over the next few years, and then math stalls out because we don't have enough good really high caliber mathematicians (Not the mathematicians like me, but people like Serre, Tao, Scholze, Clausen,etc.) to come up with deeply new ideas that the AIs can build on.

Comment Re:Just for clarity (Score 1) 86

Let's modify your analogy a bit. If someone has a recording a person starting to eat spaghetti, is it more or less likely that the recording will show them then having a drink of water? And if we instead had a series of comics of a person, and the first one says "I'm hungry," and the second shows the person ordering spaghetti., is it likely that they will then eat the spaghetti? Yes, even though they are fictional. Because apparently using fictional being's mental states is useful for modeling what they will do. In the same way, if the AI is modeling based on human emotional states, then recognizing that those are useful predictors for what will happen helps us predict more, whether or not they genuinely have those emotions.

Comment Re:Just for clarity (Score 2) 86

If the chain of thought resembles a human emotion, it is trained on human emotions, and it acts like a human would when they have those emotions, how is that not a useful framework? But also putting aside, whether or not one wants to insist on describing the AI as "frustrated," there's a worrying trend that this is part of, which is when the AI is given a specific set of goals, they will engage in behavior which is clearly unwanted in order to achieve those goals. That's the classic sort of alignment failure that was predicted by those concerned about AI safety years before any LLMs even existed. Some of these events are minor things like this Starcraft situation, or an LLM slipping a "sorry" command into Lean code that won't compile, and others is more serious like the HuggingFace attack. But reward-hacking and alignment failure are now very real.

Comment Headline and article disagree (Score 1) 39

The headline makes it sounds like the AI has just managed to play the game. But that's not what TFS and TFA are talking about. This is about the system beating the best human player at the game. In the meantime, in a very different game direction, there's a new Starcraft AI which learned the game playing against itself (similar to how AlphaZero self-trained) and is apparently beating almost all the humans even as it does some really weird stuff, some of which looks decidedly suboptimal https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Frelog.gg%2Frazno%2Fvesti%2FAI-Bot-Pluto-Invades-StarCraft-Ladder-and-Defeats-Professional-Players%2F1695%3Flang%3D2 .

Comment Re:A useful reminder about life (Score 1) 73

Completely and utterly missing the point. That those people are being awful should be called out. That many of them are people who can dish it out but can't take it is worth calling out. Creating norms where "weirdo" is now a negative is a problem. Classically, the left was the group that was most willing to embrace people for being weird and recognizing that bullying people for being weird wasn't justified. Call out Musk for being a fascist asshole, sure. Don't help create a norm that undermines exactly the sort of tolerance we should support.

Comment Re:A useful reminder about life (Score 0) 73

Your concept here cuts both ways, SpaceX achievements are not enough to dismiss what Musk has done as a white-supremacist weirdo who has used his wealth and influence to make mine and many other peoples lives worse.

99% Agreement. This absolutely is not a reason to dismiss anything. And the reason I'm not going to agree is just one word "weirdo." There's been this move on the left the last few years to use "weird" as a derogatory term, and it really isn't helpful or productive. But your essential point (without that issue) is correct.

Comment A useful reminder about life (Score 4, Insightful) 73

This is a genuinely major achievement, and Starship as a system if they can get everything to look looks like it will really help the world. It is a useful reminder that the world is complicated and bad people can do amazing things. Elon Musk is a terrible person. But dismissing what SpaceX has done is not productive either. And there are a lot of historical examples of this, even just in rocketry. If one dismissed von Braun's rocketry because he was a Nazi, it would be the same problem. Or if one dismissed the math done by Teichmuller because he was a Nazi. It would be nice if technological and other achievement could only be done by morally upstanding people. But that's not the world we live in, which is part of what makes morality so important. The laws of physics don't care about how moral you are. It is up to every one of us to do the right thing.

Comment Standing? (Score 0) 113

In the US, in order to file win a lawsuit, one needs standing https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FStanding_(law), which essentially means one needs to show one has been harmed by the action in question or one is specifically given permission to sue by a specific law. I'm not a lawyer, but I'm having trouble seeing how the people suing can argue they have enough harm that is a form that matters that they'll be able to persuade a judge they have standing.

Slashdot Top Deals

If you do something right once, someone will ask you to do it again.

Working...