Comment Re:Benchmarks (Score 1) 25
Google said already in 2019 that they moved their focus from gaming to other areas.
Google said already in 2019 that they moved their focus from gaming to other areas.
Medical research is not a finite resource. The more money you have on the field, the more people will join the field. And pretty much any research you do, can benefit other research in unexpected ways. There are even some worm researchers and yogurt researchers who have made significant contributions to medical research.
Looking at the benchmarks it looks like it is better than the previous version, so that is probably why they released it. So I don't think that it is rushed.
"One employee built an app that turned a nearly four-hour manual process into a one-second automated check, saving tens of thousands of hours annually. "
This is the reason why AI will take jobs. World is full of tasks that are dead easy to automate, but no-one has just not done it yet. AI doesn't have to be smart, it doesn't even need to be good. If it can write a script that compiles, it can automate.
In 2025 Claude had pass rate of 2% in FrontierMath. In 2026 Claude had 88%. Perhaps we should just wait a few years and see if the pass rates stop increasing?
Source:
https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fepoch.ai%2Ffrontiermath%2F...
Grades increase extrinsic motivation, but they often hurt intrinsic motivation. I've been thinking that we should replace grades with a single percentage number per subject. This number would tell that you have learned for example 12% of math. This percentage would be calculated automatically based on your answers on controlled environment (tests or class room work). Behind this number would be actual data that shows exactly the things you have not yet learned, so teaching could be targeted there. But as tests would sometimes test past knowledge also, the number might drop also.
The benefit of this single number would be that it would be very simple way to tell how much knowledge a person has and long term learning would provide better scores than studying just for the test.
"within a century (by 2030), rapid technological progress would create so much wealth and productivity that people would only need to work a 15-hour workweek"
You have to wait 4 more years before you can claim that he was wrong. And note that he didn't say that people would work 15 hour weeks, he said that would need to. I actually could do that, I could work for 15 hours only and live with that, but I have decided to work more and retire early.
Actual quote was "by 2025, AI will allow developers to complete a week’s worth of work in four days".
There are plenty of cases where AI has been used to train another AI, usually to make lighter and cheaper AI, which is in some cases better than original. For example AlphaZero was trained this way. There are some examples where humans have done similar things, one European young mathematician comes to mind for example.
No. The real issue is the double underscore. If double underscore had not been allowed in the username, this mistake would have never happened. This is an usability error in the software.
There are plenty of things they could do.
1. You can hire people to generate content. This has been often done with medical LLM training.
2. You can generate content. This can be done when training LLM to generate code or math. Outside of LLM, it has also been done in medical imaging training.
3. You can make smarter LLM that doesn't need that much training data (this is what Google Deepmind aims to do, not OpenAI, but it is an alternative for them also)
One time might be enough. Even if all AI development would stop now, as in these are the models we got and no new models are created, we could still do a ton of work with the existing models. Much more than what we currently can. We are constantly learning new ways to get better results from lesser models so AI performance would continue to improve, even if the models themselves would not improve.
Like, think of a stupid student who always gets an F in a test. What if you would let him do the test with a text book? With the text book, he might get a D. The student didn't get any smarter, nor did he learn anything, but with a small trick, the output was improved.
And even if LLM would be hype, the medical classifiers certainly are not. The usage of classifiers in real medical work is constantly increasing and we are getting results that show that they actually work. It is not hype, it is just poorly communicated to the general public.
That is not true. I did actually read the old encyclopedia in my parent's book shelf, well not the whole book, but the technology parts there as I was very interested in it. Since then I have found much better source for the information and that has been youtube. Yes, youtube is mostly unnecessary things, but as long as you can find it, it contains incredibly detailed information on niche things in a format that is much more easy to understand than it books. Books usually leave a lot of things out and same can be true for videos also, but there are also uncut videos where you can see the whole process and that is extremely valuable if you want to learn something really well.
Go visit your friends. pick a board game or D&D and play with them. Or go outside and play some basketball or anything with them. It is almost free. There are a lot of things you can do with little or no money.
While you are correct, there isn't actually any reason why it should be like that (other than human behavior). We could arrange everything so that everyone just works less, spends less and has more free time. People won't do this willingly, so you would have to use either law or taxes to force people to divide work. I am not sure if this is a smart thing to do, but it is an option.
Asynchronous inputs are at the root of our race problems. -- D. Winker and F. Prosser