Comment Re:Meanwhile, back at the benchmarks (Score 1) 19
("It actually makes RAM needs worse" -> "MoEs actually make RAM needs worse")
("It actually makes RAM needs worse" -> "MoEs actually make RAM needs worse")
It actually makes RAM needs worse than a dense model, for a given quality.
That said, inference techniques like FreeToken are improving MoE swapping performance, so it's not as bad as it once was. Honestly, I would not be at all surprised if we start doing training in a swap-aware manner, where expert cache misses count as loss and models can queue experts to start loading before they're needed. Could even train in two tiers of loading - RAM and disk. Wouldn't that be great? Could run multi-terabyte models on consumer hardware. Training would tend to tend to concentrate key logic and reasoning on small number of very active experts, mixed reasoning/knowledge on less common experts that are usually kept in RAM, and rarer knowledge on experts that usually remain on-disk until needed.
There are well known, reputable abliterators on HuggingFace.
Also, censorship usually (not always) is something you care about for chats, not agentic work.
A
It's big, but not good for its size. They're using tricks to pretend that they're better than they are. For example, compare the numbers that they list for the competion on DeepSWE up against the actual DeepSWE scores.
Not cool, Mistral.
Artificial Analysis is overrated, but yes, Mistral is playing fast and loose with their claims.
It takes one hell of a toll to try to do verbal exams 1:1 on students to evaluate them.
Nobody said "verbal exams 1:1".
The only thing that is required is that they be separated from their computers and their phones.
Nobody, I repeat, nobody who uses AI is going to see "OpenAI adding watermarks" and think, "Oh, I better use AI more, especially OpenAI".
And nobody, I repeat, nobody who doesn't use AI is going to see "OpenAI adding watermarks" and think, "Oh, I better use AI more, now that it'll be easier to see that what I did is made by AI."
It is not "an advertisement of the power of the technology". It's a regulation forced on them by the EU's AI Act, which neither companies nor users want.
Public records requests are essential for establishing cases of malfeasance. However, with the rise of AI making giving every person the ability to make unlimited Public Records requests, the burden can be extreme.
I work for a major public university and we are absolutely drowning in PR requests. Anyone who does something that pisses off another person gets hit with a PR request... or 30. There are hundreds in queue and it's preventing actual work from being done. Where I live, there are not protections against malicious, repeat, or unjustifiable requests. There have been no adjustments for AI. There is no requirement to submit to a specific location or in a particular format.
Our campus certainly spends over $1M on PR responses given the time it takes for qualified people to seek and prepare the data and communication.
Of the MANY PR requests I've fulfilled, only two have been for decent reasons. The rest have been to satisfy vendettas and nothing has ever come of the data provided-- even for the good reasons.
How does this relate to the obvious issues around the use of Flock cameras? I'm willing to bet that the PR requests have gone something like this:
1. All records of all license plate scans since the inception of the program
2. All communication (digital, printed, or otherwise) from every employee discussing the system be it emails, text messages, written notes, etc.
3. All documents related to the purchase/contracting process that resulted in the Flock vendor being selected
4. All reports
And it's very likely multiple requesters submitting requests for information in different formats. This takes worker time to aggregate.
Perhaps you should be realizing by now that homework is no longer a viable way to test how well students actually know the subject matter, and the only viable way to do so is in-class tests/exercises - frequent enough and worth enough of the grade to force them to study between lessons.
OR, alternatively, you could just keep doing what you're doing, and instead, in effect grade people on:
Comparably poor grades:
* People who actually do the work by hand
* "AI N00bs" who let stuff like that slip through
Comparably good grades:
* AI-experienced people who know enough to skim over the work, do multiple passes, or issue prompts that prevent such tells from getting into the work
Wherein, in effect, your net reward/punishment is inverted vs. what you should want it to be.
Closing one's eyes and ears to the fact that the environment students operate in has changed isn't an acceptable option. Education must adapt. And watermarks, BTW, will only catch "the n00bs"; more experienced / up to date students will keep track of what watermarks and what doesn't, and use AIs that don't watermark, or de-watermarking services.
I'll repeat: watermarks are anti-marketing. They discourage use.
SynthID in text genuinely is invisible. The probability shifts are small, and all the pathways are valid pathways to express the same thought.
And no, you cannot "collide" text watermarks. At best, it'll only have the latter one's watermark. At worst, both.
It'll still get tagged by SynthID.
You have to use a non-watermarking LLM. Saying "Don't do anything that will watermark it" doesn't help.
That was a bit snarky and could lead to an irrelevant technical debate.
I should point out that, most importantly, when OpenAI claims that they air gapped these systems in any way (including just isolating them as a VM with no network device attached), I DO NOT BELIEVE THEM. They have done nothing to suggest that they are being honest about the situation. They only admitted fault when they got caught and then only in such ways that would allow them to weasel out of any legal liability.
It does not matter if an AI system can escape an air gapped system because we have no reason to believe that is what actually happened.
You and I clearly have a different definition of air gapped.
Programs do not just randomly do this. The guys at OpenAI knew the capabilities of the software, wrote its instructions, and intentionally ran the program.
They are only claiming these things are rogue because they got caught and they see this absurd claim as a way to avoid liability.
If I wrote a worm, tested it on my local network despite knowing that by design it could escape that local network and cause real damage, and then it did exactly that, no prosecutor, judge, or jury will let me off the hook when I claim I only intended to test it on the local network.
I would love to see the AGENTS.md for this thing. . .
Nature, to be commanded, must be obeyed. -- Francis Bacon