Forgot your password?
typodupeerror

Comment Re:Meanwhile, back at the benchmarks (Score 1) 21

It actually makes RAM needs worse than a dense model, for a given quality.

That said, inference techniques like FreeToken are improving MoE swapping performance, so it's not as bad as it once was. Honestly, I would not be at all surprised if we start doing training in a swap-aware manner, where expert cache misses count as loss and models can queue experts to start loading before they're needed. Could even train in two tiers of loading - RAM and disk. Wouldn't that be great? Could run multi-terabyte models on consumer hardware. Training would tend to tend to concentrate key logic and reasoning on small number of very active experts, mixed reasoning/knowledge on less common experts that are usually kept in RAM, and rarer knowledge on experts that usually remain on-disk until needed.

Comment Re: Meanwhile, back at the benchmarks (Score 1) 21

There are well known, reputable abliterators on HuggingFace.

Also, censorship usually (not always) is something you care about for chats, not agentic work.

A .safetensors or .gguf model cannot run code unless you let it. No agentic framework = all it can do is write text. It can't hurt you directly. At worst, it can only lie to you.

Comment Re:Paywalled AI Marketing (Score 4, Interesting) 66

Nobody, I repeat, nobody who uses AI is going to see "OpenAI adding watermarks" and think, "Oh, I better use AI more, especially OpenAI".

And nobody, I repeat, nobody who doesn't use AI is going to see "OpenAI adding watermarks" and think, "Oh, I better use AI more, now that it'll be easier to see that what I did is made by AI."

It is not "an advertisement of the power of the technology". It's a regulation forced on them by the EU's AI Act, which neither companies nor users want.

Comment Re:My watermark would work better (Score 1) 66

Perhaps you should be realizing by now that homework is no longer a viable way to test how well students actually know the subject matter, and the only viable way to do so is in-class tests/exercises - frequent enough and worth enough of the grade to force them to study between lessons.

OR, alternatively, you could just keep doing what you're doing, and instead, in effect grade people on:

Comparably poor grades:
* People who actually do the work by hand
* "AI N00bs" who let stuff like that slip through

Comparably good grades:
* AI-experienced people who know enough to skim over the work, do multiple passes, or issue prompts that prevent such tells from getting into the work

Wherein, in effect, your net reward/punishment is inverted vs. what you should want it to be.

Closing one's eyes and ears to the fact that the environment students operate in has changed isn't an acceptable option. Education must adapt. And watermarks, BTW, will only catch "the n00bs"; more experienced / up to date students will keep track of what watermarks and what doesn't, and use AIs that don't watermark, or de-watermarking services.

Comment Re:So what's it good for? (Score 4, Interesting) 66

It's not hard to see how this can backfire. People treat watermarks as the Word of God. But let's say, for example, a journalist, in an article, quotes a White House press statement for something, but the White House used AI. Since they don't "measure how much a human contributed", and detect the watermark in the White House statement in the journalist's article, they'll just flag the whole journalist's article as AI.

And honestly, "not measuring how much a human contributed" is IMHO a massive flaw in general even if the author was using AI. If the person wrote an article, and then told an AI, "correct my spelling, grammar, and poor phrasing", that's IMHO entirely different from the person just telling an AI "Write this article for me", and just pasting whatever it spits out in as their own, whole cloth. Contribution assessment is, IMHO, essential for fairness. And also, eminently doable. It is perfectly technologically possible to not just see, "does this signature exist", but even get a sense of exactly what the AI contributed vs. what it didn't.

Comment Re:Very funny (Score 5, Insightful) 66

Sorry, but that's not how this works.

At the softmax, the hidden state of the last layer output is converted back to token-space / text-space. This is a high-dimensional-space to low-dimensional-space conversion, so there is usually no "single right path" in which the target concept can be represented. The softmax in effect converts the nearest token pathways into probabilities; the closer the token's position to the latent position, the higher the probability.

So you have a probability-scaled list of options. Sometimes, the token that comes after is almost 100% certain. For example:

"This artifact was found in the tomb of the boy-pharaoh Tutankhamun "

Sometimes it's not:

"Ravi was hungry, so he climbed the mango tree and picked (... a mango)"
"Ravi was hungry, so he climbed the mango tree and carefully (...picked a mango)"
"Ravi was hungry, so he climbed the mango tree and grabbed (...a mango)"

Etc. There may be a whole wide range of possibilities.

A primitive watermarking algorithm, thus, can be represented in the "even-odd" algorithm. You divide up all tokens randomly into even or odd bins. On even-numbered generated tokens, you slightly increase the odds of tokens from the "even" bin and decrease them from the "odd" bin - maybe, say, "picked" declines in odds and "grabbed" rises. On odd-numbered generations you reverse your boosting. The probability shifts aren't huge, so it's still going to say "Tutankhamun", not "TutankhJetBrains", regardless of whether it's an even or odd token. But it might flip, say, "picked" vs. "carefully" vs. "grabbed" or whatnot. With a large enough sample size - and it doesn't have to be huge - you can tell whether this token bias exists.

This is, as mentioned, a primitive algorithm, and it has vulnerabilities - but it's not hard to see how you can adapt it to more advanced algorithms that are less vulnerable to user manipulation.

So no, you can't get rid of the watermark just by cut and paste. It's embedded in the choice of wording itself.

Slashdot Top Deals

Civilization, as we know it, will end sometime this evening. See SYSNOTE tomorrow for more information.

Working...