Comment Re:Very funny (Score 1) 11
Sorry, but that's not how this works.
At the softmax, the hidden state of the last layer output is converted back to token-space / text-space. This is a high-dimensional-space to low-dimensional-space conversion, so there is usually no "single right path" in which the target concept can be represented. The softmax in effect converts the nearest token pathways into probabilities; the closer the token's position to the latent position, the higher the probability.
So you have a probability-scaled list of options. Sometimes, the token that comes after is almost 100% certain. For example:
"This artifact was found in the tomb of the boy-pharaoh Tutankhamun "
Sometimes it's not:
"Ravi was hungry, so he climbed the mango tree and picked (... a mango)"
"Ravi was hungry, so he climbed the mango tree and carefully (...picked a mango)"
"Ravi was hungry, so he climbed the mango tree and grabbed (...a mango)"
Etc. There may be a whole wide range of possibilities.
A primitive watermarking algorithm, thus, can be represented in the "even-odd" algorithm. You divide up all tokens randomly into even or odd bins. On even-numbered generated tokens, you slightly increase the odds of tokens from the "even" bin and decrease them from the "odd" bin - maybe, say, "picked" declines in odds and "grabbed" rises. On odd-numbered generations you reverse your boosting. The probability shifts aren't huge, so it's still going to say "Tutankhamun", not "TutankhJetBrains", regardless of whether it's an even or odd token. But it might flip, say, "picked" vs. "carefully" vs. "grabbed" or whatnot. With a large enough sample size - and it doesn't have to be huge - you can tell whether this token bias exists.
This is, as mentioned, a primitive algorithm, and it has vulnerabilities - but it's not hard to see how you can adapt it to more advanced algorithms that are less vulnerable to user manipulation.
So no, you can't get rid of the watermark just by cut and paste. It's embedded in the choice of wording itself.