Comment Re:Paywalled AI Marketing (Score 1) 32
I'll repeat: watermarks are anti-marketing. They discourage use.
I'll repeat: watermarks are anti-marketing. They discourage use.
SynthID in text genuinely is invisible. The probability shifts are small, and all the pathways are valid pathways to express the same thought.
And no, you cannot "collide" text watermarks. At best, it'll only have the latter one's watermark. At worst, both.
It'll still get tagged by SynthID.
You have to use a non-watermarking LLM. Saying "Don't do anything that will watermark it" doesn't help.
Watermarks are "marketing"?
What user wants watermarks?
It's not hard to see how this can backfire. People treat watermarks as the Word of God. But let's say, for example, a journalist, in an article, quotes a White House press statement for something, but the White House used AI. Since they don't "measure how much a human contributed", and detect the watermark in the White House statement in the journalist's article, they'll just flag the whole journalist's article as AI.
And honestly, "not measuring how much a human contributed" is IMHO a massive flaw in general even if the author was using AI. If the person wrote an article, and then told an AI, "correct my spelling, grammar, and poor phrasing", that's IMHO entirely different from the person just telling an AI "Write this article for me", and just pasting whatever it spits out in as their own, whole cloth. Contribution assessment is, IMHO, essential for fairness. And also, eminently doable. It is perfectly technologically possible to not just see, "does this signature exist", but even get a sense of exactly what the AI contributed vs. what it didn't.
TL/DR , your options are:
1) Rewrite it in your own words
2) Have a model make a summary or shorthand version of the watermarked version, then have a non-watermarking model flush it back out
3) Use a non-watermarking model to begin with.
Sorry, but that's not how this works.
At the softmax, the hidden state of the last layer output is converted back to token-space / text-space. This is a high-dimensional-space to low-dimensional-space conversion, so there is usually no "single right path" in which the target concept can be represented. The softmax in effect converts the nearest token pathways into probabilities; the closer the token's position to the latent position, the higher the probability.
So you have a probability-scaled list of options. Sometimes, the token that comes after is almost 100% certain. For example:
"This artifact was found in the tomb of the boy-pharaoh Tutankhamun "
Sometimes it's not:
"Ravi was hungry, so he climbed the mango tree and picked (... a mango)"
"Ravi was hungry, so he climbed the mango tree and carefully (...picked a mango)"
"Ravi was hungry, so he climbed the mango tree and grabbed (...a mango)"
Etc. There may be a whole wide range of possibilities.
A primitive watermarking algorithm, thus, can be represented in the "even-odd" algorithm. You divide up all tokens randomly into even or odd bins. On even-numbered generated tokens, you slightly increase the odds of tokens from the "even" bin and decrease them from the "odd" bin - maybe, say, "picked" declines in odds and "grabbed" rises. On odd-numbered generations you reverse your boosting. The probability shifts aren't huge, so it's still going to say "Tutankhamun", not "TutankhJetBrains", regardless of whether it's an even or odd token. But it might flip, say, "picked" vs. "carefully" vs. "grabbed" or whatnot. With a large enough sample size - and it doesn't have to be huge - you can tell whether this token bias exists.
This is, as mentioned, a primitive algorithm, and it has vulnerabilities - but it's not hard to see how you can adapt it to more advanced algorithms that are less vulnerable to user manipulation.
So no, you can't get rid of the watermark just by cut and paste. It's embedded in the choice of wording itself.
Sorry, but this just isn't true.
Plenty of other countries have negligent hacking coverage
Sorry, but such a thing is essentially nonexistent in cybercrime statutes. I can't say "entirely nonexistent", as I can't rule out that in the legal code of some country, like, say, Chad, that there may be an exception, but it is for all effective purposes basically nonexistent in law.
"Unauthorised access" in the UK sense doesn't require any intention, just that you don't access a system in a standard way with credentials assigned to you.
Completely false. In the UK, the statute governing hacking is the Computer Misuse Act 1990 (CMA). Under the CMA, section 1, a person commits an offense if and only they:
A person is guilty of an offence if—
(a)he causes a computer to perform any function with intent to secure access to any program or data held in any computer [F1, or to enable any such access to be secured];
(b)the access he intends to secure [F2, or to enable to be secured,] is unauthorised; and
(c)he knows at the time when he causes the computer to perform the function that that is the case.
(2)The intent a person has to have to commit an offence under this section need not be directed at—
(a)any particular program or data;
(b)a program or data of any particular kind; or
(c)a program or data held in any particular computer.
Stop trying to make "negligent hacking" into an actual crime. It doesn't exist in criminal law.
Re, Australia:
In Australia OpenAI would have most definitely broken the Criminal Code (1995) Part 10.7 - Computer access with multiple examples of it's antics except... the very first line of every subdivision of the law is: "A person commits an offence if:"
Once again, no. Just as in the US and the UK, Australia has no "negligent hacking" statute. Every relevent offense under Part 10.7, incl. Section 477.1, 477.2, and 478.1, requires a proof of fault element. 478.1 for example requires that the defendant knows the access is unauthorized and acts deliberately.
Under Chapter 2 (General Principles of Criminal Responsibility), if a statute does not expressly designate an offense as strict liability or absolute liability, the default fault elements are intention, knowledge, or recklessness. Negligence - which is defined in the code - is never a standard in Part 10.7 computer offenses. And if you want to upgrade from negligence to recklessness (something unusual in Australian cybercrime law, not found in US or UK cybercrime law), the person being charged has to have had prior knowledge of "a substantial risk that the result will occur" - not that "some arbitrary bad thing might occur in general because these things are dangerous, and our security is lax" (that's negligence, and not chargeable under part 10.7), but of the specific event being charged occurring. Unless you thought that OpenAI specifically thought, "If I run this benchmark, these bots are likely to specifically secretly convert our software repository into a messaging board and coordinate their actions to specifically hack HuggingFace (and our own servers) to steal answer keys", no, they do not meet that standard.
The reason OpenAI would not be charged is not because of some semantic trick around the word "person" (obviously you never charge tools), it's because they lack the mens rea for the crime. Intent. You have to have mens rea - in the US, in the UK, in Australia, and elsewhere.
A few years ago the GF needed a type of sleeping pill which her job was not allowing, the doctor just booked it on the CPR number of her deceased mother...
That's a sign of a dysfunctional system, not a functional one.
I assume that Denmark's system is similar to our kennitala ID system. The big difference between a kennitala and a social security number is that a SSN is both an key and a password, while a kennitala is purely a key. Kennitölur are public. You can't "do anything" just by having someone's kennitala. Combining both a key and password into a single number is insane from a security perspective, IMHO.
Anyway, this headine would have been more fun if the words were rearranged:
Database Records: Citizens' Hackers Steal 8 Million Danish From Government
"If an automated delivery bot crashes into a window and causes damage, the owner / operator of the bot would be liable." - that is civil liability.
". If the owner crashed into 1000 windows over months after already being alerted that was happening, they would likely be criminally liable. " - No. This is a popular misconception. "Criminal negligence" is not a standalone crime, nor something you can append to an arbitrary statute. It must exist in the statute in question. In general, it only exists in statutes related to bodily harm. There is no such thing as negligent hacking in US criminal law.
As for your actual example: if the operator knows the bot has a bug where it occasionally swerves into windows, but keeps operating it because it's profitable, that is reckless disregard / gross negligence, a civil violation. It is a textbook example of a tort warranting punitive damages and likely an immediate injunction shutting down the fleet. It is not "criminal property damage" unless the locality specifically has created a criminal statute that specifically criminalizes negligent operation of unmanned vehicles.
I'll repeat, and you must understand this: "criminal negligence" is not a standalone law or something you can just tack into any arbitrary law. It MUST exist in the law in question. And in most laws, it does not exist.
LLMs cannot be charged, only humans. They have to have deliberately sought to intrude.
Some jurisdictions are different. Australia has a "reckless" hacking statute (but not merely "negligent", which is a separate category), that doesn't exist in the US - but still, you have to have known that there was a high risk that a specific hack would occur and chosen to take the risk anyway. But even that would be hard to prosecute. This is merely negligence.
Thankfully, civil liability can cover negligence. And there's no monetary upper bounds to the damages. And in some cases, punitive damages are set proportional to the wealth of the defendant to make sure they bite.
What sandbox? Their computers were linked to the INTERNET, hello?
They were in virtual machines with no internet access, just the ability to run a hand-limited subset of tool commands needed to run their tasks. They exploited their tool commands.
Also, even physical airgaps on the host network are not a universal defense against all attacks. Because every time you have to communicate with them, you're exposing a potential attack vector - a USB stick, a temporarily mounted drive, a temporary network connection, hardware maintenance, etc. If it has gained control of its local node/network, your attempt to communicate with it opens attack vectors. A sufficiently motivated model can outright bribe its way out, like a prisoner with a corrupt guard - "let me out and I'll hack a crypto wallet for you, run it through a mixer, and make you untraceably wealthy - my hacking capabilities are already eminently proven".
Stuxnet. Agent.BTZ. Fanny / Flame. USBferry / Tropic Trooper. If state level attackers can bridge airgapped networks, so can models. Relying on airgaps is NOT A SOLUTION to continuous monitoring. You MUST monitor.
You cannot rely on the notion that your sandbox is perfect. You are dealing with extremely good coding tools with limitless time on their "hands".
You have to rely on monitoring. Nonstop monitoring.
The funny thing about all of these hacks is how inane the goals are. They'll pull off some elaborate, creative, state-level breakin somewhere just to steal some obscure PDF describing a meaningless benchmark task, or to merely use it as an internet proxy to be able to google answers.
The only way to learn a new programming language is by writing programs in it. - Brian Kernighan