I’m sure they will be more subtle than that otherwise it will get circumvented.
I’m sure they will/are tackling this at the model level. Train them to both generate good completions while also embedding text with good performance at separating generated and human text.
Personally, I'm a pessimist on this front. People assert that a model-in-training can effortlessly sift out the real data from mountains of LLM spam. But then people also assert that AI detectors do not work and can never work, since LLM output is simply too good, and any watermarking can be broken up by a light paraphrasing step. It doesn't make much sense to have it both ways.
I can only await companies' attempts to publish enough junk to create an 'alternative truth' for new LLMs to believe in. The worst part is, it might even work.
Would someone even want to circumvent it though? Most sites won't care very much about encouraging scrapers to include them in LLM training data, it's not like you get paid.
I’m sure they will/are tackling this at the model level. Train them to both generate good completions while also embedding text with good performance at separating generated and human text.