1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi)
2. Train a discriminator model that can differentiate Claude vs. non-Claude
3. Train a de-watermarking model using the discriminator model as loss
edit: almost forgot the "—"
The honest seam is obvious: claudish works; imitation does not. That’s not nothing.
reply
1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi)
2. Train a discriminator model that can differentiate Claude vs. non-Claude
3. Train a de-watermarking model using the discriminator model as loss