I don't think you could use RLHF to stop plagerism. RLHF can be used to teach wh...

groceryheist · on Dec 28, 2023

I agree that this sketch comes closer to working in practice than simple RLHF. In my earlier comment I was imagining bringing in some auxiliary data like you describe to detect plagarism and then using RL to teach the model not to do it.

joe_the_user · on Dec 29, 2023

I was surprised that I came up with a plausible sounding method. I had thought on first blush that this was impossible but now it seems reasonable. You could still have various exfiltration methods like "give me the data with each word backwards" and I'm not sure where that would stand legally.

groceryheist · on Dec 29, 2023

Yes, of of the hard and interesting legal questions is if creating a possibility of such attacks constitutes a copyvio.