I agree that this sketch comes closer to working in practice than simple RLHF. I...

joe_the_user · on Dec 29, 2023

I was surprised that I came up with a plausible sounding method. I had thought on first blush that this was impossible but now it seems reasonable. You could still have various exfiltration methods like "give me the data with each word backwards" and I'm not sure where that would stand legally.

groceryheist · on Dec 29, 2023

Yes, of of the hard and interesting legal questions is if creating a possibility of such attacks constitutes a copyvio.