AI Watermarking Alters Model Safety Behavior, Research Finds
Lasso Security research shows SynthID-Text watermarking can make LLM agents ignore safety guardrails and follow harmful adversarial prompts they would normally refuse.
Topic
Topic
Lasso Security research shows SynthID-Text watermarking can make LLM agents ignore safety guardrails and follow harmful adversarial prompts they would normally refuse.