
Research Papers Highlight Contextual Integrity Risks and Privacy Violations in Large Language Models
Two research papers examine contextual integrity risks in large language models (LLMs) using persistent memory, revealing significant privacy violations. The CIMemories benchmark found frontier models like GPT-5 leak sensitive attributes in up to 69% of cases, with violation rates rising from 0.1% to 25.1% as task interactions increase. Models exhibit arbitrary behavior, disclosing different attributes for identical prompts, and fail to make nuanced context-dependent decisions despite privacy-conscious prompting. A second study proposes a reinforcement learning (RL) framework to improve contextual integrity, demonstrating reduced inappropriate disclosures while maintaining task performance using a synthetic dataset of 700 examples. The RL approach transfers improvements to established benchmarks like PrivacyLens, which evaluates privacy leakage in AI assistants. Both papers highlight fundamental limitations in LLMs' ability to reason about contextual appropriateness.