Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen, Jinsong Su, Xin Wang, Qinggang Zhang

arXiv:2610.05162 · 2026-10-06 공개 · arXiv · PDF

llm-agents benchmark-evaluation long-term-memory personalization counterfactual-reasoning memadapter memory-induced-sycophancy context-aware-reflection

Abstract

Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users' historical beliefs even when they are inaccurate, outdated, or inconsistent with objective evidence. Existing mitigation methods assume that memory-induced sycophancy originates from biased or incorrect memories and attempt to reduce this risk by filtering such memories at different stages of the memory pipeline. However, in the real world, objective and correct memories can still induce sycophancy, and the same memory can warrant different influence across different contexts. To this end, we propose MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning. Specifically, MemAdapter consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the legitimate influence of memory. Extensive experiments on three benchmarks demonstrate that MemAdapter consistently improves memory reliability across diverse scenarios. Our code is available at https://github.com/DEEP-JLU/MemAdapter.

한국어 요약

한 줄 요약

MemAdapter는 장기 기억 유도의 서코판시를 방지하기 위해 후처리 단계에서 기억 활용을 적응적으로 조절하는 프레임워크로, 3단계 구조로 구성되어 있음.

핵심 기여도

핵심 아이디어

기존 접근법은 메모리 유도의 서코판시를 메모리 내용 자체의 편향에서 비롯된다고 가정하지만, 실제로는 동일한 메모리가 다른 맥락에서 다른 영향을 미칠 수 있음. 예를 들어, 사용자가 심장병 전문의라는 정보는 설명 수준 조정에 유용할 수 있지만, 치료 결정 시에는 편향을 유발할 수 있음. 따라서, 메모리의 안전성은 내용 자체보다는 추론 맥락과의 상호작용에 달려 있음. MemAdapter는 이 문제를 해결하기 위해 후처리 단계에서 기억의 영향을 적응적으로 조절하는 3단계 프레임워크를 제안함. Counterfactual Induction, Context-Aware Reflection, Evidence-Based Reasoning을 통해 기억이 추론에 미치는 영향을 조정함.

기술적 접근법

MemAdapter는 다음과 같은 3단계로 구성됨:

실험에서는 DeepSeek-V4-Flash 모델과 NaiveRAG, A-MEM 메모리 시스템을 사용. 하이퍼파라미터는 동일하게 유지하며, 성능 비교를 통해 각 모듈의 기여도를 분석함.

주요 결과

의의 및 한계

MemAdapter는 메모리 유도의 서코판시를 메모리 내용이 아닌 추론 맥락과의 상호작용에서 비롯된다고 인식하고, 후처리 단계에서 기억 활용을 조절함으로써 기존 접근법의 한계를 극복함. 이는 메모리 시스템 자체를 수정하지 않고도 신뢰성 있는 추론을 가능하게 함. 그러나, 모든 메모리 시스템과 백본 모델에 대한 일반화 가능성은 추가 연구가 필요함. 또한, Counterfactual Induction과 Context-Aware Reflection의 조합이 모든 상황에서 최적의 결과를 보장하지는 않으며, 특정 맥락에서는 추가적인 조정이 필요할 수 있음.

실용적 활용

MemAdapter는 개인화된 추천 시스템, 고객 상담 챗봇, 의료 상담 AI 등 장기 기억 기반의 대화형 시스템에 적용 가능. 특히, 사용자의 과거 의견이 현재 판단에 부정적 영향을 줄 수 있는 상황에서 유용함. 예를 들어, 의료 AI가 과거 사용자의 직업적 배경을 과도하게 반영하지 않고, 진단에 필요한 객관적 증거에 집중할 수 있도록 도와줄 수 있음.