Towards In-Parameter Memory Augmentation for Large Language Models

Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao, Hong Ting Tsang, Wuganjing Song, Huihao Jing, Yufei Li, Yangqiu Song

arXiv:2610.08630 · 2026-10-07 공개 · arXiv · PDF

llm-agents large-language-models in-context-learning recursive-self-improvement embedding-layer attention-layer ffn-layer in-parameter-memory

Abstract

Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. In-parameter memory offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: Parameter Placement, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and Parameter Acquisition Time, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.

한국어 요약

한 줄 요약

LLM에 사전 학습 이후 지식을 효율적으로 통합하기 위한 in-parameter memory augmentation 방법론을 두 축으로 분류하고 정리한다.

핵심 기여도

핵심 아이디어

기존 ICL 기반 메모리 확장은 사전 학습된 모델에 외부 정보를 토큰으로 주입하는 방식이지만, 이는 **context window 소모**와 **토큰 수에 비례한 계산 비용 증가**를 유발한다. IPM은 이와 달리, **추론 시점에 파라미터 공간에 메모리 객체를 조합**하여 정보를 재사용하는 방식을 제안한다. 이는 **LLM의 파라미터 공간을 메모리 매체로 활용**함으로써, **반복적인 토큰 인코딩 비용을 줄이고, 메모리 재사용을 가능하게 한다**. 핵심 통찰은, **메모리 객체 $ \phi $**가 추론 시점에 **모델 파라미터 $ \theta $**와 결합되며, 이는 **online 또는 offline 방식**으로 생성될 수 있다는 점이다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용

IPM은 **도메인별 지식 통합**, **사용자 선호 반영**, **대화 기반 에이전트의 경험 학습** 등에 적용 가능. 특히, **I/O 비용이 높은 대규모 배포 환경**에서 효율적인 메모리 관리가 필요한 경우 유용.