Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, K. Dvijotham, Thomas Steinke, J. Hayase, A. F. Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, D. Rolnick, Florian Tramèr

arXiv:2403.06634 · 2026-07-27 공개 · arXiv · PDF

transformer chatgpt black-box-models model-stealing embedding-projection palm-2 ada-model babbage-model

Abstract

We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under \$20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.

한국어 요약

한 줄 요약

생산 환경 언어 모델에서 임베딩 프로젝션 계층을 추출하는 최초의 모델 도난 공격이 제시되었다.

핵심 기여도

핵심 아이디어

기존의 모델 도난 공격은 입력 계층부터 상향식(bottom-up)으로 모델을 재구성하는 방식을 사용했으나, 본 연구는 **최종 계층**(임베딩 프로젝션)을 직접 추출하는 **하향식(top-down)** 접근법을 제안한다.

이 접근법은 언어 모델의 마지막 계층이 **은닉 차원(hidden dimension)**에서 **로짓 벡터(logit vector)**로의 **저랭크(low-rank) 매핑**이라는 사실을 활용한다. API를 통해 **목표 토큰에 대한 로짓 조정**(logit bias)을 요청함으로써, 모델의 마지막 계층을 추출할 수 있다.

이러한 공격은 모델의 **은닉 차원 크기**를 드러내어, 모델의 전체 파라미터 수와 연관성을 제공하며, 향후 공격 가능성에 대한 경고를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 **생산 환경 언어 모델**(예: GPT-4, PaLM-2)에서 **정확한 내부 정보**를 추출하는 최초의 사례로, 모델의 **보안 취약성**을 드러낸다. 특히, **API의 로짓-바이어스(logit bias)**와 **로그프로브(logprobs)** 기능이 공격에 활용된 점은, 시스템 설계 결정이 보안에 미치는 영향을 보여준다.

한편, 본 공격은 **단일 계층**만을 추출하며, 전체 모델 가중치 복원은 아직 불가능하다. 또한, OpenAI와 Google은 이미 **방어 조치**를 도입하여 공격 비용을 증가시키고 있다.

실용적 활용

본 연구는 **모델 보안 설계**, **API 정책 개선**, **모델 도난 방지 기술** 개발에 기여할 수 있다. 특히, **대규모 언어 모델의 보안 취약점**을 사전에 파악하고, **서비스 제공자**가 보다 안전한 API를 설계하는 데 활용될 수 있다.