Large Language Models for Data Annotation and Synthesis: A Survey

Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, Huan Liu

arXiv:2402.13446 · 2026-07-27 공개 · arXiv · PDF

large-language-models data-annotation annotation-synthesis llm-taxonomy annotation-assessment machine-learning-data llm-utilization data-types

Abstract

Data annotation and synthesis generally refers to the labeling or generating of raw data with relevant information, which could be used for improving the efficacy of machine learning models. The process, however, is labor-intensive and costly. The emergence of advanced Large Language Models (LLMs), exemplified by GPT-4, presents an unprecedented opportunity to automate the complicated process of data annotation and synthesis. While existing surveys have extensively covered LLM architecture, training, and general applications, we uniquely focus on their specific utility for data annotation. This survey contributes to three core aspects: LLM-Based Annotation Generation, LLM-Generated Annotations Assessment, and LLM-Generated Annotations Utilization. Furthermore, this survey includes an in-depth taxonomy of data types that LLMs can annotate, a comprehensive review of learning strategies for models utilizing LLM-generated annotations, and a detailed discussion of the primary challenges and limitations associated with using LLMs for data annotation and synthesis. Serving as a key guide, this survey aims to assist researchers and practitioners in exploring the potential of the latest LLMs for data annotation, thereby fostering future advancements in this critical field.

한국어 요약

한 줄 요약

GPT-4 기반 대형 언어 모델을 활용한 데이터 어노테이션 및 합성에 대한 체계적 서베이.

핵심 기여도

핵심 아이디어

기존 데이터 어노테이션은 수작업이 많고 비용이 높은 작업이지만, GPT-4와 같은 LLM은 이를 자동화할 수 있는 잠재력을 가진다. 본 논문은 LLM이 단순히 어노테이션 생성 도구를 넘어서, 어노테이션의 품질 평가 및 활용 전략까지 포괄하는 종합적인 프레임워크를 제시한다. 특히, $\mathcal{A}$ (annotator model)와 $\mathcal{L}$ (task learner)의 이중 모델 구조를 통해 어노테이션 생성과 활용을 분리하고, 이를 기반으로 다양한 학습 전략을 적용할 수 있다. 이는 기존 서베이들이 주로 모델 아키텍처나 특정 태스크에 집중한 것과 달리, 어노테이션 생성이라는 복잡한 과정 자체에 초점을 맞춘 점에서 차별화된다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM이 데이터 어노테이션 분야에서 기존 수작업 방식을 대체할 수 있는 잠재력을 제시하며, 어노테이션 생성, 평가, 활용의 전 과정을 체계적으로 정리한 점에서 학술적·실용적 가치가 있다. 특히, $\mathcal{A}$-$\mathcal{L}$ 이중 모델 구조는 어노테이션 생성과 활용을 분리하여 모델의 유연성과 정확도를 동시에 확보할 수 있다. 그러나 LLM 생성 어노테이션의 한계점도 명시되는데, hallucination(허구) 발생 가능성, 도메인 적응성 부족, 품질 평가의 주관성 등이 주요 문제로 지적된다. 또한, multimodal LLM은 다루지 않았다는 점에서 연구 범위의 제한이 있다.

실용적 활용

본 연구는 NLP 및 머신러닝 분야에서 데이터 어노테이션을 대규모로 필요로 하는 연구자와 엔지니어에게 실용적인 가이드를 제공한다. 특히, GPT-4, LLaMA-2 등 LLM을 활용한 어노테이션 생성 및 활용 전략은 의료, 법률, 금융 등 전문 도메인에서의 어노테이션 자동화에 적용 가능하다. 또한, 어노테이션 품질 평가 방법론은 데이터셋 개선 및 모델 학습 효율성 향상에 기여할 수 있다.