HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Heng Wang, Peng Wang, Yuyin Zhou, Cihang Xie

arXiv:2404.09990 · 2026-07-27 공개 · arXiv · PDF

foundation-models image-editing high-resolution evaluation-metrics instruction-based-editing gpt-4v dall-e-3 hq-edit

Abstract

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. The project page is https://thefllood.github.io/HQEdit_web.

한국어 요약

한 줄 요약

HQ-Edit는 GPT-4V와 DALL-E 3를 활용한 20만 개의 고해상도 이미지 편집 데이터셋으로, 기존 모델보다 편집 정확도를 12.3% 향상시킨다.

핵심 기여도

핵심 아이디어

기존 편집 데이터셋은 GPT-3와 Stable Diffusion 1.5 등 낡은 모델을 사용해 해상도와 정확도가 낮았으며, 인간 피드백에 의존해 확장성이 제한적이었다. 본 연구는 최신 기초 모델인 GPT-4V와 DALL-E 3를 활용해 편집 데이터셋을 자동 생성함으로써 이 문제를 해결한다. 특히, DALL-E 3를 통해 생성된 이미지 쌍(diptych)을 기반으로 입력-출력 이미지의 정밀한 정렬을 보장하는 후처리 과정을 도입했다. 이는 기존 Prompt-to-Prompt 방식보다 편집 영역 외부의 일관성을 향상시키는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

HQ-Edit는 기존 인간 주도 데이터셋의 한계를 극복한 첫 대규모 자동 생성 편집 데이터셋으로, GPT-4V와 DALL-E 3의 강력한 생성 능력을 활용해 편집 정확도와 일관성을 동시에 향상시킨다. 특히, Alignment와 Coherence 지표를 통해 편집 품질을 정량적으로 평가할 수 있는 새로운 기준을 제시한다. 그러나 DALL-E 3의 API 기반 접근으로 모델 가중치에 직접 접근하지 못한 점이 한계이며, 일부 후처리 과정에서 이미지 왜곡이 발생할 수 있다.

실용적 활용

HQ-Edit는 디지털 콘텐츠 제작, 광고, 영화 산업에서의 자동 이미지 편집 모델 개발에 활용 가능하다. 특히, GPT-4V와 DALL-E 3의 API를 활용한 대규모 데이터셋 생성 방식은 편집 모델의 훈련 데이터 확보를 효율화할 수 있다.