Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution

Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, Chao Dong

arXiv:2501.11561 · 2026-07-27 공개 · arXiv · PDF

large-language-models multi-modal soft-labels image-quality-assessment iqa-benchmarks fidelity-loss deqa-score score-regression

Abstract

With the rapid advancement of Multi-modal Large Language Models (MLLMs), MLLM-based Image Quality Assessment (IQA) methods have shown promising performance in linguistic quality description. However, current methods still fall short in accurately scoring image quality. In this work, we aim to leverage MLLMs to regress accurate quality scores. A key challenge is that the quality score is inherently continuous, typically modeled as a Gaussian distribution, whereas MLLMs generate discrete token outputs. This mismatch necessitates score discretization. Previous approaches discretize the mean score into a one-hot label, resulting in information loss and failing to capture inter-image relationships. We propose a distribution-based approach that discretizes the score distribution into a soft label. This method preserves the characteristics of the score distribution, achieving high accuracy and maintaining inter-image relationships. Moreover, to address dataset variation, where different IQA datasets exhibit various distributions, we introduce a fidelity loss based on Thurstone’s model. This loss captures intra-dataset relationships, facilitating co-training across multiple IQA datasets. With these designs, we develop the distribution-based Depicted image Quality Assessment model for Score regression (DeQA-Score). Experiments across multiple benchmarks show that DeQA-Score stably outperforms baselines in score regression. Also, DeQA-Score can predict the score distribution that closely aligns with human annotations. Codes and model weights have been released in https://depictqa.github.io/deqa-score/.

한국어 요약

한 줄 요약

DeQA-Score는 MLLM 기반의 정확한 이미지 품질 점수 회귀를 위해 분포 기반 소프트 라벨과 퓨이델리티 손실을 도입한 모델이다.

핵심 기여도

핵심 아이디어

기존 MLLM 기반 IQA 방법은 점수를 이산화할 때 평균 점수를 one-hot 라벨로 변환하는 방식을 사용하여 정보 손실이 발생하고, 이미지 간 관계를 포착하지 못한다. DeQA-Score는 대신 **점수 분포 자체를 이산화하여 소프트 라벨로 사용**함으로써, **점수 분포의 특성을 보존**하고, **이미지 간 관계를 유지**한다. 예를 들어, Image A와 B가 각각 다른 품질을 가졌음에도 one-hot 라벨은 동일하게 예측할 수 있지만, DeQA-Score는 분포 기반 라벨로 이 관계를 정확히 반영한다.

또한, 다양한 IQA 데이터셋 간 분포 차이를 해결하기 위해 **Thurstone 모델 기반 퓨이델리티 손실**을 도입하여, **데이터셋 내 관계를 학습**하고, **다중 데이터셋 공동 훈련**을 가능하게 한다. 이는 특히 합성 왜곡과 실제 왜곡이 혼합된 데이터셋에서 효과적이다.

기술적 접근법

주요 결과

의의 및 한계

DeQA-Score는 MLLM이 정확한 수치 점수를 회귀하는 데 기여하며, **점수 분포 예측**이라는 새로운 차원의 IQA를 제시한다. 특히, **다중 데이터셋 간 관계 학습**을 통해 **실제 세계 왜곡 유형에 대한 일반화 능력**을 향상시킨다. 또한, **인간 주관적 평가와 유사한 분포 예측**을 통해 주관적 평가와의 일관성을 높인다.

그러나, **PIPAL 데이터셋**과 같이 모델-처리된 왜곡이 포함된 데이터에서는 **다른 데이터셋 성능이 감소**하는 문제가 있다. 이는 모델-처리 왜곡이 다른 유형의 왜곡과 구조적으로 차이가 크기 때문으로, **더 다양한 왜곡 유형을 포함한 데이터셋**이 필요하다는 한계를 드러낸다.

실용적 활용

DeQA-Score는 **이미지 압축 및 전송**, **스마트폰 사진 처리**, **AI 생성 이미지 품질 평가** 등 다양한 산업 분야에서 활용 가능하다. 특히, **다양한 왜곡 유형을 포함한 대규모 데이터셋**이 필요한 연구 및 제품 개발에 적합하며, **인간 주관 평가와 유사한 수치를 제공**하여 주관적 평가와의 일관성을 유지할 수 있다.