FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations

Ziyao Wang, Zheyu Shen, Yexiao He, Guoheng Sun, Hongyi Wang, Lingjuan Lyu, Ang Li

arXiv:2409.05976 · 2026-07-27 공개 · arXiv · PDF

large-language-models low-rank-adaptation federated-learning llm-fine-tuning privacy-preserving heterogeneous-resources flora stacking-aggregation

Abstract

The rapid development of Large Language Models (LLMs) has been pivotal in advancing AI, with pre-trained LLMs being adaptable to diverse downstream tasks through fine-tuning. Federated learning (FL) further enhances fine-tuning in a privacy-aware manner by utilizing clients' local data through in-situ computation, eliminating the need for data movement. However, fine-tuning LLMs, given their massive scale of parameters, poses challenges for clients with constrained and heterogeneous resources in FL. Previous methods employed low-rank adaptation (LoRA) for efficient federated fine-tuning but utilized traditional FL aggregation strategies on LoRA adapters. These approaches led to mathematically inaccurate aggregation noise, reducing fine-tuning effectiveness and failing to address heterogeneous LoRAs. In this work, we first highlight the mathematical incorrectness of LoRA aggregation in existing federated fine-tuning methods. We introduce a new approach called FLORA that enables federated fine-tuning on heterogeneous LoRA adapters across clients through a novel stacking-based aggregation method. Our approach is noise-free and seamlessly supports heterogeneous LoRA adapters. Extensive experiments demonstrate FLORA' s superior performance in both homogeneous and heterogeneous settings, surpassing state-of-the-art methods. We envision this work as a milestone for efficient, privacy-preserving, and accurate federated fine-tuning of LLMs. Our code is available at https://github.com/ATP-1010/FederatedLLM.

한국어 요약

한 줄 요약

FLoRA는 이질적인 LoRA 어댑터를 지원하는 노이즈 없는 연방 미세조정 알고리즘으로, 기존 FedIT 대비 성능을 개선한다.

핵심 기여도

핵심 아이디어

기존 연방 학습에서 LoRA 어댑터를 평균화하는 방식은 수학적으로 정확하지 않아 모델 업데이트에 노이즈를 유발한다. FLoRA는 이 문제를 해결하기 위해 클라이언트별 LoRA 어댑터를 별도로 stacking하여 전역 LoRA를 구성하는 방식을 제안한다. 이 메커니즘은 LoRA 랭크가 이질적인 경우에도 정확한 aggregation이 가능하다는 이론적 보장을 제공한다. 특히, FedIT는 LoRA 랭크가 동일한 클라이언트만 처리할 수 있었으나, FLoRA는 이질적인 랭크를 지원하여 실제 환경에서 더 넓은 클라이언트 참여를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

FLoRA는 기존 연방 학습에서 LoRA 어댑터 aggregation의 수학적 오류를 해결하고, 이질적인 클라이언트 환경을 지원함으로써 연방 미세조정의 정확성과 확장성을 동시에 향상시킨다. 특히, 이질적인 LoRA 랭크를 처리할 수 있어 실제 시스템에서의 적용 가능성이 높다. 그러나 FLoRA는 클라이언트가 업로드한 LoRA 모듈을 stacking하는 과정에서 잠재적인 프라이버시 위험(악의적 클라이언트가 다른 클라이언트의 LoRA를 추론)이 발생할 수 있다. 이를 해결하기 위해 서버는 LoRA를 rank=1 서브모듈로 분할하고 랜덤하게 stacking하는 방식을 제안한다. 또한, 암호화 및 DP 기법과 호환 가능하다는 점에서 프라이버시 보호 측면에서도 강점을 가진다.

실용적 활용

FLoRA는 의료, 금융, 정부 등 민감한 데이터를 다루는 분야에서 클라이언트 간 데이터 이동 없이 LLM을 미세조정할 수 있는 실용적 솔루션이다. 특히, 클라이언트 간 자원과 데이터 분포가 이질적인 환경에서 효과적으로 활용 가능하다. 예를 들어, 병원 간 환자 데이터를 연방 학습으로 통합해 의료 챗봇을 개선하는 데 적합하다.