Zero-Shot Multi-Reference Personalization via MLLMs-Guided Layout Planning

Authors: Junhao Feng, Xinghang Xu, Zihao Zhang, Zhonghua Wan
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Image Personalization, Diffusion Models, Multimodal LLMs, Zero-Shot Learning.

Abstract

Though zero-shot adapters excel in image personalization, they often encounter significant challenges in multi-reference personalized generation, specifically failing to precisely adhere to the spatial layouts described in text prompts and suffering from feature leakage between reference images. To address these two challenges, we propose RIG (Regional Image-prompt Generation), a novel training-free framework. For the first challenge, leveraging Multimodal Large Language Models (MLLMs), we introduce a layout planning binder. Leveraging Chain-of-Thought (CoT) reasoning, this module infers and generates precise global layouts from text prompts, while simultaneously binding reference images to their corresponding regions. For the second, we introduce a satially decoupled diffusion mechanism that isolates feature streams during attention computation. By injecting reference features exclusively into designated regions, this mechanism effectively prevents feature interference between reference images. Extensive experiments demonstrate that RIG significantly outperforms state-of-the-art adapter methods in terms of both personalization fidelity and text-layout alignment.
📄 View Full Paper (PDF) 📋 Show Citation