Facial appearance editing plays a pivotal role in digital avatars, AR/VR systems, and personalized content creation. However, achieving identity-preserving editing from a single reference image remains a significant challenge. Existing approaches typically rely on multiple images per identity, yet still struggle with rigging accuracy and consistency in appearance. To address this limitation, we propose InstaFace, a diffusion-based framework for identity-preserving facial image generation from a single input. To enable precise control, InstaFace introduces a 3D Fusion Controller Network that integrates multiple 3DMM-derived conditionals, augmented with lightweight adjustment modules for fine-grained rigging control. These modules are optimized using a novel geometry-aware objective, 3D Morphable Reinference Discrepancy (3D-MRD), which aligns morphable reconstructions with the original conditioning. Additionally, to retain high-fidelity contextual features such as background, hair, and accessories, we incorporate a Contextual Identity Mixer, a hybrid embedding module that leverages both facial recognition and vision-language priors. Empirical evaluations demonstrate that InstaFace achieves strong identity preservation and photorealism while offering fine control over expression, pose, and lighting, using only a single reference image.
@inproceedings{khan2026instaface,
title = {InstaFace: Identity-Preserving Facial Editing with Single Image Inference},
author = {Khan, MD Wahiduzzaman and Jia, Mingshan and Zhang, Xiaolin and Yu, En and Shan, Caifeng and Musial-Gabrys, Kaska},
booktitle = {Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition (FG)},
year = {2026},
eprint = {2502.20577},
archivePrefix = {arXiv},
note = {To appear}
}