Xiaoxiao He

dblp:22/9737 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2026 DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative Models
abstract
Recent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces.
Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas
WACV1
2026 Large Sign Language Models: Toward 3D American Sign Language Translation
abstract
We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals’ virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language.
Xiaoxiao He, Di Liu 0003, Zhaoyang Xia, Chaowei Tan, Vivian Li, Bo Liu 0005, Dimitris N. Metaxas, Mubbasir Kapadia
WACV2
2025 LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
abstract
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing **Lo**w-**R**ank matrix multiplication for **V**isual **P**rompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to $6\times$ faster training times, utilizing $18\times$ fewer visual prompt parameters, and delivering a 3.1% improvement in performance.
Can Jin, Shiyu Zhao 0001, Zhenting Wang, Xiaoxiao He, Ligong Han, Tong Che, Dimitris N. Metaxas
ICLR6
2025 Camera-Based Infant Suffocation Risk Detection Via Text-to-Image Generation for Guarding Sleep Safety
abstract
Current camera-based infant monitoring mainly focuses on physiological measurement, overlooking its important semantic analysis potential for detecting accidental suffocation caused by oronasal occlusion during sleep. However, developing a robust infant suffocation risk detection model typically requires substantial labeled data, which is very difficult to obtain in real-world scenarios. To address this, we utilized the text-to-image diffusion model to generate diverse infant images depicting oronasal occlusion and non-occlusion scenarios controlled by text prompts. To ease the process of labeling, self- and semi-supervised learning algorithms are leveraged to learn the semantic information from unlabeled data with the support of minimal labeled data to train different model architectures. To evaluate the feasibility of this solution, we conducted a clinical trial in the neonatology department, which collected video data from 22 infants under various oronasal occlusion scenarios using breathable covers (e.g. clinical tissue). The clinical evaluation shows that most models trained on 25,000 generated images achieved over 90% performance on metrics of accuracy, recall, and F1-score, outperforming conventional approaches that pre-train and fine-tune the model using over 90,000 labeled task-related online images. This demonstrates the feasibility of leveraging text-to-image generated data to achieve robust camera-based infant suffocation risk detection, so as to secure the sleep safety of infants. More importantly, it beacons the potential of using text-based large-scale model to solve the general issue of scarcity of human data in artificial intelligence-based healthcare or clinical applications.
Dongmin Huang, Chuchu Liao, Jingyun Mai, Xiaoxiao He, Liping Pan, Ming Xia 0003, Huailei Lai, Xuhui Yang, Zhenlang Lin, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics4
2024 ProxEdit: Improving Tuning-Free Real Image Editing with Proximal Guidance
abstract
DDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used for enhanced editing. Null-text inversion (NTI) optimizes null embeddings to align the reconstruction and inversion trajectories with larger CFG scales, enabling real image editing with cross-attention control. Negative-prompt inversion (NPI) further offers a training-free closed-form solution of NTI. However, it may introduce artifacts and is still constrained by DDIM reconstruction quality. To overcome these limitations, we propose proximal guidance and incorporate it to NPI with cross-attention control. We enhance NPI with a regularization term and inversion guidance, which reduces artifacts while capitalizing on its training-free nature. Additionally, we extend the concepts to incorporate mutual self-attention control, enabling geometry and layout alterations in the editing process. Our method provides an efficient and straightforward approach, effectively addressing real image editing tasks with minimal computational overhead.
Ligong Han, Song Wen 0001, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen 0002, Di Liu 0003, Qilong Zhangli, Jindong Jiang, Zhaoyang Xia, Akash Srivastava, Dimitris N. Metaxas
WACV9
2023 DMCVR: Morphology-Guided Diffusion Model for 3D Cardiac Volume Reconstruction
Xiaoxiao He, Chaowei Tan, Ligong Han, Bo Liu 0005, Leon Axel, Kang Li 0004, Dimitris N. Metaxas
MICCAI (7)1
2022 TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers
Di Liu 0003, Yunhe Gao, Qilong Zhangli, Ligong Han, Xiaoxiao He, Zhaoyang Xia, Song Wen 0001, Zhennan Yan, Mu Zhou, Dimitris N. Metaxas
MICCAI (5)5
2022 Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu 0003, Xiaoxiao He, Zhaoyang Xia, Ligong Han, Yunhe Gao, Song Wen 0001, Haiming Tang, He Wang 0016, Mu Zhou, Dimitris N. Metaxas
MICCAI (4)4
2005 An Interval-Based Knowledge Model and Query Language for Temporal Information
Zhongzhi Shi, Xiaoxiao He, Lirong Qiu, Jiewen Luo
PRIMA3