Haomiao Ni

dblp:196/1975 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
10since 2021 · last 2025
0009-0004-6747-4080ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Enhancing AI-Assisted Stroke Emergency Triage with Adaptive Uncertainty Estimation
Tongan Cai, Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Sharon X. Huang, Stephen T. C. Wong
MICCAI (14)3
2025 Computer-Aided Layout Generation for Building Design: A Review
abstract
Generating realistic building layouts for automatic building design has been studied in both computer vision and architectural domains. Traditional approaches in the latter, which are based on optimization techniques or heuristic design guidelines, can synthesize desirable layouts, but usually require post-processing and involve human interaction in the design pipeline, making them costly and time-consuming. The advent of deep generative models has significantly improved the fidelity and diversity of the generated architecture layouts, reducing the workload of designers and making the process much more efficient. This paper presents a comprehensive review of three major research topics in architectural layout design and generation: floorplan layout generation, scene layout synthesis, and generation of various other formats of building layouts. For each topic, we overview the leading paradigms, categorized either by research domains (architecture or machine learning) or by user input conditions or constraints. We then introduce commonly-adopted benchmark datasets used to verify the effectiveness of the methods, as well as corresponding evaluation metrics. Finally, we identify the well-solved problems and limitations of existing approaches, and then propose promising directions for future research. This survey has an associated project which aims to maintain the resources, at https://github.com/jcliu0428/awesome-building-layout-generation.
Yuan Xue 0002, Haomiao Ni, Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang
Comput. Vis. Media3
2024 TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
abstract
Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., “a woman is drinking water.”). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a “repeat-and-slide” strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation.
Haomiao Ni, Bernhard Egger 0001, Suhas Lohit, Anoop Cherian, Ye Wang 0001, Toshiaki Koike-Akino, Sharon X. Huang, Tim K. Marks
CVPR1
2024 3D-Aware Talking-Head Video Motion Transfer
abstract
Motion transfer of talking-head videos involves generating a new video with the appearance of a subject video and the motion pattern of a driving video. Current methodologies primarily depend on a limited number of subject images and 2D representations, thereby neglecting to fully utilize the multi-view appearance features inherent in the subject video. In this paper, we propose a novel 3D-aware talking-head video motion transfer network, Head3D, which fully exploits the subject appearance information by generating a visually-interpretable 3D canonical head from the 2D subject frames with a recurrent network. A key component of our approach is a self-supervised 3D head geometry learning module, designed to predict head poses and depth maps from 2D subject video frames. This module facilitates the estimation of a 3D head in canonical space, which can then be transformed to align with driving video frames. Additionally, we employ an attention-based fusion network to combine the background and other details from subject frames with the 3D subject head to produce the synthetic target video. Our extensive experiments on two public talking-head video datasets demonstrate that Head3D outperforms both 2D and 3D prior arts in the practical cross-identity setting, with evidence showing it can be readily adapted to the pose-controllable novel view synthesis task.
Haomiao Ni, Yuan Xue 0002, Sharon X. Huang
WACV1
2023 Conditional Image-to-Video Generation with Latent Flow Diffusion Models
abstract
Conditional image-to-video (cI2V) generation aims to synthesize a new plausible video starting from an image (e.g., a person's face) and a condition (e.g., an action class label like smile). The key challenge of the cI2V task lies in the simultaneous generation of realistic spatial appearance and temporal dynamics corresponding to the given image and condition. In this paper, we propose an approach for cI2V using novel latent flow diffusion models (LFDM) that synthesize an optical flow sequence in the latent space based on the given condition to warp the given image. Compared to previous direct-synthesis-based works, our proposed LFDM can better synthesize spatial details and temporal motion by fully utilizing the spatial content of the given image and warping it in the latent space according to the generated temporally-coherent flow. The training of LFDM consists of two separate stages: (1) an unsupervised learning stage to train a latent flow auto-encoder for spatial content generation, including a flow predictor to estimate latent flow between pairs of video frames, and (2) a conditional learning stage to train a 3D-UNet-based diffusion model (DM) for temporal latent flow generation. Unlike previous DMs operating in pixel space or latent feature space that couples spatial and temporal information, the DM in our LFDM only needs to learn a low-dimensional latent flow space for motion generation, thus being more computationally efficient. We conduct comprehensive experiments on multiple datasets, where LFDM consistently outperforms prior arts. Furthermore, we show that LFDM can be easily adapted to new domains by simply finetuning the image decoder. Our code is available at https://github.com/nihaomiao/CVPR23_LFDM.
Haomiao Ni, Changhao Shi, Kai Li 0012, Sharon X. Huang, Martin Renqiang Min
CVPR1
2023 Synthetic Augmentation with Large-Scale Unconditional Pre-training
Jiarong Ye, Haomiao Ni, Sharon X. Huang, Yuan Xue 0002
MICCAI (2)2
2023 Cross-identity Video Motion Retargeting with Joint Transformation and Synthesis
abstract
In this paper, we propose a novel dual-branch Transformation-Synthesis network (TS-Net), for video motion retargeting. Given one subject video and one driving video, TS-Net can produce a new plausible video with the subject appearance of the subject video and motion pattern of the driving video. TS-Net consists of a warp-based transformation branch and a warp-free synthesis branch. The novel design of dual branches combines the strengths of deformation-grid-based transformation and warp-free generation for better identity preservation and robustness to occlusion in the synthesized videos. A mask-aware similarity module is further introduced to the transformation branch to reduce computational overhead. Experimental results on face and dance datasets show that TS-Net achieves better performance in video motion retargeting than several state-of-the-art models as well as its single-branch variants. Our code is available at https://github.com/nihaomiao/WACV23_TSNet.
Haomiao Ni, Yihao Liu 0003, Sharon X. Huang, Yuan Xue 0002
WACV1
2023 Semi-supervised body parsing and pose estimation for enhancing infant general movement assessment
Haomiao Ni, Yuan Xue 0002, Liya Ma, Qian Zhang 0081, Xiaoye Li, Sharon X. Huang
Medical Image Anal.1
2022 Asymmetry Disentanglement Network for Interpretable Acute Ischemic Stroke Infarct Segmentation in Non-contrast CT Scans
Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, Stephen T. C. Wong, James Z. Wang 0001, Sharon X. Huang
MICCAI (8)1
2022 DeepStroke: An efficient stroke screening framework for emergency rooms with multimodal adversarial deep learning
Tongan Cai, Haomiao Ni, Mingli Yu, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong
Medical Image Anal.2
2020 SiamParseNet: Joint Body Parsing and Label Propagation in Infant Movement Videos
Haomiao Ni, Yuan Xue 0002, Qian Zhang 0081, Sharon X. Huang
MICCAI (4)1
2018 Multiple Visual Fields Cascaded Convolutional Neural Network for Breast Cancer Detection
Haomiao Ni, Hong Liu 0007, Zichao Guo, Taijiao Jiang, Kuansong Wang, Yueliang Qian
PRICAI (1)1
2016 Action Recognition Based on Optimal Joint Selection and Discriminative Depth Descriptor
Haomiao Ni, Hong Liu 0007, Yueliang Qian
ACCV (2)1