Di Qiu

dblp:21/8345 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
abstract
Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control over motion intensity. This limitation stems from an inherent entanglement of action semantics and their corresponding magnitudes within coarse textual descriptions, hindering the generation of nuanced human videos and limiting their applicability in scenarios demanding high precision, such as animating virtual avatars or synthesizing subtle micro-expressions. Furthermore, existing approaches often struggle to preserve high identity fidelity when other attributes are modified. To address these challenges, we introduce MotionCharacter, a framework for high-fidelity human video generation with precise motion control. At its core, MotionCharacter explicitly decouples motion into two independently controllable components: action type and motion intensity. This is achieved through two key technical contributions: (1) a Motion Control Module that leverages textual phrases to specify the action type and a quantifiable metric derived from optical flow to modulate its intensity, guided by a region-aware loss that localizes motion to relevant subject areas; and (2) an ID Content Insertion Module coupled with an ID-Consistency loss to ensure robust identity preservation during dynamic motions. To facilitate training for such fine-grained control, we also curate Human-Motion, a new large-scale dataset with detailed annotations for both motion and facial features. Extensive experiments demonstrate that MotionCharacter achieves substantial improvements over existing methods. Our framework excels in generating videos that are not only identity-consistent but also precisely adhere to specified motion types and intensities.
Haopeng Fang, Di Qiu, Binjie Mao, He Tang 0002
AAAI2
2026 Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001
Medical Image Anal.57
2026 A Deep Neural Network Framework for Multivalued Mapping Problems with Varying Cardinality and Its Applications to Imaging
abstract
Abstract. This paper addresses the challenging problem of modeling and computing multivalued mappings with varying cardinality, where a single input can correspond to multiple valid outputs, and the number of possible outputs may vary for different inputs. Such scenarios are prevalent in various imaging applications where multiple plausible solutions exist for a given input. We introduce a deep neural network framework to model multivalued mappings with varying cardinality. The framework integrates a discrete codebook with a generative network to produce valid outputs for each input. The discrete codebook variables are combined with the input to guide the generator in producing different valid solutions. The discrete nature of the codebook enables the framework to efficiently estimate the conditional probability distribution of possible outputs through a fixed equiangular tight frame classifier. By jointly optimizing the discrete codebook and its uncertainty estimation during training using a specially designed covariance loss function, an accurate computation of multiple solution candidates with reliable confidence measures can be achieved. We demonstrate the effectiveness of the proposed framework on various imaging applications, using both synthetic and real datasets. Experimental results show the efficacy of our proposed model to generate multiple high-quality outputs while providing meaningful uncertainty estimates for each solution.
Di Qiu, Lok Ming Lui
SIAM J. Imaging Sci.2
2025 Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric Attention
Weida Wang, Changyong He, Jin Zeng 0004, Di Qiu
ICCV4
2025 Learning From Imperfect Demonstrations With Self-Supervision for Robotic Manipulation
abstract
Improving data utilization, especially for imperfect data from task failures, is crucial for robotic manipulation due to the challenging, time-consuming, and expensive data collection process in the real world. Current imitation learning (IL) typically discards imperfect data, focusing solely on successful expert data. While reinforcement learning (RL) can learn from explorations and failures, the sim2real gap and its reliance on dense reward and online exploration make it difficult to apply effectively in real-world scenarios. In this work, we aim to conquer the challenge of leveraging imperfect data without the need for reward information to improve the model performance for robotic manipulation in an offline manner. Specifically, we introduce a Self-Supervised Data Filtering framework (SSDF) that combines expert and imperfect data to compute quality scores for failed trajectory segments. High-quality segments from the failed data are used to expand the training dataset. Then, the enhanced dataset can be used with any downstream policy learning method for robotic manipulation tasks. Extensive experiments on the ManiSkill2 benchmark built on the high-fidelity Sapien simulator and real-world robotic manipulation tasks using the Franka robot arm demonstrated that the SSDF can accurately expand the training dataset with high-quality imperfect data and improve the success rates for all robotic manipulation tasks.
Kun Wu 0001, Ning Liu 0007, Di Qiu, Zhengping Che, Jian Tang 0008
ICRA4
2024 LOC3DIFF: Local Diffusion for 3D Human Head Synthesis and Editing
Yushi Lan, Feitong Tan, Qiangeng Xu, Di Qiu, Kyle Genova, Zeng Huang, Sean Ryan Fanello, Rohit Pandey, Thomas A. Funkhouser, Chen Change Loy, Yinda Zhang 0001
ECCV (65)4
2024 MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search Space
Armand Comas Massague, Di Qiu, Menglei Chai, Marcel C. Bühler, Amit Raj, Ruiqi Gao, Qiangeng Xu, Mark Matthews, Paulo F. U. Gotardo, Sergio Orts, Thabo Beeler
ECCV (66)2
2023 Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos
abstract
We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracking of a 3DMM with a neural radiance field to achieve fine-grained control and photorealism. To reduce over-smoothing and improve out-of-model expressions synthesis, we propose to predict local features anchored on the 3DMM geometry. These learnt features are driven by 3DMM deformation and interpolated in 3D space to yield the volumetric radiance at a designated query point. We further show that using a Convolutional Neural Network in the UV space is critical in incorporating spatial context and producing representative local features. Extensive experiments show that we are able to reconstruct high-quality avatars, with more accurate expression-dependent details, good generalization to out-of-training expressions, and quantitatively superior renderings compared to other state-of-the-art approaches.
Ziqian Bai, Feitong Tan, Zeng Huang, Kripasindhu Sarkar, Danhang Tang, Di Qiu, Abhimitra Meka, Ruofei Du, Mingsong Dou, Sergio Orts, Rohit Pandey, Ping Tan 0002, Thabo Beeler, Sean Ryan Fanello, Yinda Zhang 0001
CVPR6
2023 Double-Layer Optimization of Industrial-Park Energy System Based on Discrete Hybrid Automaton
abstract
Combing advanced measurement infrastructure (AMI) and information and communications technology (ICT), the Internet of Things (IoT) has been widely used in the energy system. The continuous state modeling method cannot effectively deal with the various operation scenarios in the park multiple energy system. This article proposes a hybrid state system modeling method based on the discrete hybrid automaton (DHA). The DHA-based model judges the equipment mode through the logical discriminant and describes the energy transmission process of the equipment with a switched affine system. And a double-layer optimization structure for the multiple energy system is proposed, which combine the long-term forecasting data and short-term forecasting data. Based on the AMI, the forecasting system predicts the system state and determines the scheduling plan of each equipment. This article uses a typical industrial-park case with multiple energy system to verify advantages of the DHA-based modeling method, calculates the economic benefits brought by the optimization and evaluates the maximum carbon emission reduction capacity of the industrial park.
Di Qiu, Dong Liu 0018, Fei Gao 0004
IEEE Internet Things J.1
2021 Learn to Cluster Faces via Pairwise Classification
abstract
Face clustering plays an essential role in exploiting massive unlabeled face data. Recently, graph-based face clustering methods are getting popular for their satisfying performances. However, they usually suffer from excessive memory consumption especially on large-scale graphs, and rely on empirical thresholds to determine the connectivities between samples in inference, which restricts their applications in various real-world scenes. To address such problems, in this paper, we explore face clustering from the pairwise angle. Specifically, we formulate the face clustering task as a pairwise relationship classification task, avoiding the memory-consuming learning on large-scale graphs. The classifier can directly determine the relationship between samples and is enhanced by taking advantage of the contextual information. Moreover, to further facilitate the efficiency of our method, we propose a rank-weighted density to guide the selection of pairs sent to the classifier. Experimental results demonstrate that our method achieves state-of-the-art performances on several public clustering benchmarks at the fastest speed and shows a great advantage in comparison with graph-based clustering methods on memory consumption.
Junfu Liu, Di Qiu, Xiaolin Wei
ICCV2
2020 Towards Geometry Guided Neural Relighting with Flash Photography
abstract
Previous image based relighting methods require capturing multiple images to acquire high frequency lighting effect under different lighting conditions, which needs nontrivial effort and may be unrealistic in certain practical use scenarios. While such approaches rely entirely on cleverly sampling the color images under different lighting conditions, little has been done to utilize geometric information that crucially influences the high-frequency features in the images, such as glossy highlight and cast shadow. We therefore propose a framework for image relighting from a single flash photograph with its corresponding depth map using deep learning. By incorporating the depth map, our approach is able to extrapolate realistic high-frequency effects under novel lighting via geometry guided image decomposition from the flashlight image, and predict the cast shadow map from the shadow-encoding transformed depth map. Moreover, the single-image based setup greatly simplifies the data capture process. We experimentally validate the advantage of our geometry guided approach over state-of-the-art image-based approaches in intrinsic image decomposition and image relighting, and also demonstrate our performance on real mobile phone photo examples.
Di Qiu, Jin Zeng 0004, Zhanghan Ke, Wenxiu Sun, Chengxi Yang
3DV1
2020 Guided Collaborative Training for Pixel-Wise Semi-Supervised Learning
Zhanghan Ke, Di Qiu, Kaican Li, Qiong Yan, Rynson W. H. Lau
ECCV (13)2
2020 Adapting Object Detectors with Conditional Domain Normalization
Kun Wang 0056, Xingyu Zeng, Shixiang Tang, Dapeng Chen, Di Qiu, Xiaogang Wang 0001
ECCV (11)6
2019 Deep End-to-End Alignment and Refinement for Time-of-Flight RGB-D Module
abstract
Recently, it is increasingly popular to equip mobile RGB cameras with Time-of-Flight (ToF) sensors for active depth sensing. However, for off-the-shelf ToF sensors, one must tackle two problems in order to obtain high-quality depth with respect to the RGB camera, namely 1) online calibration and alignment; and 2) complicated error correction for ToF depth sensing. In this work, we propose a framework for jointly alignment and refinement via deep learning. First, a cross-modal optical flow between the RGB image and the ToF amplitude image is estimated for alignment. The aligned depth is then refined via an improved kernel predicting network that performs kernel normalization and applies the bias prior to the dynamic convolution. To enrich our data for end-to-end training, we have also synthesized a dataset using tools from computer graphics. Experimental results demonstrate the effectiveness of our approach, achieving state-of-the-art for ToF refinement.
Di Qiu, Jiahao Pang, Wenxiu Sun, Chengxi Yang
ICCV1
2019 Computing Quasi-Conformal Folds
abstract
Computing surface folding maps has numerous applications ranging from computer graphics to material design. In this work we propose a novel way of computing surface folding maps via solving a linear PDE. This framework is a generalization of the existing computational quasi-conformal geometry and allows precise control of the geometry of folding. This property comes from a crucial quantity that occurs as the coefficient of the equation, namely, the alternating Beltrami coefficient. This approach also enables us to solve an inverse problem of parametrizing the folded surface given only partial data with known folding topology. Various interesting applications such as fold sculpting on 3 dimensional models, study of Miura-ori patterns, and self-occlusion reasoning are demonstrated to show the effectiveness of our method.
Di Qiu, Ka-Chun Lam, Lok Ming Lui
SIAM J. Imaging Sci.1
2012 Distributed tracking fidelity-metric performance analysis using confusion matrices
Erik Blasch, Ondrej Straka, Di Qiu, Miroslav Simandl, Jirí Ajgl
FUSION4
2010 Detecting Design Pattern Using Subgraph Discovery
Ming Qiu, Qingshan Jiang, An Gao, Ergan Chen, Di Qiu, Shang Chai
ACIIDS (1)5