VLDB 2026 Research / reviewers in the wild / expert
Yipeng Qin
dblp:169/5516
· DBLP profile ↗
41ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0002-1551-9126ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 3 first-author · 27 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved Cinematic-Guided Camera Language Transfer in 3D SceneabstractDirectors and cinematographers often recreate iconic scenes by replicating the underlying camera language to evoke shared aesthetic and narrative meaning. In this work, we refer to this as the task of Cinematic-Guided Camera Language Transfer, where the goal is to reproduce the cinematic camera language of a reference video clip in a new 3D scene. The pioneer work, Jaws [62], tackles this problem by adapting generic computer vision methods but fails to model the essential principles of cinematography, often leading to inaccurate framing, motion mismatches, and loss of expressive intent. To overcome these limitations, we systematically define the objectives of camera language transfer, grounding them in professional cinematography literature. Specifically, we conduct an in-depth review of cinematography literature to identify eight key cinematic features and encode them into five novel camera language losses. These losses not only guide optimization of camera parameters for effective transfer, but also serve as quantitative metrics for evaluating cinematographic fidelity. Extensive experiments demonstrate the superiority of our method. Zhuoling Jiang, Bailin Deng, Yipeng Qin |
3DV | 4 |
| 2026 | Improving Sparse IMU-based Motion Capture with Motion Label SmoothingabstractSparse Inertial Measurement Units (IMUs) based human motion capture has gained significant momentum, driven by the adaptation of fundamental AI tools such as recurrent neural networks (RNNs) and transformers that are tailored for temporal and spatial modeling. Despite these achievements, current research predominantly focuses on pipeline and architectural designs, with comparatively little attention given to regularization methods, highlighting a critical gap in developing a comprehensive AI toolkit for this task. To bridge this gap, we propose motion label smoothing, a novel method that adapts the classic label smoothing strategy from classification to the sparse IMU-based motion capture task. Specifically, we first demonstrate that a naive adaptation of label smoothing, including simply blending a uniform vector or a "uniform" motion representation (e.g., dataset-average motion or a canonical T-pose), is suboptimal; and argue that a proper adaptation requires increasing the entropy of the smoothed labels. Second, we conduct a thorough analysis of human motion labels, identifying three critical properties: 1) Temporal Smoothness, 2) Joint Correlation, and 3) Low-Frequency Dominance, and show that conventional approaches to entropy enhancement (e.g., blending Gaussian noise) are ineffective as they disrupt these properties. Finally, we propose the blend of a novel skeleton-based Perlin noise for motion label smoothing, designed to raise label entropy while satisfying motion properties. Extensive experiments applying our motion label smoothing to three state-of-the-art methods across four real-world IMU datasets demonstrate its effectiveness and robust generalization (plug-and-play) capability. Zhaorui Meng, Yangqing Hou, Anjun Chen, Shihui Guo, Yipeng Qin |
AAAI | 6 |
| 2025 | Hierarchically Controlled Deformable 3D Gaussians for Talking Head SynthesisabstractAudio-driven talking head synthesis is a critical task in digital human modeling. While recent advances using diffusion models and Neural Radiance Fields (NeRF) have improved visual quality, they often require substantial computational resources, limiting practical deployment. We present a novel framework for audio-driven talking head synthesis, namely it Hierarchically Controlled Deformable 3D Gaussians (HiCoDe), which achieves state-of-the-art performance with significantly reduced computational costs. Our key contribution is a hierarchical control strategy that effectively bridges the gap between sparse audio features and dense 3D Gaussian point clouds. Specifically, this strategy comprises two control levels: i) coarse-level control based on a 3D Morphable Model (3DMM) and ii) fine-level control using facial landmarks. Extensive experiments on the HDTF dataset and additional test sets demonstrate that our method outperforms existing approaches in visual quality, facial landmark accuracy, and audio-visual synchronization while being more computationally efficient in both training and inference. Linxuan Jiang, Chaowei Fang, Yipeng Qin, Guanbin Li |
AAAI | 5 |
| 2025 | VTON 360: High-Fidelity Virtual Try-On from Any Viewing DirectionabstractVirtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-view rendering. Specifically, we leverage the equivalence between a 3D model and its rendered multi-view 2D images, and reformulate 3D VTON as an extension of 2D VTON that ensures 3D consistent results across multiple views. To achieve this, we extend 2D VTON models to include multi-view garments and clothing-agnostic human body images as input, and propose several novel techniques to enhance them, including: i) a pseudo-3D pose representation using normal maps derived from the SMPL-X 3D human model, ii) a multi-view spatial attention mechanism that models the correlations between features from different viewing angles, and iii) a multi-view CLIP embedding that enhances the garment CLIP features used in 2D VTON with camera information. Extensive experiments on large-scale real datasets and clothing images from e-commerce platforms demonstrate the effectiveness of our approach. Project page: https://scnuhealthy.github.io/VTON360. Yuwei Ning, Yipeng Qin, Guangrun Wang, Sibei Yang, Liang Lin 0004, Guanbin Li |
CVPR | 3 |
| 2025 | LLM-driven Multimodal and Multi-Identity Listening Head GenerationabstractGenerating natural listener responses in conversational scenarios is crucial for creating engaging digital humans and avatars. Recent work has shown that large language models (LLMs) can be effectively leveraged for this task, demonstrating remarkable capabilities in generating contextually appropriate listener behaviors. However, current LLM-based methods face two critical limitations: they rely solely on speech content, overlooking other crucial communication signals, and they entangle listener identity with response generation, compromising output fidelity and generalization. In this work, we present a novel framework that addresses these limitations while maintaining the advantages of LLMs. Our approach introduces a Multimodal-LM architecture that jointly processes speech content, acoustics, and speaker emotion, capturing the full spectrum of communication cues. Additionally, we propose an identity disentanglement strategy using instance normalization and adaptive instance normalization in a VQ-VAE framework, enabling high-fidelity listening head synthesis with flexible identity control. Extensive experiments demonstrate that our method significantly outperforms existing approaches in terms of response naturalness and fidelity, while enabling effective identity control without retraining. Peiwen Lai, Weizhi Zhong, Yipeng Qin, Xiaohang Ren, Baoyuan Wang, Guanbin Li |
CVPR | 3 |
| 2025 | MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisabstractWhile data augmentation (DA) has been extensively studied in computer vision, its application to Inertial Measurement Unit (IMU) signals remains largely unexplored, despite IMUs’ growing importance in human motion analysis. In this paper, we present the first systematic study of IMU-specific data augmentation, beginning with a comprehensive analysis that identifies three fundamental properties of IMU signals: their time-series nature, inherent multimodality (rotation and acceleration) and motion-consistency characteristics. Through this analysis, we demonstrate the limitations of applying conventional time-series augmentation techniques to IMU data. We then introduce Motion-Drift Augmentation (MODA), a novel technique that simulates the natural displacement of body-worn IMUs during motion. We evaluate our approach across five diverse datasets and five deep learning settings, including i) fully-supervised, ii) semi-supervised, iii) domain adaptation, iv) domain generalization and v) few-shot learning for both Human Action Recognition (HAR) and Human Pose Estimation (HPE) tasks. Experimental results show that our proposed MODA consistently outperforms existing augmentation methods, with semi-supervised learning performance approaching state-of-the-art fully-supervised methods. Yinghao Wu, Shihui Guo, Yipeng Qin |
CVPR | 3 |
| 2025 | ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion CaptureabstractSparse inertial measurement units (IMUs) provide a portable, low-cost solution for human motion tracking but struggle with error accumulation from drift and sensor noise when estimating joint position through time-based linear acceleration integration (i.e., indirect measurement).
To address this, we propose ToF-IP, a novel 3D full-body pose estimation system that integrates Time-of-Flight (ToF) sensors with sparse IMUs.
The distinct advantage of our approach is that ToF sensors provide direct distance measurements, effectively mitigating error accumulation without relying on indirect time-based integration.
From a hardware perspective, we maintain the portability of existing solutions by attaching ToF sensors to selected IMUs with a negligible volume increase of just 3\%.
On the software side, we introduce two novel techniques to enhance multi-sensor integration: (i) a Node-Centric Data Integration strategy that leverages a Transformer encoder to explicitly model both intra-node and inter-node data integration by treating each sensing node as a token; and (ii) a Dynamic Spatial Positional Encoding scheme that encodes the continuously changing spatial positions of wearable nodes as motion-conditioned functions, enabling the model to better capture human body dynamics in the embedding space.Additionally, we contribute a 208-minute human motion dataset from 10 participants, including synchronized IMU-ToF measurements and ground-truth from optical tracking.
Extensive experiments demonstrate that our method outperforms state-of-the-art approaches such as PNP, achieving superior accuracy in tracking complex and slow motions like Tai Chi, which remains challenging for inertial-only methods. Shifan Jiang, Yangqing Hou, Chengxu Zuo, Shihui Guo, Yipeng Qin |
NeurIPS | 7 |
| 2025 | WristSketcher: Creating 2D Dynamic Sketches in AR With a Sensing WristbandabstractRestricted by the limited interaction area of native AR glasses, creating sketches is a challenge in it. Existing solutions attempt to use mobile devices (e.g., tablets) or mid-air hand gestures to expand the interactive spaces and as the 2D/3D sketching input interfaces for AR glasses. Between them, mobile devices allow for accurate sketching but are often heavy to carry. Sketching with bare hands is zero-burden but can be inaccurate due to arm instability. In addition, mid-air sketching can easily lead to social misunderstandings and its prolonged use can cause arm fatigue. In this work, we present WristSketcher, a new AR system based on a flexible sensing wristband that enables users to place multiple virtual plane canvases in the real environment and create 2D dynamic sketches based on them, featuring an almost zero-burden authoring model for accurate and comfortable sketch creation in real-world scenarios. Specifically, we streamlined the interaction space from the mid-air to the surface of a lightweight sensing wristband, and implemented AR sketching and associated interaction commands by developing a gesture recognition method based on the sensing pressure points. We designed a set of interactive gestures consisting of Long Press, Tap and Double Tap based on a heuristic study involving 26 participants. These gestures are correspondingly mapped to various command interactions using a combination of multi-touch and hotspots. Moreover, we endow our WristSketcher with the ability of animation creation, allowing it to create dynamic and expressive sketches. Experimental results demonstrate that our WristSketcher (i) recognizes users’ gesture interactions with a high accuracy of 95.9%; (ii) achieves higher sketching accuracy than Freehand sketching; (iii) achieves high user satisfaction in ease of use, usability and functionality; and (iv) shows innovation potentials in art creation, memory aids, and entertainment applications. Enting Ying, Tianyang Xiong, Gaoxiang Zhu, Ming Qiu, Yipeng Qin, Shihui Guo |
Int. J. Hum. Comput. Interact. | 5 |
| 2025 | Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion CaptureabstractIn this paper, we propose a novel dynamic calibration method for sparse inertial motion capture systems, which is the first to break the restrictive absolute static assumption in IMU calibration, i.e., the coordinate drift R G′ G and measurement offset R BS remain constant during the entire motion, thereby significantly expanding their application scenarios. Specifically, we achieve real-time estimation of R G′ G and R BS under two relaxed assumptions: i) the matrices change negligibly in a short time window; ii) the human movements/IMU readings are diverse in such a time window. Intuitively, the first assumption reduces the number of candidate matrices, and the second assumption provides diverse constraints, which greatly reduces the solution space and allows for accurate estimation of R G′ G and R BS from a short history of IMU readings in real time. To achieve this, we created synthetic datasets of paired R G′ G , R BS matrices and IMU readings, and learned their mappings using a Transformer-based model. We also designed a calibration trigger based on the diversity of IMU readings to ensure that assumption ii) is met before applying our method. To our knowledge, we are the first to achieve implicit IMU calibration (i.e., seamlessly putting IMUs into use without the need for an explicit calibration process), as well as the first to enable long-term and accurate motion capture using sparse IMUs. The code and dataset are available at https://github.com/ZuoCX1996/TIC. Chengxu Zuo, Xiangren Shi, Xinyu Yi, Feng Xu 0005, Shihui Guo, Yipeng Qin |
ACM Trans. Graph. | 10 |
| 2025 | Diverse Motion In-Betweening From Sparse Keyframes With Dual Posture StitchingabstractIn-betweening is a technique for generating transitions given start and target character states. The majority of existing works require multiple (often 10) frames as input, which are not always available. In addition, they produce results that lack diversity, which may not fulfill artists' requirements. Addressing these gaps, our work deals with a focused yet challenging problem: generating diverse and high-quality transitions given exactly two frames (only the start and target frames). To cope with this challenging scenario, we propose a bi-directional motion generation and stitching scheme which generates forward and backward transitions from the start and target frames with two adversarial autoregressive networks, respectively, and stitches them midway between the start and target frames. In contrast to stitching at the start or target frames, where the ground truth cannot be altered, there is no strict midway ground truth. Thus, our method can capitalize on this flexibility and generate high-quality and diverse transitions simultaneously. Specifically, we employ conditional variational autoencoders (CVAEs) to implement our autoregressive networks and propose a novel stitching loss to stitch the bi-directional generated motions around the midway point. Extensive experiments demonstrate that our method achieves higher motion quality and more diverse results than existing methods on the LaFAN1, Human3.6m and AMASS datasets. Tianxiang Ren, Jubo Yu, Shihui Guo, Yutao Ouyang, Zijiao Zeng, Yazhan Zhang, Yipeng Qin |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | NeRF-HuGS: Improved Neural Radiance Fields in Non-static Scenes Using Heuristics-Guided SegmentationabstractNeural Radiance Field (NeRF) has been widely recognized for its excellence in novel view synthesis and 3D scene reconstruction. However, their effectiveness is in-herently tied to the assumption of static scenes, rendering them susceptible to undesirable artifacts when confronted with transient distractors such as moving objects or shad-ows. In this work, we propose a novel paradigm, namely “Heuristics-Guided Segmentation” (HuGS), which signifi-cantly enhances the separation of static scenes from tran-sient distractors by harmoniously combining the strengths of hand-crafted heuristics and state-of-the-art segmentation models, thus significantly transcending the limitations of previous solutions. Furthermore, we delve into the metic-ulous design of heuristics, introducing a seamless fusion of Structure-from-Motion (SfM)-based heuristics and color residual heuristics, catering to a diverse range of texture profiles. Extensive experiments demonstrate the superiority and robustness of our method in mitigating transient dis-tractors for NeRFs trained in non-static scenes. Project page: https://cnhaox.github.io/NeRF-HuGS/ Yipeng Qin, Lingjie Liu, Jiangbo Lu, Guanbin Li |
CVPR | 2 |
| 2024 | Deep Generative Model based Rate-Distortion for Image Downscaling AssessmentabstractIn this paper, we propose Image Downscaling Assessment by Rate-Distortion (IDA-RD), a novel measure to quantitatively evaluate image downscaling algorithms. In contrast to image-based methods that measure the quality of downscaled images, ours is process-based that draws ideas from rate-distortion theory to measure the distortion incurred during downscaling. Our main idea is that downscaling and super-resolution (SR) can be viewed as the encoding and decoding processes in the rate-distortion model, respectively, and that a downscaling algorithm that preserves more details in the resulting low-resolution (LR) images should lead to less dis-torted high-resolution (HR) images in SR. In other words, the distortion should increase as the downscaling algorithm deteriorates. However, it is non-trivial to measure this distortion as it requires the SR algorithm to be blind and stochastic. Our key insight is that such requirements can be met by re-cent SR algorithms based on deep generative models that can find all matching HR images for a given LR image on their learned manifolds. Extensive experimental results show the effectiveness of our IDA-RD measure. Our code is available at: https://github.com/Byronliang8/Ida-Rd Yuanbang Liang, Bhavesh Garg, Paul L. Rosin, Yipeng Qin |
CVPR | 4 |
| 2024 | PICTURE: PhotorealistIC Virtual Try-on from UnconstRained dEsignsabstractIn this paper, we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types, our method allows flexible specification of style (text or image) and texture (full garment, cropped sections, or texture patches) conditions. To address the entanglement challenge when using full garment images as conditions, we develop a two-stage pipeline with explicit disentanglement of style and texture. In the first stage, we generate a human parsing map reflecting the desired style conditioned on the input. In the second stage, we composite textures onto the parsing map areas based on the texture input. To represent complex and non-stationary textures that have never been achieved in previous fashion editing works, we first propose extracting hierarchical and balanced CLIP features and applying position encoding in VTON. Experiments demonstrate superior synthesis quality and personalization enabled by our method. The flexible control over style and texture mixing brings virtual try-on to a new level of user experience for online shopping and fashion design. Shuliang Ning, Duomin Wang, Yipeng Qin, Zirong Jin, Baoyuan Wang, Xiaoguang Han 0001 |
CVPR | 3 |
| 2024 | Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear JacketabstractExisting wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing, limiting their application in everyday life. In this paper, we introduce Loose Inertial Poser, a novel motion capture solution with high wearing comfortableness, by integrating four Inertial Measurement Units (IMUs) into a loose-wear jacket. Specifically, we address the challenge of scarce loose-wear IMU training data by proposing a Secondary Motion AutoEncoder (SeMo-AE) that learns to model and synthesize the effects of secondary motion between the skin and loose clothing on IMU data. SeMo-AE is leveraged to generate a diverse synthetic dataset of loose-wear IMU data to augment training for the pose estimation network and significantly improve its accuracy. For validation, we collected a dataset with various subjects and 2 wearing styles (zipped and unzipped). Experimental results demonstrate that our approach maintains high-quality real-time posture estimation even in loose-wear scenarios. Our dataset and code are available at: https://github.com/ZuoCX1966/Loose-Inertial-Poser Chengxu Zuo, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu 0005, Yipeng Qin |
CVPR | 7 |
| 2024 | Dcctnet: Kidney Tumors Segmentation Based On Dual-Level Combination Of Cnn And TransformerabstractThe hybrid model of CNN(Convolution Neural Networks) and Transformer is a popular method in segmenting kidney images, but most existing hybrid models directly fused local features from CNN with global features from Transformer, ignoring the issue of semantic gaps between distinct features. Furthermore, feature fusion is typically performed solely at the feature level, without considering alignment at the mask (prediction map) level. To address these limitations, we propose a novel segmentation method called Dual-level Combination of CNN and Transformers Network (DCCTNet). Specifically, we select similar features from both CNN and Transformer to reduce semantic gaps at the feature level. Additionally, we further utilize the global information of the Transformers by reducing the difference between the prediction maps in the coding stage at the mask level. We evaluate DCCTNet on the KiTS19 dataset, achieving $97.3 \%$ dice score for kidneys segmentation and $81.2 \%$ dice score for kidney tumors segmentation, respectively. https://github.com/hou-bz/DCCTNet. Bingzhen Hou, Gui-mei Zhang, Huiqun Liu, Yipeng Qin, Ying Chen 0023 |
ICIP | 4 |
| 2024 | SuDA: Support-based Domain Adaptation for Sim2Real Hinge Joint Tracking with Flexible SensorsabstractFlexible sensors hold promise for human motion capture (MoCap), offering advantages such as wearability, privacy preservation, and minimal constraints on natural movement. However, existing flexible sensor-based MoCap methods rely on deep learning and necessitate large and diverse labeled datasets for training. These data typically need to be collected in MoCap studios with specialized equipment and substantial manual labor, making them difficult and expensive to obtain at scale. Thanks to the high-linearity of flexible sensors, we address this challenge by proposing a novel Sim2Real solution for hinge joint tracking based on domain adaptation, eliminating the need for labeled data yet achieving comparable accuracy to supervised learning. Our solution relies on a novel Support-based Domain Adaptation method, namely SuDA, which aligns the supports of the predictive functions rather than the instance-dependent distributions between the source and target domains. Extensive experimental results demonstrate the effectiveness of our method and its superiority overstate-of-the-art distribution-based domain adaptation methods in our task. Jiawei Fang, Haishan Song, Chengxu Zuo, Xiaoxia Gao, Xiaowei Chen 0017, Shihui Guo, Yipeng Qin |
ICML | 7 |
| 2024 | Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware SamplingabstractDespite impressive results, deep generative models require massive datasets for training, and as dataset size increases, effective evaluation metrics like precision and recall (P&R) become computationally infeasible on commodity hardware. In this paper, we address this challenge by proposing efficient P&R (eP&R) metrics that give almost identical results as the original P&R but with much lower computational costs. Specifically, we identify two redundancies in the original P&R: i) redundancy in ratio computation and ii) redundancy in manifold inside/outside identification. We find both can be effectively removed via hubness-aware sampling, which extracts representative elements from synthetic/real image samples based on their hubness values, i.e., the number of times a sample becomes a k-nearest neighbor to others in the feature space. Thanks to the insensitivity of hubness-aware sampling to exact k-nearest neighbor (k-NN) results, we further improve the efficiency of our eP&R metrics by using approximate k-NN methods. Extensive experiments show that our eP&R matches the original P&R but is far more efficient in time and space. Our code is available at: https://github.com/Byronliang8/Hubness_Precision_Recall Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICML | 4 |
| 2024 | SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground TactilityabstractEstimating 3D human poses from monocular images is an important research area with many practical applications. However, the depth ambiguity of 2D solutions limits their accuracy in actions where occlusion exits or where slight centroid shifts can result in significant 3D pose variations. In this paper, we introduce a novel multimodal approach to mitigate the depth ambiguity inherent in monocular solutions by integrating spatial-aware pressure information. We first establish a data collection system with a pressure mat and a monocular camera, and construct a large-scale multimodal human activity dataset comprising over 600,000 frames of motion data. Utilizing this dataset, we propose a pressure image reconstruction network to extract pressure priors from monocular images. Subsequently, we introduce a Transformer-based multimodal pose estimation network to combine pressure priors with monocular images, achieving a world mean per joint position error of 51.6mm, outperforming state-of-the-art methods. Extensive experiments demonstrate the effectiveness of our multimodal 3D human pose estimation method across various actions and joints, highlighting the significance of spatial-aware pressure in improving the accuracy of monocular-vision-based methods. Our dataset is available at: https://github.com/LishuangZhan/SATPose. Lishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo, Boyu Gao 0003, Yipeng Qin |
ACM Multimedia | 6 |
| 2024 | Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and ModulationabstractTransformer models excel at capturing long-range dependencies in sequential data, but lack explicit mechanisms to leverage structural patterns inherent in fixed-length input sequences.
In this paper, we propose a novel sequence structure learning and modulation approach that endows Transformers with the ability to model and utilize such fixed-sequence structural properties for improved performance on inertial pose estimation tasks.
Specifically, our method introduces a Sequence Structure Module (SSM) that utilizes structural information of fixed-length inertial sensor readings to adjust the input features of transformers.
Such structural information can either be acquired by learning or specified based on users' prior knowledge.
To justify the prospect of our approach, we show that i) injecting spatial structural information of IMUs/joints learned from data improves accuracy, while ii) injecting temporal structural information based on smooth priors reduces jitter (i.e., improves steadiness), in a spatial-temporal transformer solution for inertial pose estimation.
Extensive experiments across multiple benchmark datasets demonstrate the superiority of our approach against state-of-the-art methods and has the potential to advance the design of the transformer architecture for fixed-length sequences. Yinghao Wu, Chaoran Wang, Shihui Guo, Yipeng Qin |
NeurIPS | 5 |
| 2024 | Universal Semi-supervised Model Adaptation via Collaborative Consistency TrainingabstractIn this paper, we introduce a realistic and challenging domain adaptation problem called Universal Semi-supervised Model Adaptation (USMA), which i) requires only a pre-trained source model, ii) allows the source and target domain to have different label sets, i.e., they share a common label set and hold their own private label set, and iii) requires only a few labeled samples in each class of the target domain. To address USMA, we propose a collaborative consistency training framework that regularizes the prediction consistency between two models, i.e., a pre-trained source model and its variant pre-trained with target data only, and combines their complementary strengths to learn a more powerful model. The rationale of our framework stems from the observation that the source model performs better on common categories than the target-only model, while on target-private categories, the target-only model performs better. We also propose a two-perspective, i.e., sample-wise and class-wise, consistency regularization to improve the training. Experimental results demonstrate the effectiveness of our method on several benchmark datasets. Zizheng Yan, Yushuang Wu, Yipeng Qin, Xiaoguang Han 0001, Shuguang Cui, Guanbin Li |
WACV | 3 |
| 2024 | Exploration and Exploitation of Unlabeled Data for Open-Set Semi-supervised Learning
Ganlong Zhao, Guanbin Li, Yipeng Qin, Zhenhua Chai, Xiaolin Wei, Liang Lin 0004, Yizhou Yu |
Int. J. Comput. Vis. | 3 |
| 2024 | Full-body Human Motion Reconstruction with Sparse Joint Tracking Using Flexible SensorsabstractHuman motion tracking is a fundamental building block for various applications including computer animation, human-computer interaction, healthcare, and so on. To reduce the burden of wearing multiple sensors, human motion prediction from sparse sensor inputs has become a hot topic in human motion tracking. However, such predictions are non-trivial as (i) the widely adopted data-driven approaches can easily collapse to average poses, and (ii) the predicted motions contain unnatural jitters. In this work, we address the aforementioned issues by proposing a novel framework which can accurately predict the human joint moving angles from the signals of only four flexible sensors, thereby achieving the tracking of human joints in multi-degrees of freedom. Specifically, we mitigate the collapse to average poses by implementing the model with a Bi-LSTM neural network that makes full use of short-time sequence information; we reduce jitters by adding a median pooling layer to the network, which smooths consecutive motions. Although being bio-compatible and ideal for improving the wearing experience, the flexible sensors are prone to aging which increases prediction errors. Observing that the aging of flexible sensors usually results in drifts of their resistance ranges, we further propose a novel dynamic calibration technique to rescale sensor ranges, which further improves the prediction accuracy. Experimental results show that our method achieves a low and stable tracking error of 4.51 degrees across different motion types with only four sensors. Xiaowei Chen 0017, Lishuang Zhan, Shihui Guo, Qunsheng Ruan, Guoliang Luo, Minghong Liao, Yipeng Qin |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2023 | Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentabstractAudio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance Fields), thus suffering from the trade-offs between interpretability and expressive power, hence between controllability and quality of the results. In this work, we break these trade-offs with our novel parametric implicit face representation and propose a novel audio-driven facial reenactment framework that is both controllable and can generate high-quality talking heads. Specifically, our parametric implicit representation parameterizes the implicit representation with interpretable parameters of 3D face models, thereby taking the best of both explicit and implicit methods. In addition, we propose several new techniques to improve the three components of our framework, including i) incorporating contextual information into the audio-to-expression parameters encoding; ii) using conditional image synthesis to parameterize the implicit representation and implementing it with an innovative tri-plane structure for efficient learning; iii) formulating facial reenactment as a conditional image inpainting problem and proposing a novel data augmentation technique to improve model generalizability. Extensive experiments demonstrate that our method can generate more realistic results than previous methods with greater fidelity to the identities and talking styles of speakers. Ricong Huang, Peiwen Lai, Yipeng Qin, Guanbin Li |
CVPR | 3 |
| 2023 | Improved Distribution Matching for Dataset CondensationabstractDataset Condensation aims to condense a large dataset into a smaller one while maintaining its ability to train a well-performing model, thus reducing the storage cost and training effort in deep learning applications. However, conventional dataset condensation methods are optimization-oriented and condense the dataset by performing gradient or parameter matching during model optimization, which is computationally intensive even on small datasets and models. In this paper, we propose a novel dataset condensation method based on distribution matching, which is more efficient and promising. Specifically, we identify two important shortcomings of naive distribution matching (i.e., imbalanced feature numbers and unvalidated embeddings for distance computation) and address them with three novel techniques (i.e., partitioning and expansion augmentation, efficient and enriched model sampling, and class-aware distribution regularization). Our simple yet effective method outperforms most previous optimization-oriented methods with much fewer computational resources, thereby scaling data condensation to larger datasets and models. Extensive experiments demonstrate the effectiveness of our method. Codes are available at https://github.com/uitrbn/IDM Ganlong Zhao, Guanbin Li, Yipeng Qin, Yizhou Yu |
CVPR | 3 |
| 2023 | Feature Proliferation - the "Cancer" in StyleGAN and its TreatmentsabstractDespite the success of StyleGAN in image synthesis, the images it synthesizes are not always perfect and the well-known truncation trick has become a standard post-processing technique for StyleGAN to synthesize high quality images. Although effective, it has long been noted that the truncation trick tends to reduce the diversity of synthesized images and unnecessarily sacrifices many distinct image features. To address this issue, in this paper, we first delve into the StyleGAN image synthesis mechanism and discover an important phenomenon, namely Feature Proliferation, which demonstrates how specific features reproduce with forward propagation. Then, we show how the occurrence of Feature Proliferation results in StyleGAN image artifacts. As an analogy, we refer to it as the "cancer" in StyleGAN from its proliferating and malignant nature. Finally, we propose a novel feature rescaling method that identifies and modulates risky features to mitigate feature proliferation. Thanks to our discovery of Feature Proliferation, the proposed feature rescaling method is less destructive and retains more useful image features than the truncation trick, as it is more fine-grained and works in a lower-level feature space rather than a high-level latent space. Experimental results justify the validity of our claims and the effectiveness of the proposed feature rescaling method. Our code is available at https://github.com/songc42/Feature-proliferation. Shuang Song 0012, Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICCV | 5 |
| 2023 | Self-Adaptive Motion Tracking against On-body Displacement of Flexible SensorsabstractFlexible sensors are promising for ubiquitous sensing of human status due to their flexibility and easy integration as wearable systems. However, on-body displacement of sensors is inevitable since the device cannot be firmly worn at a fixed position across different sessions. This displacement issue causes complicated patterns and significant challenges to subsequent machine learning algorithms. Our work proposes a novel self-adaptive motion tracking network to address this challenge. Our network consists of three novel components: i) a light-weight learnable Affine Transformation layer whose parameters can be tuned to efficiently adapt to unknown displacements; ii) a Fourier-encoded LSTM network for better pattern identification; iii) a novel sequence discrepancy loss equipped with auxiliary regressors for unsupervised tuning of Affine Transformation parameters. Chengxu Zuo, Jiawei Fang, Shihui Guo, Yipeng Qin |
NeurIPS | 4 |
| 2023 | Computational Design of Wiring Layout on Tight Suits with Minimal Motion ResistanceabstractAn increasing number of electronics are directly embedded on the clothing to monitor human status (e.g., skeletal motion) or provide haptic feedback. A specific challenge to prototype and fabricate such a clothing is to design the wiring layout, while minimizing the intervention to human motion. We address this challenge by formulating the topological optimization problem on the clothing surface as a deformation-weighted Steiner tree problem on a 3D clothing mesh. Our method proposed an energy function for minimizing strain energy in the wiring area under different motions, regularized by its total length. We built the physical prototype to verify the effectiveness of our method and conducted user study with participants of both design experts and smart cloth users. On three types of commercial products of smart clothing, the optimized layout design reduced wire strain energy by an average of 77% among 248 actions compared to baseline design, and 18% over the expert design. Kai Wang 0107, Yinping Zheng, Da Zhou, Shihui Guo, Yipeng Qin, Xiaohu Guo |
SIGGRAPH Asia | 6 |
| 2023 | Reduced-Reference Quality Assessment of Point Clouds via Content-Oriented Saliency ProjectionabstractMany dense 3D point clouds have been exploited to represent visual objects instead of traditional images or videos. To evaluate the perceptual quality of various point clouds, in this letter, we propose a novel and efficient Reduced-Reference quality metric for point clouds, which is based on Content-oriented sAliency Projection (RR-CAP). Specifically, we make the first attempt to simplify reference and distorted point clouds into projected saliency maps with a downsampling operation. Through this process, we tackle the issue of transmitting large-volume original point clouds to end-users for quality assessment. Then, motivated by the characteristics of the human visual system (HVS), the objective quality scores of distorted point clouds are produced by combining content-oriented similarity and statistical correlation measurements. Finally, extensive experiments are conducted on SJTU-PCQA and WPC databases. The experiment results demonstrate that our proposed algorithm outperforms existing reduced-reference and no-reference quality metrics, and significantly reduces the performance gap between state-of-the-art full-reference quality assessment methods. In addition, we show the performance variation of each proposed technical component by ablation tests. Wei Zhou 0021, Guanghui Yue 0001, Ruizeng Zhang, Yipeng Qin, Hantao Liu |
IEEE Signal Process. Lett. | 4 |
| 2022 | Centrality and Consistency: Two-Stage Clean Samples Identification for Learning with Instance-Dependent Noisy Labels
Ganlong Zhao, Guanbin Li, Yipeng Qin, Feng Liu 0036, Yizhou Yu |
ECCV (25) | 3 |
| 2022 | Exploring and Exploiting Hubness Priors for High-Quality GAN Latent SamplingabstractDespite the extensive studies on Generative Adversarial Networks (GANs), how to reliably sample high-quality images from their latent spaces remains an under-explored topic. In this paper, we propose a novel GAN latent sampling method by exploring and exploiting the hubness priors of GAN latent distributions. Our key insight is that the high dimensionality of the GAN latent space will inevitably lead to the emergence of hub latents that usually have much larger sampling densities than other latents in the latent space. As a result, these hub latents are better trained and thus contribute more to the synthesis of high-quality images. Unlike the a posterior "cherry-picking", our method is highly efficient as it is an a priori method that identifies high-quality latents before the synthesis of images. Furthermore, we show that the well-known but purely empirical truncation trick is a naive approximation to the central clustering effect of hub latents, which not only uncovers the rationale of the truncation trick, but also indicates the superiority and fundamentality of our method. Extensive experimental results demonstrate the effectiveness of the proposed method. Our code is available at: https://github.com/Byronliang8/HubnessGANSampling. Yuanbang Liang, Jing Wu 0004, Yukun Lai, Yipeng Qin |
ICML | 4 |
| 2022 | Multi-level Consistency Learning for Semi-supervised Domain AdaptationabstractSemi-supervised domain adaptation (SSDA) aims to apply knowledge learned from a fully labeled source domain to a scarcely labeled target domain. In this paper, we propose a Multi-level Consistency Learning (MCL) framework for SSDA. Specifically, our MCL regularizes the consistency of different views of target domain samples at three levels: (i) at inter-domain level, we robustly and accurately align the source and target domains using a prototype-based optimal transport method that utilizes the pros and cons of different views of target samples; (ii) at intra-domain level, we facilitate the learning of both discriminative and compact target feature representations by proposing a novel class-wise contrastive clustering loss; (iii) at sample level, we follow standard practice and improve the prediction accuracy by conducting a consistency-based self-training. Empirically, we verified the effectiveness of our MCL framework on three popular SSDA benchmarks, i.e., VisDA2017, DomainNet, and Office-Home datasets, and the experimental results demonstrate that our MCL framework achieves the state-of-the-art performance. Zizheng Yan, Yushuang Wu, Guanbin Li, Yipeng Qin, Xiaoguang Han 0001, Shuguang Cui |
IJCAI | 4 |
| 2022 | Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution PriorsabstractA key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degradations (\eg, downsampling, noise and compression). Most previous works restore such missing details in the image space. To cope with the high diversity of natural images, they either rely on the unstable GANs that are difficult to train and prone to artifacts, or resort to explicit references from high-resolution (HR) images that are usually unavailable. In this work, we propose Feature Matching SR (FeMaSR), which restores realistic HR images in a much more compact feature space. Unlike image-space methods, our FeMaSR restores HR images by matching distorted LR image features to their distortion-free HR counterparts in our pretrained HR priors, and decoding the matched features to obtain realistic HR images. Specifically, our HR priors contain a discrete feature codebook and its associated decoder, which are pretrained on HR images with a Vector Quantized Generative Adversarial Network (VQGAN). Notably, we incorporate a novel semantic regularization in VQGAN to improve the quality of reconstructed images. For the feature matching, we first extract LR features with an LR encoder consisting of several Swin Transformer blocks and then follow a simple nearest neighbour strategy to match them with the pretrained codebook. In particular, we equip the LR encoder with residual shortcut connections to the decoder, which is critical to the optimization of feature matching loss and also helps to complement the possible feature matching errors.Experimental results show that our approach produces more realistic HR images than previous methods. Code will be made publicly available. Chaofeng Chen, Yipeng Qin, Xiaoming Li 0001, Xiaoguang Han 0001, Shihui Guo |
ACM Multimedia | 3 |
| 2022 | PVSeRF: Joint Pixel-, Voxel- and Surface-Aligned Radiance Field for Single-Image Novel View SynthesisabstractWe present PVSeRF, a learning framework that reconstructs neural radiance fields from single-view RGB images, for novel view synthesis. Previous solutions, such as pixelNeRF, rely only on pixel-aligned features and suffer from feature ambiguity issues. As a result, they struggle with the disentanglement of geometry and appearance, leading to implausible geometries and blurry results. To address this challenge, we propose to incorporate explicit geometry reasoning and combine it with pixel-aligned features for radiance field prediction. Specifically, in addition to pixel-aligned features, we further constrain the radiance field learning to be conditioned on i) voxel-aligned features learned from a coarse volumetric grid and ii) fine surface-aligned features extracted from a regressed point cloud. We show that the introduction of such geometry-aware features helps to achieve a better disentanglement between appearance and geometry, i.e. recovering more accurate geometries and synthesizing higher quality images of novel views. Extensive experiments against state-of-the-art methods on ShapeNet benchmarks demonstrate the superiority of our approach for single-image novel view synthesis. Xianggang Yu, Jiapeng Tang, Yipeng Qin, Chenghong Li, Xiaoguang Han 0001, Linchao Bao, Shuguang Cui |
ACM Multimedia | 3 |
| 2021 | Pixel-level Intra-domain Adaptation for Semantic SegmentationabstractRecent advances in unsupervised domain adaptation have achieved remarkable performance on semantic segmentation tasks. Despite such progress, existing works mainly focus on bridging the inter-domain gaps between the source and target domain, while only few of them noticed the intra-domain gaps within the target data. In this work, we propose a pixel-level intra-domain adaptation approach to reduce the intra-domain gaps within the target data. Compared with image-level methods, ours treats each pixel as an instance, which adapts the segmentation model at a more fine-grained level. Specifically, we first conduct the inter-domain adaptation between the source and target domain; Then, we separate the pixels in target images into the easy and hard subdomains; Finally, we propose a pixel-level adversarial training strategy to adapt a segmentation network from the easy to the hard subdomain. Moreover, we show that the segmentation accuracy can be further improved by incorporating a continuous indexing technique in the adversarial training. Experimental results show the effectiveness of our method against existing state-of-the-art approaches. Zizheng Yan, Xianggang Yu, Yipeng Qin, Yushuang Wu, Xiaoguang Han 0001, Shuguang Cui |
ACM Multimedia | 3 |
| 2021 | Human posture tracking with flexible sensors for motion recognitionabstractAbstract The integration of conventional clothes with flexible electronics is a promising solution as a future‐generation computing platform. However, the problem of user authentication on this novel platform is still underexplored. This work uses flexible sensors to track human posture and achieves the goal of user authentication. We capture human movement pattern by four stretch sensors around the shoulder and one on the elbow. We introduce the long short‐term memory fully convolutional network (LSTM‐FCN), which directly takes noisy and sparse sensor data as input and verifies its consistency with the user's predefined movement patterns. The method can identify a user by matching movement patterns even if there are large intrapersonal variations. The authentication accuracy of LSTM‐FCN reaches 98.0%, which is 10.7% and 6.5% higher than that of dynamic time warping and dynamic time warping dependent. Xiaowei Chen 0017, Yong Ma 0005, Shihui Guo, Yipeng Qin, Minghong Liao |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Image2StyleGAN++: How to Edit the Embedded Images?abstractWe propose Image2StyleGAN++, a flexible image editing framework with many applications. Our framework extends the recent Image2StyleGAN in three ways. First, we introduce noise optimization as a complement to the W+ latent space embedding. Our noise optimization can restore high frequency features in images and thus significantly improves the quality of reconstructed images, e.g. a big increase of PSNR from 20 dB to 45 dB. Second, we extend the global W+ latent space embedding to enable local embeddings. Third, we combine embedding with activation tensor manipulation to perform high quality local edits along with global semantic edits on images. Such edits motivate various high quality image editing applications, e.g. image reconstruction, image inpainting, image crossover, local style transfer, image editing using scribbles, and attribute level feature transfer. Examples of the edited images are shown across the paper for visual inspection. Rameen Abdal, Yipeng Qin, Peter Wonka |
CVPR | 2 |
| 2020 | SEAN: Image Synthesis With Semantic Region-Adaptive NormalizationabstractWe propose semantic region-adaptive normalization (SEAN), a simple but effective building block for Generative Adversarial Networks conditioned on segmentation masks that describe the semantic regions in the desired output image. Using SEAN normalization, we can build a network architecture that can control the style of each semantic region individually, e.g., we can specify one style reference image per region. SEAN is better suited to encode, transfer, and synthesize style than the best previous method in terms of reconstruction quality, variability, and visual quality. We evaluate SEAN on multiple datasets and report better quantitative metrics (e.g. FID, PSNR) than the current state of the art. SEAN also pushes the frontier of interactive image editing. We can interactively edit images by changing segmentation masks or the style for any given region. We can also interpolate styles from two reference images per region. Peihao Zhu 0001, Rameen Abdal, Yipeng Qin, Peter Wonka |
CVPR | 3 |
| 2020 | How Does Lipschitz Regularization Influence GAN Training?
Yipeng Qin, Niloy J. Mitra, Peter Wonka |
ECCV (16) | 1 |
| 2019 | Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?abstractWe propose an efficient algorithm to embed a given image into the latent space of StyleGAN. This embedding enables semantic image editing operations that can be applied to existing photographs. Taking the StyleGAN trained on the FFHD dataset as an example, we show results for image morphing, style transfer, and expression transfer. Studying the results of the embedding algorithm provides valuable insights into the structure of the StyleGAN latent space. We propose a set of experiments to test what class of images can be embedded, how they are embedded, what latent space is suitable for embedding, and if the embedding is semantically meaningful. Rameen Abdal, Yipeng Qin, Peter Wonka |
ICCV | 2 |
| 2017 | Fast and Memory-Efficient Voronoi Diagram Construction on Triangle MeshesabstractAbstract Geodesic based Voronoi diagrams play an important role in many applications of computer graphics. Constructing such Voronoi diagrams usually resorts to exact geodesics. However, exact geodesic computation always consumes lots of time and memory, which has become the bottleneck of constructing geodesic based Voronoi diagrams. In this paper, we propose the window‐VTP algorithm, which can effectively reduce redundant computation and save memory. As a result, constructing Voronoi diagrams using the proposed window‐VTP algorithm runs 3–8 times faster than Liu et al.'s method [ LCT11 ], 1.2 times faster than its FWP‐MMP variant and more importantly uses 10–70 times less memory than both of them. Yipeng Qin, Hongchuan Yu, Jian J. Zhang 0001 |
Comput. Graph. Forum | 1 |
| 2016 | Fast and exact discrete geodesic computation based on triangle-oriented wavefront propagationabstractComputing discrete geodesic distance over triangle meshes is one of the fundamental problems in computational geometry and computer graphics. In this problem, an effective window pruning strategy can significantly affect the actual running time. Due to its importance, we conduct an in-depth study of window pruning operations in this paper, and produce an exhaustive list of scenarios where one window can make another window partially or completely redundant. To identify a maximal number of redundant windows using such pairwise cross checking, we propose a set of procedures to synchronize local window propagation within the same triangle by simultaneously propagating a collection of windows from one triangle edge to its two opposite edges. On the basis of such synchronized window propagation, we design a new geodesic computation algorithm based on a triangle-oriented region growing scheme. Our geodesic algorithm can remove most of the redundant windows at the earliest possible stage, thus significantly reducing computational cost and memory usage at later stages. In addition, by adopting triangles instead of windows as the primitive in propagation management, our algorithm significantly cuts down the data management overhead. As a result, it runs 4--15 times faster than MMP and ICH algorithms, 2-4 times faster than FWP-MMP and FWP-CH algorithms, and also incurs the least memory usage. Yipeng Qin, Xiaoguang Han 0001, Hongchuan Yu, Yizhou Yu, Jian J. Zhang 0001 |
ACM Trans. Graph. | 1 |