Peng Dai 0002

dblp:08/3547-2 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
12since 2021 · last 2024
0000-0002-3015-7485ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang
ECCV (38)6
2024 GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
Yuxuan Mu, Xinxin Zuo, Chuan Guo 0002, Juwei Lu, Songcen Xu, Peng Dai 0002, Youliang Yan, Li Cheng 0001
ECCV (79)8
2024 Generative Human Motion Stylization in Latent Space
abstract
Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction and infusion. Building upon this, we present a novel \textit{generative} model that produces diverse stylization results of a single motion (latent) code. During training, a motion code is decomposed into two coding components: a deterministic content code, and a probabilistic style code adhering to a prior distribution; then a generator massages the random combination of content and style codes to reconstruct the corresponding motion codes. Our approach is versatile, allowing the learning of probabilistic style space from either style labeled or unlabeled motions, providing notable flexibility in stylization as well. In inference, users can opt to stylize a motion using style cues from a reference motion or a label. Even in the absence of explicit style input, our model facilitates novel re-stylization by sampling from the unconditional style prior distribution. Experimental results show that our proposed stylization models, despite their lightweight design, outperform the state-of-the-arts in style reeanactment, content preservation, and generalization across various applications and settings.
Chuan Guo 0002, Yuxuan Mu, Xinxin Zuo, Peng Dai 0002, Youliang Yan, Juwei Lu, Li Cheng 0001
ICLR4
2023 CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language Retrieval
abstract
Pre-training a vision-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vision-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capability of a large model to a small one while maintaining the accuracy, which has achieved remarkable success in natural language processing. However, it faces many problems when applying KD to the multi-modality applications. In this paper, we propose a novel knowledge distillation method, named CLIPPING11In this paper, CLIPPING means cutting something to make it smaller through distilling., where the plentiful knowledge of a large teacher model that has been fine-tuned for video-language tasks with the powerful pre-trained CLIP can be effectively transferred to a small student only at the fine-tuning stage. Especially, a new layer-wise alignment with the student as the base is proposed for knowledge distillation of the intermediate layers in CLIPPING, which enables the student's layers to be the bases of the teacher, and thus allows the student to fully absorb the knowledge of the teacher. CLIPPING with MobileViT-v2 as the vision encoder without any vision-language pre-training achieves 88.1%-95.3% of the performance of its teacher on three video-language retrieval benchmarks, with its vision encoder being 19.5x smaller. CLIPPING also significantly outperforms a state-of-the-art small baseline (ALL-in-one-B) on the MSR-VTT dataset, obtaining relatively 7.4% performance gain, with 29% fewer parameters and 86.9% fewer flops. Moreover, CLIPPING is comparable or even superior to many large pre-training models.
Renjing Pei, Jianzhuang Liu, Weimian Li, Songcen Xu, Peng Dai 0002, Juwei Lu, Youliang Yan
CVPR6
2023 HiVLP: Hierarchical Interactive Video-Language Pre-Training
abstract
Video-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and video-text pairs, which significantly outperforms pre-training with only the video-text pairs. Besides, existing methods usually model the cross-modal interaction using cross-attention between single-scale visual tokens and textual tokens. These visual features are either of low resolutions lacking fine-grained information, or of high resolutions without high-level semantics. To address the issue, we propose Hierarchical interactive Video-Language Pre-training (HiVLP) that efficiently uses a hierarchical visual feature group for multi-modal cross-attention during pre-training. In the hierarchical framework, low-resolution features are learned with focus on more global high-level semantic information, while high-resolution features carry fine-grained details. As a result, HiVLP has the ability to effectively learn both the global and fine-grained representations to achieve better alignment between video and text inputs. Furthermore, we design a hierarchical multi-scale vision contrastive loss for self-supervised learning to boost the interaction between them. Experimental results show that HiVLP establishes new state-of-the-art results in three downstream tasks, text-video retrieval, video-text retrieval, and video captioning.
Jianzhuang Liu, Renjing Pei, Songcen Xu, Peng Dai 0002, Juwei Lu, Weimian Li, Youliang Yan
ICCV5
2023 Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild
abstract
This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependent texture, and a diffuse UV texture. Subsequently, users can either manually edit the UV or provide a prompt for the automatic generation of a new 3D-consistent texture. To achieve high-quality 3D texture generation, we propose a structure-aware score distillation sampling method to optimize a neural UV texture based on user-defined text and empower an image diffusion model with 3D-consistent generation capability. Furthermore, we introduce a few-view resampling training method and utilize a super-resolution model to obtain refined high-resolution UV textures (2048$\times$2048) for 3D texturing. Extensive experiments collectively validate the superior performance of Decorate3D in retexturing real-world 3D objects. Project page: https://decorate3d.github.io/Decorate3D/.
Xinxin Zuo, Peng Dai 0002, Juwei Lu, Li Cheng 0001, Youliang Yan, Songcen Xu
NeurIPS3
2022 Self-Supervised Spatiotemporal Representation Learning by Exploiting Video Continuity
abstract
Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the \textit{video continuity}, to obtain supervision signals for self-supervised representation learning. Specifically, we formulate three novel continuity-related pretext tasks, i.e. continuity justification, discontinuity localization, and missing section approximation, that jointly supervise a shared backbone for video representation learning. This self-supervision approach, termed as Continuity Perception Network (CPNet), solves the three tasks altogether and encourages the backbone network to learn local and long-ranged motion and context representations. It outperforms prior arts on multiple downstream tasks, such as action recognition, video retrieval, and action localization. Additionally, the video continuity can be complementary to other coarse-grained video properties for representation learning, and integrating the proposed pretext task to prior arts can yield much performance gains.
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai 0002, Juwei Lu, Yang Wang 0003
AAAI5
2022 Decompose the Sounds and Pixels, Recompose the Events
abstract
In this paper, we propose a framework centering around a novel architecture called the Event Decomposition Recomposition Network (EDRNet) to tackle the Audio-Visual Event (AVE) localization problem in the supervised and weakly supervised settings. AVEs in the real world exhibit common unraveling patterns (termed as Event Progress Checkpoints(EPC)), which humans can perceive through the cooperation of their auditory and visual senses. Unlike earlier methods which attempt to recognize entire event sequences, the EDRNet models EPCs and inter-EPC relationships using stacked temporal convolutions. Based on the postulation that EPC representations are theoretically consistent for an event category, we introduce the State Machine Based Video Fusion, a novel augmentation technique that blends source videos using different EPC template sequences. Additionally, we design a new loss function called the Land-Shore-Sea loss to compactify continuous foreground and background representations. Lastly, to alleviate the issue of confusing events during weak supervision, we propose a prediction stabilization method called Bag to Instance Label Correction. Experiments on the AVE dataset show that our collective framework outperforms the state-of-the-art by a sizable margin.
Varshanth R. Rao, Md Ibrahim Khalil, Haoda Li, Peng Dai 0002, Juwei Lu
AAAI4
2022 Dual Perspective Network for Audio-Visual Event Localization
Varshanth R. Rao, Md Ibrahim Khalil, Haoda Li, Peng Dai 0002, Juwei Lu
ECCV (34)4
2021 Boosting the Generalization Capability in Cross-Domain Few-shot Learning via Noise-enhanced Supervised Autoencoder
abstract
State of the art (SOTA) few-shot learning (FSL) methods suffer significant performance drop in the presence of domain differences between source and target datasets. The strong discrimination ability on the source dataset does not necessarily translate to high classification accuracy on the target dataset. In this work, we address this cross-domain few-shot learning (CDFSL) problem by boosting the generalization capability of the model. Specifically, we teach the model to capture broader variations of the feature distributions with a novel noise-enhanced supervised autoencoder (NSAE). NSAE trains the model by jointly reconstructing inputs and predicting the labels of inputs as well as their reconstructed pairs. Theoretical analysis based on intra-class correlation (ICC) shows that the feature embeddings learned from NSAE have stronger discrimination and generalization abilities in the target domain. We also take advantage of NSAE structure and propose a two-step fine-tuning procedure that achieves better adaption and improves classification performance in the target domain. Extensive experiments and ablation studies are conducted to demonstrate the effectiveness of the proposed method. Experimental results show that our proposed method consistently outperforms SOTA methods under various conditions.
Hanwen Liang, Peng Dai 0002, Juwei Lu
ICCV3
2021 Class Semantics-based Attention for Action Detection
Deepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li, Peng Dai 0002, Juwei Lu
ICCV5
2021 Structural fragmentation in scene graphs
Varshanth Rao, Peng Dai 0002, Sidharth Singla
Knowl. Based Syst.2
2020 Weight Excitation: Built-in Attention Mechanisms in Convolutional Neural Networks
Niamul Quader, Md Mafijul Islam Bhuiyan, Juwei Lu, Peng Dai 0002, Wei Li 0002
ECCV (30)4
2020 Towards Efficient Coarse-to-Fine Networks for Action and Gesture Recognition
Niamul Quader, Juwei Lu, Peng Dai 0002, Wei Li 0002
ECCV (30)3
2017 Healthy Cognitive Aging: A Hybrid Random Vector Functional-Link Model for the Analysis of Alzheimer's Disease
abstract
Alzheimer's disease (AD) is a genetically complex neurodegenerative disease, which leads to irreversible brain damage, severe cognitive problems and ultimately death. A number of clinical trials and study initiatives have been set up to investigate AD pathology, leading to large amounts of high dimensional heterogeneous data (biomarkers) for analysis. This paper focuses on combining clinical features from different modalities, including medical imaging, cerebrospinal fluid (CSF), etc., to diagnose AD and predict potential progression. Due to privacy and legal issues involved with clinical research, the study cohort (number of patients) is relatively small, compared to thousands of available biomarkers (predictors). We propose a hybrid pathological analysis model, which integrates manifold learning and Random Vector functional-link network (RVFL) so as to achieve better ability to extract discriminant information with limited training materials. Furthermore, we model (current and future) cognitive healthiness as a regression problem about age. By comparing the difference between predicted age and actual age, we manage to show statistical differences between different pathological stages. Verification tests are conducted based on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database. Extensive comparison is made against different machine learning algorithms, i.e. Support Vector Machine (SVM), Random Forest (RF), Decision Tree and Multilayer Perceptron (MLP). Experimental results show that our proposed algorithm achieves better results than the comparison targets, which indicates promising robustness for practical clinical implementation.
Peng Dai 0002, Femida Gwadry-Sridhar, Michael Bauer 0003, Michael Borrie, Xue Teng
AAAI1
2016 Bagging Ensembles for the Diagnosis and Prognostication of Alzheimer's Disease
abstract
Alzheimer's disease (AD) is a chronic neurodegenerative disease, which involves the degeneration of various brain functions, resulting in memory loss, cognitive disorder and death. Large amounts of multivariate heterogeneous medical test data are available for the analysis of brain deterioration. How to measure the deterioration remains a challenging problem. In this study, we first investigate how different regions of the human brain change as the patient develops AD. Correlation analysis and feature ranking are performed based on the feature vectors from different stages of the pathologic process in Alzheimer disease. Then, an automatic diagnosis system is presented, which is based on a hybrid manifold learning for feature embedding and the bootstrap aggregating (Bagging) algorithm for classification.We investigate two different tasks, i.e. diagnosis and progression prediction. Extensive comparison is made against Support Vector Machines (SVM), Random Forest (RF), Decision Tree (DT) and Random Subspace (RS) methods. Experimental results show that our proposed algorithm yields superior results when compared to the other methods, suggesting promising robustness for possible clinical applications.
Peng Dai 0002, Femida Gwadry-Sridhar, Michael Bauer 0003, Michael Borrie
AAAI1
2016 Manifold Learning for Multivariate Variable-Length Sequences With an Application to Similarity Search
abstract
Multivariate variable-length sequence data are becoming ubiquitous with the technological advancement in mobile devices and sensor networks. Such data are difficult to compare, visualize, and analyze due to the nonmetric nature of data sequence similarity measures. In this paper, we propose a general manifold learning framework for arbitrary-length multivariate data sequences driven by similarity/distance (parameter) learning in both the original data sequence space and the learned manifold. Our proposed algorithm transforms the data sequences in a nonmetric data sequence space into feature vectors in a manifold that preserves the data sequence space structure. In particular, the feature vectors in the manifold representing similar data sequences remain close to one another and far from the feature points corresponding to dissimilar data sequences. To achieve this objective, we assume a semisupervised setting where we have knowledge about whether some of data sequences are similar or dissimilar, called the instance-level constraints. Using this information, one learns the similarity measure for the data sequence space and the distance measures for the manifold. Moreover, we describe an approach to handle the similarity search problem given user-defined instance level constraints in the learned manifold using a consensus voting scheme. Experimental results on both synthetic data and real tropical cyclone sequence data are presented to demonstrate the feasibility of our manifold learning framework and the robustness of performing similarity search in the learned manifold.
Shen-Shyang Ho, Peng Dai 0002, Frank Rudzicz
IEEE Trans. Neural Networks Learn. Syst.2
2015 A hybrid manifold learning algorithm for the diagnosis and prognostication of Alzheimer's disease
Peng Dai 0002, Femida Gwadry-Sridhar, Michael Bauer 0003, Michael Borrie
AMIA1
2015 Sequential behavior prediction based on hybrid similarity and cross-user activity transfer
Peng Dai 0002, Shen-Shyang Ho, Frank Rudzicz
Knowl. Based Syst.1
2015 2D Psychoacoustic modeling of equivalent masking for automatic speech recognition
Peng Dai 0002, Frank Rudzicz, Ing Yann Soon, Alex Mihailidis, Huijun Ding
Signal Process.1
2015 Objective measures for quality assessment of noise-suppressed speech
Huijun Ding, Tan Lee, Ing Yann Soon, Chai Kiat Yeo, Peng Dai 0002, Guo Dan
Speech Commun.5
2014 A Smartphone User Activity Prediction Framework Utilizing Partial Repetitive and Landmark Behaviors
abstract
In this paper, we propose a general smartphone user activity prediction framework utilizing the general concept of partial repetitive behavior (instead of the stronger periodicity condition) for similarity scoring and the landmark behaviors (representative behaviors to identify groups of similar behavior vectors). Prediction of the next-day(s) behavior is based on a weighted sum of the most similar behavior vectors related to the landmark behavior of the next-day(s) behavior. These behavior vectors are selected based on the likely partial repetition of the next-day behavior and similarity in the eigen behavior feature space. Our proposed prediction algorithm allows one to categorically quantify the frequency of a target behavior, such as no behavior, normal behavior, and high frequency behavior, or other more refined categorization based on user preference. Extensive experiments are carried out using the Nokia Mobile Data Challenge (MDC) dataset to demonstrate the feasibility of our proposed approach and its generality using arbitrary call activity, voice call activity, short message activity, media consumption, and apps usage data types.
Peng Dai 0002, Shen-Shyang Ho
MDM (1)1
2013 Robust speech recognition by using spectral subtraction with noise peak shifting
abstract
In this study, a novel technique that recovers the temporal structure of speech power spectrum is proposed. The histogram of average speech log power spectrum shows that the contamination of noise leads to the shift of noise peak, which in return degrades the performance of speech recognition systems. A two‐step scheme is proposed to weaken the noise effects by first reducing the noise variance and then shifting the noise mean. The proposed algorithm consists of two parts, two‐dimensional smoothing and controlled noise subtraction, which leads to the name SNS. The proposed algorithm manages to solve the speech probability distribution function discontinuity problem caused by traditional spectral subtraction series algorithms. In contrast to the clean speech estimation methods, the proposed algorithm does not need a prior speech/noise statistical model, which makes it simple but effective. The effectiveness of the proposed filter is tested using the AURORA2 database. Very promising results are obtained, 88.59% for noisy speech (average from signal‐to‐noise ratio 0–20 dB). Comparison is made against eight state‐of‐the‐art speech recognition algorithms. Overall the proposed algorithm produces significant improvements over the comparison targets.
Peng Dai 0002, Ing Yann Soon
IET Signal Process.1
2013 An improved model of masking effects for robust speech recognition system
Peng Dai 0002, Ing Yann Soon
Speech Commun.1
2012 A temporal frequency warped (TFW) 2D psychoacoustic filter for robust speech recognition system
Peng Dai 0002, Ing Yann Soon
Speech Commun.1
2011 A temporal warped 2D psychoacoustic modeling for robust speech recognition system
Peng Dai 0002, Ing Yann Soon
Speech Commun.1