Peng Chang 0002

dblp:55/1-2 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement
abstract
Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS key-point movements, which often lead to artifacts like blurry hands and distorted fingers. In response to these challenges, we present the Implicit Motion-Audio Entanglement (IMAE) method for co-speech gesture video generation. IMAE strengthens audio control by entangling implicit motion parameters, including pose and expression, with audio inputs. Our method utilizes a two-branch framework that combines an audio-to-motion generation branch with a video diffusion branch, enabling realistic gesture generation without requiring additional inputs during inference. To improve training efficiency, we propose a two-stage slow-fast training strategy that balances memory constraints while facilitating the learning of meaningful gestures from long frame sequences. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple metrics. Project Page.
Xinjie Li 0002, Ziyi Chen 0005, Xinlu Yu, Iek-Heng Chu, Peng Chang 0002, Jing Xiao 0006
CVPR5
2025 Data-Free Knowledge Distillation with Diffusion Models
abstract
Recently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any access to training data. Although diffusion models are adept at synthesizing high-fidelity photorealistic images across various domains, existing methods cannot be easiliy implemented to DFKD. To bridge that gap, this paper proposes a novel approach based on diffusion models, DiffDFKD. Specifically, DiffDFKD involves targeted optimizations in two key areas. Firstly, DiffDFKD utilizes valuable information from teacher models to guide the pre-trained diffusion models’ data synthesis, generating datasets that mirror the training data distribution and effectively bridge domain gaps. Secondly, to reduce computational burdens, DiffDFKD introduces Latent CutMix Augmentation, an efficient technique, to enhance the diversity of diffusion model-generated images for DFKD while preserving key attributes for effective knowledge transfer. Extensive experiments validate the efficacy of DiffDFKD, yielding state-of-the-art results exceeding existing DFKD approaches.We release our code at https://github.com/xhqi0109/DiffDFKD.
Xiaohua Qi, Renda Li, Qiang Ling 0001, Jun Yu 0001, Ziyi Chen 0005, Peng Chang 0002, Jing Xiao 0006
ICME7
2025 SyncDiff: Diffusion-Based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization
abstract
Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image fidelity. Recent studies demonstrate that GAN-based and diffusion-based models achieve state-of-the-art (SOTA) performance on this task, with diffusion-based models achieving superior image fidelity but experiencing lower synchronization compared to their GAN-based counterparts. To this end, we propose SYNcDIFF, a simple yet effective approach to improve diffusion-based models using a temporal pose frame with information bottleneck and facial-informative audio features extracted from AVHuBERT, as conditioning input into the diffusion process. We evaluate SYNcDIFF on two canonical talking head datasets, LRS2 and LRS3 for direct comparison with other SOTA models. Experiments on LRS2/LRS3 datasets show that SYNcDIFF achieves a synchronization score 27.7%/62.3% relatively higher than previous diffusion-based methods, while preserving their high-fidelity characteristics.
Xulin Fan, Heting Gao, Ziyi Chen 0005, Peng Chang 0002, Mark Hasegawa-Johnson
WACV4
2024 Co-speech Gesture Video Generation with 3D Human Meshes
Aniruddha Mahapatra, Renda Li, Ziyi Chen 0005, Boyang Ding, Shoulei Wang, Jun-Yan Zhu, Peng Chang 0002, Jing Xiao 0006
ECCV (89)8
2024 Dialogue Cross-Enhanced Central Engagement Attention Model for Real-Time Engagement Estimation
Jun Yu 0001, Keda Lu, Ji Zhao 0020, Zhihong Wei, Iek-Heng Chu, Peng Chang 0002
IJCAI6
2023 A CTC Alignment-Based Non-Autoregressive Transformer for End-to-End Automatic Speech Recognition
abstract
Recently, end-to-end models have been widely used in automatic speech recognition (ASR) systems. Two of the most representative approaches are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. Autoregressive transformers, variants of AED, adopt an autoregressive mechanism for token generation and thus are relatively slow during inference. In this paper, we present a comprehensive study of a CTC Alignment-based Single-Step Non-Autoregressive Transformer (CASS-NAT) for end-to-end ASR. In CASS-NAT, word embeddings in the autoregressive transformer (AT) are substituted with token-level acoustic embeddings (TAE) that are extracted from encoder outputs with the acoustical boundary information offered by the CTC alignment. TAE can be obtained in parallel, resulting in a parallel generation of output tokens. During training, Viterbi-alignment is used for TAE generation, and multiple training strategies are further explored to improve the word error rate (WER) performance. During inference, an error-based alignment sampling method is investigated in depth to reduce the alignment mismatch in the training and testing processes. Experimental results show that the CASS-NAT has a WER that is close to AT on various ASR tasks, while providing a$\sim$24x inference speedup. With and without self-supervised learning, we achieve new state-of-the-art results for non-autoregressive models on several datasets. We also analyze the behavior of the CASS-NAT decoder to explain why it can perform similarly to AT. We find that TAEs have similar functionality to word embeddings for grammatical structures, which might indicate the possibility of learning some semantic information from TAEs without a language model.
Ruchao Fan, Peng Chang 0002, Abeer Alwan
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
abstract
Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous efforts typically only model one aspect (e.g., accuracy) at one granularity (e.g., at the phoneme-level). In this work, we explore modeling multi-aspect pronunciation assessment at multiple granularities. Specifically, we train a Goodness Of Pronunciation feature-based Transformer (GOPT) with multi-task learning. Experiments show that GOPT achieves the best results on speechocean762 with a public automatic speech recognition (ASR) acoustic model trained on Librispeech.
Yuan Gong 0001, Ziyi Chen 0005, Iek-Heng Chu, Peng Chang 0002, James R. Glass
ICASSP4
2021 Leveraging Large-Scale Weakly Labeled Data for Semi-Supervised Mass Detection in Mammograms
abstract
Mammographic mass detection is an integral part of a computer-aided diagnosis system. Annotating a large number of mammograms at pixel-level in order to train a mass detection model in a fully supervised fashion is costly and time-consuming. This paper presents a novel self-training framework for semi-supervised mass detection with soft image-level labels generated from diagnosis reports by Mammo-RoBERTa, a RoBERTa-based natural language processing model fine-tuned on the fully labeled data and associated mammography reports. Starting with a fully supervised model trained on the data with pixel-level masks, the proposed framework iteratively refines the model itself using the entire weakly labeled data (image-level soft label) in a self-training fashion. A novel sample selection strategy is proposed to identify those most informative samples for each iteration, based on the current model output and the soft labels of the weakly labeled data. A soft cross-entropy loss and a soft focal loss are also designed to serve as the image-level and pixel-level classification loss respectively. Our experiment results show that the proposed semi-supervised framework can improve the mass detection accuracy on top of the supervised baseline, and outperforms the previous state-of-the-art semi-supervised approaches with weakly labeled data, in some cases by a large margin.
Yuxing Tang, Zhenjie Cao, Zongcheng Ji, Jing Xiao 0006, Peng Chang 0002
CVPR10
2021 CASS-NAT: CTC Alignment-Based Single Step Non-Autoregressive Transformer for Speech Recognition
abstract
We propose a CTC alignment-based single step non-autoregressive transformer (CASS-NAT) for speech recognition. Specifically, the CTC alignment contains the information of (a) the number of tokens for decoder input, and (b) the time span of acoustics for each token. The information are used to extract acoustic representation for each token in parallel, referred to as token-level acoustic embedding which substitutes the word embedding in autoregressive transformer (AT) to achieve parallel generation in decoder. During inference, an error-based alignment sampling method is proposed to be applied to the CTC output space, reducing the WER and retaining the parallelism as well. Experimental results show that the proposed method achieves WERs of 3.8%/9.1% on Librispeech test clean/other dataset without an external LM, and a CER of 5.8% on Aishell1 Mandarin corpus, respectively1. Compared to the AT baseline, the CASS-NAT has a performance reduction on WER, but is 51.2x faster in terms of RTF. When decoding with an oracle CTC alignment, the lower bound of WER without LM reaches 2.3% on the test-clean set, indicating the potential of the proposed method.
Ruchao Fan, Peng Chang 0002, Jing Xiao 0006
ICASSP3
2021 Extending Pronunciation Dictionary with Automatically Detected Word Mispronunciations to Improve PAII's System for Interspeech 2021 Non-Native Child English Close Track ASR Challenge
Peng Chang 0002, Jing Xiao 0006
Interspeech2
2021 An Improved Single Step Non-Autoregressive Transformer for Automatic Speech Recognition
abstract
Non-autoregressive mechanisms can significantly decrease inference time for speech transformers, especially when the single step variant is applied. Previous work on CTC alignment-based single step non-autoregressive transformer (CASS-NAT) has shown a large real time factor (RTF) improvement over autoregressive transformers (AT). In this work, we propose several methods to improve the accuracy of the end-to-end CASS-NAT, followed by performance analyses. First, convolution augmented self-attention blocks are applied to both the encoder and decoder modules. Second, we propose to expand the trigger mask (acoustic boundary) for each token to increase the robustness of CTC alignments. In addition, iterated loss functions are used to enhance the gradient update of low-layer parameters. Without using an external language model, the WERs of the improved CASS-NAT, when using the three methods, are 3.1%/7.2% on Librispeech test clean/other sets and the CER is 5.4% on the Aishell1 test set, achieving a 7%~21% relative WER/CER improvement. For the analyses, we plot attention weight distributions in the decoders to visualize the relationships between token-level acoustic embeddings. When the acoustic embeddings are visualized, we find that they have a similar behavior to word embeddings, which explains why the improved CASS-NAT performs similarly to AT.
Ruchao Fan, Peng Chang 0002, Jing Xiao 0006, Abeer Alwan
Interspeech3
2021 Supervised Contrastive Pre-training forMammographic Triage Screening Models
Zhenjie Cao, Yuxing Tang, Jing Xiao 0006, Peng Chang 0002
MICCAI (7)8
2021 BI-RADS Classification of Calcification on Mammograms
Yuxing Tang, Zhenjie Cao, Jing Xiao 0006, Peng Chang 0002
MICCAI (7)7
2021 Radar Object Detection Using Data Merging, Enhancement and Fusion
abstract
Compared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system.
Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002
ICMR6
2021 MommiNet-v2: Mammographic multi-view mass identification networks
abstract
Many existing approaches for mammogram analysis are based on single view. Some recent DNN-based multi-view approaches can perform either bilateral or ipsilateral analysis, while in practice, radiologists use both to achieve the best clinical outcome. MommiNet is the first DNN-based tri-view mass identification approach, which can simultaneously perform bilateral and ipsilateral analysis of mammographic images, and in turn, can fully emulate the radiologists' reading practice. In this paper, we present MommiNet-v2, with improved network architecture and performance. Novel high-resolution network (HRNet)-based architectures are proposed to learn the symmetry and geometry constraints, to fully aggregate the information from all views for accurate mass detection. A multi-task learning scheme is adopted to incorporate both Breast Imaging-Reporting and Data System (BI-RADS) and biopsy information to train a mass malignancy classification network. Extensive experiments have been conducted on the public DDSM (Digital Database for Screening Mammography) dataset and our in-house dataset, and state-of-the-art results have been achieved in terms of mass detection accuracy. Satisfactory mass malignancy classification result has also been obtained on our in-house dataset.
Zhenjie Cao, Yuxing Tang, Xiaohui Lin 0010, Rushan Ouyang, Mingxiang Wu, Jing Xiao 0006, Lingyun Huang, Shibin Wu, Peng Chang 0002
Medical Image Anal.12
2020 MABEL: An AI-Powered Mammographic Breast Lesion Diagnostic System
abstract
Mammography plays an essential role in early detection of breast cancer. Interpreting mammography is a professional task that requires well-trained radiologists with longtime clinical experience. In this paper, we present MABEL, an artificial intelligence-powered system to assist doctors for breast cancer screening and diagnosis in mammograms, in order to reduce their workloads and accelerate the diagnostic process. Our system smoothly integrates our upgraded lesion identification models, provides a doctor-oriented annotation tool and web interface, and can communicate with Picture Archiving and Communication System (PACS) in our collaborative hospital. Our lesion identification performance is evaluated on both public and in-house datasets, in which mass detection has achieved state-of-the-art accuracy in the single-view manner. The overall high satisfaction from doctors of our system is also demonstrated.
Zhenjie Cao, Peng Chang 0002, Shibin Wu, Lingyun Huang, Wei Xu 0007, Jing Xiao 0006, Mingxiang Wu
HealthCom4
2020 MommiNet: Mammographic Multi-view Mass Identification Networks
Zhenjie Cao, Jing Xiao 0006, Lingyun Huang, Shibin Wu, Peng Chang 0002
MICCAI (6)9
2002 Extract highlights from baseball game video with hidden Markov models
abstract
We describe a statistical method to detect highlights in a baseball game video. The input video is first segmented into scene shots, within which the camera motion is continuous. Our approach is based on the observations that (1) most highlights in baseball games are composed of certain types of scene shots and (2) those scene shots exhibit special transition context in time. To exploit those two observations, we first build statistical models for each type of scene shots with products of histograms, and then for each type of highlight a hidden Markov model is learned to represent the context of transition in the time domain. A probabilistic model can be obtained by combining the two, which is used for highlight detection and classification. Satisfactory results have been achieved on initial experimental results.
Peng Chang 0002, Yihong Gong
ICIP (1)1