VLDB 2026 Research / reviewers in the wild / expert
Changchong Sheng
dblp:304/1219
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Diffusion-Based Camouflaged Object Detection via Asynchronous Denoising and Linear AttentionabstractCamouflaged object detection (COD) aims to precisely locate and identify objects concealed within their surrounding backgrounds, a task that has been significantly advanced by recent powerful diffusion models. Nevertheless, existing diffusion-based COD methods demonstrate suboptimal performance in inference speed, attributable to two key factors: time-consuming iterative sampling processes and the quadratic complexity introduced by self-attention mechanisms. To address these limitations, we propose Fast CamoDiff, a novel and efficient diffusion-based model for COD. To tackle the computational overhead, we incorporate an asynchronous denoising paradigm leveraging dynamic encoding and early termination, which significantly reduces computational costs during the sampling process. Additionally, we introduce the linear attention mechanism from state space models (SSM) into the diffusion process to achieve linear computational complexity. Meanwhile, to address the side effect of compromising fine-grained feature preservation through information compression, we design a novel bidirectional pooling layer that enhances the model's capability of preserving detailed features without sacrificing computational benefits. Extensive experimental results on four public datasets demonstrate that Fast-CamoDiff achieves superior detection accuracy with only 36.5M parameters (67.4% reduction) and 28 FPS (2× speedup) compared to previous state-of-the-art diffusion-based models. The source code will be available at https://github.com/wty-team/diff-ssm. Xinghua Xu, Changchong Sheng, Shaohua Qiu, Li Liu 0002, Denghua Guo |
IEEE Trans. Multim. | 3 |
| 2024 | Deep Learning for Visual Speech Analysis: A SurveyabstractVisual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper will present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions. Changchong Sheng, Gangyao Kuang, Liang Bai 0003, Chenping Hou, Yulan Guo, Xin Xu 0001, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Importance-Aware Information Bottleneck Learning Paradigm for Lip ReadingabstractLip reading is the task of decoding text from speakers' mouth movements. Numerous deep learning-based methods have been proposed to address this task. However, these existing deep lip reading models suffer from poor generalization due to overfitting the training data. To resolve this issue, we present a novel learning paradigm that aims to improve the interpretability and generalization of lip reading models. In specific, a Variational Temporal Mask (VTM) module is customized to automatically analyze the importance of frame-level features. Furthermore, the prediction consistency constraints of global information and local temporal important features are introduced to strengthen the model generalization. We evaluate the novel learning paradigm with multiple lip reading baseline models on the LRW and LRW-1000 datasets. Experiments show that the proposed framework significantly improves the generalization performance and interpretability of lip reading models. Changchong Sheng, Li Liu 0002, Wanxia Deng, Liang Bai 0003, Zhong Liu 0002, Songyang Lao, Gangyao Kuang, Matti Pietikäinen |
IEEE Trans. Multim. | 1 |
| 2023 | Sim2Word: Explaining Similarity with Representative Attribute Words via Counterfactual ExplanationsabstractRecently, we have witnessed substantial success using the deep neural network in many tasks. Although there still exist concerns about the explainability of decision making, it is beneficial for users to discern the defects in the deployed deep models. Existing explainable models either provide the image-level visualization of attention weights or generate textual descriptions as post hoc justifications. Different from existing models, in this article we propose a new interpretation method that explains the image similarity models by salience maps and attribute words. Our interpretation model contains visual salience maps generation and the counterfactual explanation generation. The former has two branches: global identity relevant region discovery and multi-attribute semantic region discovery. The first branch aims to capture the visual evidence supporting the similarity score, which is achieved by computing counterfactual feature maps. The second branch aims to discover semantic regions supporting different attributes, which helps to understand which attributes in an image might change the similarity score. Then, by fusing visual evidence from two branches, we can obtain the salience maps indicating important response evidence. The latter will generate the attribute words that best explain the similarity using the proposed erasing model. The effectiveness of our model is evaluated on the classical face verification task. Experiments conducted on two benchmarks—VGGFace2 and Celeb-A—demonstrate that our model can provide convincing interpretable explanations for the similarity. Moreover, our algorithm can be applied to evidential learning cases, such as finding the most characteristic attributes in a set of face images, and we verify its effectiveness on the VGGFace2 dataset. Ruoyu Chen 0001, Jingzhi Li 0002, Hua Zhang 0008, Changchong Sheng, Li Liu 0002, Xiaochun Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Boosting Lip Reading with a Multi-View Fusion NetworkabstractLip reading aims to decode speech information by analyzing lip movement without involving audio. Numerous deep learning based methods are proposed to address this task. Generally, most existing methods extract visual features only based on the lip appearance, while ignoring the shape dynamic information of the lip region. Motivated by this, we propose a Multi-View Fusion Network (MVFN), which can extract more discriminative visual representations by incorporating appearance and shape information. Besides, a novel adaptive graph convolutional network model called Adaptive Spatial Graph Model(ASGM) is proposed to learn lip spatial topology and lip shape dynamics automatically. Experiments on LRW (word-level) and OuluVS2 (phrase-level) clearly show that the proposed method significantly outperforms the baseline methods by a large margin and achieves state-of-the-art performance. Xueyi Zhang 0001, Jinping Sui, Changchong Sheng, Wanxia Deng, Li Liu 0002 |
ICME | 4 |
| 2022 | Adaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip ReadingabstractThe goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical flow information of videos. However, these methods do not fully exploit the characteristics of lip motion. In addition to appearance and optical flow, the mouth contour deformation usually conveys significant information that is complementary to others. However, the modeling of dynamic mouth contour has received little attention than that of appearance and optical flow. In this work, we propose a novel model of dynamic mouth contours called Adaptive Semantic-Spatio-Temporal Graph Convolution Network (ASST-GCN), to go beyond previous methods by automatically learning both the spatial and temporal information from videos. To combine the complementary information from appearance and mouth contour, a two-stream visual front-end network is proposed. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art lip reading methods on several large-scale lip reading benchmarks. Changchong Sheng, Xinzhong Zhu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 1 |
| 2021 | Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingabstractThe goal of this work is to learn discriminative visual representations for lip reading without access to manual text annotation. Recent advances in cross-modal self-supervised learning have shown that the corresponding audio can serve as a supervisory signal to learn effective visual representations for lip reading. However, existing methods only exploit the natural synchronization of the video and the corresponding audio. We find that both video and audio are actually composed of speech-related information, identity-related information, and modal information. To make the visual representations (i) more discriminative for lip reading and (ii) indiscriminate with respect to the identities and modals, we propose a novel self-supervised learning framework called Adversarial Dual-Contrast Self-Supervised Learning (ADC-SSL), to go beyond previous methods by explicitly forcing the visual representations disentangled from speech-unrelated information. Experimental results clearly show that the proposed method outperforms state-of-the-art cross-modal self-supervised baselines by a large margin. Besides, ADC-SSL can outperform its supervised counterpart without any finetune. Changchong Sheng, Matti Pietikäinen, Qi Tian 0001, Li Liu 0002 |
ACM Multimedia | 1 |