VLDB 2026 Research / reviewers in the wild / expert
Chengyou Wang
dblp:23/2651 · also Cheng-You Wang
· DBLP profile ↗
13ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-0901-2492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional AnnotationabstractThe development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline. Longhao Li, Zhao Guo, Hongjie Chen 0001, Yuhang Dai, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Hui Bu, Jie Li 0001, Jian Kang 0006, Ruibin Yuan, Ziya Zhou, Wei Xue 0002, Lei Xie 0001 |
AAAI | 8 |
| 2026 | CF2M-Net: Cross-feature fusion and memory-constraint network for video anomaly detection
Qiming Ma, Chengyou Wang, Xiao Zhou 0002 |
Inf. Sci. | 2 |
| 2026 | GECS-Net: Two-View Correspondence Learning via Geometric Embedding and Cross-Stage ConsensusabstractTwo-view correspondence learning is fundamental to geometric computer vision. However, existing methods often fail in complex scenarios due to inadequate spatial modeling and stage-wise inconsistency. We propose geometric embedding and cross-stage-network (GECS-Net), a robust framework that unifies geometry-aware embedding with a progressive cross-stage consensus strategy. Specifically, GECS-Net synergizes a geometric embedding attention (GEA) module using rotary positional embeddings (RoPE) to encode relative spatial constraints, and a cross-stage graph feature extraction (CSGFE) module to enforce structural consensus via sparse graph neural networks. Evaluations on YFCC100M and SUN3D show that GECS-Net significantly outperforms state-of-the-art methods, achieving an 82.2% F-score on YFCC100M and improving mAP@5$^\circ$by 8.29% in unknown outdoor scenes, demonstrating superior robustness and generalization. Chengyou Wang |
IEEE Signal Process. Lett. | 3 |
| 2026 | Multi-scale cross-attention network for enhanced face forgery detection
Chengyou Wang, Xiao Zhou 0002 |
Vis. Comput. | 2 |
| 2025 | Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
Mingchen Shao, Xinfa Zhu, Chengyou Wang, Bingshen Mu, Danming Xie, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2025 | Multi-modality complementary learning network with cross-modality interaction and adaptive fusion for face forgery detection
Chunyin Shi, Chengyou Wang, Xiao Zhou 0002, Zhiliang Qin |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | MSER-Net: Multi-stage edge refinement network for deepfake detectionabstractToday, the malicious creation and dissemination of deepfake content has reached a level that threatens personal privacy and even social stability. To address this issue, many researchers have focused on deepfake detection tasks. However, many approaches rely solely on single spatial domain information, with training models directly extracting global artefacts from images for detection. Such approaches frequently lack accuracy and are vulnerable to interference. Meanwhile, forgery traces typically appear at the edges of tampered regions, and these subtle edge inconsistencies serve as effective cues for detecting deepfakes. To address these issues, we propose the multi-stage edge refinement network (MSER-Net), which uses VMamba as its backbone network and incorporates channel residual image (CRI) to detect unnatural colour anomalies introduced during the deepfake process. First, the designed multiscale edge enhancement module (MSEEM) processes low-level features from both branches, obtaining multi-scale edge information using the Sobel operator and highlighting subtle forgery traces. In addition, we use the detail-aware interaction module (DAIM) to extract complementary information from spatial and residual features, which improves edge artefacts. Finally, we propose the multidimensional attention fusion module (MAFM), which improves dual-branch features from multiple perspectives using the attention mechanism, resulting in more effective fusion. Experimental results on multiple public datasets show that our method outperforms state-of-the-art detection approaches and exhibits the best robustness against most post-processing attacks. The code is available at: https://github.com/ypzhang123/MSER-Net . Chengyou Wang, Xiao Zhou 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Robust and Compatible Video Watermarking via Spatio-Temporal Enhancement and Multiscale Pyramid AttentionabstractDeep learning-based video watermarking is shown to be effective in improving robustness. However, existing methods neglect the enhancement of long-distance spatio-temporal features and the representation of the inter-frame difference and the intra-frame difference, which lead to poor robustness against H.264 compression and low compatibility with high-definition (HD) and full high-definition (FHD) videos for copyright protection, respectively. To address these issues, we propose a robust and compatible video watermarking network (RC-VWN) based on spatio-temporal enhancement and multiscale pyramid attention. For robustness, RC-VWN extracts long-distance spatio-temporal features using a central difference 3D U-Net and enhances them through multiscale spatio-temporal fusion, which alleviates the loss of the watermark caused by attacks through the association of long-distance spatio-temporal features. Then, the simulated compression network is developed to simulate H.264 compression with high-accuracy, which guides the decoder to recover the watermark accurately. For compatibility, a multiscale pyramid attention is designed to represent the intra-frame difference and the inter-frame difference effectively. Experimental results demonstrate that RC-VWN outperforms the state-of-the-art methods with higher robustness and imperceptibility under quantitative evaluation and visual quality. Furthermore, RC-VWN exhibits high compatibility with various videos, including HD and FHD videos, ensuring effective copyright protection. Chengyou Wang, Xiao Zhou 0002, Zhiliang Qin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MSG: Robust Multimodal Remote Sensing Image Matching Using Side Window Gaussian SpaceabstractRobust and accurate image matching and registration are foundational tasks for numerous applications. However, current methods often fail when dealing with multimodal remote sensing images (MRSI) that exhibit significant spatial geometric differences (SGD) and nonlinear radiometric differences (NRD). To address these challenges, this paper proposes a novel MRSI matching method: matching using side window Gaussian space (MSG). MSG leverages an intuitive concept that human visual perception relies heavily on salient features at image edges for precise matching. Specifically, the proposed method: (1) constructs a multi-scale side window Gaussian filter scale space (MSGSS) that preserves edge information at different scales while blurring the image; (2) enhances the repeatability of keypoints by performing corner detection on edge maps; (3) increases descriptor robustness by using second-order gradients combined with steerable filtering; (4) further utilizes a two-stage matching strategy within a constrained search space and designs a new distance, making full use of densely distributed edge keypoints. Quantitative and qualitative experiments conducted on 5 datasets spanning 957 image pairs across 9 multimodal types demonstrate that MSG outperforms nine advanced algorithms (six feature-based methods: SIFT, OS-SIFT, RIFT, CoFSM, HOWP, POS-GIFT; and three deep learning-based methods: SuperPoint+SuperGlue, LoFTR, ReDFeat). The results indicate that MSG achieved a number of correct matches (NCM) much higher than the compared algorithms, with the highest success rate (SR), lowest RMSE and good time efficiency, while achieving both scale and rotation invariance. Codes are available at https://github.com/ZCYla/MSG. Chongyue Zheng, Shanshan Li 0003, Chengyou Wang, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Copy-move image forgery detection based on evolving circular domains coverageabstractAbstract The aim of this paper is to improve the accuracy of copy-move forgery detection (CMFD) in image forensics by proposing a novel scheme and the main contribution is evolving circular domains coverage (ECDC) algorithm. The proposed scheme integrates both block-based and keypoint-based forgery detection methods. Firstly, the speed-up robust feature (SURF) in log-polar space and the scale invariant feature transform (SIFT) are extracted from an entire image. Secondly, generalized 2 nearest neighbor (g2NN) is employed to get massive matched pairs. Then, random sample consensus (RANSAC) algorithm is employed to filter out mismatched pairs, thus allowing rough localization of counterfeit areas. To present these forgery areas more accurately, we propose the efficient and accurate ECDC algorithm to present them. This algorithm can find satisfactory threshold areas by extracting block features from jointly evolving circular domains, which are centered on matched pairs. Finally, morphological operation is applied to refine the detected forgery areas. Experimental results indicate that the proposed CMFD scheme can achieve better detection performance under various attacks compared with other state-of-the-art CMFD schemes. Shilin Lu, Xinghong Hu, Chengyou Wang, Shulu Han, Yuejia Han |
Multim. Tools Appl. | 3 |
| 2009 | All phase biorthogonal transform and its application in JPEG-like image compression
Zheng-Xin Hou, Chengyou Wang, Ai-Ping Yang |
Signal Process. Image Commun. | 2 |
| 2006 | On-Line Vibration Source Detection of Running Trains Based on Acceleration MeasurementabstractTo ensure safety of railway operation, it is important to regularly check railway conditions such as deformation of the rails. To monitor rail deformation, this paper presents a method for detecting sources of vibrations a running train on-line by measuring accelerations, which include the train bogie's lateral acceleration, and the crossbeam's lateral and vertical accelerations. A series of detection algorithms including peak-peak value entropy comparison, weighted correlation coefficients comparison etc. are proposed in the method, according to different characters of vibrations from train itself and rail deformation. To eliminate the vibration due to the train itself, the algorithm employs the peak-peak value entropy comparison. To identify the order of the vibrations between crossbeam and bogie, a weighted correlation coefficient is applied. Weight center and maximum position are used to detect at last. The algorithms were implemented on a passenger train using ARM processor and real experiments were conducted on the train on the railway between Shenyang and Dalian in China. The experiments demonstrated that the proposed method can produce satisfactory results. Chengyou Wang, Qiugen Xiao, Hua Liang, Xuanping Cai, Yun-Hui Liu 0001 |
IROS | 1 |
| 2006 | Uncalibrated visual servoing of robots using a depth-independent interaction matrixabstractThis paper presents a new adaptive controller for image-based dynamic control of a robot manipulator using a fixed camera whose intrinsic and extrinsic parameters are not known. To map the visual signals onto the joints of the robot manipulator, this paper proposes a depth-independent interaction matrix, which differs from the traditional interaction matrix in that it does not depend on the depths of the feature points. Using the depth-independent interaction matrix makes the unknown camera parameters appear linearly in the closed-loop dynamics so that a new algorithm is developed to estimate their values on-line. This adaptive algorithm combines the Slotine-Li method with on-line minimization of the errors between the real and estimated projections of the feature points on the image plane. Based on the nonlinear robot dynamics, we prove asymptotic convergence of the image errors to zero by the Lyapunov theory. Experiments have been conducted to verify the performance of the proposed controller. The results demonstrated good convergence of the image errors. Yun-Hui Liu 0001, Hesheng Wang 0001, Chengyou Wang, Kinkwan Lam |
IEEE Trans. Robotics | 3 |