Dawei Luo

dblp:78/644 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2025 CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
abstract
With the swift progression of image generation technology, the widespread emergence of facial deepfakes poses significant challenges to the field of security, thus amplifying the urgent need for effective deepfake detection. Existing techniques for face forgery detection can broadly be categorized into two primary groups: visual-based methods and multimodal approaches. The former often lacks clear explanations for forgery details, while the latter, which merges visual and linguistic modalities, is more prone to the issue of hallucinations.To address these shortcomings, we introduce a visual detail enhanced self-correction framework, designated CorrDetail, for interpretable face forgery detection. CorrDetail is meticulously designed to rectify authentic forgery details when provided with error-guided questioning, with the aim of fostering the ability to uncover forgery details rather than yielding hallucinated responses. Additionally, to bolster the reliability of its findings, a visual fine-grained detail enhancement module is incorporated, supplying CorrDetail with more precise visual forgery details. Ultimately, a fusion decision strategy is devised to further augment the model's discriminative capacity in handling extreme samples, through the integration of visual information compensation and model bias reduction. Experimental results demonstrate that CorrDetail not only achieves state-of-the-art performance compared to the latest methodologies but also excels in accurately identifying forged details, all while exhibiting robust generalization capabilities.
Binjia Zhou, Hengrui Lou, Lizhe Chen, Dawei Luo, Jie Lei 0002, Zunlei Feng, Yijun Bei
IJCAI5
2023 Multi-Task Sub-Band Network For Deep Residual Echo Suppression
abstract
This paper introduces the SWANT team’s entry to the ICASSP 2023 AEC Challenge. We submit a system that cascades a linear filter with a neural post-filter. Particularly, we adopt sub-band processing to handle full-band signals and shape the network with multi-task learning, where dual signal voice activity detection (DSVAD) and echo estimation are adopted as auxiliary tasks. Moreover, we particularly improve the time frequency convolution module (TFCM) to increase the receptive field using small convolution kernels. Finally, our system has ranked 4th in ICASSP 2023 AEC Challenge Non-personalized track.
Jiayao Sun, Dawei Luo, Zhaoxia Li, Yukai Jv
ICASSP2
2022 Uformer: A Unet Based Dilated Complex & Real Dual-Path Conformer Network for Simultaneous Speech Enhancement and Dereverberation
abstract
Complex spectrum and magnitude are considered as two major features of speech enhancement and dereverberation. Traditional approaches always treat these two features separately, ignoring their underlying relationship. In this paper, we propose Uformer, a Unet based dilated complex & real dual-path conformer network in both complex and magnitude domain for simultaneous speech enhancement and dereverberation. We exploit time attention (TA) and dilated convolution (DC) to leverage local and global contextual information and frequency attention (FA) to model dimensional information. These three sub-modules contained in the proposed dilated complex & real dual-path conformer module effectively improve the speech enhancement and dereverberation performance. Furthermore, hybrid encoder and decoder are adopted to simultaneously model the complex spectrum and magnitude and promote the information interaction between two domains. Encoder decoder attention is also applied to enhance the interaction between encoder and decoder. Our experimental results outperform all SOTA time and complex domain models objectively and subjectively. Specifically, Uformer reaches 3.6032 DNSMOS on the blind test set of Interspeech 2021 DNS Challenge, which outperforms all top-performed models. We also carry out ablation experiments to tease apart all proposed submodules that are most important.
Yihui Fu, Jingdong Li, Dawei Luo, Shubo Lv, Yukai Jv, Lei Xie 0001
ICASSP4
2022 The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO Convolutional Recurrent Network for Multi Channel Speech Enhancement and Speech Recognition
abstract
This paper described the PCG-AIID system for L3DAS22 challenge in Task 1: 3D speech enhancement in office reverberant environment. We proposed a two-stage framework to address multi-channel speech denoising and dereverberation. In the first stage, a multiple input and multiple out-put (MIMO) network is applied to remove background noise while maintaining the spatial characteristics of multi-channel signals. In the second stage, a multiple input and single out-put (MISO) network is applied to enhance the speech from desired direction and post-filtering. As a result, our system ranked 3rd place in ICASSP2022 L3DAS22 challenge and significantly outperforms the baseline system, while achieving 3.2% WER and 0.972 STOI on the blind test-set.
Jingdong Li, Dawei Luo, Guohui Cui, Zhaoxia Li
ICASSP3
2022 A Lipreading Model Based on Fine-Grained Global Synergy of Lip Movement
abstract
Lipreading is a type of speech recognition based on visual information. It is instructive to design a lipreading model according to the lip movement law. Algorithms in the field of computer vision cannot fully satisfy the characteristics of lipreading, and direct use does not necessarily improve the performance of lipreading. In this paper, we propose that lipreading has fine-grained global synergy by comparing other computer vision tasks and analyzing lip muscle motion patterns. To address this feature, we propose a tailored model and name it Fine-Grained Global Synergy Lipreading (FGSLip). Our model aims to make features synergistic to improve lipreading performance. We introduce global features to represent the overall characteristics of the lip, and local features to learn coarse-grained and fine-grained correlations between features. Then, diffusion and fusion methods are used to make the local features and global features synergistic. Based on the above, several different feature extraction structures are constructed to demonstrate the fine-grained global synergy of lipreading. To verify the effectiveness of the proposed model, extensive experiments are conducted on the laboratory record dataset ICSLR and the public dataset CMLR, and the experimental results show that the proposed method can effectively improve the accuracy of lipreading.
Baosheng Sun, Dongliang Xie, Dawei Luo, Xiaojie Yin
ICTAI3
2021 Densely Connected Multi-Stage Model with Channel Wise Subband Feature for Real-Time Speech Enhancement
abstract
Research on single channel speech enhancement (SE) has a long tradition, but two main practical problems still remain unsolved. Firstly, it’s hard to balance between enhancement quality and computational efficiency, and low-latency always brings loss of quality. Secondly, enhancement in specific scenarios, such as singing and emotional speech, is also an intricate problem of conventional methods. In this paper, we propose a computationally efficient real-time speech enhancement network with densely connected multi-stage structures, which progressively enhances the channel-wise subband speech. The enhanced speech from earlier stage is used to guide the processing of deeper stage in order to obtain coarse to fine estimations. Besides, supervision is applied to all intermediate results in order to stabilize training and accelerate convergence. Moreover, an adaptive fine-tune step is utilized with some small datasets of specific scenarios, which achieves superb improvement under corresponding scenes. As a result, the proposed method achieves promising performance improvements in terms of speech quality and demonstrates robustness in complex scenarios. We submitt the proposed method to the deep noise suppression (DNS) challenge 2021, real-time denoising track, which was held by Microsoft. In the subjective evaluation, our system outperforms DNS-Challenge baseline by 0.14 points in terms of mean opinion score (MOS).
Jingdong Li, Dawei Luo, Zhaoxia Li, Guohui Cui, Wenqi Tang, Wei Chen 0071
ICASSP2
2020 Few-Shot Object Detection by Second-Order Pooling
Shan Zhang 0002, Dawei Luo, Lei Wang 0001, Piotr Koniusz
ACCV (4)2
2017 Sequential DOA estimation method for multi-group coherent signals
Jinqiang Wei, Xu Xu 0003, Dawei Luo, Zhongfu Ye
Signal Process.3
1984 FORMANAGER: An Office Forms Management System
abstract
The form has become an important abstraction for data management in an office application environment.Structured office forms present data to users in an easily understood and easily manipulated manner.In this paper we classify forms systems in terms of three dimensions: data structuring, user interfaces, and programming interfaces.Current forms systems are analyzed under these dimensions.We have designed a comprehensive forms management system, FORMANAGER, that includes facilities for form specification, form processing, and form control.The system transforms data from a relational database into a hierarchical data structure which defines the form.The design and algorithms for implementation of the system are described, and future extensions to enhance the capabilities of forms systems are proposed.
Alan R. Hevner, Zhongzhi Shi, Dawei Luo
ACM Trans. Inf. Syst.4
1981 Form Operation By Example: A Language For Office Information Processing
abstract
In this paper we introduce a high level nonprocedural language Form Operation by Example (FOBE) to manipulate forms in office systems. The form data model is selected as the basis for the user interface. The idea of query-by-example is applied to forms. A precise semantics definition of the language is given. FOBE has a predetermined structure. It combines the advantages of procedural and nonprocedural languages. We believe that it is user-friendly and sufficiently powerful for office environments.
Dawei Luo
SIGMOD Conference1