EDBT 2026 Demo / reviewers in the wild / expert
Jian Zhou 0006
dblp:97/97-6
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-6509-5520ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementabstractAlthough the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the performance of SE, many modules are stacked onto SE, resulting in increased model complexity that limits the application of SE. To address these problems, we proposed a dual-path network based on compressed frequency using Mamba. First, we extract amplitude and phase information through parallel dual branches. This approach leverages structured complex spectra to implicitly capture phase information and solves the compensation effect by decoupling amplitude and phase, and the network incorporates an interaction module to suppress unnecessary parts and recover missing components from the other branch. Second, to reduce network complexity, the network introduces a band-split strategy to compress the frequency dimension. To further reduce complexity while maintaining good performance, we designed a Mamba-based module that models the time and frequency dimensions under linear complexity. Finally, compared to baselines, our model achieves an average 8.3 times reduction in computational complexity while maintaining superior performance. Furthermore, it achieves a 25 times reduction in complexity compared to transformer-based models. Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 0001, Jian Zhou 0006, Chengshi Zheng, Zhao Lv |
AAAI | 5 |
| 2025 | M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker ExtractionabstractThe brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet. Cunhang Fan, Jian Zhou 0006, Zexu Pan, Youdian Gao, Xiaoke Yang, Zhengqi Wen, Zhao Lv |
IJCAI | 3 |
| 2025 | ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention DetectionabstractAuditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at:https://github.com/fchest/ListenNet. Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv |
IJCAI | 6 |
| 2025 | MHANet: Multi-scale Hybrid Attention Network for Auditory Attention DetectionabstractAuditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale contextual information within EEG signals, limiting their ability to capture long-short range spatiotemporal dependencies simultaneously. To address these issues, this paper proposes a multi-scale hybrid attention network (MHANet) for AAD, which consists of the multi-scale hybrid attention (MHA) module and the spatiotemporal convolution (STC) module. Specifically, MHA combines channel attention and multi-scale temporal and global attention mechanisms. This effectively extracts multi-scale temporal patterns within EEG signals and captures long-short range spatiotemporal dependencies simultaneously. To further improve the performance of AAD, STC utilizes temporal and spatial convolutions to aggregate expressive spatiotemporal representations. Experimental results show that the proposed MHANet achieves state-of-the-art performance with fewer trainable parameters across three datasets, 3 times lower than that of the most advanced model. Code is available at: https://github.com/fchest/MHANet. Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv |
IJCAI | 6 |
| 2025 | REB-former: RWKV-enhanced E-branchformer for Speech Recognition
Wang Xiang, Jian Zhou 0006, Cunhang Fan, Zhao Lv |
INTERSPEECH | 3 |
| 2025 | SSF-DST: A Spectro-Spatial Features Enhanced Deep Spatiotemporal Network for EEG-Based Auditory Attention Detection
Xiaoke Yang, Jian Zhou 0006, Zhao Lv, Cunhang Fan |
INTERSPEECH | 3 |
| 2025 | DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention DetectionabstractAuditory attention detection (AAD) aims to identify the attended speaker in multi-talker environments by analyzing brain activity recorded through neural monitoring techniques. Recent AAD approaches have achieved great progress in improving detection accuracy. However, they still face challenges in capturing complex spatio-temporal dependencies and high-order nonlinear relationships across brain regions. To address these challenges, this paper proposes DHGCN, a dual hypergraph convolutional network that integrates a hypergraph modeling module, a dual-branch hypergraph learning (DHGL) module, and a feature fusion module. Specifically, the hypergraph modeling module constructs spatial and temporal hypergraphs from EEG signals, enabling the representation of high-order relationships among channels and time points. The DHGL module comprises two parallel branches: a spatial branch that learns high-order spatial dependencies across EEG channels, and a temporal branch that captures complex temporal dependencies. Each branch uses its corresponding hypergraph structure, which is established during the modeling phase. The feature fusion module then aggregates spatial and temporal representations from both branches to support robust auditory attention classification. Extensive experiments on multiple benchmark datasets demonstrate that DHGCN consistently outperforms state-of-the-art AAD models. It achieves superior classification performance while reducing the trainable parameters count by over 50% compared to the state-of-the-art models. Code is available at: https://github.com/nobody1219/DHGCN.git. Jian Zhou 0006, Yingjie Xie, Cunhang Fan, Zhao Lv |
ACM Multimedia | 1 |
| 2025 | Enhancing bone-conducted speech with spectrum similarity metric in adversarial learning
Jian Zhou 0006, Wenming Zheng, Hon Keung Kwan |
Speech Commun. | 2 |
| 2025 | Controllable Multi-Speaker Emotional Speech Synthesis With an Emotion Representation of High Generalization CapabilityabstractThe aim of multi-speaker emotional speech synthesis is to generate speech for a designated speaker in a desired emotional state. The task is challenging due to the presence of speech variations, such as noise, content, and timbre, which can obstruct emotion extraction and transfer. This paper proposes a new approach to performing multi-speaker emotional speech synthesis. The proposed method, which is based on a seq2seq synthesizer, integrates emotion embedding as a conditioned variable to convey exact emotional information from reference audio to the synthesized speech. To boost emotion representation capability, we utilize a three-dimensional acoustic feature as input. And an emotion generalization module with adaptive instance normalization (AdaIN) is proposed to obtain emotion embedding with high generalization ability, which also results in improved controllability. The derived emotion embedding from the generalization module can be readily conditioned by affine parameters, allowing for control both the emotion category and the emotion intensity of synthesized speech. Various emotional speech synthesis experimental results of the propposed method demonstrate its state-of-the-art performance in multi-speaker emotional speech synthesis, coupled with its advantage of high emotion controllability. Jian Zhou 0006, Wenming Zheng, Hon Keung Kwan |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Dynamic Ensemble Teacher-Student Distillation Framework for Light-Weight Fake Audio DetectionabstractIn recent years, fake audio detection (FAD) has made great progress, and lightweight is important to achieve fast and reliable audio authenticity verification on resource-limited devices. However, most of the researchers ignore lightweight when improving the performance of FAD. To develop the application of FAD for small-end devices, this paper proposes a novel light-weight network named Light-ECA2Net. Given that networks with different depths have different abilities in capturing fake speech artifacts, this paper proposes a dynamic ensemble teacher-student distillation framework to fully transfer distillation knowledge. The dynamic ensemble distillation is divided into two aspects. First, we adopt one-to-one feature mapping to perceive the multidimensional feature knowledge and dynamically adjust every dimension feature weight by using ground truth labels, which can enable students to receive feature knowledge efficiently. Secondly, different network layers also have their strengths of predicting, further dynamically predicting weight can improve the learning ability of the student. Experimental results on the ASVspoof 2019 LA and PA datasets show that compared to the baseline, our system further improves performance by reducing the model complexity by 45%. Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Jian Zhou 0006, Zhao Lv |
IEEE Signal Process. Lett. | 4 |
| 2024 | Multi-Level Information Aggregation Based Graph Attention Networks Towards Fake Speech DetectionabstractIt is widely acknowledged that distinguishing genuine speech from spoofed speech encompasses various subbands and temporal segments within speech signals. However, prevailing spoofing detection methods tend to oversimplify the relationships between these cues by employing linear models. In this paper, we introduce a multi-level information aggregation Graph Attention Networks (MiaGATs) to generate highly discriminative features for fake speech detection (FSD). In MiaGATs, each subband and temporal segment of a speech signal is represented as distinct nodes. MiaGATs incorporates channel information aggregation within each node to effectively harness the unique spectral and temporal characteristics during the feature encoding stage. In particular, MiaGATs address the interactions between nodes through indirect node aggregation and integrates both indirect and direct node aggregation by max-pooling operation. Experimental results on ASVspoof2019 and ASVspoof2021 LA databases show significant relative improvement compared to the current state-of-the-art. In comparison to the leading integrated spectro-temporal graph attention networks, MiaGATs gains an impressive performance improvement in various conditions, underscoring MiaGATs's position as a new benchmark in spoofing detection performance. Jian Zhou 0006, Yong Li 0032, Cunhang Fan, Hon Keung Kwan |
IEEE Signal Process. Lett. | 1 |
| 2014 | A consistent pixel-wise blur measure for partially blurred imagesabstractDespite numerous efforts on blur measurement of partially blurred images, there still lacks an effective blur measure that is both pixel-wise and locally sharp consistent. The paper proposes a novel method with two contributions to overcome this limitation: 1) A new pixel-based blur metric, Multi-resolution Singular Value (MSV), which leverages the average singular value of high frequency bands to measure the blur of each pixel, and 2) a locally continuous strategy, maximum-likelihood estimation (MLE) based refinement, that ensures local continuity by imposing the local sharp consistency on pixel blur in a local correcting process. Experimental results show that our method is effective to smoothly measure the partially blurred images without local discontinuity. Xianyong Fang, Yanwen Guo 0001, Christian Jacquemin, Jian Zhou 0006, Shanchun Huang |
ICIP | 5 |
| 2014 | Unsupervised learning of phonemes of whispered speech in a noisy environment based on convolutive non-negative matrix factorization
Jian Zhou 0006, Ruiyu Liang, Li Zhao 0003, Cairong Zou |
Inf. Sci. | 1 |
| 2013 | Fast window fusion using fuzzy equivalence relation
Xianyong Fang, Jian Zhou 0006 |
Pattern Recognit. Lett. | 3 |