Qu Yang

dblp:225/1980 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
abstract
Yue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Xin He, Qu Yang, Qingxuan Jiang, Fanghua Ye, Juntao Li, Zhaopeng Tu, Xiaolong Li, Liefeng Bo, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yue Wang 0039, Ruotian Ma, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Qu Yang, Qingxuan Jiang, Fanghua Ye 0004, Juntao Li 0005, Zhaopeng Tu, Liefeng Bo, Min Zhang 0005
ACL (1)10
2026 Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research
abstract
Deep Research systems have revolutionized how LLMs solve complex questions through iterative reasoning and evidence gathering. However, current systems remain fundamentally constrained to textual web data, overlooking the vast knowledge embedded in multimodal documents: scientific papers, technical reports, and financial documents where critical information exists in figures, tables, charts, and equations. Processing such documents demands sophisticated parsing to preserve visual semantics, intelligent chunking to maintain structural coherence, and adaptive retrieval across modalities, which are capabilities absent in existing systems. In response, we present Doc-Researcher, a unified system that bridges this gap through three integrated components: (i) deep multimodal parsing that preserves layout structure and visual semantics while creating multi-granular representations from chunk to document level, (ii) systematic retrieval architecture supporting text-only, vision-only, and hybrid paradigms with dynamic granularity selection, and (iii) iterative multi-agent workflows that decompose complex queries, progressively accumulate evidence, and synthesize comprehensive answers across documents and modalities. To enable rigorous evaluation, we introduce M4DocBench, the first benchmark for Multi-modal, Multi-hop, Multi-document, and Multi-turn deep research. Featuring 158 expert-annotated questions with complete evidence chains across 304 documents, M4DocBench tests capabilities that existing benchmarks cannot assess. Experiments demonstrate that Doc-Researcher achieves 50.6% accuracy, 3.4× better than state-of-the-art baselines, validating that effective document research requires not just better retrieval, but fundamentally deep parsing that preserve multimodal integrity and support iterative research. Our work establishes a new paradigm for conducting deep research on multimodal document collections.
Kuicai Dong, Shurui Huang, Fangda Ye, Dexun Li, Qu Yang, Gang Wang 0056, Yichao Wang 0002, Chen Zhang 0003, Yong Liu 0020
WWW8
2026 Synergy of Sight and Semantics: Holistic Visual Understanding With CLIP
abstract
Holistic Visual Understanding (HVU), encompassing tasks like intention recognition, emotion analysis, scene understanding, and content moderation, necessitates integrating low-level visual perception ('sight') with high-level semantic reasoning ('semantics'). While large Vision-Language Models (VLMs) like CLIP offer powerful representations, their inherent 'sight' bias limits their direct application to these semantically rich tasks. Our prior work, IntCLIP, addressed Multi-label Intention Understanding (MIU) using a dual-branch architecture but faced challenges with label generation instability (Hierarchical Class Integration - HCI) and limited feature interaction (unidirectional Sight-assisted Aggregation). This paper introduces an enhanced framework that significantly extends IntCLIP to tackle the broader HVU challenge. We propose Semantic Label Refinement (SLR), an iterative, metric-guided process leveraging Large Language Models (LLMs) and quantitative evaluation within the CLIP embedding space to generate stable, optimized semantic labels. We also introduce a novel bidirectional attention mechanism (Symmetric Aggregation) that enables balanced, mutual refinement between sight and semantic feature maps. By evaluating on a comprehensive benchmark spanning MIU, Image Emotion Recognition, Indoor Scene Recognition, and Visual Content Moderation, we demonstrate that our framework not only advances the state-of-the-art in MIU but also achieves superior performance across diverse HVU tasks. This framework provides a unified and robust solution for synergizing sight and semantics, pushing towards more human-like visual intelligence. Code is available at https://github.com/yan9qu/PAMI25-HVU.
Qu Yang, Mang Ye, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models
Jian Liang 0003, Wenke Huang 0003, Guancheng Wan, Qu Yang, Mang Ye
CVPR4
2025 Uncertain Multimodal Intention and Emotion Understanding in the Wild
abstract
Understanding intention and emotion from social media poses unique challenges due to the inherent uncertainty in multimodal data, where posts often contain incomplete or missing modalities. While this uncertainty reflects real-world scenarios, it remains underexplored within the computer vision community, particularly in conjunction with the intrinsic relationship between emotion and intention. To address these challenges, we introduce the Multimodal IntentioN and Emotion Understanding in the Wild (MINE) dataset, comprising over 20,000 topic-specific social media posts with natural modality variations across text, image, video, and audio. MINE is distinctively constructed to capture both the uncertain nature of multimodal data and the implicit correlations between intentions and emotions, providing extensive annotations for both aspects. To tackle these scenarios, we propose the Bridging Emotion-Intention via Implicit Label Reasoning (BEAR) framework. BEAR consists of two key components: a BEIFormer that leverages emotion-intention correlations, and a Modality Asynchronous Prompt that handles modality uncertainty. Experiments show that BEAR outperforms existing methods in processing uncertain multimodal data while effectively mining emotion-intention relationships for social media content understanding.
Qu Yang, Qinghongya Shi, Tongxin Wang, Mang Ye
CVPR1
2025 Adaptive Re-calibration Learning for Balanced Multimodal Intention Recognition
abstract
Multimodal Intention Recognition (MIR) plays a critical role in applications such as intelligent assistants, service robots, and autonomous systems. However, in real-world settings, different modalities often vary significantly in informativeness, reliability, and noise levels. This leads to modality imbalance, where models tend to over-rely on dominant modalities, thereby limiting generalization and robustness. While existing methods attempt to alleviate this issue at either the sample or model level, most overlook its multi-level nature. To address this, we propose Adaptive Re-calibration Learning (ARL), a novel dual-path framework that models modality importance from both sample-wise and structural perspectives. ARL incorporates two key mechanisms: Contribution-Inverse Sample Calibration (CISC), which dynamically masks overly dominant modalities at the sample level to encourage attention to underutilized ones; and Weighted Encoder Calibration (WEC), which adjusts encoder weights based on global modality contributions to prevent overfitting. Experimental results on multiple MIR benchmarks demonstrate that ARL significantly outperforms existing methods in both accuracy and robustness, particularly under noisy or modality-degraded conditions.
Qu Yang, Xiyang Li, Mang Ye
NeurIPS1
2025 An intelligent framework based on optimized variational mode decomposition and temporal convolutional network: Applications to stock index multi-step forecasting
Yuanyuan Yu, Dongsheng Dai, Qu Yang
Expert Syst. Appl.3
2025 Toward Ultralow-Power Neuromorphic Speech Enhancement With Spiking-FullSubNet
abstract
Speech enhancement (SE) is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved SE performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultralow-power SE system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and subband fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive subband modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multiscale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art (SOTA) methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (algorithmic track), opening up a myriad of opportunities for ultralow-power SE at the edge. Our source code and model checkpoints are publicly available at github.com/haoxiangsnr/spiking-fullsubnet.
Chenxiang Ma, Qu Yang, Jibin Wu, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.3
2024 TC-LIF: A Two-Compartment Spiking Neuron Model for Long-Term Sequential Modelling
abstract
The identification of sensory cues associated with potential opportunities and dangers is frequently complicated by unrelated events that separate useful cues by long delays. As a result, it remains a challenging task for state-of-the-art spiking neural networks (SNNs) to establish long-term temporal dependency between distant cues. To address this challenge, we propose a novel biologically inspired Two-Compartment Leaky Integrate-and-Fire spiking neuron model, dubbed TC-LIF. The proposed model incorporates carefully designed somatic and dendritic compartments that are tailored to facilitate learning long-term temporal dependencies. Furthermore, the theoretical analysis is provided to validate the effectiveness of TC-LIF in propagating error gradients over an extended temporal duration. Our experimental results, on a diverse range of temporal classification tasks, demonstrate superior temporal classification capability, rapid training convergence, and high energy efficiency of the proposed TC-LIF model. Therefore, this work opens up a myriad of opportunities for solving challenging temporal processing tasks on emerging neuromorphic computing systems. Our code is publicly available at https://github.com/ZhangShimin1/TC-LIF.
Qu Yang, Chenxiang Ma, Jibin Wu, Haizhou Li 0001, Kay Chen Tan
AAAI2
2024 Synergy of Sight and Semantics: Visual Intention Understanding with CLIP
Qu Yang, Mang Ye, Dacheng Tao
ECCV (11)1
2024 SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
abstract
Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achieve noise robustness and often require large models for high performance. This paper introduces a novel SNN-based VAD model, referred to as sVAD, which features an auditory encoder with an SNN-based attention mechanism. Particularly, it provides effective auditory feature representation through SincNet and 1D convolution, and improves noise robustness with attention mechanisms. The classifier utilizes Spiking Recurrent Neural Networks (sRNN) to exploit temporal speech information. Experimental results demonstrate that our sVAD achieves remarkable noise robustness and meanwhile maintains low power consumption and a small footprint, making it a promising solution for real-world VAD applications.
Qu Yang, Qianhui Liu, Meng Ge, Zeyang Song, Haizhou Li 0001
ICASSP1
2024 ED-sKWS: Early-Decision Spiking Neural Networks for Rapid, and Energy-Efficient Keyword Spotting
Zeyang Song, Qianhui Liu, Qu Yang, Yizhou Peng, Haizhou Li 0001
INTERSPEECH3
2024 A Hybrid Neural Coding Approach for Pattern Recognition With Spiking Neural Networks
abstract
Recently, brain-inspired spiking neural networks (SNNs) have demonstrated promising capabilities in solving pattern recognition tasks. However, these SNNs are grounded on homogeneous neurons that utilize a uniform neural coding for information representation. Given that each neural coding scheme possesses its own merits and drawbacks, these SNNs encounter challenges in achieving optimal performance such as accuracy, response time, efficiency, and robustness, all of which are crucial for practical applications. In this study, we argue that SNN architectures should be holistically designed to incorporate heterogeneous coding schemes. As an initial exploration in this direction, we propose a hybrid neural coding and learning framework, which encompasses a neural coding zoo with diverse neural coding schemes discovered in neuroscience. Additionally, it incorporates a flexible neural coding assignment strategy to accommodate task-specific requirements, along with novel layer-wise learning methods to effectively implement hybrid coding SNNs. We demonstrate the superiority of the proposed framework on image classification and sound localization tasks. Specifically, the proposed hybrid coding SNNs achieve comparable accuracy to state-of-the-art SNNs, while exhibiting significantly reduced inference latency and energy consumption, as well as high noise robustness. This study yields valuable insights into hybrid neural coding designs, paving the way for developing high-performance neuromorphic systems.
Qu Yang, Jibin Wu, Haizhou Li 0001, Kay Chen Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Composed Image Retrieval via Cross Relation Network With Hierarchical Aggregation Transformer
abstract
Composing Text and Image to Image Retrieval (CTI-IR) aims at finding the target image, which matches the query image visually along with the query text semantically. However, existing works ignore the fact that the reference text usually serves multiple functions, e.g., modification and auxiliary. To address this issue, we put forth a unified solution, namely Hierarchical Aggregation Transformer incorporated with Cross Relation Network (CRN). CRN unifies modification and relevance manner in a single framework. This configuration shows broader applicability, enabling us to model both modification and auxiliary text or their combination in triplet relationships simultaneously. Specifically, CRN includes: 1) Cross Relation Network comprehensively captures the relationships of various composed retrieval scenarios caused by two different query text types, allowing a unified retrieval model to designate adaptive combination strategies for flexible applicability; 2) Hierarchical Aggregation Transformer aggregates top-down features with Multi-layer Perceptron (MLP) to overcome the limitations of edge information loss in a window-based multi-stage Transformer. Extensive experiments demonstrate the superiority of the proposed CRN over all three fashion-domain datasets. Code is available at github.com/yan9qu/crn.
Qu Yang, Mang Ye, Zhaohui Cai, Kehua Su, Bo Du 0001
IEEE Trans. Image Process.1
2022 Knowledge distillation for In-memory keyword spotting model
Zeyang Song, Qi Liu 0005, Qu Yang, Haizhou Li 0001
INTERSPEECH3
2022 Deep residual spiking neural network for keyword spotting in low-resource settings
Qu Yang, Qi Liu 0005, Haizhou Li 0001
INTERSPEECH1
2022 Training Spiking Neural Networks with Local Tandem Learning
abstract
Spiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a generalized learning rule, termed Local Tandem Learning (LTL). The LTL rule follows the teacher-student learning approach by mimicking the intermediate feature representations of a pre-trained ANN. By decoupling the learning of network layers and leveraging highly informative supervisor signals, we demonstrate rapid network convergence within five training epochs on the CIFAR-10 dataset while having low computational complexity. Our experimental results have also shown that the SNNs thus trained can achieve comparable accuracies to their teacher ANNs on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. Moreover, the proposed LTL rule is hardware friendly. It can be easily implemented on-chip to perform fast parameter calibration and provide robustness against the notorious device non-ideality issues. It, therefore, opens up a myriad of opportunities for training and deployment of SNN on ultra-low-power mixed-signal neuromorphic computing chips.
Qu Yang, Jibin Wu, Malu Zhang, Yansong Chua, Xinchao Wang, Haizhou Li 0001
NeurIPS1
2021 Rethinking Benchmarks for Neuromorphic Learning Algorithms
abstract
We rely on benchmarking datasets to monitor the research progress. However, recent studies have cast doubts on the effectiveness of current neuromorphic benchmarking datasets; and the debate remains largely unsettled. In this paper, we assess the richness and usefulness of temporal information embedded in these benchmarking datasets for SNN decision making. To this end, we propose a segregated spatio-temporal learning framework that allows us to selectively control the information flow along both spatial and temporal directions during feedforward and backward propagation. Leveraging on this framework, we conduct a comprehensive study on seven widely used neuromorphic audio and vision datasets. Our findings are threefold. First, the existing neuromorphic benchmarks only make limited contributions in highlighting the temporal processing capability of spiking neurons. Second, the temporal credit assignment is redundant for tasks that only require short-range temporal dependency. Third, we recommend the neuromorphic research community to develop novel benchmarks that require both short-range and long-range temporal dependencies. Such appropriate benchmark datasets would be helpful in guiding the development of powerful SNN-based learning algorithms and computational models.
Qu Yang, Jibin Wu, Haizhou Li 0001
IJCNN1
2019 Deep Spiking Neural Network with Spike Count based Learning Rule
abstract
Deep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training.
Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001
IJCNN4