Weilong Huang

dblp:180/2944 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Privacy-preserving for user-uploaded images and text in Vision-Language Models
Zixiang Liu, Chi Chen 0001, Shuguang Yuan 0003, Weilong Huang, Xiaojie Zhu, Peizhuo Lv
Comput. Secur.4
2026 High-resolution image deraining via dual-branch features interaction and fusion
Weilong Huang, Jiaxue Mei, Tao Yan 0001, Yinghui Wang 0001, Xiaojun Chang
Neural Networks1
2026 MDeRainNet: An Efficient Macro-pixel Image Rain Removal Network
abstract
Since raining weather always degrades image quality and poses significant challenges to most computer vision-based intelligent systems, image de-raining has been a hot research topic in computer vision community. Fortunately, in a rainy Light Field (LF) image, background obscured by rain streaks in one sub-view may be visible in the other sub-views, and implicit depth information and recorded 4D structural information may benefit rain streak detection and removal. However, existing LF image rain removal methods either do not fully exploit the global correlations of 4D LF data or only utilize partial sub-views (i.e., under-utilization of the rich angular information), resulting in sub-optimal rain removal performance and no-equally good quality for all de-rained sub-views. In this article, we propose an efficient neural network, called MDeRainNet , for rain streak removal from LF images. The proposed network adopts a multi-scale encoder–decoder architecture, which directly works on Macro-pixel Images (MPIs) for improving the rain removal performance. To fully model the global correlation between the spatial information and the angular information, we propose an Extended Spatial-angular Interaction (ESAI) module to merge the two types of information, in which a simple and effective Transformer-based Spatial-angular Interaction Attention (SAIA) block is also proposed for modeling long-range geometric correlations and making full use of the angular information. Furthermore, to improve the generalization performance of our network on real-world rainy scenes, we propose a novel semi-supervised learning framework for our MDeRainNet , which utilizes multi-level KL loss to bridge the domain gap between features of synthetic and that of real-world rain streaks and introduces colored-residue image-guided contrastive regularization to reconstruct rain-free images. Extensive experiments conducted on both synthetic and real-world Light Field Images (LFIs) demonstrate that our method outperforms the state-of-the-art methods both quantitatively and qualitatively.
Tao Yan 0001, Weilong Huang, Weijiang He, Cihang Wei, Xiangjie Zhu, Yinghui Wang 0001, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Multi-scale Attention and Adaptive Self-attention for Occlusion-Aware Depth Estimation in Light Field
Zhineng Zhang, Tao Yan 0001, Jinsheng Liu, Weilong Huang
CGI (2)6
2025 DDYOLO: Efficient Rainy Scene Object Detection with Deraining Guidance
abstract
Recently, deep learning-based object detection has been widely used in autonomous driving and intelligent security. However, distinct features of targets would be contaminated due to domain shift and unexpected noise in adverse weather conditions, which may lead to poor performance in object detection. Several methods preprocess rain images with complex image deraining network before detection network to improve the detection performance, but fail to balance the time and detection accuracy. In this paper, we propose an end-to-end hybrid network called DDYOLO, which combines an image deraining sub-network and an object detection sub-network to utilize derained features for improving detection results with low time cost. Specifically, to ensure domain consistency between deraining and detection features while reducing redundant computations, both sub-networks share the same backbone network during features encoding. In the decoding process, we propose the Semantic Information Enhancement Block (SIEB) to supplement the missing semantic information layer by layer from the clean image features obtained by our proposed image deraining sub-network to the detection features. Moreover, we propose a rain layer extraction sub-network composed of the External Attention Module (EAM) to exploit the residual background information from degradation layers to assist background restoration and object detection. Extensive results demonstrate that our DDYOLO performs comparably with state-of-the-art methods while reducing inference time for each single-frame. Our code and dataset will be available at https://github.com/YT3DVision/DDYOLO.
Tao Yan 0001, Jinsheng Liu, Zhineng Zhang, Weilong Huang
IJCNN6
2025 GCGP-YOLO: Global-Local and Channel Grouping Perception Network for Small Object Detection based on YOLOv8
abstract
Due to the wide field of view, small object sizes, and high object densities in Unmanned Aerial Vehicle (UAV) imagery, conventional object detection networks often struggle to effectively extract and perceive the features of small objects. This leads to misclassification and localization errors for small objects. To address these challenges, we propose GCGP-YOLO, a novel network based on YOLOv8n that improves small object detection performance for aerial image through Global-Local and Channel Grouping Perception. Specifically, we propose a Channel Grouping Perception Module (CGPM) to model the global context information while extracting local features. Subsequently, the implicit semantic relations of objects in the global contextual information are employed to guide the network to enhance the perception of small objects. Meanwhile, by combining large separable kernel convolutions, we propose Long and Short Dependency Pooling (LSDP), a method that captures long-range dependencies without introducing additional computational cost or parameters, thereby enriching the fine-grained information extracted by the network. In addition, we construct Linear Deformable Convolution Block (LDCB) to fit the diverse shapes of small objects, enabling the network to focus more on small objects during multi-scale feature fusion. Finally, the experimental results on the VisDrone2019 and TinyPerson datasets demonstrate that our network outperforms the baseline models with lower computational cost, achieving significant improvements of 8.30% and 5.68% in mAP50, respectively. The code and models are available at https://github.com/YT3DVision/GCGP-YOLO.
Jinsheng Liu, Tao Yan 0001, Zhineng Zhang, Weilong Huang
IJCNN6
2024 A Practical Online Multichannel Dereverberation Approach with Data-Reuse Technique
abstract
One of the most effective online dereverberation algorithms is the weighted prediction error (WPE) method and its improved version, switching WPE (SwWPE). This paper proposes a practical online dereverberation approach to improve SwWPE by introducing a datareuse technique, DR-SwWPE; we then show analytically that DR-SwWPE is equivalent to a SwWPE followed by a post-filter to suppress the residual reverberation. Secondly, we compare this suppression technique and two explicit post-filtering schemes to SwWPE using both simulated data and real-world recordings. Experimental results show that DR-SwWPE out-performs SwWPE with little additional computational cost.
Weilong Huang, Jinwei Feng, W. Bastiaan Kleijn
ICASSP1
2023 Neural Optimization Of Geometry And Fixed Beamformer For Linear Microphone Arrays
abstract
Fixed beamforming based on uniform linear microphone arrays often suffers from non-optimal performance for broadband signals. This paper addresses the issue by jointly optimizing the array geometry and spatial filters through a neural network based model. The model, composed of two feed forward neural networks, is optimized in an end-to-end manner. It satisfies the distortionless constraint in the look direction. Experimental results show that the proposed model outperforms the previous state-of-the-art fixed beamformer with overall better scores. Moreover, the proposed model can control the tradeoff between Directivity Factor (DF) and White Noise Gain (WNG) in a flexible way.
Longfei Yan 0001, Weilong Huang, W. Bastiaan Kleijn, Thushara D. Abhayapala
ICASSP2
2022 M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
abstract
Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologies. Speaker diarization and multi-speaker automatic speech recognition in meeting scenarios have attracted much attention recently. However, the lack of large public meeting data has been a major obstacle for advancement of the field. Therefore, we make available the AliMeeting corpus, which consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone. Each meeting session is composed of 2-4 speakers with different speaker overlap ratio, recorded in meeting rooms with different size. Along with the dataset, we launch the ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) with two tracks, namely speaker diarization and multi-speaker ASR, aiming to provide a common testbed for meeting rich transcription and promote reproducible research in this field. In this paper we provide a detailed introduction of the AliMeeting dateset, challenge rules, evaluation methods and baseline systems.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Lei Xie 0001, Zhihao Du, Weilong Huang, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP7
2022 Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
abstract
The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP7
2021 A PDRVIO Loosely coupled Indoor Positioning System via Robust Particle Filter
abstract
In recent years, the Visual Inertial Odometry (VIO) technology has attracted attention as a support technology that improves medical experience and management efficiency. However, the performance of most VIO systems will drop drastically when the light intensity changes significantly or there are few texture features from the images. This paper designs a visual-inertial fusion-based navigation Indoor Positioning system to deal with the challenging scenario. It loosely coupled an inertial sensor-based pedestrian dead reckoning (PDR) model with the VIO model via a robust particle filter. The state estimation of the particle filter is based on the PDR model. The VIO model is used for the measurements of the particle filter. It compensates the gross errors of the VIO with a visual error propagation model which is established according to the posterior observation residuals of visual feature points. It is verified through experiments that the PDR/VIO fusion indoor positioning system based on the robust particle filter implemented in this paper has improved positioning accuracy and strong ability to deal with complex scenes.
Xinwei Hu, Weilong Huang, Lingxiang Zheng, Ao Peng, Huiru Zheng, Haiying Wang 0001
BIBM4
2021 A Real-Time Speaker Diarization System Based on Spatial Spectrum
abstract
In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in speaker diarization tasks: (1) to segment and separate overlapping speech from two speakers; (2) to estimate the number of speakers when participants may enter or leave the conversation at any time; (3) to provide accurate speaker identification on short text-independent utterances; (4) to track down speakers movement during the conversation; (5) to detect speaker change incidence real-time. First, a differential directional microphone array-based approach is exploited to capture the target speakers’ voice in far-field adverse environment. Second, an online speaker-location joint clustering approach is proposed to keep track of speaker location. Third, an instant speaker number detector is developed to trigger the mechanism that separates overlapped speech. The results suggest that our system effectively incorporates spatial information and achieves significant gains.
Weilong Huang, Xianliang Wang, Hongbin Suo, Jinwei Feng, Zhijie Yan
ICASSP2
2021 Minimum-Norm Differential Beamforming for Linear Array with Directional Microphones
Weilong Huang, Jinwei Feng
Interspeech1
2021 Real-Time Multi-Channel Speech Enhancement Based on Neural Network Masking with Attention Model
Weilong Huang, Weiguang Chen, Jinwei Feng
Interspeech2
2021 Investigation of Spatial-Acoustic Features for Overlapping Speech Detection in Multiparty Meetings
Shiliang Zhang, Weilong Huang, Hongbin Suo, Jinwei Feng, Zhijie Yan
Interspeech3
2020 Differential Beamforming for Uniform Circular Array with Directional Microphones
Weilong Huang, Jinwei Feng
INTERSPEECH1
2018 Cohort Analysis with Ease
abstract
The tremendous volume of user behavior records generated in various domains provides data analysts new opportunities to mine valuable insights into user behavior. Cohort analysis, which aims to find user behavioral trends hidden in time series, is one of the most commonly used techniques. Since traditional database systems suffer from both operability and efficiency when processing cohort analysis queries, we proposed COHANA, a query processing system specialized for cohort analysis. In order to make COHANA easy-to-use, we present a comprehensive and powerful tool in this demo, covering the major use cases in cohort analysis with intuitive and accessible operations. Analysts can easily adapt COHANA to their own use with provided visualizations which can help verify their analysis assumptions and inconspicuous trends hidden in user behavior data.
Zhongle Xie, Qingchao Cai, Gene Yan Ooi, Weilong Huang, Beng Chin Ooi
SIGMOD Conference5
2017 QALink: Enriching Text Documents with Relevant Q&A Site Contents
abstract
With rapid development of Q&A sites such as Quora and StackExchange, high quality question-answer pairs have been produced by users. These Q&A contents cover a wide range of topics, and they are useful for users to resolve queries and obtain new knowledge. Meanwhile, when people are reading digital documents, they may encounter reading problems such as lack of background information and unclear illustration of concepts. We believe that Q&A sites offer high-quality contents which can serve as rich supplements to digital documents. In this paper, we devise a rigorous formulation of the novel text enrichment problem, and design an end-to-end system named QALink which assigns the most relevant Q&A contents to the corresponding section of the document. We first present a new segmentation approach to model each document with a hierarchical structure. Based on the hierarchy, queries are constructed to retrieve and rank related question-answer pairs. Both syntactical and semantic features are adopted in our system. The empirical evaluation results indicate that QALink is able to effectively enrich text documents with relevant Q&A contents to help people better understand the documents.
Weilong Huang, Qi Liu 0049, Anthony K. H. Tung, Xiaoli Wang 0002, Jisong Yang
CIKM2
2016 A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometries
abstract
An increasing number of spatial filtering approaches requires narrowband direction-of-arrival (DOA) estimates. State-of-the-art (SOA) estimators such as root-MUSIC and ESPRIT are computationally complex and can be used only with specific array geometries. In this work, a low complexity DOA estimator is proposed that can be applied to arbitrary array geometries. The DOA is estimated by minimizing the weighted error between the observed and expected inter-microphone phase differences. The complexity of the proposed DOA estimator is significantly lower compared to that of the SOA estimators while providing a similar performance.
Oliver Thiergart, Weilong Huang, Emanuël A. P. Habets
ICASSP2