VLDB 2026 Research / reviewers in the wild / expert
Yunlong Huang
dblp:192/2987
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Segment Anything Model adaptation framework for battery visual inspection under complex radiographic imaging conditions
Chun Cao, Tianyu Wang 0007, Shiyu Lu, Yunlong Huang, Xi Vincent Wang |
Pattern Recognit. | 5 |
| 2025 | PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelabstractTransformers have significantly advanced the field of 3D human pose estimation (HPE). However, existing transformer-based methods primarily use self-attention mechanisms for spatio-temporal modeling, leading to a quadratic complexity, unidirectional modeling of spatio-temporal relationships, and insufficient learning of spatial-temporal correlations. Recently, the Mamba architecture, utilizing the state space model (SSM), has exhibited superior long-range modeling capabilities in a variety of vision tasks with linear complexity. In this paper, we propose PoseMamba, a novel purely SSM-based approach with linear complexity for 3D human pose estimation in monocular video. Specifically, we propose a bidirectional global-local spatio-temporal SSM block that comprehensively models human joint relations within individual frames as well as temporal correlations across frames. Within this bidirectional global-local spatio-temporal SSM block, we introduce a reordering strategy to enhance the local modeling capability of the SSM. This strategy provides a more logical geometric scanning order and integrates it with the global SSM, resulting in a combined global-local spatial scan. We have quantitatively and qualitatively evaluated our approach using two benchmark datasets: Human3.6M and MPI-INF-3DHP. Extensive experiments demonstrate that PoseMamba achieves state-of-the-art performance on both datasets while maintaining a smaller model size and reducing computational costs. Yunlong Huang, Junshuo Liu, Ke Xian, Robert C. Qiu |
AAAI | 1 |
| 2025 | A Two-Stage Electrode Refinement Method with Denoising Diffusion for Automated Segmentation and Overhang Analysis on Battery X-ray ImagesabstractThe increasing demand for high-performance Lithium-Ion Batteries (LIBs) in Electric Vehicles (EVs) underscores the need for high-quality manufacturing. A key challenge is accurately identifying electrodes and measuring anode overhang during the winding process, which directly affects battery performance and safety. Recently, Segment Anything Model (SAM) has been explored to automate data annotation for training small yet specialized segmentation models like Mask R-CNN to infer on the edge. However, SAM struggles to capture sub-pixel boundaries of thin and elongated electrodes due to fixed patch size, lack of high-frequency spatial features, and difficulty of deep fine-tuning. This leads to fragmented masks and irregular boundaries, degrading the annotation quality. On the other hand, existing mask refinement methods, designed primarily for natural images, are not directly applicable for the dense electrodes and low signal-to-noise conditions in industrial X-ray images. To bridge the research gap, a novel local-global two-stage refinement method is proposed in this paper with denoising diffusion steps. The local stage repairs breaks and refines boundary precision at electrode level, while the global stage ensures overall structural consistency across the entire battery cell. Artifact detection and secondary local refinement further improves the mask quality. Experiment results on real-world LIB data collected demonstrate the proposed method outperforms other refinement techniques with visible advancement over coarse masks, enabling more accurate overhang analysis. These findings suggest the potential to reduce manual annotation efforts, paving the way for more scalable and effective LIB manufacturing inspection. Shiyu Lu, Chun Cao, Mian Li 0001, Yunlong Huang, Songhua Zhang |
IECON | 5 |
| 2025 | TRIS-HAR: Transmissive Reconfigurable Intelligent Surfaces-Assisted Human Activity Recognition Using State Space ModelsabstractHuman activity recognition (HAR) using radio frequency (RF) signals has attracted increasing interest due to its non-intrusive and privacy-preserving nature. However, traditional systems often suffer from multipath fading, environmental noise, and limited spatial diversity, particularly in through-the-wall scenarios. In this paper, we propose TRIS-HAR, a novel HAR system that integrates a transmissive reconfigurable intelligent surface (TRIS) with an advanced dual-stream state space model, Human intelligence Mamba (HiMamba). The TRIS actively reshapes the propagation environment by constructing deterministic quasi-line-of-sight (QLoS) paths across obstacles, significantly improving channel state information (CSI) quality. Complementing this, HiMamba leverages a lightweight structured state space architecture to jointly model temporal and spectral dynamics, enabling robust activity recognition under non-line-of-sight conditions. Extensive experiments on both public and real-world datasets demonstrate that TRIS-HAR improves recognition accuracy from 85.00% to 98.06% and maintains strong generalizability across environments. The model is also deployed on a CPU-based edge device, achieving real-time inference at 108 FPS with minimal memory cost. This work establishes a co-designed hardware-algorithm framework for RF-based HAR, offering a scalable and deployable solution for smart homes, healthcare, and next-generation pervasive sensing applications. Junshuo Liu, Yunlong Huang, Rujing Xiong, Tiebin Mi, Robert C. Qiu |
IEEE Internet Things J. | 2 |
| 2024 | TRGR: Transmissive RIS-aided Gait Recognition Through WallsabstractGait recognition with radio frequency (RF) signals enables many potential applications requiring accurate identification. However, current systems require individuals to be within a line-of-sight (LOS) environment and struggle with low signal-to-noise ratio (SNR) when signals traverse concrete and thick walls. To address these challenges, we present TRGR, a novel transmissive reconfigurable intelligent surface (RIS)-aided gait recognition system. TRGR can recognize human identities through walls using only the magnitude measurements of channel state information (CSI) from a pair of transceivers. Specifically, by leveraging transmissive RIS alongside a configuration alternating optimization algorithm, TRGR enhances wall penetration and signal quality, enabling accurate gait recognition. Furthermore, a residual convolution network (RCNN) is proposed as the backbone network to learn robust human information. Experimental results confirm the efficacy of transmissive RIS, highlighting the significant potential of transmissive RIS in enhancing RF-based gait recognition systems. Extensive experiment results show that TRGR achieves an average accuracy of 97.88% in identifying persons when signals traverse concrete walls, demonstrating the effectiveness and robustness of TRGR. Yunlong Huang, Junshuo Liu, Tiebin Mi, Robert C. Qiu |
GLOBECOM | 1 |
| 2024 | TRTAR: Transmissive RIS-Assisted Through-the-Wall Human Activity RecognitionabstractDevice-free human activity recognition plays a pivotal role in wireless sensing. However, current systems often fail to accommodate signal transmission through walls or necessitate dedicated noise removal algorithms. To overcome these limitations, we introduce TRTAR: a device-free passive human activity recognition system integrated with a transmissive reconfigurable intelligent surface (RIS). TRTAR eliminates the necessity for dedicated devices or noise removal algorithms, while specifically addressing signal propagation through walls. Unlike existing approaches, TRTAR solely employs a transmissive RIS at the transmitter or receiver without modifying the inherent hardware structure. Experimental results demonstrate that TRTAR attains an average accuracy of 98.13% when signals traverse concrete walls. Junshuo Liu, Yunlong Huang, Rujing Xiong, Robert C. Qiu |
WCNC | 2 |
| 2024 | RISAR: Reconfigurable Intelligent Surfaces-Assisted Human Activity Recognition With Commercial Wi-Fi DevicesabstractHuman activity recognition (HAR) is crucial in smart homes, security, and healthcare. Existing systems are limited by insufficient spatial diversity due to the constrained number of antennas. Additionally, challenges in noise reduction and feature extraction from sensing data, particularly channel state information (CSI), affect recognition performance. This study introduces a reconfigurable intelligent surface (RIS)-assisted passive HAR (RISAR) method compatible with commercial Wi-Fi devices. RISAR leverages RIS to enhance the spatial diversity of Wi-Fi signals, capturing a broader range of spatial information. A novel high-dimensional factor model based on random matrix theory is proposed to improve noise reduction and feature extraction in the temporal domain. Furthermore, a dual-stream spatiotemporal attention network model is developed to assign variable weights to different characteristics and sequences, mimicking human cognitive processes in prioritizing essential information. Experimental results demonstrate that RISAR significantly outperforms existing HAR methods in both accuracy and efficiency, achieving an average accuracy of 97.26%. These findings highlight RISAR’s adaptability and potential as a robust activity recognition solution in real-world environments. Junshuo Liu, Tiebin Mi, Yunlong Huang, Rujing Xiong, Robert C. Qiu |
IEEE Internet Things J. | 4 |
| 2021 | OMP-Based Channel Estimation without Prior Information for Underwater Acoustic OFDM SystemsabstractA crucial prerequisite for orthogonal matching pursuit (OMP), a widely-used channel estimation method in underwater acoustic (UWA) orthogonal frequency division multiplexing (OFDM) communication systems, is the determination of a termination condition. However, the appropriate condition, which is commonly considered equal to the physical sparsity of the UWA channel, actually dramatically varies with the suffered noise, thus possibly leading to extremely unstable estimation performance. Existing OMP-based algorithms attempt to solve this problem by elaborately adjusting iteration numbers to balance the proportion of genuine channel taps and noise in the reconstructed signal based on noise levels, which inevitably increases the dependency on the prior information, i.e., signal-to-noise ratio (SNR). In order to overcome this challenge, an intuitive idea is eliminating the influence of noise to restore the originally sparse signal before implementing the standard OMP, naturally avoiding the variation of termination conditions. Considering the powerful ability of deep learning, we imitate and elegantly modify the feed-forward denoising convolution neural network (DnCNN), one of the most typical neural networks for image denoising, to develop our prior-information-free denoising OMP (DnOMP) algorithm with a constant iteration number. Simulation results validate that, compared to the standard OMP with the dynamic termination condition, the DnOMP can reduce the normalized mean square error (NMSE) by 39.47%. Donghong Ouyang, Yuzhou Li 0001, Zhizhan Wang, Chengcai Wang, Yunlong Huang |
GLOBECOM | 5 |
| 2021 | Adaptive F-FFT Demodulation for ICI Mitigation in Differential Underwater Acoustic OFDM SystemsabstractThis paper addresses the problem of frequency-domain inter-carrier interference (ICI) mitigation for differential orthogonal frequency-division multiplexing (OFDM) systems. The classical fractional fast Fourier transform (F-FFT), adopting the fixed sampling interval, would suffer from the limited accu-racy of ICI mitigation and low adaptability in dynamic Doppler spread. To target the above challenges, we propose an adaptive fractional Fourier transform (A-FFT) demodulation method, in which an estimation algorithm based on the coordinate descent approach is designed to compute the fiducial frequency offset without increasing pilots. By means of compensating ICI at fractions of the fiducial frequency offset adapted to the time-varying Doppler shift, the A-FFT has the capability of tracking Doppler fluctuations over the underwater acoustic channels, thus extending the application range of frequency-domain ICI mitigation. Simulation results show that the A-FFT is significantly superior to the existing classical methods, the partial fast Fourier transform (P-FFT) and the F-FFT, for both medium and high Doppler factors and large carrier numbers in terms of the mean squared error (MSE). Numerically, the MSE of the A-FFT is reduced by 39.88% - 72.14% compared to that of the F-FFT with the input signal-to-noise ratio ranging from 10 dB to 30 dB at a Doppler factor of 2.5 x 10−4and a carrier number of 1024, while the P-FFT even cannot work well. Jihui Qiu, Yuzhou Li 0001, Yunlong Huang, Lingyu Gu |
GLOBECOM | 3 |
| 2021 | Inter-Carrier Interference Mitigation for Differentially Coherent Detection in Underwater Acoustic OFDM SystemsabstractSuppressing the inter-carrier interference (ICI) is crucial for differentially coherent detection in underwater acoustic (UWA) orthogonal frequency division multiplexing (OFDM) systems due to the fact that the UWA channel is inherently violently Doppler-shifted. In this paper, we propose a new ICI suppression method, referred to as the partially-shifted fast Fourier transform (PS-FFT), which eliminates the ICI from both the time and frequency domains. Specifically, the PS-FFT first divides the received signal in the entire block duration into several short non-overlapping ones to reduce the channel variation in the time domain. It then applies the Fourier transform at several predefined frequencies to the received signal in each of these intervals to compensate Doppler shifts in the frequency domain. Finally, it weightedly combines the multiple demodulator outputs at each carrier as one output for symbol detection, with the combiner weights being solved by the stochastic gradient algorithm. Simulation results show that the PS-FFT dramatically outperforms the existing classical methods, the partial fast Fourier transform (P-FFT) and the fractional fast Fourier transform (F-FFT), for both medium and high Doppler factors and large carrier numbers in terms of the mean squared error (MSE). Numerically, the MSE of the PS-FFT is reduced by 61.83% – 84.89% compared to that of the F-FFT when the input signal-to-noise ratio (SNR) at the receiver ranges from 10 dB to 30 dB at a Doppler factor of 3 × 10−4and a carrier number of 1024 where the P-FFT even cannot work. Yunlong Huang, Yuzhou Li 0001 |
ICC | 1 |
| 2020 | CHOpinionMiner: An unsupervised system for Chinese opinion target extractionabstractSummary Opinion target extraction (OTE) is an important task in fine‐grained sentiment analysis field, which focuses on the identification of the targets of users' opinions or sentiments from online reviews. Existing approaches to OTE are mainly based on unsupervised rule‐based methods or supervised machine learning methods. However, the latter needs a large number of labeled samples to train their models and lack of labeled corpus limits the research progress on OTE of Chinese. In this paper, we proposed a novel unsupervised and domain independent system called CHOpinionMiner. First, noun phrases are extracted as candidate opinion targets based on phrase structure grammar, then the syntactic paths of the candidate opinion targets and the opinion words in one opinionated sentence are extracted. After that, the legal paths, or rather, the opinion targets‐opinion lexicon pairs that are syntactically associated are selected according to the syntactic rules. Finally, not only the formal targets but also the orientations are obtained. Experiments on data sets consisting of microblog topics and product reviews demonstrate that our approach outperforms the existing state‐of‐the‐art methods. Minghu Jiang, Yunlong Huang, Peijun Qiu |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | EPAN: Effective parts attention network for scene text recognitionabstractFor most previous attention-based scene text recognition methods, images are transformed into high-level feature vectors that form a feature map with height equal to one. Such vectors may contain unnecessary noise that limits recognition performance. To address this issue, in this paper, we propose the effective parts attention network (EPAN) which can attentively highlight the character region for more precise recognition. EPAN consists of a text image encoder and character effective parts decoder (CEPD), and it is end-to-end trainable. The former separates the high-dimensional feature map into one-dimensional vectors row-by-row, which are connected to a bidirectional long short term memory unit to encode contextual information. Subsequently, the CEPD transforms the vectors using a novel glimpse network at each time step to roughly determine the position of the characters. Then the CEPD uses a refinement network to generate a mask to gradually localize the precise position of important parts of the current character. Experiments were conducted on various benchmarks, including IIIT5K-Words, Street View Text, ICDAR 2003, ICDAR 2013, CUTE80, Street View Text Perspective, and ICDAR 2015, which demonstrated that the proposed EPAN method significantly outperformed or was comparable to existing methods in terms of lexicon-free word accuracy. Additionally, substantial qualitative results further demonstrated the robustness of our method. Yunlong Huang, Zenghui Sun, Canjie Luo |
Neurocomputing | 1 |
| 2019 | Attention After Attention: Reading Text in the Wild with Cross AttentionabstractRecent methods mostly regarded scene text recognition as a sequence-to-sequence problem. These methods roughly transform the image into a feature sequence and use the algorithms for sequence-to-sequence problem like CTC or attention to decode the characters. However, text in images is distributed in a two-dimensional (2D) space and roughly converting the features of text into a feature sequence may introduce extra noise, especially if the text is irregular. In this paper, we propose a novel framework named cross attention network, which learns to attend to local features of a 2D feature map corresponding to individual characters. The network contains two 1D attention networks, which operates harmoniously in two directions. Thus, one of the attention modules vertically attends to the features corresponding to the whole text of 2D features and the other horizontal module selects the local features to decode individual characters. Extensive experiments are performed on various regular benchmarks, including SVT, ICDAR2003, ICDAR2013, and IIIT5K-Words, which demonstrate that the proposed model either outperforms or is comparable to all previous methods. Moreover, the model is evaluated on irregular benchmarks including SVT-Perspective, CUTE80 and ICDAR 2015. The performance on irregular benchmarks shows the robustness of our model. Yunlong Huang, Canjie Luo, Qingxiang Lin, Weiying Zhou |
ICDAR | 1 |