Huilin Zhu

dblp:28/11106 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 RegenTrack: Distance-Adaptive Regeneration Pool Matching for Drone-Based Crowd Tracking
abstract
Drone-based crowd tracking remains challenging due to low object distinctiveness, high crowd density, frequent occlusions, and the difficulty of maintaining continuous and precise localization from aerial views. To address these challenges, we proposeRegenTrack, a tracking framework that integrates distance-adaptive fusion with Regeneration Pool (RegenPool) matching. RegenTrack dynamically balances appearance and motion cues for trajectory-object association. Appearance cues are extracted from multi-frame fused trajectory and object features, while motion cues are derived by comparing predicted and observed positions via a motion network. Unmatched trajectories are temporarily stored in RegenPool and re-matched in subsequent frames to mitigate object loss caused by occlusion. Meanwhile, RegenPool evaluates unmatched objects through multi-frame observations to determine whether they should be promoted to new trajectories. In addition, a two-stage cascade matching strategy is employed to further enhance trajectory continuity and stability. Experiments on DRONECROWD, DRONEBIRD, and CROHD demonstrate that RegenTrack achieves notable improvements in both tracking accuracy and reliability for drone-based crowd scenarios. The code will be released at https://github.com/Zebrabeast/RegenTrack.
Jingling Yuan, Huilin Zhu, Jinqiao Wang, Xian Zhong
IEEE Trans. Circuits Syst. Video Technol.4
2025 Synergistic Integration of Cross-Spatial Learning for Lightweight Crack Detection
abstract
Efficient crack segmentation is crucial for engineering surface inspection, especially on edge devices where both accuracy and computational efficiency are essential. To address the challenges posed by crack directionality and blurred edges while enhancing performance, we propose a lightweight segmentation model, SCCU-Net, based on cross-spatial synergistic learning and coordinate awareness. The model integrates coordinate information and perceptual sets into a hybrid attention mechanism, significantly boosting segmentation accuracy. We introduce a Mapping Attention Gate (MAG), which utilizes gating signals from fine-grained features to guide cross-spatial learning, alongside an Adaptive Skip Fusion (ASF) to ensure smooth feature integration while optimizing computational resources. Extensive experiments on six benchmark datasets demonstrate that SCCU-Net consistently outperforms existing lightweight models, setting a new benchmark for crack detection on edge devices. Code is available at https://github.com/lsyyy20000830/SCCU-Net.
Senyao Li, Jingling Yuan, Huilin Zhu, Xian Zhong
ICASSP3
2025 Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection
abstract
Maritime object detection is essential for navigation safety, surveillance, and autonomous operations, yet constrained by two key challenges: the scarcity of annotated maritime data and poor generalization across various maritime attributes (e.g., object category, viewpoint, location, and imaging environment). To address these challenges, we propose Neptune-X, a data-centric generative-selection framework that enhances training effectiveness by leveraging synthetic data generation with task-aware sample selection. From the generation perspective, we develop X-to-Maritime, a multi-modality-conditioned generative model that synthesizes diverse and realistic maritime scenes. A key component is the Bidirectional Object-Water Attention module, which captures boundary interactions between objects and their aquatic surroundings to improve visual fidelity. To further improve downstream tasking performance, we propose Attribute-correlated Active Sampling, which dynamically selects synthetic samples based on their task relevance. To support robust benchmarking, we construct the Maritime Generation Dataset, the first dataset tailored for generative maritime learning, encompassing a wide range of semantic conditions. Extensive experiments demonstrate that our approach sets a new benchmark in maritime scene synthesis, significantly improving detection accuracy, particularly in challenging and previously underrepresented settings. The code is available at https://github.com/gy65896/Neptune-X.
Yu Guo 0008, Shengfeng He, Yuxu Lu, Haonan An 0001, Yihang Tao, Huilin Zhu, Jingxian Liu, Yuguang Fang
NeurIPS6
2025 Clothing Purification with Causality Meets Vision-Language Pretraining Models
Zhengwei Yang 0001, Huilin Zhu, Nan Lei, Basura Fernando, Zheng Wang 0007
Int. J. Comput. Vis.2
2025 Multi-Granularity Distribution Alignment for Cross-Domain Crowd Counting
abstract
Unsupervised domain adaptation enables the transfer of knowledge from a labeled source domain to an unlabeled target domain, and its application in crowd counting is gaining momentum. Current methods typically align distributions across domains to address inter-domain disparities at a global level. However, these methods often struggle with significant intra-domain gaps caused by domain-agnostic factors such as density, surveillance angles, and scale, leading to inaccurate alignment and unnecessary computational burdens, especially in large-scale training scenarios. To address these challenges, we propose the Multi-Granularity Optimal Transport (MGOT) distribution alignment framework, which aligns domain-agnostic factors across domains at different granularities. The motivation behind multi-granularity is to capture fine-grained domain-agnostic variations within domains. Our method proceeds in three phases: first, clustering coarse-grained features based on intra-domain similarity; second, aligning the granular clusters using an optimal transport framework and constructing a mapping from cluster centers to finer patch levels between domains; and third, re-weighting the aligned distribution for model refinement in domain adaptation. Extensive experiments across twelve cross-domain benchmarks show that our method outperforms existing state-of-the-art methods in adaptive crowd counting. The code will be available at https://github.com/HopooLinZ/MGOT.
Xian Zhong, Lingyue Qiu, Huilin Zhu, Jingling Yuan, Shengfeng He, Zheng Wang 0007
IEEE Trans. Image Process.3
2024 OneRestore: A Universal Restoration Framework for Composite Degradation
Yu Guo 0008, Yuan Gao 0015, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He
ECCV (19)4
2024 Zero-Shot Object Counting with Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Yu Guo 0008, Zheng Wang 0007, Xian Zhong, Shengfeng He
ECCV (5)1
2024 DenseTrack: Drone-Based Crowd Tracking via Density-Aware Motion-Appearance Synergy
abstract
Drone-based crowd tracking faces difficulties in accurately identifying and monitoring objects from an aerial perspective, largely due to their small size and close proximity to each other, which complicates both localization and tracking. To address these challenges, we present the Density-aware Tracking (DenseTrack) framework. DenseTrack capitalizes on crowd counting to precisely determine object locations, blending visual and motion cues to improve the tracking of small-scale objects. It specifically addresses the problem of cross-frame motion to enhance tracking accuracy and dependability. DenseTrack employs crowd density estimates as anchors for exact object localization within video frames. These estimates are merged with motion and position information from the tracking network, with motion offsets serving as key tracking cues. Moreover, DenseTrack enhances the ability to distinguish small-scale objects using insights from the visual-language model, integrating appearance with motion cues. The framework utilizes the Hungarian algorithm to ensure the accurate matching of individuals across frames. Demonstrated on DroneCrowd dataset, our approach exhibits superior performance, confirming its effectiveness in scenarios captured by drones. Our code will be available at: https://github.com/Zebrabeast/DenseTrack.
Huilin Zhu, Jingling Yuan, Guangli Xiang, Xian Zhong, Shengfeng He
ACM Multimedia2
2024 Find Gold in Sand: Fine-Grained Similarity Mining for Domain-Adaptive Crowd Counting
abstract
The domain shift of crowd scenes significantly hinders the application of crowd counting models in open scenarios. Although domain adaptation methods for crowd counting have bridged this gap to some extent, they ignore one of the significant causes of domain shift, which is the inter-domain data distribution bias. We discover that there exists a connection between the known and unknown distribution, which can be utilized by similarity mining to address the domain shift. However, there are still challenges related to insufficient and inaccurate similarity mining. In this article, a novel Fine-grained Inter-domain Similarity Mining (FSIM) framework is proposed. To comprehensively explore the similar distributions between source and target domains, we propose a Multi-scale Distribution Alignment (MDA) module based on diffusion retrieval. To enhance the reliability of inter- domain similarity mining, we propose a Multi-retrieval Refinement (MR) module based on evidence theory, which serves as an uncertainty measurement method. Eventually, to eliminate the data distribution bias, we perform model retraining using a similar distribution. Extensive experiments conducted on five standard crowd counting benchmarks, SHA, SHB, QNRF, NWPU, and JHU-CROWD++, show that the proposed FSIM has strong generalizability.
Huilin Zhu, Jingling Yuan, Xian Zhong, Zheng Wang 0007
IEEE Trans. Multim.1
2023 Striking a Balance: Unsupervised Cross-Domain Crowd Counting via Knowledge Diffusion
abstract
Supervised crowd counting relies on manual labeling, which is costly and time-consuming. This led to an increased interest in unsupervised methods. However, there is a significant domain gap issue in unsupervised methods, which is manifested by a model trained on one dataset serving dramatic performance drops when being transferred to another. This phenomenon can be attributed to the diverse domain knowledge making it difficult for the unsupervised models to transfer between general (e.g., similar distribution) and domain-specific (e.g., unique density, perspective, illumination, etc.) knowledge, leading to knowledge bias. Existing methods focus on exploring distinguishable relationships and establishing connections between the source and target domains. However, the similar knowledge transfer cannot perfectly simulate the contents of the target domain, leading to the model's inability to generalize to domain-specific knowledge. In this paper, we propose a Self-awareness Knowledge Diffusion method (SaKnD) that leverages the self-knowledge without establishing cross-domain knowledge relationships, which aims to balance the knowledge bias between general and domain-specific knowledge. Specifically, we propose a strategy to evaluate the uncertainty and consistency to define the clueless and informed areas, which determine the location and orientation of knowledge diffusion. These clueless areas serve as domain-specific knowledge that needs to be optimized, and these informed areas serve as general knowledge across domains. Extensive experiments on three standard crowd-counting benchmarks, ShanghaiTech PartA, ShanghaiTech PartB, and UCF_QNRF, show that the proposed SaKnD achieves state-of-the-art performance.
Haiyang Xie, Zhengwei Yang 0001, Huilin Zhu, Zheng Wang 0007
ACM Multimedia3
2023 DAOT: Domain-Agnostically Aligned Optimal Transport for Domain-Adaptive Crowd Counting
abstract
Domain adaptation is commonly employed in crowd counting to bridge the domain gaps between different datasets. However, existing domain adaptation methods tend to focus on inter-dataset differences while overlooking the intra-differences within the same dataset, leading to additional learning ambiguities. These domain-agnostic factors,e.g., density, surveillance perspective, and scale, can cause significant in-domain variations, and the misalignment of these factors across domains can lead to a drop in performance in cross-domain crowd counting. To address this issue, we propose a Domain-agnostically Aligned Optimal Transport (DAOT) strategy that aligns domain-agnostic factors between domains. The DAOT consists of three steps. First, individual-level differences in domain-agnostic factors are measured using structural similarity (SSIM). Second, the optimal transfer (OT) strategy is employed to smooth out these differences and find the optimal domain-to-domain misalignment, with outlier individuals removed via a virtual "dustbin'' column. Third, knowledge is transferred based on the aligned domain-agnostic factors, and the model is retrained for domain adaptation to bridge the gap across domains. We conduct extensive experiments on five standard crowd-counting benchmarks and demonstrate that the proposed method has strong generalizability across diverse datasets. Our code will be available at: https://github.com/HopooLinZ/DAOT/.
Huilin Zhu, Jingling Yuan, Xian Zhong, Zhengwei Yang 0001, Zheng Wang 0007, Shengfeng He
ACM Multimedia1
2023 Computer-Aided Autism Spectrum Disorder Diagnosis With Behavior Signal Processing
abstract
Behavioral observation plays an essential role in the diagnosis of Autism Spectrum Disorder (ASD) by analyzing children's atypical patterns in social activities (e.g., impaired social interaction, restricted interests, and repetitive behavior). To date, this process still heavily relies on the questionnaire survey, clinical observation, or retrospective video analysis, leading to high demand for professionals with massive labor costs. This article proposes a standardized platform for stimulating, gathering, analyzing, modeling, and interpreting human behavioral data in the application of computer-aided ASD diagnosis. By a structured assessment process, the proposed system can automatically evaluate children's multiple social interaction skills using the captured audio-visual data and provide the final diagnostic suggestions. We collect a multimodal behavioral database of 95 participants (71 children with ASD and 24 age-matched typical controls) in a real clinic environment, the Third Affiliated Hospital of Sun Yat-sen University, China. On the clinical database, our proposed computer-aided ASD diagnosis system obtains an accuracy of 88.42% for identifying ASD children with an average age of 24 months, representing a performance comparable to top-level human experts. As a unified and replicable solution, it has good potential to be promoted to less developed areas with limited high-quality medical resources.
Ming Cheng 0005, Yixiang Xie, Yueran Pan, Xiao Li 0048, Chengyan Yu, Dong Zhang 0002, Xiaoqian Huang, Cong You, Yuanyuan Zou 0003, Yuchong Liu, Fengjing Liang, Huilin Zhu, Chun Tang, Hongzhu Deng, Xiaobing Zou, Ming Li 0026
IEEE Trans. Affect. Comput.16
2022 Fine-Grained Fragment Diffusion for Cross Domain Crowd Counting
abstract
Deep learning improves the performance of crowd counting, but model migration remains a tricky challenge. Due to the reliance on training data and inherent domain shift, model application to unseen scenarios is tough. To facilitate the problem, this paper proposes a cross-domain Fine-Grained Fragment Diffusion model (FGFD) that explores feature-level fine-grained similarities of crowd distributions between different fragments to bridge the cross-domain gap (content-level coarse-grained dissimilarities). Specifically, we obtain features of fragments in both source and target domains, and then perform the alignment of the crowd distribution across different domains. With the assistance of the diffusion of crowd distribution, it is able to label unseen domain fragments and make source domain close to target domain, which is fed back to the model to reduce the domain discrepancy. By monitoring the distribution alignment, the distribution perception model is updated, then the performance of distribution alignment is improved. During the model inference, the gap between different domains is gradually alleviated. Multiple sets of migration experiments show that the proposed method achieves competitive results with other state-of-the-art domain-transfer methods.
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Xian Zhong, Zheng Wang 0007
ACM Multimedia1
2021 A Semi-Supervised Classification Method of Apicomplexan Parasites and Host Cell using Contrastive Learning Strategy
abstract
A common shortfall of supervised learning for medical imaging is the greedy need for human annotations, which is often expensive and time-consuming to obtain. This paper proposes a semi-supervised classification method for three kinds of apicomplexan parasites and non-infected host cells microscopic images, which uses a small number of labeled data and a large number of unlabeled data for training. There are two challenges in microscopic image recognition. The first is that salient structures of the microscopic images are more fuzzy and intricate than natural images’ on a real-world scale. The second is that insignificant textures, like background staining, lightness, and contrast level, vary a lot in samples from different clinical scenarios. To address these challenges, we aim to learn a distinguishable and appearance-invariant representation by contrastive learning strategy. On one hand, macroscopic images, which share similar shape characteristics in morphology, are introduced to contrast for structure enhancement. On the other hand, different appearance transformations, including color distortion and flittering, are utilized to contrast for texture elimination. In the case where only 1% of microscopic images are labeled, the proposed method reaches an accuracy of 94.90% in a generalized testing set.
Yanni Ren, Hangyu Deng, Hao Jiang 0028, Huilin Zhu, Jinglu Hu
SMC4
2021 Establishing A Hybrid Pieceswise Linear Model for Air Quality Prediction Based Missingness Challenges
abstract
Air pollution has threatened people’s health. It is urgent for the government to strengthen and improve the ability of air pollution monitoring. This paper proposes a winner-take-all (WTA) autoencoder-based piecewise linear model for imputing air quality prediction under the missing data scenario. The main idea consists of two parts. Firstly, overcomplete WTA stacked denoising autoencoders (SDAEs) are proposed to handle missing data, which play two roles: 1) devise the multiple imputation strategy to fill in missing values; 2) generate a set of binary gate control sequences to construct the sophisticated partitioning. Besides, renewed teacher signals are updated based on clustering information in the trained SDAEs to improve the accuracy of filling in missing samples. Secondly, the piecewise linear model is then proposed by the generated set of gate signals using the information from the feature layers of SDAEs. By using a quasi-linear kernel based on the trained gating mechanism, our piecewise air quality predictor is finally identified in the exact same way as support vector regression. The proposed modeling method is applied to real air quality datasets to show that it has led to greater performance than traditional models.
Huilin Zhu, Yanni Ren, Jinglu Hu
SMC1
2019 An automated assessment framework for atypical prosody and stereotyped idiosyncratic phrases related to autism spectrum disorder
Ming Li 0026, Dengke Tang, Junlin Zeng, Tianyan Zhou, Huilin Zhu, Biyuan Chen, Xiaobing Zou
Comput. Speech Lang.5
2015 Area Spectral Efficiency and Energy Efficiency Analysis in Downlink Massive MIMO Systems
abstract
We consider the downlink multi-user multi-cell massive MIMO systems, assuming that the number of antennas at base station (BS) and the number of users are large. Our system model accounts for channel estimation, pilot contamination, and uniformly random user location distribution. We derive the approximation of area spectral efficiency (ASE) with regularized zero-forcing (RZF) precoding technique which are proven to be accurate via simulation results. With a realistic power consumption model considering not only transmit power but also the fundamental power for operating the circuit at transmitter and receiver, we analyze the performance of area energy efficiency (AEE). Finally, based on the proposed power consumption model, we determine the optimal number of antennas at BS aimed at maximizing AEE when transmit power is given.
Yuanxue Xin, Dongming Wang 0002, Jiamin Li 0001, Huilin Zhu, Jiangzhou Wang, Xiaohu You 0001
VTC Fall4