VLDB 2026 Research / reviewers in the wild / expert
Jie Hu 0041
dblp:90/5064-41
· DBLP profile ↗
13ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-3296-5459ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Lightweight and Robust MCMT Tracking: Dual-Retrieval Knowledge Distillation and Scene-Aware FusionabstractMulti-Camera Multi-Target Tracking (MCMT) aims to achieve robust identity association of objects across cameras under diverse and challenging real-world conditions, such as varying viewpoints and occlusions. Current mainstream approaches often rely on deploying large-scale models with extensive feature extraction capabilities. However, the high computational demands of these models make them impractical for real-world MCMT scenarios where efficiency is critical. To address these challenges, we propose a novel Dual-Retrieval Knowledge Distillation (DRKD) framework, which enhances the student’s feature representation learning by leveraging multi-teacher guidance and dual-retrieval optimization. Unlike traditional knowledge distillation (KD) methods focusing primarily on feature alignment, DRKD introduces the Cross Triplet Loss, which optimizes the feature space for cross-camera identity association by enhancing intra-class compactness and inter-class separability. This dual-retrieval optimization ensures that the student model learns from the teacher’s feature representations and develops strong retrieval capabilities, which are crucial for robust identity association in MCMT. Additionally, we present the Dynamic Feature Fusion (DFF) module, which integrates short-term and historical features to balance the trade-off between responsiveness in single-camera tracking (SCT) and stability in multi-camera tracking (MCT). To further improve adaptability, we design Scene-Aware Regularization, which dynamically adjusts feature contributions in DFF based on temporal gaps and appearance variations. Extensive experiments on the HST and AI City Challenge S02 datasets demonstrate the effectiveness of our approach. The proposed DRKD framework with DFF achieves state-of-the-art performance, with IDF1 scores of 67.13 and MOTA scores of 60.50, while maintaining high inference speeds of up to 62 FPS. These results highlight the ability of DRKD to combine accuracy, robustness, and efficiency, making it a promising solution for real-world MCMT applications. Sixian Chan 0001, Xiaoxiang Chen, Wei Wang 0307, Jiafa Mao, Jie Hu 0041 |
IJCNN | 6 |
| 2025 | Deformable Blur Sensing and Regression Analysis ReID Feature Fusion for Multitarget Multicamera Tracking Systems in Highway ScenariosabstractIn highway scenarios, the rapid motion of vehicles can cause deformation and blur in camera footage, significantly affecting the accuracy of vehicle detection and re-identification (ReID) in multitarget multicamera tracking (MTMCT) systems. To address this issue, this article develops the deformable and blur sensing and regression analysis ReID feature fusion MTMCT system (DSRF). First, a deformable and blur sensing detection module (DFB) in DSRF is designed to overcome the limitations of cameras in capturing fast-moving objects, thereby accurately detecting vehicles moving at high speeds on highways. Then, a regression-based ReID feature fusion algorithm (RARF) in DSRF is proposed, which enhances ReID features by modeling the relationship between vehicle motion and its features, thereby better associating the detected vehicles in consecutive frames into trajectories and establishing intertrajectory relationships. Finally, extensive experiments are conducted on the highway surveillance traffic (HST) dataset developed by our team and the public dataset (CityFlow). Promising results are achieved, validating the effectiveness of our proposed method. Sixian Chan 0001, Shenghao Ni, Jie Hu 0041, Tinglong Tang, Xiaolong Zhou 0001, Pengyi Hao |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Adaptive Target-Oriented TrackingabstractThe current one-stream tracking pipelines are early relation modeling in feature extraction. However, insufficient discrimination may result in ambiguous relation modeling during early feature extraction. Moreover, the non-target information occupies most of the search image, rendering most relation modeling futile. To tackle the above issues, we propose tracking via learning adaptive target-oriented representation, named ATOTrack . We design an Untied positional encoding to mark the template token and the search region token separately, which reduces the confused relationship between the template and the search region. Besides, we introduce an Auto-Mask Learner to decouple the target and non-target information in the search region. Interestingly, the Auto-Mask Learner can self-learn and mask the ineffective information to interpret adaptive target-oriented representation. Extensive experiments demonstrate that ATOTrack is superior to existing methods, which achieves the state-of-the-art performance on six tracking benchmarks. In particular, ATOTrack establishes a new record on AViST with 57% AO. The code and models will be released as soon. Sixian Chan 0001, Xianpeng Zeng, Zhoujian Wu, Yu Wang 0254, Xiaolong Zhou 0001, Tinglong Tang, Jie Hu 0041 |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2024 | Modeling Label Correlations with Latent Context for Multi-label Recognition
Quan Cui, Ruoxi Deng, Jie Hu 0041, Guodao Zhang |
ECCV (33) | 4 |
| 2024 | Advancing Semantic Edge Detection through Cross-Modal Knowledge LearningabstractSemantic edge detection (SED) is pivotal for the precise demarcation of object boundaries, yet it faces ongoing challenges due to the prevalence of low-quality labels in current methods. In this paper, we present a novel solution to bolster SED through the encoding of both language and image data. Distinct from antecedent language-driven techniques, which predominantly utilize static elements such as dataset labels, our method taps into the dynamic language content that details the objects in each image and their interrelations. By encoding this varied input, we generate integrated features that utilize semantic insights to refine the high-level image features and the ultimate mask representations. This advancement improves the quality of these features and elevates SED performance. Experimental evaluation on benchmark datasets, including SBD and Cityscape, showcases the efficacy of our method, achieving leading ODS F-scores of 79.0 and 76.0, respectively. Our approach signifies a notable advancement in SED technology by seamlessly integrating multimodal textual information, embracing both static and dynamic aspects. Ruoxi Deng, Caixia Zhou, Jie Hu 0041 |
ACM Multimedia | 6 |
| 2024 | An efficient and real-time steel surface defect detection method based on single-stage detection algorithm
Qiqi Miao, Suqiang Li, Sixian Chan 0001, Jie Hu 0041, Cong Bai |
Multim. Tools Appl. | 6 |
| 2024 | Separable Spatial-Temporal Patch-Tensor Pair Completion for Infrared Small Target DetectionabstractThe infrared small target detection (IRSTD) task presents significant challenges due to low signal-to-clutter ratio, complicated background, and strong interferences. While tensor theory has shown promise in detection performance, three issues regarding damaged tensor construction, inaccurate tensor models, and high computation complexity remain. This study addresses these issues by introducing an independent spatial-temporal perspective, and proposes a fast and separable spatial-temporal tensor completion model. A new tensor structure named separable spatial-temporal patch-tensor pair (SSPP) is conceived to alleviate the dilemma of maintaining neighborhood structure and temporal consistency when constructing image tensors. By treating spatial and temporal dimensions as independent, SSPP enables flexible distribution hypotheses and representations in each dimension. Two tensor models are devised: the spatial model focuses on target enhancement in the spatial dimension, while the temporal one concentrates on suppressing strong interference in the temporal dimension. A long-term memory regularization is further introduced to the temporal model for target movement perception, enhancing its robustness to interferences. By combining these tensor models and employing a coarse-to-fine detection strategy, our method offers an effective solution for IRSTD. Extensive experiments on practical datasets have demonstrated the superiority of the proposed method in terms of target enhancement, background suppression, and detection efficiency. Chaoqun Xia, Shuhan Chen, Risheng Huang, Jie Hu 0041 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Diverse-Feature Collaborative Progressive Learning for Visible-Infrared Person Re-IdentificationabstractVisible–infrared person reidentification (VI-ReID) aims to search for pedestrian identities in different spectra. The major challenge is the modality differences between infrared and visible images for the VI-ReID task. Existing approaches try to design networks based on a single-stage training strategy to extract features. However, they often excessively rely on a particular feature, such as modality-specific features or modality-independent features, and overlook the significance of the diverse features obtained by combining them. To address this problem, we propose a diverse-feature collaborative progressive learning network (DCPLNet) for VI-ReID in this article. With the benefit of diverse information, our DCPLNet can effectively learn informative representations for reducing the modality differences. Specifically, we propose a novel three-stage progressive learning strategy (t-PLS) to progressively learn diverse features. For the proposed t-PLS, we design a contour feature enhancement module to mine human contour features and raise a perceptual contour feature loss for supervised feature extraction. Finally, we advance a batch adaptation module to establish feature links between samples. Extensive experiments on SYSU-MM01, RegDB, and LLCM datasets demonstrate that our proposed model performs better than most state-of-the-art methods. Sixian Chan 0001, Weihao Meng, Cong Bai, Jie Hu 0041, Shenyong Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Auxiliary Feature Fusion and Noise Suppression for HOI DetectionabstractIn recent years, one-stage HOI (Human–Object Interaction) detection methods tend to divide the original task into multiple sub-tasks by using a multi-branch network structure. However, there is no sufficient attention to information communication between these branches. The inference approach in the cascaded structure is singular, while fully parallel methods will disrupt the associations between different pieces of information. Besides, noise interference may occur during the fusion of different features and thus affect the detection performance. To address these issues, this article proposes a one-stage three-branch parallel HOI detection method, which treats HOI as three separate sub-tasks (human detection, object detection, and interaction detection) and leverages three distinct reasoning relationships to generate richer relational information. Firstly , an auxiliary feature fusion (AFF) module is introduced, which integrates features originally extracted independently to form fused features enriched with supplementary information. This approach strengthens communication between branches in the network while handling the three sub-tasks concurrently, thereby facilitating the exchange of more contextual information. Secondly , to mitigate noise interference generated during the fusion process, a fusion noise suppression (FNS) module is introduced, which effectively suppresses noise and enhances the model’s performance in interaction detection tasks. Finally , experiments are conducted on two major benchmark datasets, and experimental results show that our HOI detection method is superior to previous methods. Also, ablation studies confirm the effectiveness of all the components in our proposed method. Sixian Chan 0001, Xianpeng Zeng, Jie Hu 0041, Cong Bai |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | The Improvement of Road Driving Safety Guided by Visual Inattentional BlindnessabstractThe computational modeling of human visual attention has received much attention in recent decades. In advanced industrial applications, it has been demonstrated that computational visual attention models (CVAMs) can predict visual attention very similarly to human visual attention. However, it is controversial whether the driver’s eye fixation location (EFL) or the predicted eye fixation location of computational visual attention models is more reliable and helpful for actual driving. To address this issue, an open database of videos taken under the most common 18 driving conditions in everyday driving has been established. In experiments using this database, expert drivers found that it was not sufficient for drivers to rely on only one of the two EFLs. Based on this finding, a hybrid EFL recommendation strategy is proposed for improving driving safety. By extracting visual characteristics from human dynamic vision, the performance of the proposed recommendation method demonstrates its potential value in these collected driving tasks. In addition, the visual comfort of driving is further addressed to enhance the safety of driving. From the results of experiments on 108 driving video clips taken of the most common 18 real driving conditions, it is confirmed that the proposed EFL recommendation achieves an experience rating of driving comfort between 88.1 and 92.7 out of 100. Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002, Jie Hu 0041 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Adaptive Data Structure Regularized Multiclass Discriminative Feature SelectionabstractFeature selection (FS), which aims to identify the most informative subset of input features, is an important approach to dimensionality reduction. In this article, a novel FS framework is proposed for both unsupervised and semisupervised scenarios. To make efficient use of data distribution to evaluate features, the framework combines data structure learning (as referred to as data distribution modeling) and FS in a unified formulation such that the data structure learning improves the results of FS and vice versa. Moreover, two types of data structures, namely the soft and hard data structures, are learned and used in the proposed FS framework. The soft data structure refers to the pairwise weights among data samples, and the hard data structure refers to the estimated labels obtained from clustering or semisupervised classification. Both of these data structures are naturally formulated as regularization terms in the proposed framework. In the optimization process, the soft and hard data structures are learned from data represented by the selected features, and then, the most informative features are reselected by referring to the data structures. In this way, the framework uses the interactions between data structure learning and FS to select the most discriminative and informative features. Following the proposed framework, a new semisupervised FS (SSFS) method is derived and studied in depth. Experiments on real-world data sets demonstrate the effectiveness of the proposed method. Mingyu Fan, Xiaoqin Zhang 0002, Jie Hu 0041, Nannan Gu, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Multiple Lane Detection via Combining Complementary Structural ConstraintsabstractMany studies have been conducted on single lane detection, but multi-lane detection is rarely addressed. The latter is more advantageous for applications such as autonomous navigation, unmanned vehicles, departure warning, and cruise control. In this paper, we propose a novel and robust multiple lane detection algorithm based on the road structure information, which contains five complementary constraints: length constraint, parallel constraint, distribution constraint, pair constraint and uniform width constraint. All the five constraints are incorporated into a Hough transform (HT) based unified framework to select lane candidates. Nearly 99% of the false alarm candidates in HT space can be removed. Moreover, a dynamic programming strategy is proposed to find the most rational solutions among the remaining candidates. This strategy can effectively deal with combination complexity and interferences introduced by multi-lane detection. Experimental results on the benchmark dataset and other collected data demonstrate that the proposed method can outperform the state-of-the-art approaches in both accuracy and efficiency. Sheng Luo 0003, Xiaoqin Zhang 0002, Jie Hu 0041, Jinghua Xu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Semi-Supervised Dictionary Learning Based on Atom Graph RegularizationabstractIn this paper, we propose a novel unified optimization framework for semi-supervised dictionary learning, which optimizes a graph Laplacian component and the dictionary simultaneously. In the framework, the graph Laplacian is defined on the atoms and the corresponding sparse codings. Since the atoms are more concise and representative than the original training samples, the constructed graph Laplacian can not only effectively capture the manifold structure of training samples, but also be more robust to noise and outliers. Moreover, the dictionary and the graph Laplacian can facilitate each other during the learning iterations. We derive an efficient algorithm by combining the block coordinate descent method with the alternating direction method of multipliers to solve the unified optimization problem. Extensive experimental evaluation on several challenging datasets demonstrates the superior performance of the proposed method. Xiaoqin Zhang 0002, Di Wang 0008, Jie Hu 0041, Nannan Gu, Tianhao Wang 0006 |
IEEE BigData | 4 |