VLDB 2026 Research / reviewers in the wild / expert
Shoudong Han
dblp:55/7737
· DBLP profile ↗
24ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-0572-4748ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focusing on temporal feature variations: Improving the adaptability of Re-identification in multi-object tracking
Mengyu He, Shoudong Han, Chaoyue Li |
Neurocomputing | 2 |
| 2025 | Head-Dominant Enhancement With Local Count for Better Human Detection in CrowdsabstractIn crowded scenes, it is difficult to extract discriminating human features due to occlusion. Some human detectors have improved this issue by introducing head detection. However, a complex problem still exists in associating full-body detection with its corresponding head detection. Instead of learning the association, we propose a Head-dominant Enhancement Module (HDEM) that uses the full-body proposal to regress the head bounding box. To embed the head information into the target human feature, we further propose a Consistent Weighted (CW) loss. Additionally, existing Non-Maximum Suppression (NMS) algorithms do not consider density changes with the selection of detection boxes, which leads to false and missed detection. Similar to human visual habits in occluded scenarios, we propose a Count-aware Dynamic Threshold Module (CADTM) that utilizes head context information to predict local count, which is associated with crowd density. CADTM can solve the inherent defects of Greedy-NMS in crowded scenes by adjusting the Intersection over Union (IoU) threshold dynamically. Ultimately, through the combination of HDEM and CADTM, we achieve state-of-the-art performance on CrowdHuman with a small computational cost. Our method achieves 4.6% AP gains, 2.2% MR-2 gains, and 3.2% JI gains over a Cascade R-CNN baseline. Furthermore, the proposed method is flexible and can be used with most proposal-based detection frameworks and various IoU-based NMS. Note to Practitioners—The motivation for this study arises from a prevalent issue encountered in human detection applications, particularly in densely populated areas such as shopping malls, streets, and subway stations, where occlusion poses a significant challenge. Employing a generic object detector results in numerous missed detections, thereby significantly compromising the overall performance of the detector. This research proposes a convenient plugin that achieves significant performance improvements at a very small cost. To validate its efficacy, our proposed method is thoroughly evaluated on proposal-based detection frameworks. The experimental results demonstrate the robustness of our approach and its ability to adapt to diverse crowded scenarios. Shoudong Han, Huilin Ding, Zhiling Han |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | DfTrack: Deconfused Data Association Framework for Multi-Object TrackingabstractAccurate data association plays a crucial role in Multi-Object Tracking (MOT) as it helps reduce confusion such as identity switches and assignment errors. However, many existing advanced methods often overlook the diversity among trajectories and the ambiguity and conflicts present in various types of cues. Consequently, when performing simple global data association, confusion arises between detections, trajectories, and associations. To address this problem, we propose a simple, versatile, and highly interpretable Deconfused Data Association Framework (DDAF). DDAF decomposes the traditional association problem into multiple sub-problems using a series of non-learnable modules, and selectively resolves confusion in each sub-problem by strategically utilizing new cues. Building upon DDAF, we design a powerful multi-object tracker named DfTrack, which specifically targets confusion in MOT. Furthermore, we discuss different specific implementations of DDAF to tackle challenging environments characterized by low frame rate, camera motion, and cross-domain scenarios. Correspondingly, we also develop several variants of DfTrack, demonstrating the remarkable scalability and adaptability of DDAF. Extensive experiments conducted on the MOT17, MOT20, and DanceTrack datasets demonstrate that DDAF significantly outperforms simple global association methods, and its variants can adapt to various challenging environments. Furthermore, DfTrack achieves state-of-the-art performance on multiple datasets, with HOTA of 65.2%, 63.9%, and 64.4% on MOT17, MOT20, and DanceTrack, respectively. The DfTrack-Hybrid variant further improves the performance on this basis. These results validate that our DDAF can effectively decompose and resolve various confusion in global association without any learning cost. Shoudong Han, Mengyu He, Yuhao Wei |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | DeconfuseTrack: Dealing with Confusion for Multi-Object TrackingabstractAccurate data association is crucial in reducing confusion, such as ID switches and assignment errors, in multi-object tracking (MOT). However, existing advanced methods often overlook the diversity among trajectories and the am-biguity and conflicts present in motion and appearance cues, leading to confusion among detections, trajectories, and associations when performing simple global data association. To address this issue, we propose a simple, versatile, and highly interpretable data association approach called De-composed Data Association (DDA). DDA decomposes the traditional association problem into multiple sub-problems using a series of non-learning-based modules and selectively addresses the confusion in each sub-problem by incorporating targeted exploitation of new cues. Additionally, we introduce Occlusion-aware Non-Maximum Suppression (ONMS) to retain more occluded detections, thereby increasing op-portunities for association with trajectories and indirectly reducing the confusion caused by missed detections. Finally, based on DDA and ONMS, we design a powerful multi-object tracker named DeconfuseTrack, specifically focused on resolving confusion in MOT. Extensive experiments conducted on the MOT17 and MOT20 datasets demonstrate that our proposed DDA and ONMS significantly enhance the performance of several popular trackers. Moreover, De-confuseTrack achieves state-of-the-art performance on the MOT17 and MOT20 test sets, significantly outperforms the baseline tracker ByteTrack in metrics such as HOTA, IDF1, AssA. This validates that our tracking design effectively reduces confusion caused by simple global association. Shoudong Han, Mengyu He, Yuhao Wei |
CVPR | 2 |
| 2024 | Focus and imagine: Occlusion suppression and repairing transformer for occluded person re-identification
Shoudong Han, Donghaisheng Liu, Delie Ming |
Neurocomputing | 2 |
| 2023 | Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language RepresentationabstractAlthough existing multi-object tracking (MOT) algorithms have obtained competitive performance on various benchmarks, almost all of them train and validate models on the same domain. The domain generalization problem of MOT is hardly studied. To bridge this gap, we first draw the observation that the high-level information contained in natural language is domain invariant to different tracking domains. Based on this observation, we propose to introduce natural language representation into visual MOT models for boosting the domain generalization ability. However, it is infeasible to label every tracking target with a textual description. To tackle this problem, we design two modules, namely visual context prompting (VCP) and visual-language mixing (VLM). Specifically, VCP generates visual prompts based on the input frames. VLM joints the information in the generated visual prompts and the textual prompts from a pre-defined Trackbook to obtain instance-level pseudo textual description, which is domain invariant to different tracking scenes. Through training models on MOT17 and validating them on MOT20, we observe that the pseudo textual descriptions generated by our proposed modules improve the generalization performance of query-based trackers by large margins. En Yu, Zhuoling Li, Shoudong Han, Wenbing Tao |
AAAI | 6 |
| 2023 | Focus On Details: Online Multi-Object Tracking with Diverse Fine-Grained RepresentationabstractDiscriminative representation is essential to keep a unique identifier for each target in Multiple object tracking (MOT). Some recent MOT methods extract features of the bounding box region or the center point as identity embeddings. However, when targets are occluded, these coarse-grained global representations become unreliable. To this end, we propose exploring diverse fine-grained representation, which describes appearance comprehensively from global and local perspectives. This fine-grained represen-tation requires high feature resolution and precise semantic information. To effectively alleviate the semantic misalignment caused by indiscriminate contextual information aggregation, Flow Alignment FPN (FAFPN) is proposed for multi-scale feature alignment aggregation. It generates semantic flow among feature maps from different resolutions to transform their pixel positions. Furthermore, we present a Multi-head Part Mask Generator (MPMG) to extract fine-grained representation based on the aligned feature maps. Multiple parallel branches of MPMG allow it to focus on different parts of targets to generate local masks without label supervision. The diverse details in target masks facilitate fine-grained representation. Eventually, benefiting from a Shuffle-Group Sampling (SGS) training strategy with positive and negative samples balanced, we achieve state-of-the-art performance on MOT17 and MOT20 test sets. Even on DanceTrack, where the appearance of targets is extremely similar, our method significantly outperforms Byte-Track by 5.0% on HOTA and 5.6% on IDF1. Extensive experiments have proved that diverse fine- grained representation makes Re-ID great again in MOT. Shoudong Han, Huilin Ding, Faquan Wang |
CVPR | 2 |
| 2023 | Spatial complementary and self-repair learning for occluded person re-identification
Shoudong Han, Donghaisheng Liu, Delie Ming |
Neurocomputing | 1 |
| 2023 | RelationTrack: Relation-Aware Multiple Object Tracking With Decoupled RepresentationabstractExisting online multiple object tracking (MOT) algorithms often consist of two subtasks, detection and re-identification (ReID). In order to enhance the inference speed and reduce the complexity, current methods commonly integrate these double subtasks into a unified framework. Nevertheless, detection and ReID demand diverse features. This issue results in an optimization contradiction during the training procedure. With the target of alleviating this contradiction, we devise a module named Global Context Disentangling (GCD) that decouples the learned representation into detection-specific and ReID-specific embeddings. As such, this module provides an implicit manner to balance the different requirements of these two subtasks. Moreover, we observe that preceding MOT methods typically leverage local information to associate the detected targets and neglect to consider the global semantic relation. To resolve this limitation, we develop a module, referred to as Guided Transformer Encoder (GTE), by combining the powerful reasoning ability of Transformer encoder and deformable attention. Unlike previous works, GTE avoids analyzing all the pixels and only attends to capture the relation between query nodes and a few self-adaptively selected key samples. Therefore, it is computationally efficient. Extensive experiments have been conducted on the MOT16, MOT17 and MOT20 benchmarks to demonstrate the superiority of the proposed MOT framework, namely RelationTrack. The experimental results indicate that RelationTrack has surpassed preceding methods significantly and established a new state-of-the-art performance, e.g., IDF1 of 70.5% and MOTA of 67.2% on MOT20. En Yu, Zhuoling Li, Shoudong Han |
IEEE Trans. Multim. | 3 |
| 2022 | Towards Discriminative Representation: Multi-view Trajectory Contrastive Learning for Online Multi-object TrackingabstractDiscriminative representation is crucial for the association step in multi-object tracking. Recent work mainly utilizes features in single or neighboring frames for constructing metric loss and empowering networks to extract representation of targets. Although this strategy is effective, it fails to fully exploit the information contained in a whole trajectory. To this end, we propose a strategy, namely multi-view trajectory contrastive learning, in which each trajectory is represented as a center vector. By maintaining all the vectors in a dynamically updated memory bank, a trajectory-level contrastive loss is devised to explore the inter-frame information in the whole trajectories. Besides, in this strategy, each target is represented as multiple adaptively selected keypoints rather than a pre-defined anchor or center. This design allows the network to generate richer representation from multiple views of the same target, which can better characterize occluded objects. Additionally, in the inference stage, a similarity-guided feature fusion strategy is developed for further boosting the quality of the trajectory representation. Extensive experiments have been conducted on MOTChallenge to verify the effectiveness of the proposed techniques. The experimental results indicate that our method has surpassed preceding trackers and established new state-of-the-art performance. En Yu, Zhuoling Li, Shoudong Han |
CVPR | 3 |
| 2022 | MAT: Motion-aware multi-object tracking
Shoudong Han, Piao Huang, En Yu, Donghaisheng Liu, Xiaofeng Pan |
Neurocomputing | 1 |
| 2022 | Foreground-guided textural-focused person re-identification
Donghaisheng Liu, Shoudong Han, Chenfei Xia, Jun Zhao 0007 |
Neurocomputing | 2 |
| 2022 | Toward Efficiently Evaluating the Robustness of Deep Neural Networks in IoT Systems: A GAN-Based MethodabstractIntelligent Internet of Things (IoT) systems based on deep neural networks (DNNs) have been widely deployed in the real world. However, DNNs are found to be vulnerable to adversarial examples, which raises people’s concerns about intelligent IoT systems’ reliability and security. Testing and evaluating the robustness of IoT systems become necessary and essential. Recently, various attacks and strategies have been proposed, but the efficiency problem remains unsolved properly. Existing methods are either computationally extensive or time consuming, which is not applicable in practice. In this article, we propose a novel framework, called attack-inspired generative adversarial networks (AI-GAN) to generate adversarial examples conditionally. Once trained, it can generate adversarial perturbations efficiently given input images and target classes. We apply AI-GAN on different data sets in white-box settings, black-box settings, and targeted models protected by state-of-the-art defenses. Through extensive experiments, AI-GAN achieves high attack success rates, outperforming existing methods, and reduces generation time significantly. Moreover, for the first time, AI-GAN successfully scales to complex data sets, e.g., CIFAR-100 and ImageNet, with about 90% success rates among all classes. Jun Zhao 0007, Jinlin Zhu, Shoudong Han, Jiefeng Chen 0001, Bo Li 0026, Alex Chichung Kot |
IEEE Internet Things J. | 4 |
| 2022 | Poisson kernel: Avoiding self-smoothing in graph convolutional networks
Ziqing Yang 0004, Shoudong Han, Jun Zhao 0007 |
Pattern Recognit. | 2 |
| 2021 | AI-GAN: Attack-Inspired Generation of Adversarial ExamplesabstractDeep neural networks (DNNs) are vulnerable to adversarial examples, which are crafted by adding imperceptible perturbations to inputs. Recently different attacks and strategies have been proposed, but how to generate adversarial examples perceptually realistic and more efficiently remains unsolved. This paper proposes a novel framework called Attack-Inspired GAN (AI-GAN), where a generator, a discriminator, and an attacker are trained jointly. Once trained, it can generate adversarial perturbations efficiently given input images and target classes. Through extensive experiments on several popular datasets e.g., MNIST and CFAR-10, AI-GAN achieves high attack success rates and reduces generation time significantly in various settings. Moreover, for the first time, AI-GAN successfully scales to complicated datasets e.g., CFAR-100 with around 90% success rates among all classes. Jun Zhao 0007, Jinlin Zhu, Shoudong Han, Jiefeng Chen 0001, Bo Li 0026, Alex Chichung Kot |
ICIP | 4 |
| 2019 | Deformed landmark fitting for sequential faces
Shoudong Han, Ziqing Yang 0004 |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Visual object tracking via enhanced structural correlation filter
Kai Chen 0023, Wenbing Tao, Shoudong Han |
Inf. Sci. | 3 |
| 2017 | Color-texture cosegmentation based on nonlinear compact multi-scale structure tensor and TV-flow
Shoudong Han, Wenbing Tao |
Signal Process. | 1 |
| 2015 | Combining pixel-level and patch-level information for segmentation
Tao Wang 0020, Zexuan Ji, Quan-Sen Sun, Shoudong Han |
Neurocomputing | 4 |
| 2015 | Image segmentation based on weighting boundary information via graph cut
Tao Wang 0020, Zexuan Ji, Quan-Sen Sun, Qiang Chen 0004, Shoudong Han |
J. Vis. Commun. Image Represent. | 5 |
| 2013 | Multilayer graph cuts based unsupervised color-texture image segmentation using multivariate mixed student's t-distribution and regional credibility merging
Shoudong Han, Tianjiang Wang, Wenbing Tao, Xue-Cheng Tai |
Pattern Recognit. | 2 |
| 2011 | Texture segmentation using independent-scale component-wise Riemannian-covariance Gaussian mixture model in KL measure based multi-scale nonlinear structure tensor space
Shoudong Han, Wenbing Tao, Xianglin Wu |
Pattern Recognit. | 1 |
| 2010 | Fast image segmentation based on multilevel banded closed-form method
Shoudong Han, Wenbing Tao, Xianglin Wu, Xue-Cheng Tai, Tianjiang Wang |
Pattern Recognit. Lett. | 1 |
| 2009 | Image Segmentation Based on GrabCut Framework Integrating Multiscale Nonlinear Structure TensorabstractIn this paper, we propose an interactive color natural image segmentation method. The method integrates color feature with multiscale nonlinear structure tensor texture (MSNST) feature and then uses GrabCut method to obtain the segmentations. The MSNST feature is used to describe the texture feature of an image and integrated into GrabCut framework to overcome the problem of the scale difference of textured images. In addition, we extend the Gaussian Mixture Model (GMM) to MSNST feature and GMM based on MSNST is constructed to describe the energy function so that the texture feature can be suitably integrated into GrabCut framework and fused with the color feature to achieve the more superior image segmentation performance than the original GrabCut method. For easier implementation and more efficient computation, the symmetric KL divergence is chosen to produce the estimates of the tensor statistics instead of the Riemannian structure of the space of tensor. The Conjugate norm was employed using Locality Preserving Projections (LPP) technique as the distance measure in the color space for more discriminating power. An adaptive fusing strategy is presented to effectively adjust the mixing factor so that the color and MSNST texture features are efficiently integrated to achieve more robust segmentation performance. Last, an iteration convergence criterion is proposed to reduce the time of the iteration of GrabCut algorithm dramatically with satisfied segmentation accuracy. Experiments using synthesis texture images and real natural scene images demonstrate the superior performance of our proposed method. Shoudong Han, Wenbing Tao, Xue-Cheng Tai, Xianglin Wu |
IEEE Trans. Image Process. | 1 |