EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Zhang 0003
dblp:93/4664-3
· DBLP profile ↗
30ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0003-4854-3736ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeHub: Learning Hub-Resistant Representations for Text-Video RetrievalabstractText-video retrieval has achieved remarkable progress through fine-grained cross-modal alignment. However, the hubness problem, where a small subset of samples disproportionately dominate nearest-neighbor rankings for many queries, remains a fundamental obstacle in high-dimensional embedding spaces. Existing approaches primarily rely on post-hoc normalization during inference, making them sensitive to prior data distributions and vulnerable under domain shift. To address this limitation, we propose DeHub, a training-phase framework that learns hub-resistant representations by explicitly suppressing hubness during optimization through a dual-stage design. In the deterministic stage, a Hierarchical Subspace Decomposition (HiSD) module performs in-batch PCA to extract latent semantic subspaces, enriching representation granularity and reducing embedding collapse toward central hubs. A hubness weighting loss further identifies high-centrality samples and adaptively increases their contrastive penalties, mitigating semantically relevant hubs. In the probabilistic stage, text-video pairs are modeled as Gaussian distributions rather than fixed points, and a boundary distance loss enforces extremum-based separation at distributional boundaries, effectively repelling semantically irrelevant hubs. Extensive experiments on five benchmarks demonstrate consistent improvements over state-of-the-art methods. Moreover, cross-domain evaluations confirm that training-phase hubness suppression leads to stronger generalization under distribution shift. Anjun Jia, Xiaowei Zhang 0003 |
ICMR | 2 |
| 2026 | ASCFormer: An Adaptive Structure-Aware Cascaded Transformer for 3D Object Detectionabstract3D object detection has achieved significant progress in outdoor LiDAR point clouds, however, the inherent irregularity and varying sparsity distribution of point occupancy present a key challenge. Existing transformer-based 3D detectors often treat all tokens within the attention window as equally important, regardless of varying sparsity, which not only fails to address the disparities between the varying beam densities but also results in increased memory and computational costs. In this work, we propose an adaptive structure-aware cascaded transformer (ASCFormer) that dynamically captures density-insensitive multiscale structure features to model long-range dependencies via cascaded learning. Our ASCFormer detector includes an adaptive structure-aware token learning module that embeds voxel-level foreground probability and grid-level local density into the grid tokens to enhance structural perception capability. Moreover, we integrate these factors to compute significance scores, which are then utilized in inverse transform sampling to select a subset of multiscale tokens with varying receptive field sizes. To improve the training convergence of the window-based transformer in 3D voxel space, we employ cascaded learning via cross-stage attention to enhance the feature representation capability and refine the localization precision of 3D bounding boxes. This design of structure-aware reweighting effectively enhances the cascade paradigm, making to more adaptable to the varying sparsity distribution of point clouds. Extensive experiments on the KITTI and Waymo Open datasets demonstrate that the proposed ASCFormer detector achieves exceptional performance compared with state-of-the-art 3D object detection methods. The source code is publicly available at https://github.com/Xinglong-Li1/ASCFormer. Xiaowei Zhang 0003, Xinglong Li, Mingliang Zhou 0001, Min Gan, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | 3D Human Pose Estimation with Spatio-Temporal Topological Decoupling
Xiaolong Qin, Xiaowei Zhang 0003 |
IEEE Big Data | 2 |
| 2025 | Knowledge Distilled Group Prompts Learning for HOI Detection with Large Vision-Language ModelsabstractLarge vision-language models (VLMs) have significantly advanced human-object interaction (HOI) detection. However, existing VLM-based HOI detectors primarily rely on simple text prompt paradigms, specifically in relation to knowledge hallucination, with limited exploration of the intrinsic attributes or extrinsic context. In this paper, we propose a knowledge distilled group prompts learning method for HOI detection, termed GPL-HOI, which transfer knowledge from vision-language models via group prompts and knowledge distillation. Specifically, we design visual-textual group prompts by combining scene-aware, region-aware, and pose-aware prompt to guide knowledge transfer from VLMs. Additionally, we introduce a cross-modal group distillation module,which aligns the semantic features of both the vision and text models via KL divergence, encouraging the visual encoder to generate similar probability distributions to the text encoder through the learnable prompts. Extensive experiments demonstrate that our method surpasses state-of-the-art approaches in both conventional and zero-shot settings, achieving improvements of +2.04 mAP and +1.84 mAP on HICO-DET, respectively. Code will be available at https://github.com/hxqstree/GPL-HOI. Xiaoqian Han, Guanglin Niu, Mingliang Zhou 0001, Xiaowei Zhang 0003 |
ICME | 4 |
| 2025 | Exploring Hierarchical Multi-modal Prompts Learning for HOI Detection via Hybrid LearningabstractLarge Visual-Language Models (VLMs) have shown great potential in providing interaction priors for Human-Object Interaction (HOI) detection. However, existing VLM-based HOI detectors mainly focus on text prompt learning and have yet to fully explore the effectiveness of multimodal prompt learning. Additionally, due to annotation sparsity, interaction head generate a large number of candidate human-object pairs, but only a small subset is supervised. To address these challenges, we propose a hierarchical multimodal prompts learning approach for HOI detection via hybrid learning, called HMMP. Specifically, we introduce hierarchical visual-textual prompts designed to build multimodal prompts across three levels to guide knowledge transfer from VLMs, incorporating spatial priors to further explore multi-level visual information. Furthermore, we introduce a hybrid learning method that combines HOI triplets with probabilistic soft label supervision to explore implicit relationships between interactions, enhancing the model’s generalization ability without the need for additional data. Our proposed HMMP achieves state-of-the-art performance on both the HICO-DET and V-COCO datasets, in both standard and zero-shot settings. Xiaoqian Han, Xiaowei Zhang 0003 |
IJCNN | 2 |
| 2025 | Pose Guided Cross-Modal Invariant Feature Learning for Visible-Infrared Person Re-identification
TongJiaHao Teng, Xiaowei Zhang 0003 |
PRCV (17) | 3 |
| 2025 | Consistency-aware unsupervised label learning for cross-domain person re-identification
Yanbing Geng, Yongjian Lian, Fangshu Cui, Xiaowei Zhang 0003, Mingliang Zhou 0001, Geao Zhang |
Multim. Tools Appl. | 4 |
| 2025 | Geometry-Guided Point Generation for 3D Object DetectionabstractPoint cloud completion 3D object detectors effectively tackle the challenge of incomplete shapes in sparse point clouds by generating pseudo points to improve detection performance. However, the absence of guidance provided by the heatmap information and the geometric shape information renders the precise recovery of object shapes an arduous task. To this end, we propose a Geometry-guided Point Generation for 3D Object Detection, named GgPG. Specifically, we first design a 3D heatmap auxiliary supervision subnetwork to enhance the quality of object proposals by capturing the actual size and position of the object within the 3D heatmap representation. Moreover, we introduce a density-aware point generation module that employs Kernel Density Estimation (KDE) to embed the point density into the grid point's feature representation, thereby enabling the completion of more precise object shapes. Our GgPG achieves progressive performance in both Waymo and KITTI benchmarks, notably GgPG outperforms PGRCNN by +1.02$\%$, +1.18$\%$, and +0.56$\%$on the vehicle, pedestrian, and cyclist under LEVEL$\_$2 mAPH classes on Waymo Open Dataset, respectively. Mingliang Zhou 0001, Guanglin Niu, Xiaowei Zhang 0003 |
IEEE Signal Process. Lett. | 5 |
| 2024 | HT-SSPG:Hierarchical Transformers for Semantic Surface Point Generation in 3D Object Detection
Wenhao Kong, Xiaowei Zhang 0003 |
ACCV (7) | 2 |
| 2024 | Deformable Shape-Aware Point Generation for 3D Object Detection
Xiaowei Zhang 0003 |
ACCV (10) | 2 |
| 2024 | Hierarchical bi-directional conceptual interaction for text-video retrieval
Wenpeng Han, Guanglin Niu, Mingliang Zhou 0001, Xiaowei Zhang 0003 |
Multim. Syst. | 4 |
| 2024 | Dark knowledge association guided hashing for unsupervised cross-modal retrieval
Han Kang, Xiaowei Zhang 0003, Wenpeng Han, Mingliang Zhou 0001 |
Multim. Syst. | 2 |
| 2024 | Matching Multi-Scale Feature Sets in Vision Transformer for Few-Shot ClassificationabstractRecently, Transformer-based few-shot classification methods are widely exploited. However, they only leverage feature information at a single scale, resulting in weak feature representations, which cannot fully capture the rich information contained in a limited number of images regarding diverse objects with different scales, even those belonging to the same category. To mitigate this issue, we propose a multi-scale feature sets matching scheme in vision Transformer for few-shot classification, and name it FSViT, which can sufficiently extract discriminative features from the few number of labeled support examples. Concretely, we establish a patch-based multi-scale feature representation based on the feature extractors of FSViT, where we introduce an attention-aware grid pooling operation to merge adjacent patches with various scales to obtain multi-scale feature sets. Moreover, we devise a multi-scale patch matching metric to aggregate the measurement of similarity over the multi-scale feature sets for few-shot classification. Extensive experiments demonstrate the effectiveness of the proposed FSViT in both 1-shot and 5-shot scenarios on standard single-domain and cross-domain few-shot classification, especially improving the state-of-the-art recognition accuracy by 1.27% and 1.33% on average on the Mini-ImageNet and CFAIR-FS datasets, respectively. The code of FSViT is available athttps://github.com/codeshop715/FSViT. Mingchen Song, Fengqin Yao, Guoqiang Zhong 0001, Zhong Ji, Xiaowei Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Educational Pattern Guided Self-knowledge Distillation for Siamese Visual Tracking
Xiaowei Zhang 0003 |
ICONIP (14) | 2 |
| 2023 | CasFormer: Cascaded Transformer Based on Dynamic Voxel Pyramid for 3D Object Detection from Point Clouds
Xinglong Li, Xiaowei Zhang 0003 |
PRCV (3) | 2 |
| 2023 | Object Centric Body Part Attention Network for Human-Object Interaction Detection
Xiaowei Zhang 0003 |
PRCV (12) | 2 |
| 2023 | KTPose: Keypoint-Based Tokens in Vision Transformer for Human Pose EstimationabstractTransformers have made remarkable progress on human pose estimation in recent years, however, vision tokens are all in a fixed position, a property unsuitable for unknown human deformation. In this paper, we propose KTPose, a novel keypoint-based tokens in Vision Transformer for human pose estimation, which includes an instance-aware keypoint head and a keypoint refinement with transformer. To address the limb deformation issue, the instance-aware keypoint head is devised to capture the discriminative features dynamically based on the coarse localized keypoints. Further, we propose the multi-granularity vision tokens, in which each keypoint is explicitly embedded as a token to simultaneously learn spatial dependencies and constraint relationships from vision transformer for human pose estimation. Extensive experiments are carried out on two benchmark datasets, which demonstrate that KTPose outperforms state-of-the-art methods and achieves 76.6AP (?1.06%) and 75.7AP (?0.93%) on COCO validation and test-dev sets, respectively. This is accomplished with a smaller computational footprint when compared to the current mainstream transformer-based methods. Code is publicly available11https://github.com/WINGS-999/KTPose. Xiaowei Zhang 0003 |
SMC | 3 |
| 2022 | Heterogeneous Interactive Learning Network for Unsupervised Cross-Modal Retrieval
Yuanchao Zheng, Xiaowei Zhang 0003 |
ACCV (4) | 2 |
| 2022 | Mask-Guided Self-Distillation For Visual TrackingabstractRecently, Siamese-based visual trackers has been dramatically improved on performance, however, excessive parameters of tracking networks make models seriously hinder the practical deployment on edge devices. In this work, we present mask-guided self-distillation(MGSD) to compress the models of Siamese-based visual trackers, which enables Siamese-based visual trackers to capture crucial knowledge for effecting the performance of tracking. Specifically, MGSD consists of mask-guided semantic features self-distillation and decoupled tracking-head self-distillation, which discards the redundant convolution parameters in feature-dependency and task-dependency. It is worth noting that the proposed method compress exponentially the models of Siamese-based trackers under the condition of keeping even improving the performance of tracking. Extensive comprehensive experiments on tracking benchmarks including OTB2015, UAV123, LaSOT demonstrate that our proposed method can compress model size of Siamese-based visual trackers to 50% while achieving state-of-the-art performance. Our source code is available at: https://github.com/xl0312/MGSD. Luming Li, Chenglizhao Chen, Xiaowei Zhang 0003 |
ICME | 3 |
| 2022 | Relation-Guided Dual Hash Network for Unsupervised Cross-Modal Retrieval
Yuanchao Zheng, Xiaowei Zhang 0003 |
ICONIP (3) | 3 |
| 2022 | Multi-stream Feature Aggregation Network for 3D Object Detection in Point CloudabstractIn the recent 3D object detection methods for point clouds, the combination of point-based methods and voxel-based methods is gradually becoming a trend. Point-based methods retain the accurate position and pose information in the raw points and voxel-based methods get multi-scale structure information through the 3D backbone. However, because of the sparsity and irregularity of point clouds, both representations ignore the context information, which is important for the detection of sparse and small objects. To solve this problem, we propose a multi-stream feature aggregation network to extract features from three representations of the point cloud for object detection. Specifically, we exploit multi-stream features extracted from point, voxel, and perspective view (PV) respectively on a parallel way, where the complementary information between different perspectives can be used to enrich the feature representations, especially for the perspective view containing rich semantic context information. Secondly, to eliminate redundant information and better exploit the correlation between different feature representations, we design an attention-based multi-stream feature fusion module (MSFF) to combine features from three information streams. Besides, we introduce a new voxel RoI pooling with the self-attention in the second refinement stage, which can further strengthen the connection between local features in the proposal to obtain accurate classification and localization predictions. Our method achieves progressive results on the KITTI dataset, especially in the cyclist category, which improves the baseline significantly by 5.56%, 4.73%, 5.16% AP in the test set for easy, moderate, and hard levels respectively. Code will be available at https://github.com/june2678/MR F. Yingjie Hou, Xiaowei Zhang 0003 |
SMC | 2 |
| 2022 | Heterogeneous Interactive Attention Network for Human ParsingabstractAs a fine-grained semantic segmentation task, human parsing has attracted extensive attention in computer vision. However, without the assistance of heterogeneous information, it is difficult to obtain detailed human parsing directly. At present, although some studies have introduced heterogeneous data to guide human parsing task, such as pose estimation and edge prediction, the correlations between these heterogeneous data has not been effectively utilized. To avoid the distribution gap among heterogeneous data, we proposed a Heterogeneous Interactive Attention Network (HIANet), in which we exploit the attention between heterogeneous data to capture long-distance context dependence. And the supplementary cues with plentiful interaction can mutually guide multi-source features to correct their respective prediction errors, further refine the result of human parsing. Extensive experiments on three human body parsing datasets are conducted, especially on the LIP dataset, where the mean accuracy and mean Intersection-over-Union of the proposed HIANet are improved by 2.86% and 3.90% compared with PGECNet, respectively. Our code has been made available at https://github.com/wangwenjiawj/HIANet. Xiaowei Zhang 0003 |
SMC | 3 |
| 2022 | Cluster-aware Diversity Samples Mining for Unsupervised Person Re-IdentificationabstractRecently, state-of-the-art unsupervised re-ID methods train the neural network by calculating cluster-level or instance-level contrastive loss for learning discriminative features with unlabeled data. However, due to the divergence of the individual cluster, the previous methods did not fully utilize inherent feature of clustered samples for contrastive learning, which merged an unreliable instance into a wrong cluster by simply using cluster centroid or hard instance. To solve this issue, we propose a novel Cluster-aware Diversity Samples Mining (CDSM) framework based on the compactness and independence of individual cluster to generate diverse samples for updating memory dictionary of each cluster, so as to reduce the effect of noisy labels and improve the robustness of the model performance. Significantly, the proposed Cluster-aware Diversity Samples Mining method gradually creates more reliable clusters to generate more robust pseudo labels by refining the memory which is of central importance to our outstanding performance. Extensive experiments demonstrate that the proposed CDSM framework achieves performances of 85.6%, 73.7% and 31.0% in mAP on Market1501, DukeMTMC-reID and MSMT17, respectively. Code is available at https://github.com/colinzhaoxp/CDSM. Xinpeng Zhao 0001, Xiao Dou, Xiaowei Zhang 0003 |
SMC | 3 |
| 2022 | Disentangling classification and regression in Siamese-based network for visual trackingabstractSummary Siamese‐based trackers have made great progress in visual tracking community, however, the shared structure of network between classification and regression tasks limits the ability of the trackers to obtain more robust classification prediction and more accurate regression prediction. In this paper, we propose an effective visual tracking framework (named Siamese Disentangled Tracking‐Head, SiamDTH), which disentangles classification and regression in Siamese‐based network for visual tracking from two aspects: feature decoupling and differentiated tracking‐head. First of all, we gather the features of receptive fields with different scales and ratios, and decouple the correlation features through two different styles of feature fusion mode for classification and regression respectively. Moreover, we design the differentiated tracking‐head structure in the sibling head for discriminately handling the parallel classification and regression tasks on visual tracking. Extensive experiments on visual tracking benchmarks including VOT2018, VOT2019 and OTB100 demonstrate that our proposed SiamDTH achieves state‐of‐the‐art performance with a considerable real‐time speed. Our source code is available at: https://github.com/xl0312/SiamDTH . Xiaowei Zhang 0003, Luming Li, Hong Liu 0013 |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Rule-Guided Compositional Representation Learning on Knowledge GraphsabstractRepresentation learning on a knowledge graph (KG) is to embed entities and relations of a KG into low-dimensional continuous vector spaces. Early KG embedding methods only pay attention to structured information encoded in triples, which would cause limited performance due to the structure sparseness of KGs. Some recent attempts consider paths information to expand the structure of KGs but lack explainability in the process of obtaining the path representations. In this paper, we propose a novel Rule and Path-based Joint Embedding (RPJE) scheme, which takes full advantage of the explainability and accuracy of logic rules, the generalization of KG embedding as well as the supplementary semantic structure of paths. Specifically, logic rules of different lengths (the number of relations in rule body) in the form of Horn clauses are first mined from the KG and elaborately encoded for representation learning. Then, the rules of length 2 are applied to compose paths accurately while the rules of length 1 are explicitly employed to create semantic associations among relations and constrain relation embeddings. Moreover, the confidence level of each rule is also considered in optimization to guarantee the availability of applying the rule to representation learning. Extensive experimental results illustrate that RPJE outperforms other state-of-the-art baselines on KG completion task, which also demonstrate the superiority of utilizing logic rules as well as paths for improving the accuracy and explainability of representation learning. Guanglin Niu, Yongfei Zhang, Bo Li 0006, Peng Cui 0001, Si Liu 0001, Xiaowei Zhang 0003 |
AAAI | 7 |
| 2020 | Improved Robust Video Saliency Detection Based on Long-Term Spatial-Temporal InformationabstractThis paper proposes to utilize supervised deep convolutional neural networks to take full advantage of the long-term spatial-temporal information in order to improve the video saliency detection performance. The conventional methods, which use the temporally neighbored frames solely, could easily encounter transient failure cases when the spatial-temporal saliency clues are less-trustworthy for a long period. To tackle the aforementioned limitation, we plan to identify those beyond-scope frames with trustworthy long-term saliency clues first and then align it with the current problem domain for an improved video saliency detection. Chenglizhao Chen, Guotao Wang 0004, Chong Peng 0001, Xiaowei Zhang 0003, Hong Qin 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Too Far to See? Not Really! - Pedestrian Detection With Scale-Aware Localization PolicyabstractA major bottleneck of pedestrian detection lies on the sharp performance deterioration in the presence of small-size pedestrians that are relatively far from the camera. Motivated by the observation that pedestrians of disparate spatial scales exhibit distinct visual appearances, we propose in this paper an active pedestrian detector that explicitly operates over multiple-layer neuronal representations of the input still image. More specifically, convolutional neural nets, such as ResNet and faster R-CNNs, are exploited to provide a rich and discriminative hierarchy of feature representations, as well as initial pedestrian proposals. Here each pedestrian observation of distinct size could be best characterized in terms of the ResNet feature representation at a certain layer of the hierarchy. Meanwhile, initial pedestrian proposals are attained by the faster R-CNNs techniques, i.e., region proposal network and follow-up region of interesting pooling layer employed right after the specific ResNet convolutional layer of interest, to produce joint predictions on the bounding-box proposals' locations and categories (i.e., pedestrian or not). This is engaged as an input to our active detector, where for each initial pedestrian proposal, a sequence of coordinate transformation actions is carried out to determine its proper x-y 2D location and the layer of feature representation, or eventually terminated as being background. Empirically our approach is demonstrated to produce overall lower detection errors on widely used benchmarks, and it works particularly well with far-scale pedestrians. For example, compared with 60.51% log-average miss rate of the state-of-the-art MS-CNN for far-scale pedestrians (those below 80 pixels in bounding-box height) of the Caltech benchmark, the miss rate of our approach is 41.85%, with a notable reduction of 18.66%. Xiaowei Zhang 0003, Li Cheng 0001, Bo Li 0006, Hai-Miao Hu |
IEEE Trans. Image Process. | 1 |
| 2017 | Scale-aware hierarchical loss: A multipath RPN for multi-scale pedestrian detectionabstractPedestrians with different spatial scales exhibiting dramatically differences, the serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection. Considering the local feature differences for multi-scale pedestrians, a scale-aware multipath region proposal network is exploited to improve the recall rate, which is divided into several branches to generate a proper object proposal for target with specific scale range. Moreover, motivated by the visual semantic concepts of different convolutional layers, a scale-aware hierarchical loss model is introduced to minimize the error rate for pedestrians with different scales, in which the hierarchical features of higher convolutional layers are jointed to calculate a multi-task loss to learn scale-aware weighting of multipath region proposal network for each object proposal. Finally, compared to state-of-the-art methods, experimental results on the challenging ETH and Caltech benchmark show the superiority of the proposed method for large variance in instance scales. Xiaowei Zhang 0003, Bo Li 0006, Hai-Miao Hu |
VCIP | 1 |
| 2015 | Pedestrian detection based on hierarchical co-occurrence model for occlusion handling
Xiaowei Zhang 0003, Hai-Miao Hu, Bo Li 0006 |
Neurocomputing | 1 |
| 2015 | Joint global-local information pedestrian detection algorithm for outdoor video surveillance
Hai-Miao Hu, Xiaowei Zhang 0003, Bo Li 0006 |
J. Vis. Commun. Image Represent. | 2 |