EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Zhang 0015
dblp:58/4582-15
· DBLP profile ↗
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-2999-3291ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I2V-Adapter: Fast adapting image pre-trained models for video correspondence
Hannan Lu, Xinyu Zhang 0015, Zhi Tian, Xiaohe Wu, Wangmeng Zuo, Jingdong Wang 0001 |
Pattern Recognit. | 2 |
| 2026 | FastPillars: A Deployment-Friendly Pillar-Based 3D DetectorabstractThe deployment of 3D detectors strikes one of the major challenges in real-world self-driving scenarios. Existing BEV-based (i.e., Bird Eye View) detectors favor sparse convolutions (known as SPConv) to speed up training and inference, which puts a hard barrier for deployment, especially for on-device applications. In this paper, in order to tackle the challenge of efficient 3D object detection from an industry perspective, we devise a deployment-friendly pillar-based 3D detector, termed FastPillars. Specifically, aiming to compensate the geometric information loss of pillar encoding. First, we design a novel lightweight Max-and-Attention Pillar Encoding (MAPE) module specially for enhancing small objects. Second, we propose a simple yet effective backbone design for pillar-based 3D detection, enhancing pillar representations. We construct FastPillars based on these designs, achieving high performance and low latency without SPConv. Extensive experiments on two large-scale datasets demonstrate the effectiveness and efficiency of FastPillars for on-device 3D detection regarding both performance and speed. Specifically, FastPillars delivers real-time state-of-the-art accuracy on Waymo Open Dataset with 1.8 × speed up and 3.8 mAPH/L2 improvement over CenterPoint (SPConv-based). Code will be opened soon in: https://github.com/StiphyJay/FastPillars. Sifan Zhou, Xinyu Zhang 0015, Xiangxiang Chu, Bo Zhang 0046, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Lenna: Language Enhanced Reasoning Detection AssistantabstractWith the fast-paced development of multimodal large language models (MLLMs), we can now converse with AI systems in natural languages to understand images. However, the reasoning power and world knowledge embedded in the large language models have been much less investigated and exploited for image perception tasks. In this paper, we propose Lenna, a Language enhanced reasoning detection assistant, which utilizes the robust multimodal feature representation of MLLMs, while preserving location information for detection. This is achieved by incorporating an additionaltoken in the MLLM vocabulary that is free of explicit semantic context but serves as a prompt for the detector to identify the corresponding position. To evaluate the reasoning capability of Lenna, we construct a ReasonDet dataset to measure its performance on reasoning-based detection. Remarkably, Lenna demonstrates outstanding performance on ReasonDet and comes with significantly low training costs. It also incurs minimal transferring overhead when extended to other tasks. Fei Wei, Xinyu Zhang 0015, Bo Zhang 0046, Xiangxiang Chu |
ICASSP | 2 |
| 2024 | LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object DetectionabstractDue to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted in 2D vision tasks. However, applying it directly to 3D lidar-based tasks inevitably leads to performance degradation. As a remedy, we propose an effective PTQ method called LiDAR-PTQ, which is particularly curated for 3D lidar detection (both SPConv-based and SPConv-free). Our LiDAR-PTQ features three main components, (1) a sparsity-based calibration method to determine the initialization of quantization parameters, (2) an adaptive rounding-to-nearest operation to minimize the layerwise reconstruction error, (3) a Task-guided Global Positive Loss (TGPL) to reduce the disparity between the final predictions before and after quantization. Extensive experiments demonstrate that our LiDAR-PTQ can achieve state-of-the-art quantization performance when applied to CenterPoint (both Pillar-based and Voxel-based). To our knowledge, for the very first time in lidar-based 3D detection tasks, the PTQ INT8 model's accuracy is almost the same as the FP32 model while enjoying 3X inference speedup. Moreover, our LiDAR-PTQ is cost-effective being 6X faster than the quantization-aware training method. The code will be released. Sifan Zhou, Liang Li 0003, Xinyu Zhang 0015, Bo Zhang 0046, Shipeng Bai, Xiaobo Lu, Xiangxiang Chu |
ICLR | 3 |
| 2024 | STAT: Multi-Object Tracking Based on Spatio-Temporal Topological ConstraintsabstractThe mainstream tracking-by-detection paradigm for multi-object tracking generally conducts detection first, followed by Re-IDentification (Re-ID) and motion estimation. The associations between the predicted boxes and existing tracks are then performed via visual and motion association. However, challenges such as irregular motion patterns, similar appearances, and frequent occlusions often arise, making object tracking a nontrivial task. In this article, we propose a multi-object tracker based on Spatio-TemporAl Topological (STAT) constraints to address the above issues. More specifically, we design the Feature Adaptive Association Module (FAAM) to establish the association between motion and appearance regionally, completing a complementary combination of appearance and motion features. Among these, the Appearance Feature Update Module (AFUM) is proposed to manage the appearance updates of tracked objects by imposing constraints based on the spatial locations and the degree of object occlusion, while temporal consistency is adopted to smooth the appearance states of tracks to mitigate the accumulation of appearance noise. Moreover, the Robust Motion Tracking Module (RMTM) is established to reduce the impact of irregular motions and certain unreliable detection results. The proposed module includes a higher weighted momentum term to accommodate the excessive motion amplitude and considers low-confidence boxes accompanied by the stage-wise association strategy for high-confidence boxes. Extensive experiments on DanceTrack and benchmark MOT datasets verify the effectiveness of our STAT tracker, especially the state-of-the-art results on DanceTrack, which is characterized by irregular motion and indistinguishable appearance attributes. Junjie Zhang 0002, Xinyu Zhang 0015, Chenggang Yan 0001, Dan Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationabstractThe pre-training task is indispensable for the text-to-image person re-identification (T2I-ReID) task. However, there are two underlying inconsistencies between these two tasks that may impact the performance: i) Data inconsistency. A large domain gap exists between the generic images/texts used in public pre-trained models and the specific person data in the T2I-ReID task. This gap is especially severe for texts, as general textual data are usually unable to describe specific people in fine-grained detail. ii) Training inconsistency. The processes of pre-training of images and texts are independent, despite cross-modality learning being critical to T2I-ReID. To address the above issues, we present a new unified pre-training pipeline (UniPT) designed specifically for the T2I-ReID task. We first build a large-scale text-labeled person dataset "LUPerson-T", in which pseudo-textual descriptions of images are automatically generated by the CLIP paradigm using a divide-conquer-combine strategy. Benefiting from this dataset, we then utilize a simple vision-and-language pre-training framework to explicitly align the feature space of the image and text modalities during pre-training. In this way, the pre-training task and the T2I-ReID task are made consistent with each other on both data and training levels. Without the need for any bells and whistles, our UniPT achieves competitive Rank-1 accuracy of, i.e., 68.50%, 60.09%, and 51.85% on CUHK-PEDES, ICFG-PEDES and RSTPReid, respectively. Both the LUPerson-T dataset and code are available at https://github.com/ZhiyinShao-H/UniPT. Zhiyin Shao, Xinyu Zhang 0015, Changxing Ding, Jian Wang 0066, Jingdong Wang 0001 |
ICCV | 2 |
| 2023 | HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionabstractModel pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation. Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001 |
NeurIPS | 2 |
| 2023 | A Real-Time Memory Updating Strategy for Unsupervised Person Re-IdentificationabstractRecently, clustering-based methods have been the dominant solution for unsupervised person re-identification (ReID). Memory-based contrastive learning is widely used for its effectiveness in unsupervised representation learning. However, we find that the inaccurate cluster proxies and the momentum updating strategy do harm to the contrastive learning system. In this paper, we propose a real-time memory updating strategy (RTMem) to update the cluster centroid with a randomly sampled instance feature in the current mini-batch without momentum. Compared to the method that calculates the mean feature vectors as the cluster centroid and updating it with momentum, RTMem enables the features to be up-to-date for each cluster. Based on RTMem, we propose two contrastive losses, i.e., sample-to-instance and sample-to-cluster, to align the relationships between samples to each cluster and to all outliers not belonging to any other clusters. On the one hand, sample-to-instance loss explores the sample relationships of the whole dataset to enhance the capability of density-based clustering algorithm, which relies on similarity measurement for the instance-level images. On the other hand, with pseudo-labels generated by the density-based clustering algorithm, sample-to-cluster loss enforces the sample to be close to its cluster proxy while being far from other proxies. With the simple RTMem contrastive learning strategy, the performance of the corresponding baseline is improved by 9.3% on Market-1501 dataset. Our method consistently outperforms state-of-the-art unsupervised learning person ReID methods on three benchmark datasets. Code is made available at:https://github.com/PRIS-CV/RTMem. Junhui Yin, Xinyu Zhang 0015, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Implicit Sample Extension for Unsupervised Person Re-IdentificationabstractMost existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters substantially hampers the Re-ID accuracy. Due to the limited samples in each identity, we suppose there may lack some underlying information to well reveal the accurate clusters. To discover these information, we propose an Implicit Sample Extension (ISE) method to generate what we call support samples around the cluster boundaries. Specifically, we generate support samples from actual samples and their neighbouring clusters in the embedding space through a progressive linear interpolation (PLI) strategy. PLI controls the generation with two critical factors, i.e., 1) the direction from the actual sample towards its K-nearest clusters and 2) the degree for mixing up the context information from the K-nearest clusters. Meanwhile, given the support samples, ISE further uses a label-preserving loss to pull them towards their corresponding actual samples, so as to compact each cluster. Consequently, ISE reduces the “sub and mixed” clustering errors, thus improving the Re-ID performance. Extensive experiments demonstrate that the proposed method is effective and achieves state-of-the-art performance for unsupervised person Re-ID. Code is available at: https://github.com/PaddlePaddle/PaddleClas. Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Errui Ding, Qinfeng Shi, Zhaoxiang Zhang 0001, Jingdong Wang 0001 |
CVPR | 1 |
| 2022 | UFO: Unified Feature Optimization
Teng Xi, Yifan Sun 0003, Deli Yu, Bi Li 0005, Nan Peng, Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
ECCV (26) | 7 |
| 2022 | Self-Guided Hard Negative Generation for Unsupervised Person Re-IdentificationabstractRecent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative samples, playing important role in training reID models, are significantly reduced. To alleviate this problem, we propose a self-guided hard negative generation method for unsupervised person re-ID. Specifically, a joint framework is developed which incorporates a hard negative generation network (HNGN) and a re-ID network. To continuously generate harder negative samples to provide effective supervisions in the contrastive learning, the two networks are alternately trained in an adversarial manner to improve each other, where the reID network guides HNGN to generate challenging data and HNGN enforces the re-ID network to enhance discrimination ability. During inference, the performance of re-ID network is improved without introducing any extra parameters. Extensive experiments demonstrate that the proposed method significantly outperforms a strong baseline and also achieves better results than state-of-the-art methods. Zhigang Wang 0002, Jian Wang 0066, Xinyu Zhang 0015, Errui Ding, Jingdong Wang 0001, Zhaoxiang Zhang 0001 |
IJCAI | 4 |
| 2022 | Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationabstractText-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal gaps. Existing works usually ignore the difference in feature granularity between the two modalities, i.e., the visual features are usually fine-grained while textual features are coarse, which is mainly responsible for the large inter-modal gaps. In this paper, we propose an end-to-end framework based on transformers to learn granularity-unified representations for both modalities, denoted as LGUR. LGUR framework contains two modules: a Dictionary-based Granularity Alignment (DGA) module and a Prototype-based Granularity Unification (PGU) module. In DGA, in order to align the granularities of two modalities, we introduce a Multi-modality Shared Dictionary (MSD) to reconstruct both visual and textual features. Besides, DGA has two important factors, i.e., the cross-modality guidance and the foreground-centric reconstruction, to facilitate the optimization of MSD. In PGU, we adopt a set of shared and learnable prototypes as the queries to extract diverse and semantically aligned features for both modalities in the granularity-unified feature space, which further promotes the ReID performance. Comprehensive experiments show that our LGUR consistently outperforms state-of-the-arts by large margins on both CUHK-PEDES and ICFG-PEDES datasets. Code will be released at https://github.com/ZhiyinShao-H/LGUR. Zhiyin Shao, Xinyu Zhang 0015, Zhifeng Lin, Jian Wang 0066, Changxing Ding |
ACM Multimedia | 2 |
| 2022 | Part-Guided Attention Learning for Vehicle Instance RetrievalabstractVehicle instance retrieval (IR) often requires one to recognize the fine-grained visual differences between vehicles. Besides the holistic appearance of vehicles which is easily affected by the viewpoint variation and distortion, vehicle parts also provide crucial cues to differentiate near-identical vehicles. Motivated by these observations, we introduce aPart-Guided Attention Network(PGAN) to pinpoint the prominent part regions and effectively combine the global and local information for discriminative feature learning. PGAN first detects the locations of different part components and salient regions regardless of the vehicle identity, which serves as thebottom-up attentionto narrow down the possible searching regions. To estimate the importance of detected parts, we propose aPart Attention Module(PAM) to adaptively locate the most discriminative regions with high-attention weights and suppress the distraction of irrelevant parts with relatively low weights. The PAM is guided by the identification loss and therefore providestop-down attentionthat enables attention to be calculated at the level of car parts and other salient regions. Finally, we aggregate the global appearance and local features together to improve the feature performance further. The PGAN combines part-guided bottom-up and top-down attention, global and local visual features in an end-to-end framework. Extensive experiments demonstrate that the proposed method achieves new state-of-the-art vehicle IR performance on four large-scale benchmark datasets.1 Xinyu Zhang 0015, Rufeng Zhang, Jiewei Cao, Dong Gong, Mingyu You, Chunhua Shen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Diverse Knowledge Distillation for End-to-End Person SearchabstractPerson search aims to localize and identify a specific person from a gallery of images. Recent methods can be categorized into two groups, i.e., two-step and end-to-end approaches. The former views person search as two independent tasks and achieves dominant results using separately trained person detection and re-identification (Re-ID) models. The latter performs person search in an end-to-end fashion. Although the end-to-end approaches yield higher inference efficiency, they largely lag behind those two-step counterparts in terms of accuracy. In this paper, we argue that the gap between the two kinds of methods is mainly caused by the Re-ID sub-networks of end-to-end methods. To this end, we propose a simple yet strong end-to-end network with diverse knowledge distillation to break the bottleneck. We also design a spatial-invariant augmentation to assist model to be invariant to inaccurate detection results. Experimental results on the CUHK-SYSU and PRW datasets demonstrate the superiority of our method against existing approaches -- it achieves on par accuracy with state-of-the-art two-step methods while maintaining high efficiency due to the single joint model. Code is available at: https://git.io/DKD-PersonSearch. Xinyu Zhang 0015, Jiawang Bian, Chunhua Shen, Mingyu You |
AAAI | 1 |
| 2019 | Self-Training With Progressive Augmentation for Unsupervised Cross-Domain Person Re-IdentificationabstractPerson re-identification (Re-ID) has achieved great improvement with deep learning and a large amount of labelled training data. However, it remains a challenging task for adapting a model trained in a source domain of labelled data to a target domain of only unlabelled data available. In this work, we develop a self-training method with progressive augmentation framework (PAST) to promote the model performance progressively on the target dataset. Specially, our PAST framework consists of two stages, namely, conservative stage and promoting stage. The conservative stage captures the local structure of target-domain data points with triplet-based loss functions, leading to improved feature representations. The promoting stage continuously optimizes the network by appending a changeable classification layer to the last layer of the model, enabling the use of global information about the data distribution. Importantly, we propose a new self-training strategy that progressively augments the model capability by adopting conservative and promoting stages alternately. Furthermore, to improve the reliability of selected triplet samples, we introduce a ranking-based triplet loss in the conservative stage, which is a label-free objective function based on the similarities between data pairs. Experiments demonstrate that the proposed method achieves state-of-the-art person Re-ID performance under the unsupervised cross-domain setting. Xinyu Zhang 0015, Jiewei Cao, Chunhua Shen, Mingyu You |
ICCV | 1 |
| 2018 | An Extended Filtered Channel Framework for Pedestrian DetectionabstractPedestrian detection is an important example of object detection and has attracted much attention. Many works have shown that good image features provide high detection accuracy, and a few works have investigated enhancing low-level features (e.g., gradient and color features) using a filtered layer (i.e., convolutional layer) to obtain enhanced features or filtered channel features. To investigate whether these features are saturated, this paper adopts the concept of filtered channel features and strengthens them by adding more convolutional layers. Acting as convolution kernels, multilayer filters are applied to low-level features to obtain the extended filtered channel features, providing a powerful feature extractor with multilayer transformation for pedestrian detection. The proposed extended filtered channel framework (ExtFCF) achieves competitive performance on widely used benchmark datasets (Caltech, INRIA, and KITTI datasets), using only histogram of oriented gradient (HOG) and CIE-LUV [a color space composing of luminance (L) and two chrominance (UV) components by International Commission on Illumination] color features (HOG+LUV) as low-level features. One representative ExtFCF implementation achieves the best result compared with the current best traditional pedestrian detection methods on the Caltech dataset. Mingyu You, Yubin Zhang, Chunhua Shen, Xinyu Zhang 0015 |
IEEE Trans. Intell. Transp. Syst. | 4 |