Shiyu Xuan

dblp:252/0070 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-9950-6025ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spike Camera Optical Flow Estimation Based on Continuous Spike Streams
abstract
Spike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting binary spike streams. Optical flow estimation aims to estimate pixel-level correspondences between different moments, describing motion information along time, which is a key task of spike camera. High-quality optical flow is important since motion information is a foundation for analyzing spikes. However, extracting stable light-intensity information from spikes is difficult due to the randomness of binary spikes. Besides, the continuity of spikes can offer contextual information for optical flow. In this paper, we propose a network Spike2Flow++ to estimate optical flow for spike camera. In Spike2Flow++, we propose a differential of spike firing time (DSFT) to represent information in binary spikes. Moreover, we propose a dual DSFT representation and a dual correlation construction to extract stable light-intensity information for reliable correlations. To use the continuity of spikes as motion contextual information, we propose a joint correlation decoding (JCD) that jointly estimates a series of flow fields. To adaptively fuse different motions in JCD, we propose a global motion bank aggregation to construct an information bank for all motions and adaptively extract contexts from the bank for each iteration during recurrent decoding of each motion. To train and evaluate our network, we construct a real scene with spikes and flow++ (RSSF++) based on real-world scenes. Experiments demonstrate that our Spike2Flow++ achieves state-of-the-art performance on RSSF++, photo-realistic high-speed motion (PHM), and real-captured data.
Rui Zhao 0010, Ruiqin Xiong, Dongkai Wang, Shiyu Xuan, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Robust object detection in adverse weather with feature decorrelation via independence learning
Shiyu Xuan, Zechao Li
Pattern Recognit.2
2026 Object Detection under Low-light Conditions via Degradation Learning Driven by Foundation Models
abstract
The previous approach tackled object detection challenges in low-light scenes by training on images captured in such conditions. However, the limited availability of annotated data in such environments has impeded the development of specialized detectors. As a result, current detectors continue to struggle with images degraded by low illumination. To overcome these limitations, a novel framework for low-light image generation is proposed. These generated images alter only the illumination conditions while preserving the original content, thereby enabling fine-tuning of object detectors. The framework incorporates a trainable degradation module and integrates two frozen Foundation Models: CLIP and a low-light image enhancement network. Under the guidance of CLIP, the degradation module transfers low-light characteristics from textual descriptions to well-lit images by aligning embeddings in a shared vision-language space. This process effectively simulates realistic low-light conditions. With the assistance of the low-light enhancement network, the generated low-light images are recovered. By enforcing similarity between these enhanced versions and the original well-lit counterparts in both the spatial and frequency domains, the content of the generated low-light images is effectively preserved. Extensive experiments on real-world low-light datasets show that fine-tuning detectors using the generated low-light images significantly improves detection performance, even when only well-lit images are available. This validates the effectiveness of the proposed method.
Shiyu Xuan, Zechao Li
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Generalizable Object Keypoint Localization from Generative Priors
abstract
Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performance and generalization capability. Instead of relying on large scale training data, this work tackles this challenge by exploiting the rich priors from large generative models. We propose a data-efficient generalizable localization method named GenLoc. GenLoc extracts the generative priors from a pre-trained image generation model by calculating the correlation map between image latent feature and condition embedding. Those priors are hence optimized with our proposed heatmap expectation loss to perform object keypoint localization. Benefited by the rich knowledge of generative priors in understanding of object semantics and structures, GenLoc achieves superior performance on various object keypoint localization benchmarks. It shows more substantial performance enhancements in cross-domain, few-shot and zero-shot evaluation settings, e.g., getting 20%+ AP enhancement over CLAMP [43] in various zero-shot settings.
Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan, Hao Chen 0061, Shiliang Zhang
CVPR4
2025 Incremental Model Enhancement via Memory-based Contrastive Learning
Shiyu Xuan, Ming Yang 0007, Shiliang Zhang
Int. J. Comput. Vis.1
2024 Decoupled Optimisation for Long-Tailed Visual Recognition
abstract
When training on a long-tailed dataset, conventional learning algorithms tend to exhibit a bias towards classes with a larger sample size. Our investigation has revealed that this biased learning tendency originates from the model parameters, which are trained to disproportionately contribute to the classes characterised by their sample size (e.g., many, medium, and few classes). To balance the overall parameter contribution across all classes, we investigate the importance of each model parameter to the learning of different class groups, and propose a multistage parameter Decouple and Optimisation (DO) framework that decouples parameters into different groups with each group learning a specific portion of classes. To optimise the parameter learning, we apply different training objectives with a collaborative optimisation step to learn complementary information about each class group. Extensive experiments on long-tailed datasets, including CIFAR100, Places-LT, ImageNet-LT, and iNaturaList 2018, show that our framework achieves competitive performance compared to the state-of-the-art.
Cong Cong 0001, Shiyu Xuan, Sidong Liu, Shiliang Zhang, Maurice Pagnucco, Yang Song 0001
AAAI2
2024 Decoupled Contrastive Learning for Long-Tailed Recognition
abstract
Supervised Contrastive Loss (SCL) is popular in visual representation learning. Given an anchor image, SCL pulls two types of positive samples, i.e., its augmentation and other images from the same class together, while pushes negative images apart to optimize the learned embedding. In the scenario of long-tailed recognition, where the number of samples in each class is imbalanced, treating two types of positive samples equally leads to the biased optimization for intra-category distance. In addition, similarity relationship among negative samples, that are ignored by SCL, also presents meaningful semantic cues. To improve the performance on long-tailed recognition, this paper addresses those two issues of SCL by decoupling the training objective. Specifically, it decouples two types of positives in SCL and optimizes their relations toward different objectives to alleviate the influence of the imbalanced dataset. We further propose a patch-based self distillation to transfer knowledge from head to tail classes to relieve the under-representation of tail classes. It uses patch-based features to mine shared visual patterns among different instances and leverages a self distillation procedure to transfer such knowledge. Experiments on different long-tailed classification benchmarks demonstrate the superiority of our method. For instance, it achieves the 57.7% top-1 accuracy on the ImageNet-LT dataset. Combined with the ensemble-based method, the performance can be further boosted to 59.7%, which substantially outperforms many recent works. Our code will be released.
Shiyu Xuan, Shiliang Zhang
AAAI1
2024 LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
abstract
The capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more gen-eral model, this work studies keypoint localization from a different perspective by reasoning locations based on key-piont clues in text descriptions. We propose LocLLM, the first Large-Language Model (LLM) based keypoint local-ization model that takes images and text instructions as in-puts and outputs the desired keypoint coordinates. LocLLM leverages the strong reasoning capability of LLM and clues of keypoint type, location, and relationship in textual de-scriptions for keypoint localization. To effectively tune Lo-cLLM, we construct localization-based instruction conver-sations to connect keypoint description with corresponding coordinates in input image, and fine-tune the whole model in a parameter-efficient training pipeline. LocLLM shows remarkable performance on standard 2D/3D keypoint lo-calization benchmarks. Moreover, incorporating language clues into the localization makes LocLLM show superior flexibility and generalizable capability in cross dataset key-point localization, and even detecting novel type of key-points unseen during training††Project page: https://github.com/kennethwdk/LocLLM.
Dongkai Wang, Shiyu Xuan, Shiliang Zhang
CVPR2
2024 Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
abstract
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless, their performance in fine-grained image understanding tasks is still limited. To address this issue, this paper proposes a new framework to enhance the fine-grained image understanding abilities of MLLMs. Specifically, we present a new method for constructing the instruction tuning dataset at a low cost by leveraging annotations in existing datasets. A self-consistent bootstrapping method is also introduced to extend existing dense object annotations into high-quality referring-expression-bounding-box pairs. These methods enable the generation of high-quality instruction data which includes a wide range of fundamental abilities essential for fine-grained image perception. Moreover, we argue that the visual encoder should be tuned during instruction tuning to mitigate the gap between full image perception and fine-grained image perception. Experimental results demonstrate the superior performance of our method. For instance, our model exhibits a 5.2% accuracy improvement over Qwen-VL on GQA and surpasses the accuracy of Kosmos-2 by 24.7% on RefCOCO_val. We have also attained the top rank on the leaderboard of MM-Bench. This promising performance is achieved by training on only publicly available data, making it easily reproducible. The models, datasets, and codes are publicly available at https://github.com/SY-Xuan/Pink.
Shiyu Xuan, Qingpei Guo, Ming Yang 0007, Shiliang Zhang
CVPR1
2024 Intra-Inter Domain Similarity for Unsupervised Person Re-Identification
abstract
Most of unsupervised person Re-Identification (ReID) works produce pseudo-labels by measuring the feature similarity without considering the domain discrepancy among cameras, leading to degraded accuracy in pseudo-label computation across cameras. This paper targets to address this challenge by decomposing the similarity computation into two stages, i.e., the intra-domain and inter-domain computations, respectively. The intra-domain similarity directly leverages CNN features learned within each camera, hence generates pseudo-labels on different cameras to train the ReID model in a multi-branch network. The inter-domain similarity considers the classification scores of each sample on different cameras as a new feature vector. This new feature effectively alleviates the domain discrepancy among cameras and generates more reliable pseudo-labels. We further propose the Instance and Camera Style Normalization (ICSN) to enhance the robustness to domain discrepancy. ICSN alleviates the intra-camera variations by adaptively learning a combination of instance and batch normalization. ICSN also boosts the robustness to inter-camera variations through TNorm which converts the original style of features into target styles. The proposed method achieves competitive performance on multiple datasets under fully unsupervised, intra-camera supervised and domain generalization settings, e.g., it achieves rank-1 accuracy of 64.4% on the MSMT17 dataset, outperforming the recent unsupervised methods by 20+%.
Shiyu Xuan, Shiliang Zhang
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Adapting Vision-Language Models via Learning to Inject Knowledge
abstract
Pre-trained vision-language models (VLM) such as CLIP, have demonstrated impressive zero-shot performance on various vision tasks. Trained on millions or even billions of image-text pairs, the text encoder has memorized a substantial amount of appearance knowledge. Such knowledge in VLM is usually leveraged by learning specific task-oriented prompts, which may limit its performance in unseen tasks. This paper proposes a new knowledge injection framework to pursue a generalizable adaption of VLM to downstream vision tasks. Instead of learning task-specific prompts, we extract task-agnostic knowledge features, and insert them into features of input images or texts. The fused features hence gain better discriminative capability and robustness to intra-category variances. Those knowledge features are generated by inputting learnable prompt sentences into text encoder of VLM, and extracting its multi-layer features. A new knowledge injection module (KIM) is proposed to refine text features or visual features using knowledge features. This knowledge injection framework enables both modalities to benefit from the rich knowledge memorized in the text encoder. Experiments show that our method outperforms recently proposed methods under few-shot learning, base-to-new classes generalization, cross-dataset transfer, and domain generalization settings. For instance, it outperforms CoOp by 4.5% under the few-shot learning setting, and CoCoOp by 4.4% under the base-to-new classes generalization setting. Our code will be released.
Shiyu Xuan, Ming Yang 0007, Shiliang Zhang
IEEE Trans. Image Process.1
2022 SatSOT: A Benchmark Dataset for Satellite Video Single Object Tracking
abstract
By imaging a specific area continuously, satellite video shows excellent capability in various applications such as surveillance and traffic management. Although object tracking has made significant progress in recent years, development in satellite object tracking is limited by the lack of open-source satellite datasets. It is thus essential to establish a satellite video object-tracking benchmark to fill the gap and advance the research. In this work, we present SatSOT, the first densely annotated satellite video single object-tracking benchmark dataset. SatSOT consists of 105 sequences with 27664 frames, 11 attributes, and four categories of typical moving targets in satellite videos: car, plane, ship, and train. Based on the proposed dataset and the significant challenges in satellite video object tracking, such as small targets, background interference, and severe occlusion, detailed evaluation and analysis are performed on 15 among the best and most representative tracking algorithms, which provides a basis for further research on satellite video object tracking.
Manqi Zhao, Shengyang Li, Shiyu Xuan, Longxuan Kou, Shuai Gong
IEEE Trans. Geosci. Remote. Sens.3
2021 Intra-Inter Camera Similarity for Unsupervised Person Re-Identification
abstract
Most of unsupervised person Re-Identification (Re-ID) works produce pseudo-labels by measuring the feature similarity without considering the distribution discrepancy among cameras, leading to degraded accuracy in label computation across cameras. This paper targets to address this challenge by studying a novel intra-inter camera similarity for pseudo-label generation. We decompose the sample similarity computation into two stage, i.e., the intra-camera and inter-camera computations, respectively. The intra-camera computation directly leverages the CNN features for similarity computation within each camera. Pseudo-labels generated on different cameras train the re-id model in a multi-branch network. The second stage considers the classification scores of each sample on different cameras as a new feature vector. This new feature effectively alleviates the distribution discrepancy among cameras and generates more reliable pseudo-labels. We hence train our re-id model in two stages with intra-camera and inter-camera pseudo-labels, respectively. This simple intra-inter camera similarity produces surprisingly good performance on multiple datasets, e.g., achieves rank-1 accuracy of 89.5% on the Market1501 dataset, outperforming the recent unsupervised works by 9+%, and is comparable with the latest transfer learning works that leverage extra annotations.
Shiyu Xuan, Shiliang Zhang
CVPR1
2021 Rotation adaptive correlation filter for moving object tracking in satellite videos
Shiyu Xuan, Shengyang Li, Zifei Zhao, Wanfeng Zhang, Hong Tan, Gui-Song Xia, Yanfeng Gu
Neurocomputing1
2021 Siamese networks with distractor-reduction method for long-term visual object tracking
abstract
Many trackers which divide the tracking process into two stages have recently been proposed to solve the problem of long-term tracking. Their outstanding performance makes them become one of the mainstream algorithms of long-term tracking. To further improve the performance of two-stage tracking algorithms, some improvements are proposed in this paper. (a) A hard negative mining method is proposed. It can optimize the training process of the verification network and bridge the gap between the two sub-networks. (b) The architecture of the verification network is designed as a Siamese structure; therefore, the semantic ambiguity in classification can be alleviated. Extensive experiments performed on benchmarks demonstrate that the proposed approach significantly outperforms the state-of-the-art methods, yielding 7% relative gain in the VOT2018-LT dataset and 14.2% relative gain in the OxUvA dataset.
Shiyu Xuan, Shengyang Li, Zifei Zhao, Longxuan Kou, Gui-Song Xia
Pattern Recognit.1
2020 Object Tracking in Satellite Videos by Improved Correlation Filters With Motion Estimations
abstract
As a new method of Earth observation, video satellite is capable of monitoring specific events on the Earth's surface continuously by providing high-temporal resolution remote sensing images. The video observations enable a variety of new satellite applications such as object tracking and road traffic monitoring. In this article, we address the problem of fast object tracking in satellite videos, by developing a novel tracking algorithm based on correlation filters embedded with motion estimations. Based on the kernelized correlation filter (KCF), the proposed algorithm provides the following improvements: 1) proposing a novel motion estimation (ME) algorithm by combining the Kalman filter and motion trajectory averaging and mitigating the boundary effects of KCF by using this ME algorithm and 2) solving the problem of tracking failure when a moving object is partially or completely occluded. The experimental results demonstrate that our algorithm can track the moving object in satellite videos with 95% accuracy.
Shiyu Xuan, Shengyang Li, Mingfei Han 0002, Xue Wan, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.1
2019 The Multi-task Fully Convolutional Siamese Network with Correlation Filter Layer for Real-Time Visual Tracking
Shiyu Xuan, Shengyang Li, Zifei Zhao, Mingfei Han 0002
PRCV (1)1