Xuehui Yu

dblp:243/8603 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Seeing Is Believing: Grounding Long-Video Understanding in Spatio-Temporal Visual Evidence
abstract
Although Vision Language Models (VLMs) have excelled at image and video understanding, applying them to hour-long videos is held back by two interrelated challenges: exorbitant computational expense and a qualitative breakdown in long-term temporal reasoning. Thus, models tend to generate answers based on speculation instead of solid visual facts, causing both factually incorrect and plausible hallucinations. This problem is compounded by current benchmarks that, by only emphasizing final answers, lack an effective mechanism to check whether reasoning is substantiated by specific visual evidence. This makes it hard to differentiate between true understanding and pretend comprehension, inhibiting targeted model refinement. To address these interrelated challenges of model fragility and evaluation weakness, we adopt a twofold strategy. First, we present EV²-Bench, a large-scale benchmark that breaks new ground by an evaluation paradigm built upon spatio-temporal visual evidence, forcing models to justify answers with checkable hints. Second, we put forward DynamicSelect, an adaptive token compression system that efficiently condenses salient information by a dynamic semantic selector and a hierarchical compression strategy. Comprehensive experiments demonstrate that DynamicSelect significantly outperforms the baselines on EV²-Bench as well as other public benchmarks. Our study offers not only a more effective approach to long-video understanding but also a more stringent evaluation paradigm, indicating the way toward more robust models.
Zhaoyang Wei, Guohua Gao, Yanchao Hao, Wenchao Ding 0007, Shizhu He, Xuehui Yu
AAAI9
2026 SAPNet++: Evolving Point-Prompted Instance Segmentation With Semantic and Spatial Awareness
abstract
Single-point annotation is increasingly prominent in visual tasks for labeling cost reduction. However, it challenges tasks requiring high precision, such as the point-prompted instance segmentation (PPIS) task, which aims to estimate precise masks using single-point prompts to train a segmentation network. Due to the constraints of point annotations, granularity ambiguity and boundary uncertainty arise i.e., the difficulty distinguishing between different levels of detail (e.g., whole object vs. parts) and the challenge of precisely delineating object boundaries. Previous works have usually inherited the paradigm of mask generation along with proposal selection to achieve PPIS. However, proposal selection relies solely on category information, failing to resolve the ambiguity of different granularity. Furthermore, mask generators offer only finite discrete solutions that often deviate from actual masks, particularly at boundaries. To address these issues, we propose the Semantic-Aware Point-Prompted Instance Segmentation Network (SAPNet). It integrates Point Distance Guidance and Box Mining Strategy to tackle group and local issues caused by the point's granularity ambiguity. Additionally, we incorporate completeness scores within proposals to add spatial granularity awareness, enhancing multiple instance learning (MIL) in proposal selection termed S-MIL. The Multi-level Affinity Refinement conveys pixel and semantic clues, narrowing boundary uncertainty during mask refinement. These modules culminate in SAPNet++, mitigating point prompt's granularity ambiguity and boundary uncertainty and significantly improving segmentation performance. Extensive experiments on four challenging datasets validate the effectiveness of our methods, highlighting the potential to advance PPIS.
Zhaoyang Wei, Xumeng Han, Xuehui Yu, Xue Yang 0005, Guorong Li, Zhenjun Han, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Boosting Segment Anything Model Towards Open-Vocabulary Learning
abstract
The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in various domains, its primary limitation lies in the inability to grasp object semantics. In this paper, we present Sambor to seamlessly integrate SAM with the open-vocabulary object detector in an end-to-end framework. While retaining all the remarkable capabilities inherent to SAM, we boost it to detect arbitrary objects from human inputs like category names or reference expressions. Building upon the SAM image encoder, we introduce a novel SideFormer module designed to acquire SAM features adept at perceiving objects and inject comprehensive semantic information for recognition. In addition, we devise an Open-set RPN that leverages SAM proposals to assist in finding potential objects. Consequently, Sambor enables the open-vocabulary detector to equally focus on generalizing both localization and classification sub-tasks. Our approach demonstrates superior zero-shot performance across benchmarks, including COCO and LVIS, proving highly competitive against previous state-of-the-art methods. We aspire for this work to serve as a meaningful endeavor in endowing SAM to recognize diverse object categories and advancing open-vocabulary learning with the support of vision foundation models.
Xumeng Han, Longhui Wei, Xuehui Yu, Zhiyang Dou, Kuiran Wang, Yingfei Sun, Zhenjun Han, Qi Tian 0001
AAAI3
2025 Segment Any-Quality Images with Generative Latent Space Enhancement
abstract
Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose Gle-SAM, which utilizes Generative Latent space Enhancement to boost robustness on low-quality images, thus enabling generalization across various image qualities. Specifically, we adapt the concept of latent diffusion to SAM-based segmentation frameworks and perform the generative diffusion process in the latent space of SAM to reconstruct high-quality representation, thereby improving segmentation. Additionally, we introduce two techniques to improve compatibility between the pre-trained diffusion model and the segmentation framework. Our method can be applied to pre-trained SAM and SAM2 with only minimal additional learnable parameters, allowing for efficient optimization. We also construct the LQSeg dataset with a greater diversity of degradation types and levels for training and evaluating the model. Extensive experiments demonstrate that GleSAM significantly improves segmentation robustness on complex degradations while maintaining generalization to clear images. Furthermore, GleSAM also performs well on unseen degradations, underscoring the versatility of our approach and dataset.
Guangqian Guo, Xuehui Yu, Yaoxing Wang, Shan Gao 0003
CVPR3
2025 SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation
abstract
The Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes them for versatile segmentation. Specifically, given a set of classes (in texts) and a set of SAM patches, the Type-I prompt judges whether a SAM patch aligns with a text label, and the Type-II prompt judges whether two SAM patches with the same text label also belong to the same instance. To decrease the complexity in dealing with a large number of semantic classes and patches, we establish a unified framework that calculates the affinity between (semantic and instance) queries and SAM patches, and then merges patches with high affinity to the query. Experiments show that SAM-CP achieves semantic, instance, and panoptic segmentation in both open and closed domains. In particular, it achieves state-of-the-art performance in open-vocabulary segmentation. Our research offers a novel and generalized methodology for equipping vision foundation models like SAM with multi-grained semantic perception abilities. Codes are released on https://github.com/ucas-vg/SAM-CP.
Pengfei Chen 0004, Lingxi Xie, Xinyue Huo, Xuehui Yu, Xiaopeng Zhang 0008, Yingfei Sun, Zhenjun Han, Qi Tian 0001
ICLR4
2025 Causal discovery based on hierarchical reinforcement learning
Jingchi Jiang, Rujia Shen, Chao Zhao 0002, Yi Guan, Xuehui Yu, Xuelian Fu
Expert Syst. Appl.5
2025 P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Kuiran Wang, Guorong Li, Lingxi Xie, Zhenjun Han, Jianbin Jiao
Int. J. Comput. Vis.2
2025 ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han
Pattern Recognit.2
2024 Semantic-aware SAM for Point-Prompted Instance Segmentation
abstract
Single-point annotation in visual tasks, with the goal of minimizing labelling costs, is becoming increasingly prominent in research. Recently, visual foundation models, such as Segment Anything (SAM), have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation performance. However, SAM's class-agnostic output and high confidence in local segmentation introduce semantic ambiguity, posing a challenge for precise category-specific segmentation. In this paper, we introduce a cost-effective category-specific segmenter using SAM. To tackle this challenge, we have devised a Semantic-Aware Instance Segmentation Network (SAPNet) that integrates Multiple Instance Learning (MIL) with matching capability and SAM with point prompts. SAPNet strategically selects the most representative mask proposals generated by SAM to supervise segmentation, with a specific focus on object category information. Moreover, we introduce the Point Distance Guidance and Box Mining Strategy to mitigate inherent challenges: group and local issues in weakly supervised segmentation. These strategies serve to further enhance the overall segmentation performance. The experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed SAPNet, emphasizing its semantic matching capabilities and its potential to advance point-prompted instance segmentation. The code is available at https://github.com/zhaoyangwei123/SAPNet.
Zhaoyang Wei, Pengfei Chen 0004, Xuehui Yu, Guorong Li, Jianbin Jiao, Zhenjun Han
CVPR3
2024 P2Seg: Pointly-supervised Segmentation via Mutual Distillation
abstract
Point-level Supervised Instance Segmentation (PSIS) aims to enhance the applicability and scalability of instance segmentation by utilizing low-cost yet instance-informative annotations. Existing PSIS methods usually rely on positional information to distinguish objects, but predicting precise boundaries remains challenging due to the lack of contour annotations. Nevertheless, weakly supervised semantic segmentation methods are proficient in utilizing intra-class feature consistency to capture the boundary contours of the same semantic regions. In this paper, we design a Mutual Distillation Module (MDM) to leverage the complementary strengths of both instance position and semantic information and achieve accurate instance-level object perception. The MDM consists of Semantic to Instance (S2I) and Istance to Semantic (I2S). S2I is guided by the precise boundaries of semantic regions to learn the association between annotated points and instance contours. I2S leverages discriminative relationships between instances to facilitate the differentiation of various objects within the semantic map. Extensive experiments substantiate the efficacy of MDM in fostering the synergy between instance and semantic information, consequently improving the quality of instance-level object representations. Our method achieves 55.7 mAP50 and 17.6 mAP on the PASCAL VOC and MS COCO datasets, significantly outperforming recent PSIS methods and several box-supervised instance segmentation competitors.
Xuehui Yu, Xumeng Han, Wenwen Yu, Zhixun Huang, Jianbin Jiao, Zhenjun Han
ICLR2
2024 Skill-aware Mutual Information Optimisation for Zero-shot Generalisation in Reinforcement Learning
abstract
Meta-Reinforcement Learning (Meta-RL) agents can struggle to operate across tasks with varying environmental features that require different optimal skills (i.e., different modes of behaviour). Using context encoders based on contrastive learning to enhance the generalisability of Meta-RL agents is now widely studied but faces challenges such as the requirement for a large sample size, also referred to as the $\log$-$K$ curse. To improve RL generalisation to different tasks, we first introduce Skill-aware Mutual Information (SaMI), an optimisation objective that aids in distinguishing context embeddings according to skills, thereby equipping RL agents with the ability to identify and execute different skills across tasks. We then propose Skill-aware Noise Contrastive Estimation (SaNCE), a $K$-sample estimator used to optimise the SaMI objective. We provide a framework for equipping an RL agent with SaNCE in practice and conduct experimental validation on modified MuJoCo and Panda-gym benchmarks. We empirically find that RL agents that learn by maximising SaMI achieve substantially improved zero-shot generalisation to unseen tasks. Additionally, the context encoder trained with SaNCE demonstrates greater robustness to a reduction in the number of available samples, thus possessing the potential to overcome the $\log$-$K$ curse.
Xuehui Yu, Mhairi Dunion, Stefano V. Albrecht
NeurIPS1
2024 An interactive food recommendation system using reinforcement learning
Liangliang Liu 0002, Yi Guan, Rujia Shen, Guowei Zheng, Xuelian Fu, Xuehui Yu, Jingchi Jiang
Expert Syst. Appl.7
2024 CPR++: Object Localization via Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance due to the inconsistency of annotated points. Existing POL heavily rely on strict annotation rules, which are difficult to define and apply, to handle the problem. In this study, we propose coarse point refinement (CPR), which to our best knowledge is the first attempt to alleviate semantic variance from an algorithmic perspective. CPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point. Furthermore, We design a sampling region estimation module to dynamically compute a sampling region for each object and use a cascaded structure to achieve end-to-end optimization. We further integrate a variance regularization into the structure to concentrate the predicted scores, yielding CPR++. We observe that CPR++ can obtain scale information and further reduce the semantic variance in a global region, thus guaranteeing high-performance object localization. Extensive experiments on four challenging datasets validate the effectiveness of both CPR and CPR++. We hope our work can inspire more research on designing algorithms rather than annotation rules to address the semantic variance problem in POL.
Xuehui Yu, Pengfei Chen 0004, Kuiran Wang, Xumeng Han, Guorong Li, Zhenjun Han, Qixiang Ye, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Save the Tiny, Save the All: Hierarchical Activation Network for Tiny Object Detection
abstract
Tiny object detection (TOD) remains a challenging problem due to the extremely small size and weak feature presentations of tiny objects. Many effective methods have improved the detection of small objects below$32\times 32$pixels to some extent, but the performance is still poor for the tiny objects below$16\times 16$pixels. In this paper, we find that the aliasing between the features and object scales, namely feature-scale-aliasing, leads to the misalignment between feature subspaces and detection subspaces, and thus results in the interference of features, especially for tiny objects. To alleviate this, we propose a Hierarchical Activation (HA) method to obtain scale-specific feature subspaces by activating object features at different scales hierarchically. To this end, we design a Scale-Guided Feature Activation (SGFA) to decompose the original object-aliasing feature spaces into a group of scale-specific feature subspaces by scale-guided activation maps. Then, Scale-Specific Feature re-Coupling (SSFC) is used to enhance the feature subspaces by adaptively aggregating the feature subspaces from different groups. In addition, we propose to complement the scale-specific detailed information by a designed Detailed Information Compensation (DIC) method. Implementing HA, a multi-scale keypoint-based detector is constructed to improve the tiny object detection, referred to as Hierarchical Activation Network (HANet). Extensive experiments are carried out on three tiny object detection datasets, e.g., TinyPerson, AI-TOD, and TinyCOCO. Our HANet achieves 58.45%$AP_{50}^{all}$, 22.1%$AP$, and 15.76%$AP$on TinyPerson, AI-TOD, and TinyCOCO, respectively, showing a significant performance gain over the competitors.
Guangqian Guo, Pengfei Chen 0004, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao 0003
IEEE Trans. Circuits Syst. Video Technol.3
2023 Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes
abstract
Object detection via inaccurate bounding boxes supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly depends on category information, to select and refine a low-quality box. Those methods suffer from object drift, group prediction and part domination problems without exploring spatial information. In this paper, we heuristically propose a Spatial Self-Distillation based Object Detector (SSD-Det) to mine spatial information to refine the inaccurate box in a self-distillation fashion. SSD-Det utilizes a Spatial Position Self-Distillation (SPSD) module to exploit spatial information and an interactive structure to combine spatial information and category information, thus constructing a high-quality proposal bag. To further improve the selection procedure, a Spatial Identity Self-Distillation (SISD) module is introduced in SSD-Det to obtain spatial confidence to help select the best proposals. Experiments on MS-COCO and VOC datasets with noisy box annotation verify our method’s effectiveness and achieve state-of-the-art performance. The code is available at https://github.com/ucas-vg/PointTinyBenchmark/tree/SSD-Det.
Pengfei Chen 0004, Xuehui Yu, Guorong Li, Zhenjun Han, Jianbin Jiao
ICCV3
2023 DECAF: An interpretable deep cascading framework for ICU mortality prediction
Jingchi Jiang, Xuehui Yu, Boran Wang, Linjiang Ma, Yi Guan
Artif. Intell. Medicine2
2023 ARLPE: A meta reinforcement learning framework for glucose regulation in type 1 diabetics
Xuehui Yu, Yi Guan, Lian Yan, Shulang Li, Xuelian Fu, Jingchi Jiang
Expert Syst. Appl.1
2023 Rethinking Sampling Strategies for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling.
Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han
IEEE Trans. Image Process.2
2023 Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking
abstract
Unmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV.
Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han
IEEE Trans. Multim.4
2022 Contextual Policy Transfer in Meta-Reinforcement Learning via Active Learning
Jingchi Jiang, Lian Yan, Xuehui Yu, Yi Guan
WISA3
2022 Unified Fine-Grained Biomedical Entity Recognition as a Combination of Boundary Detection and Sequence Generation
abstract
Biomedical Named Entity Recognition (BioNER) is a critical component of biomedical information extraction. NER is more challenging in the biomedical domain because of fine-grained entity types and more common nested and discontinuous entity forms. However, none of the BioNER datasets contains a large amount of all three entity forms, including flat, nested, and discontinuous. Not to mention that there is a unified BioNER model for dealing with the above three entity forms simultaneously. Methods in the public domain only focus on identifying text spans and ignore distinguishing fine-grained entity types. To address these issues, we propose a unified framework based on our own BioNER dataset CCNER, which innovatively models the BioNER task as a combination of boundary recognition and sequence generation. CCNER is a comprehensive and fine-grained BioNER dataset, where the proportion of discontinuous, nested, and flat entities in the dataset is 8.9%, 52.6%, and 38.5%, respectively. Meanwhile, it includes five fine-grained entity types. Our proposed framework includes two modules which are boundary detection and entity generation. In the boundary detection module, we propose a sample-based span representation method to determine fine-grained entity boundaries better. Finally, we conduct experiments on four datasets and achieve competitive results1.1Code is available at https://github.com/lx-hit/BioNER.
Yang Yang 0137, Mingchen Ye, Yi Guan, Xuehui Yu, Jingchi Jiang
BIBM5
2022 Causal Coupled Mechanisms: A Control Method with Cooperation and Competition for Complex System
abstract
Complex systems are ubiquitous in the real world and tend to have complicated and poorly understood dynamics. For their control issues, the challenge is to guarantee accuracy, robustness, and generalization in such bloated and troubled environments. Fortunately, a complex system can be divided into multiple modular structures that human cognition appears to exploit. Inspired by this cognition, a novel control method, Causal Coupled Mechanisms (CCMs), is proposed that explores the cooperation in division and competition in combination. Our method employs the theory of hierarchical reinforcement learning (HRL), in which 1) the high-level policy with competitive awareness divides the whole complex system into multiple functional mechanisms, and 2) the low-level policy finishes the control task of each mechanism. Specifically for cooperation, a cascade control module helps the series operation of CCMs, and a forward coupled reasoning module is used to recover the coupling information lost in the division process. On both synthetic systems and a real-world biological regulatory system, the CCM method achieves robust and state-of-the-art control results even with unpredictable random noise. Moreover, generalization results show that reusing prepared specialized CCMs helps to perform well in environments with different confounders and dynamics.
Xuehui Yu, Yi Guan, Xinmiao Yu, Jingchi Jiang
BIBM1
2022 Object Localization under Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily reply on accurate keypoint annotations which are difficult to define. In this study, we propose a POL method using coarse point annotations, relaxing the supervision signals from accurate key points to freely spotted points. To this end, we propose a coarse point refinement (CPR) approach, which to our best knowledge is the first attempt to alleviate semantic variance from the perspective of algorithm. CPR constructs point bags, selects semantic-correlated points, and produces semantic center points through multiple instance learning (MIL). In this way, CPR defines a weakly supervised evolution procedure, which ensures training high-performance object localizer under coarse point supervision. Experimental results on COCO, DOTA and our proposed SeaPerson dataset validate the effectiveness of the CPR approach. The dataset and code will be available at https://github.com/ucas-vg/PointTinyBenchmark/
Xuehui Yu, Pengfei Chen 0004, Najmul Hassan, Guorong Li, Junchi Yan, Humphrey Shi, Qixiang Ye, Zhenjun Han
CVPR1
2022 Point-to-Box Network for Accurate Object Detection via Single Point Supervision
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Najmul Hassan, Kai Wang 0058, Jiachen Li 0003, Jian Zhao 0006, Humphrey Shi, Zhenjun Han, Qixiang Ye
ECCV (9)2
2022 Research and Implementation of the Early-late Gate Bit Synchronization Algorithm in the Non-coherent Spread Spectrum System
abstract
The data clock and the pseudocode clock of the non-coherent spread spectrum system are not homologous, and the phase relationship between them is not fixed, which may have the problem of random data flipping during coherent integration and affect the bit decision. Thus, the bit synchronization loop becomes a key technology for bit timing recovery. In this paper, an improved bit synchronization algorithm based on the Early-Late Gate is proposed. A bit-synchronous loop is constructed to identify the phase using the Early-Late integration results of two adjacent data codes, which can effectively improve the communication performance under random data flipping conditions. The bit clock can be dynamically adjusted according to the loop tracking result so that the integration start moment is aligned with the data bit start edge. This enables full bit coherent integration to improve the signal-to-noise ratio and thus reduce the demodulation bit error rate. The effectiveness and correctness of this method is validated by MATLAB simulations and board tests. For highly dynamic signals, the loop can be bit-synchronized within 30ms. In addition, for a signal with an Eb/N0 of 10.97dB, the bit error rate is 0 (better than 10-6) when running at a rate of 8ksps for 125s, which is in general agreement with the theoretical value.
Guangyu Tian, Wenbing Cai, Hancheng Zhao, Xuehui Yu
ISNCC6
2021 SM+: Refined Scale Match for Tiny Person Detection
abstract
Detecting tiny objects (e.g., less than 20 × 20 pixels) in large-scale images is an important yet open problem. Modern CNN-based detectors are challenged by the scale mismatch between the dataset for network pre-training and the target dataset for detector training. In this paper, we investigate the scale alignment between pre-training and target datasets, and propose a new refined Scale Match method (termed SM+) for tiny person detection. SM+ improves the scale match from image level to instance level, and effectively promotes the similarity between pre-training and target dataset. Moreover, considering SM+ possibly destroys the image structure, a new probabilistic structure inpainting (PSI) method is proposed for the background processing. Experiments conducted across various detectors show that SM+ noticeably improves the performance on TinyPerson, and outperforms the state-of-the-art detectors with a significant margin.
Xuehui Yu, Xiaoke Peng, Yuqi Gong, Zhenjun Han
ICASSP2
2021 BFConv: Improving Convolutional Neural Networks with Butterfly Convolution
Dengjie Yang, Xuehui Yu, Yi Sun 0004, Fuzhen Zhuang, Qing He 0003, Shiwei Ye
ICONIP (4)2
2021 Effective Fusion Factor in FPN for Tiny Object Detection
abstract
FPN-based detectors have made significant progress in general object detection, e.g., MS COCO and PASCAL VOC. However, these detectors fail in certain application scenarios, e.g., tiny object detection. In this paper, we argue that the top-down connections between adjacent layers in FPN bring two-side influences for tiny object detection, not only positive. We propose a novel concept, fusion factor, to control information that deep layers deliver to shallow layers, for adapting FPN to tiny object detection. After series of experiments and analysis, we explore how to estimate an effective value of fusion factor for a particular dataset by a statistical method. The estimation is dependent on the number of objects distributed in each layer. Comprehensive experiments are conducted on tiny object detection datasets, e.g., TinyPerson and Tiny CityPersons. Our results show that when configuring FPN with a proper fusion factor, the network is able to achieve significant performance gains over the baseline on tiny object detection datasets. Codes and models will be released.
Yuqi Gong, Xuehui Yu, Yao Ding 0006, Xiaoke Peng, Jian Zhao 0006, Zhenjun Han
WACV2
2020 Scale Match for Tiny Person Detection
abstract
Visual object detection has achieved unprecedented advance with the rise of deep convolutional neural networks. However, detecting tiny objects (for example tiny persons less than 20 pixels) in large-scale images remains not well investigated. The extremely small objects raise a grand challenge about feature representation while the massive and complex backgrounds aggregate the risk of false alarms. In this paper, we introduce a new benchmark, referred to as TinyPerson, opening up a promising direction for tiny object detection in a long distance and with massive backgrounds. We experimentally find that the scale mismatch between the dataset for network pre-training and the dataset for detector learning could deteriorate the feature representation and the detectors. Accordingly, we propose a simple yet effective Scale Match approach to align the object scales between the two datasets for favorable tiny-object representation. Experiments show the significant performance gain of our proposed approach over state-of-the-art detectors, and the challenging aspects of TinyPerson related to real-world scenarios. The TinyPerson benchmark and the code for our approach will be publicly available1.
Xuehui Yu, Yuqi Gong, Qixiang Ye, Zhenjun Han
WACV1
2019 High performance person re-identification via a boosting ranking ensemble
Zhaoju Li, Zhenjun Han, Junliang Xing, Qixiang Ye, Xuehui Yu, Jianbin Jiao
Pattern Recognit.5