EDBT 2026 Demo / reviewers in the wild / expert
Xiaohan Yu 0001
dblp:25/613-1
· DBLP profile ↗
66ranked-venue papers
11as first author
61since 2021 · last 2026
0000-0001-6186-0520ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 8 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SparseSurf: Sparse-View 3D Gaussian Splatting for Surface ReconstructionabstractRecent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this challenge by employing flattened Gaussian primitives to better fit surface geometry, combined with depth regularization to alleviate geometric ambiguities under limited viewpoints. Nevertheless, the increased anisotropy inherent in flattened Gaussians exacerbates overfitting in sparse-view scenarios, hindering accurate surface fitting and degrading novel view synthesis performance. In this paper, we propose SparseSurf, a method that reconstructs more accurate and detailed surfaces while preserving high-quality novel view rendering. Our key insight is to introduce Stereo Geometry-Texture Alignment, which bridges rendering quality and geometry estimation, thereby jointly enhancing both surface reconstruction and view synthesis. In addition, we present a Pseudo-Feature Enhanced Geometry Consistency that enforces multi-view geometric consistency by incorporating both training and unseen views, effectively mitigating overfitting caused by sparse supervision. Extensive experiments on the DTU, BlendedMVS, and Mip-NeRF360 datasets demonstrate that our method achieves the state-of-the-art performance. Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Haonan Luo 0002, Xiao Bai 0001 |
AAAI | 4 |
| 2026 | CurvLoc: Surface Curvature Prompted Gaussian Splatting for Visual Localization
Jiahe Li 0007, Botao Jiang, Zihang Wang 0002, Xiaohan Yu 0001, Xiao Bai 0001, Haonan Luo 0002 |
Int. J. Comput. Vis. | 6 |
| 2026 | Dual-label noise filtering for weakly supervised person search
Huadong Lin, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026 |
Neurocomputing | 3 |
| 2026 | DNGaussian++: Improving Sparse-View Gaussian Radiance Fields With Depth NormalizationabstractSynthesizing novel views from sparse views has achieved impressive advances with radiance fields, yet prevailing methods suffer from high consumption or insufficient refinement capability. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian Splatting, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the remarkable advancement of recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Although DNGaussian shows impressive performance, its patch-wise regularization obscures the inconsistency in cross-patch errors. Additionally, primitives can still be irreversibly trapped in local minima under sparse views, even if depth regularization is applied. In this paper, we propose an extended version, DNGaussian++. First, a Geometry Instance Regularizer is developed to enable depth regularization for continuous consistency by exploiting reliable instance-level depth cues. Leveraging the depth gradient guidance, we then propose a Depth-Guided Geometry Reorganization to address the aforementioned local minima problem with high representation efficiency. Extensive experiments show that DNGaussian++ exhibits state-of-the-art performance in multiple datasets and scenarios with high efficiency, and the broad applicability and effectiveness are verified on various backbones and tasks. Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Lin Gu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Ask and focus more: Question-prompt uncertainty allocation for dual-controllable video captioning
Shuqin Chen, Xingrui Yang 0003, Xiaohan Yu 0001, Xian Zhong |
Pattern Recognit. | 5 |
| 2026 | From temporal thumbnail to semantics: Debiasing multi-view action recognition
Zixian Zhu, Wenxuan Liu 0008, Xu Wang 0015, Bingyi Liu, Xiaohan Yu 0001, Xian Zhong |
Pattern Recognit. | 6 |
| 2026 | LGD: Leveraging generative descriptions for zero-shot referring image segmentationabstractZero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of aligning and matching semantics across visual and textual modalities without training. Previous works address this challenge by utilizing Vision-Language Models and mask proposal networks for region-text matching. However, this paradigm may lead to incorrect target localization due to the inherent ambiguity and diversity of free-form referring expressions. To alleviate this issue, we present LGD (Leveraging Generative Descriptions), a framework that utilizes the advanced language generation capabilities of Multi-Modal Large Language Models to enhance region-text matching performance in Vision-Language Models. Specifically, we first design two kinds of prompts, the attribute prompt and the surrounding prompt, to guide the Multi-Modal Large Language Models in generating descriptions related to the crucial attributes of the referent object and the details of surrounding objects, referred to as attribute description and surrounding description, respectively. Secondly, three visual-text matching scores are introduced to evaluate the similarity between instance-level visual features and textual features, which determines the mask most associated with the referring expression. The proposed method achieves new state-of-the-art performance on three public datasets RefCOCO, RefCOCO+ and RefCOCOg, with maximum improvements of 9.97 % in oIoU and 11.29 % in mIoU compared to previous methods Jiachen Li 0002, Qing Xie 0002, Renshu Gu, Jinyu Xu 0001, Yongjian Liu, Xiaohan Yu 0001 |
Pattern Recognit. | 6 |
| 2026 | SpectralKAN: Weighted Activation Distribution Kolmogorov-Arnold Network for Hyperspectral Image Change Detection
Xiaohan Yu 0001, Yongsheng Gao 0001, Jianjun Sha, Jian Wang 0138, Shiyong Yan, Yonggang Zhang 0001, Lianru Gao |
Pattern Recognit. | 2 |
| 2026 | ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 6 |
| 2026 | Refined generation-based framework for consistent and reliable visual question answering
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang |
Pattern Recognit. | 6 |
| 2026 | Exploring personalized federated learning from a distribution-based perspectiveabstractPersonalized federated learning (PFL) is a promising technique for tackling data heterogeneity in federated learning systems. Recently, Bayesian neural networks (BNNs) have been introduced into the PFL framework to enable uncertainty quantification and improve performance in data-scarce settings. Despite these advantages, existing BNN-based PFL methods face two key challenges in practical applications. First, in real-world scenarios, client heterogeneity often arises in the form of group-wise variation, which cannot be adequately captured by a single shared distribution as assumed in prior work. Second, existing methods rely on deterministic or stochastic approximation techniques for posterior inference, which lead to substantial computational and memory overhead, hindering their scalability and deployment. To address these limitations, we propose DBFed, a novel BNN-based PFL framework from a distribution-based perspective. DBFed introduces group-specific distributions to better model the structural heterogeneity commonly observed in federated settings. Moreover, DBFed employs a rank-1 parameterization technique to map uncertainty from the weight space to a low-dimensional subspace, significantly reducing the computational and memory overhead. Theoretically, we establish the effectiveness of the rank-1 parameterization approach. Empirically, extensive experiments on diverse datasets demonstrate that DBFed consistently outperforms alternative PFL baselines in a heterogeneous setting. Tianhao Yu, Kheng Cher Yeo, Sami Azam, Xiaohan Yu 0001, Hui Chen 0026, Xianxun Zhu |
Pattern Recognit. | 4 |
| 2026 | PBDD: A Prompt-Based Learning Approach for Few-Shot Social Media Depression DetectionabstractAutomated detection of depressive moods from social media holds great promise for early mental health intervention, yet existing multimodal approaches typically require large quantities of annotated data and extensive feature engineering, impeding their deployment in real‐world settings where labels are scarce. To address this challenge, we propose prompt‐based depression detection (PBDD), a novel prompt‐based few‐shot learning framework that leverages frozen pretrained language and vision models to identify depression indicators from paired text‐image posts without fine‐tuning. Our method begins with rigorous data cleaning and sampling to construct a high‐quality few‐shot dataset, then encodes text via a masked language model and images via a self‐supervised rotation‐prediction task to capture deep semantic cues. Multimodal representations are seamlessly fused into a unified prompt template containing a [MASK] token, enabling the pre‐trained model to infer depressive states by language completion. Extensive experiments on both large‐scale and 1 % few‐shot subsets demonstrate that PBDD consistently outperforms state‐of‐the‐art baselines, achieving significant gains in accuracy and Macro‐F1. These results validate the effectiveness and scalability of our framework for depression detection under severe label scarcity, offering a practical solution for real‐time mental health monitoring in social media environments. Rui Wang 0034, Heyang Feng, Erik Cambria, Kaize Shi, Xiaohan Yu 0001, Xuhui Fan 0001, Xianxun Zhu |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Improving Medical Visual Representation Learning With Pathological-Level Cross-Modal Alignment and Correlation ExplorationabstractLearning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential for transferring acquired knowledge to various downstream medical tasks. Previous works have predominantly focused on instance-wise or token-wise cross-modal alignment, often neglecting the importance of pathological-level consistency. This paper presents a novel framework PLACE that promotes the Pathological-Level Alignment and enriches the fine-grained details via Correlation Exploration without additional human annotations. Specifically, we propose a novel pathological-level cross-modal alignment (PCMA) approach to maximize the consistency of pathology observations from both images and reports. To facilitate this, a Visual Pathology Observation Extractor is introduced to extract visual pathological observation representations from localized tokens. The PCMA module operates independently of any external disease annotations, enhancing the generalizability and robustness of our methods. Furthermore, we design a proxy task that enforces the model to identify correlations among image patches, thereby enriching the fine-grained details crucial for various downstream tasks. Experimental results demonstrate that our proposed framework achieves new state-of-the-art performance on multiple downstream tasks, including classification, image-to-text retrieval, semantic segmentation, object detection and report generation. Jun Wang 0121, Lixing Zhu, Xiaohan Yu 0001, Abhir Bhalerao, Yulan He 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Visual Perturbation for Text-Based Person SearchabstractText-based person search aims at locating a person described by natural language in uncropped scene images. Recent works for TBPS mainly focus on aligning multi-granularity vision and language representations, neglecting a key discrepancy between training and inference where the former learns to unify vision and language features where the visual side covers all clues described by language, yet the latter matches image-text pairs where the images may capture only part of the described clues due to perturbations such as occlusions, background clutters and misaligned boundaries. To alleviate this issue, we present ViPer: a Visual Perturbation network that learns to match language descriptions with perturbed visual clues. On top of a CLIP-driven baseline, we design three visual perturbation modules: (1) Spatial ViPer that varies person proposals and produces visual features with misaligned boundaries, (2) Attentive ViPer that estimates visual attention on the fly and manipulates attentive visual tokens within a proposal to produce global features under visual perturbations, and (3) Fine-grained ViPer that learns to recover masked visual clues from detailed language descriptions to encourage matching language features with perturbed visual features at the fine granularity. This overall framework thus simulates real-world scenarios at the training stage to minimize the discrepancy and improve the generalization ability of the model. Experimental results demonstrate that the proposed method clearly surpasses previous TBPS methods on the PRW-TBPS and CUHK-SYSU-TBPS datasets. Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001 |
AAAI | 2 |
| 2025 | Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain semantic-related visual features, which may cause overfitting on seen classes with a limited number of training images. This paper proposes a novel visual and semantic prompt collaboration framework, which utilizes prompt tuning techniques for efficient feature adaptation. Specifically, we design a visual prompt to integrate the visual information for discriminative feature learning and a semantic prompt to integrate the semantic formation for visual-semantic alignment. To achieve effective prompt information integration, we further design a weak prompt fusion mechanism for the shallow layers and a strong prompt fusion mechanism for the deep layers in the network. Through the collaboration of visual and semantic prompts, we can obtain discriminative semantic-related features for generalized zero-shot image recognition. Extensive experiments demonstrate that our framework consistently achieves favorable performance in both conventional zero-shot learning and generalized zero-shot learning benchmarks compared to other state-of-the-art methods. Huajie Jiang, Zhengxian Li, Xiaohan Yu 0001, Yongli Hu, Jian Yang 0001, Yuankai Qi |
CVPR | 3 |
| 2025 | UGotMe: An Embodied System for Affective Human-Robot InteractionabstractEquipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affective human-robot interaction. However, enabling current vision-aware multimodal emotion recognition models for affective human-robot interaction in the real-world raises embodiment challenges: addressing the environmental noise issue and meeting real-time requirements. First, in multi-party conversation scenarios, the noises inherited in the visual observation of the robot, which may come from either 1) distracting objects in the scene or 2) inactive speakers appearing in the field of view of the robot, hinder the models from extracting emotional cues from vision inputs. Secondly, real-time response, a desired feature for an interactive system, is also challenging to achieve. To tackle both challenges, we introduce an affective human-robot interaction system called UGotMe designed specifically for multiparty conversations. Two denoising strategies are proposed and incorporated into the system to solve the first issue. Specifically, to filter out distracting objects in the scene, we propose extracting face images of the speakers from the raw images and introduce a customized active face extraction strategy to rule out inactive speakers. As for the second issue, we employ efficient data transmission from the robot to the local server to improve real-time response capability. We deploy UGotMe on a human robot named Ameca to validate its real-time inference capabilities in practical scenarios. Videos demonstrating real-world deployment are available at https://lipzh5.github.io/HumanoidVLE/ Pei-Zhen Li, Longbing Cao, Xiao-Ming Wu 0002, Xiaohan Yu 0001 |
ICRA | 4 |
| 2025 | Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained ModelsabstractContinual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in general continual learning, C-UFG remains an underexplored issue. This work establishes the first comprehensive C-UFG benchmark using massive soy leaf data. By analyzing recent pre-trained model (PTM) based continual learning methods on the proposed benchmark, we propose two simple yet effective PTM-based methods to boost the performance of VC-UFG and HC-UFG, respectively. On top of those, we integrate the two methods into one unified framework and propose the first unified model, Unic, that is capable of tackling the C-UFG problem where VC-UFG and HC-UFG co-exist in a single continual learning sequence. To understand the effectiveness of the proposed methods, we first evaluate the models on VC-UFG and HC-UFG challenges and then test the proposed Unic on a unified C-UFG challenge. Experimental results demonstrate the proposed methods achieve superior performance for C-UFG. The code is available at https://github.com/PatrickZad/unicufg. Pengcheng Zhang 0003, Xiaohan Yu 0001, Meiying Gu, Yongsheng Gao 0001, Xiao Bai 0001 |
IJCAI | 2 |
| 2025 | View-aware Decomposition and Unification for Fast Ground-to-Aerial Person SearchabstractGround-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall person search performance when training the model in a view-agnostic way. To address this, we propose a view-aware decomposition and unification (VADU) framework for ground-to-aerial person search. Specifically, we decompose the person search model to learn view-oriented modules for image feature encoding and person proposal generation. The data sampling and retrieval feature learning are also composed to cope with the decomposed model. This decomposition improves both person detection and discriminative feature learning within each view. On top of the decomposition, we propose view-aware unification to produce unified cross-view person features. Cross-view prototypical contrastive learning is introduced to enhance the unification between different views, enhancing model robustness to retrieve a target person in cameras of a different view. As the decomposed parts of the model are deployed on different devices for inference, this overall framework adds no extra computation cost in real-world applications. Extensive experiments demonstrate that the proposed method achieves superior person search performance and guarantees the efficiency of inference. The source code is available at https://github.com/QFWang-11/vadu. Qifei Wang, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Yongsheng Gao 0001 |
IROS | 3 |
| 2025 | Contrastive Lie Algebra Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself via contrastive learning has demonstrated great progress towards learning discriminative representation. Yet forcingly dividing highly similar categories at the representation level may over-guide the learned feature space, leading to overfitting in the ultra-FGVC tasks. To this end, this paper introduces CLA-Net, a novel contrastive Lie algebra learning framework to address this fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images at a second-order representation level. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLA-Net demonstrates strong performance on eight publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at: https://github.com/zichengpan/CLA-NET. Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0002, Qin Zhang 0011, Yongsheng Gao 0001 |
ACM Multimedia | 1 |
| 2025 | GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface ReconstructionabstractReconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based framework that explores and extends the under-investigated potential of sparse voxels for achieving accurate, detailed, and complete surface reconstruction. As strengths, sparse voxels support preserving the coverage completeness and geometric clarity, while corresponding challenges also arise from absent scene constraints and locality in surface refinement. To ensure correct scene convergence, we first propose a Voxel-Uncertainty Depth Constraint that maximizes the effect of monocular depth cues while presenting a voxel-oriented uncertainty to avoid quality degradation, enabling effective and robust scene constraints yet preserving highly accurate geometries. Subsequently, Sparse Voxel Surface Regularization is designed to enhance geometric consistency for tiny voxels and facilitate the voxel-based formation of sharp and accurate surfaces. Extensive experiments demonstrate our superior performance compared to existing methods across diverse challenging scenarios, excelling in geometric accuracy, detail preservation, and reconstruction completeness while maintaining high efficiency. Code is available at https://github.com/Fictionarry/GeoSVR. Jiahe Li 0007, Youmin Zhang 0005, Xiao Bai 0001, Xiaohan Yu 0001, Lin Gu 0003 |
NeurIPS | 6 |
| 2025 | Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian SplattingabstractWe present Eve3D, a novel framework for dense 3D reconstruction based on 3D
Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors
derived from pre-trained vision models, Eve3D fully leverages these priors by
jointly optimizing both them and the 3DGS backbone. This joint optimization
creates a mutually reinforcing cycle: the priors enhance the quality of 3DGS, which
in turn refines the priors, further improving the reconstruction. Additionally, Eve3D
introduces a novel optimization step based on bundle adjustment, overcoming the
limitations of the highly local supervision in standard 3DGS pipelines. Eve3D
achieves state-of-the-art results in surface reconstruction and novel view synthesis
on the Tanks & Temples, DTU, and Mip-NeRF360 datasets. while retaining fast
convergence, highlighting an unprecedented trade-off between accuracy and speed. Youmin Zhang 0005, Fabio Tosi, Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Matteo Poggi |
NeurIPS | 6 |
| 2025 | Gate-ViT: Gated Vision Transformer for Fine-Grained Visual Classification
Kanqi Wang, Peiyu Wang, Qin Zhang 0011, Yang Zhao 0019, Xiaohan Yu 0001 |
PAKDD (3) | 7 |
| 2025 | Uniformity and deformation: A benchmark for multi-fish real-time tracking in the farming
Jinze Huang, Xiaohan Yu 0001, Dong An 0001, Xin Ning 0001, Jincun Liu, Prayag Tiwari |
Expert Syst. Appl. | 2 |
| 2025 | Fully Decoupled End-to-End Person Search: An Approach without Conflicting Objectives
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Edwin R. Hancock |
Int. J. Comput. Vis. | 2 |
| 2025 | A Generative Pretrained Transformer for Semi-Supervised Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSIs-CD) often faces the challenge of limited sample sizes, and labeling data is both time-consuming and labor-intensive. Foundation models leverage extensive unlabeled data for self-supervised generative pre-training, allowing the model to learn rich data representations. However, few models have been specifically designed for HSIs, and existing methods often rely on pre-training datasets that are limited to data from a small number of satellite sensors. This limitation affects generalization, especially when there are significant differences between data from different sensors. Moreover, the difference map (DMP) of bi-temporal HSIs is often used as input to the networks. While the DMP-based approach reduces FLOPs, it may lead to information loss compared to dual-branch networks. In this letter, we propose a mini-patch-based generative pre-trained spectral-spatial transformer (GPSST) for semi-supervised HSIs-CD. We begin by collecting public HSIs datasets and dividing them into thousands of patches. Each patch is then split into spectral-spatial tokens, with a portion of these tokens masked and used as input for the GPSST. We then design a spectral-spatial masked autoencoder (MAE) as the backbone of GPSST for self-supervised generative learning. Finally, we fine-tune the GPSST encoder using a small number of labeled patches and design a principal component analysis (PCA) branch to compensate for the information loss caused by the DMP. Our experiments demonstrate that GPSST outperforms existing methods, achieving superior accuracy in HSIs-CD. Jianjun Sha, Xiaohan Yu 0001, Yongsheng Gao 0001, Yonggang Zhang 0001, Xianhui Rong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow EstimationabstractWith advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning. Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Dynamic and static mutual fitting for action recognition
Wenxuan Liu 0008, Xuemei Jia, Xian Zhong, Kui Jiang, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 5 |
| 2025 | Overcoming learning bias via Prototypical Feature Compensation for source-free domain adaptationabstractThe focus of Source-free Unsupervised Domain Adaptation (SFUDA) is to effectively transfer a well-trained model from the source domain to an unlabelled target domain. During the target domain adaptation , the source domain data is no longer accessible. Prevalent methodologies attempt to synchronize the data distributions between the source and target domains, utilizing pseudo-labels to impart categorical information, which has made some progress in improving the model’s performance. However, performance impairments persist due to the introduction of learning bias from the source model and the impact of noisy pseudo-labels generated for the target domain. In this research, we reveal that the central cause for feature misalignment during domain transition is the learning bias, which is generated by the discrepancy of information between source and target domain data . The source domain data may contain distinguishable features that do not appear on the target domain, which causes the pre-trained source model to fail to work during domain adaptation. To overcome the information discrepancy, we propose a Prototypical Feature Compensation (PFC) Network. The network extracts representative feature maps of the source domain. Then use them to minimize the discrepancy information in the target domain feature maps. This mechanism facilitates feature alignment across different domains, allowing the model to generate more accurate categorical data through pseudo-labelling. The experimental results and ablation studies demonstrate exceptional performance on three SFUDA datasets and provide evidence of the proposed PFC method’s ability to adjust the feature distribution of both source and target domain data, ensuring their overlap in the latent space. Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2025 | ALStereo: Active learning for stereo matching
Jiahe Li 0007, Meiying Gu, Xiaohan Yu 0001, Xiao Bai 0001, Edwin R. Hancock |
Pattern Recognit. | 4 |
| 2025 | Session-Guided Attention in Continuous Learning With Few SamplesabstractFew-shot class-incremental learning (FSCIL) aims to learn from a sequence of incremental data sessions with a limited number of samples in each class. The main issues it encounters are the risk of forgetting previously learned data when introducing new data classes, as well as not being able to adapt the old model to new data due to limited training samples. Existing state-of-the-art solutions normally utilize pre-trained models with fixed backbone parameters to avoid forgetting old knowledge. While this strategy preserves previously learned features, the fixed nature of the backbone limits the model's ability to learn optimal representations for unseen classes, which compromises performance on new class increments. In this paper, we propose a novel SEssion-Guided Attention framework (SEGA) to tackle this challenge. SEGA exploits the class relationships within each incremental session by assessing how test samples relate to class prototypes. This allows accurate incremental session identification for test data, leading to more precise classifications. In addition, an attention module is introduced for each incremental session to further utilize the feature from the fixed backbone. As the session of the testing image is determined, we can fine-tune the feature with the corresponding attention module to better cluster the sample within the selected session. Our approach adopts the fixed backbone strategy to avoid forgetting the old knowledge while achieving novel data adaptation. Experimental results on three FSCIL datasets consistently demonstrate the superior adaptability of the proposed SEGA framework in FSCIL tasks. The code is available at: https://github.com/zichengpan/SEGA. Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | DyCR: A Dynamic Clustering and Recovering Network for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims to continually learn novel data with limited samples. One of the major challenges is the catastrophic forgetting problem of old knowledge while training the model on new data. To alleviate this problem, recent state-of-the-art methods adopt a well-trained static network with fixed parameters at incremental learning stages to maintain old knowledge. These methods suffer from the poor adaptation of the old model with new knowledge. In this work, a dynamic clustering and recovering network (DyCR) is proposed to tackle the adaptation problem and effectively mitigate the forgetting phenomena on FSCIL tasks. Unlike static FSCIL methods, the proposed DyCR network is dynamic and trainable during the incremental learning stages, which makes the network capable of learning new features and better adapting to novel data. To address the forgetting problem and improve the model performance, a novel orthogonal decomposition mechanism is developed to split the feature embeddings into context and category information. The context part is preserved and utilized to recover old class features in future incremental learning stages, which can mitigate the forgetting problem with a much smaller size of data than saving the raw exemplars. The category part is used to optimize the feature embedding space by moving different classes of samples far apart and squeezing the sample distances within the same classes during the training stage. Experiments show that the DyCR network outperforms existing methods on four benchmark datasets. The code is available at: https://github.com/zichengpan/DyCR. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Self-Supervised Lie Algebra Representation Learning via Optimal Canonical MetricabstractLearning discriminative representation with limited training samples is emerging as an important yet challenging visual categorization task. While prior work has shown that incorporating self-supervised learning can improve performance, we found that the direct use of canonical metric in a Lie group is theoretically incorrect. In this article, we prove that a valid optimization measurement should be a canonical metric on Lie algebra. Based on the theoretical finding, this article introduces a novel self-supervised Lie algebra network (SLA-Net) representation learning framework. Via minimizing canonical metric distance between target and predicted Lie algebra representation within a computationally convenient vector space, SLA-Net avoids computing nontrivial geodesic (locally length-minimizing curve) metric on a manifold (curved space). By simultaneously optimizing a single set of parameters shared by self-supervised learning and supervised classification, the proposed SLA-Net gains improved generalization capability. Comprehensive evaluation results on eight public datasets show the effectiveness of SLA-Net for visual categorization with limited samples. Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0019, Yongsheng Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | PViT: Pooling Vision Transformer for Active Trachoma Image Classification
Mulugeta Shitie Zewudie, Shengwu Xiong 0001, Xiaohan Yu 0001, Aminu Onimisi Abdulsalami |
ADMA (4) | 3 |
| 2024 | EIANet: A Novel Domain Adaptation Approach to Maximize Class Distinction with Neural Collapse Principles
Zicheng Pan, Xiaohan Yu 0001, Yongsheng Gao 0001 |
BMVC | 2 |
| 2024 | Robust Synthetic-to-Real Transfer for Stereo MatchingabstractWith advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have investigated the robustness after fine-tuning them in real-world scenarios, during which the domain generalization ability can be seriously degraded. In this paper, we explore fine-tuning stereo matching networks without compromising their robustness to unseen domains. Our motivation stems from comparing Ground Truth (GT) versus Pseudo Label (PL) for fine-tuning: GT degrades, but PL preserves the domain generalization ability. Empirically, we find the difference between GT and PL implies valuable information that can regularize networks during fine-tuning. We also propose a framework to utilize this difference for fine-tuning, consisting of a frozen Teacher, an exponential moving average (EMA) Teacher, and a Student network. The core idea is to utilize the EMA Teacher to measure what the Student has learned and dynamically improve GT and PL for fine-tuning. We integrate our framework with state-of-the-art networks and evaluate its effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the domain generalization ability during fine-tuning. Code is available at: https://github.com/jiaw-z/DKT-Stereo. Jiahe Li 0007, Lei Huang 0015, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001 |
CVPR | 4 |
| 2024 | CoR-GS: Sparse-View 3D Gaussian Splatting via Co-regularization
Jiahe Li 0007, Xiaohan Yu 0001, Lei Huang 0015, Lin Gu 0003, Xiao Bai 0001 |
ECCV (1) | 3 |
| 2024 | Occluded Person Retrieval with Hierarchical Feature OptimizationabstractOccluded person retrieval aims to match images from occluded pedestrians. It pushes forward progress of person retrieval towards applications in real-world scenarios, thus attracting increasing attention in recent years. A key challenge is to learn discriminative representation within limited informative regions due to obstacle or pedestrian occlusion. To that end, we propose a hierarchical feature optimization model (HFO) that jointly optimizes image-level, object-level and part-level features for improved occluded person retrieval. A hierarchical discriminative feature grouping (HDFG) module is developed to generate hierarchical object/part masks for comprehensive feature extraction. Via learning a set of part prototypes, HDFG localizes hierarchical informative object/parts by grouping intermediate feature vectors based on their similarity to these prototypes. The proposed HFO is trained in an end-to-end manner using only identity labels, making it a practical solution for occluded person retrieval. We verify the effectiveness of the proposed method on three challenging occluded datasets and two holistic datasets, i.e., Occluded-DukeMTMC, Occluded-REID, P-DukeMTMC-reID, Market1501, and DukeMTMC-reID. Extensive experiments and ablation studies demonstrate superior or comparable performance of the proposed method over the state-of-the-art methods. The code is available at https://github.com/Patrickzad/HFO. Yang Zhao 0019, Pengcheng Zhang 0003, Xiaohan Yu 0001, Zhibin Liao, Johan Verjans, Xiao Bai 0001 |
FG | 3 |
| 2024 | Prompting Continual Person SearchabstractThe development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, existing person search models still lack the ability to continually learn from increasing real-world data and adaptively process input from different domains. To this end, this work introduces the continual person search task that sequentially learns on multiple domains and then performs person search on all seen domains. This requires balancing the stability and plasticity of the model to continually learn new knowledge without catastrophic forgetting. For this, we propose a Prompt-based Continual Person Search (PoPS) model in this paper. First, we design a compositional person search transformer to construct an effective pre-trained transformer without exhaustive pre-training from scratch on large-scale person search data. This serves as the fundamental for prompt-based continual learning. On top of that, we design a domain incremental prompt pool with a diverse attribute matching module. For each domain, we independently learn a set of prompts to encode the domain-oriented knowledge. Meanwhile, we jointly learn a group of diverse attribute projections and prototype embeddings to capture discriminative domain attributes. By matching an input image with the learned attributes across domains, the learned prompts can be properly selected for model inference. Extensive experiments are conducted to validate the proposed method for continual person search. The source code is available at https://github.com/PatrickZad/PoPS. Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001 |
ACM Multimedia | 2 |
| 2024 | Consistent prototype contrastive learning for weakly supervised person search
Huadong Lin, Xiaohan Yu 0001, Pengcheng Zhang 0003, Xiao Bai 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Adaptive feature selection for active trachoma image classificationabstractTrachoma is a neglected tropical eye disease caused by ocular strains of Chlamydia trachomatis, which affects millions of people worldwide. To examine the eye for signs of active trachoma, healthcare providers typically look for clusters of five or more follicles on the conjunctiva of the upper eyelid for the follicular inflammatory trachoma stage. However, it is also possible to find individual follicles scattered throughout the conjunctiva, particularly in mild or early-stage trachoma cases. Additionally, the datasets are photographic images collected in the field that can be high-dimensional and may contain large amounts of redundant information. We propose integrating novel attention-based feature extraction and feature selection techniques to address these challenges. First, we present the Lambda layer within the Convolutional Block Attention Module (L-CBAM) to normalize attention weights and improve the feature extraction process. Second, we introduce an adaptive mechanism, Adaptive Beta Hill Climbing (AβHC) with Social Ski-Driver (SSD), which adjusts the exploration-exploitation trade-off during the search process, allowing for better exploration of the search space and more efficient convergence toward an optimal feature subset. We then use the multilayer perceptron (MLP) classifier to produce final classification results using selected subsets. We evaluated the proposed approach on active trachoma inverted eyelid images and obtained accuracy scores of 93.3% with only 19.7% of the selected features, surpassing many of the algorithms used for comparison. Our proposed method has demonstrated excellent performance compared to recent works utilizing the same datasets. Mulugeta Shitie Zewudie, Shengwu Xiong 0001, Xiaohan Yu 0001, Xiaoyu O. Wu, Moges Ahmed Mehamed |
Knowl. Based Syst. | 3 |
| 2024 | Learning From Human Attention for Attribute-Assisted Visual RecognitionabstractWith prior knowledge of seen objects, humans have a remarkable ability to recognize novel objects using shared and distinct local attributes. This is significant for the challenging tasks of zero-shot learning (ZSL) and fine-grained visual classification (FGVC), where the discriminative attributes of objects have played an important role. Inspired by human visual attention, neural networks have widely exploited the attention mechanism to learn the locally discriminative attributes for challenging tasks. Though greatly promoted the development of these fields, existing works mainly focus on learning the region embeddings of different attribute features and neglect the importance of discriminative attribute localization. It is also unclear whether the learned attention truly matches the real human attention. To tackle this problem, this paper proposes to employ real human gaze data for visual recognition networks to learn from human attention. Specifically, we design a unified Attribute Attention Network (A$^{2}$Net) that learns from human attention for both ZSL and FGVC tasks. The overall model consists of an attribute attention branch and a baseline classification network. On top of the image feature maps provided by the baseline classification network, the attribute attention branch employs attribute prototypes to produce attribute attention maps and attribute features. The attribute attention maps are converted to gaze-like attentions to be aligned with real human gaze attention. To guarantee the effectiveness of attribute feature learning, we further align the extracted attribute features with attribute-defined class embeddings. To facilitate learning from human gaze attention for the visual recognition problems, we design a bird classification game to collect real human gaze data using the CUB dataset via an eye-tracker device. Experiments on ZSL and FGVC tasks without/with real human gaze data validate the benefits and accuracy of our proposed model. This work supports the promising benefits of collecting human gaze datasets and automatic gaze estimation algorithms learning from human attention for high-level computer vision tasks. Xiao Bai 0001, Pengcheng Zhang 0003, Xiaohan Yu 0001, Edwin R. Hancock, Jun Zhou 0001, Lin Gu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Pseudo-set Frequency Refinement architecture for fine-grained few-shot class-incremental learningabstractFew-shot class-incremental learning was introduced to solve the model adaptation problem for new incremental classes with only a few examples while still remaining effective for old data. Although recent state-of-the-art methods make some progress in improving system robustness on common datasets, they fail to work on fine-grained datasets where inter-class differences are small. The problem is mainly caused by: (1) the overlapping of new data and old data in the feature space during incremental learning, which means old samples can be falsely classified as newly introduced classes and induce catastrophic forgetting phenomena; (2) lacking discriminative feature learning ability to identify fine-grained objects. In this paper, a novel Pseudo-set Frequency Refinement (PFR) architecture is proposed to tackle these problems. We design a pseudo-set training strategy to mimic the incremental learning scenarios so that the model can better adapt to novel data in future incremental sessions. Furthermore, separate adaptation tasks are developed by utilizing frequency-based information to refine the original features and address the above challenging problems. More specifically, the high and low-frequency components of the images are employed to enrich the discriminative feature analysis ability and incremental learning ability of the model respectively. The refined features are used to perform inter-class and inter-set analyses. Extensive experiments show that the proposed method consistently outperforms the state-of-the-art methods on four fine-grained datasets. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2024 | Joint discriminative representation learning for end-to-end person search
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026, Xin Ning 0001 |
Pattern Recognit. | 2 |
| 2024 | Towards effective person search with deep learning: A survey from systematic perspective
Pengcheng Zhang 0003, Xiaohan Yu 0001, Chen Wang 0026, Xin Ning 0001, Xiao Bai 0001 |
Pattern Recognit. | 2 |
| 2024 | ICLR: Instance Credibility-Based Label Refinement for label noisy person re-identification
Xian Zhong, Xuemei Jia, Wenxin Huang, Wenxuan Liu 0008, Shuaipeng Su, Xiaohan Yu 0001, Mang Ye |
Pattern Recognit. | 7 |
| 2023 | CLE-ViT: Contrastive Learning Encoded Transformer for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself, e.g., predicting the rotation of a given image, has demonstrated great progress towards learning discriminative representation. Yet none of these works consider explicit supervision for learning mutual information at instance level. To this end, this paper introduces CLE-ViT, a novel contrastive learning encoded transformer, to address the fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLE-ViT demonstrates strong performance on 7 publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at https://github.com/Markin-Wang/CLEViT. Xiaohan Yu 0001, Jun Wang 0121, Yongsheng Gao 0001 |
IJCAI | 1 |
| 2023 | SSFE-Net: Self-Supervised Feature Enhancement for Ultra-Fine-Grained Few-Shot Class Incremental LearningabstractUltra-Fine-Grained Visual Categorization (ultra-FGVC) has become a popular problem due to its great real-world potential for classifying the same or closely related species with very similar layouts. However, there present many challenges for the existing ultra-FGVC methods, firstly there are always not enough samples in the existing ultraFGVC datasets based on which the models can easily get overfitting. Secondly, in practice, we are likely to find new species that we have not seen before and need to add them to existing models, which is known as incremental learning. The existing methods solve these problems by Few-Shot Class Incremental Learning (FSCIL), but the main challenge of the FSCIL models on ultra-FGVC tasks lies in their inferior discrimination detection ability since they usually use low-capacity networks to extract features, which leads to insufficient discriminative details extraction from ultrafine-grained images. In this paper, a self-supervised feature enhancement for the few-shot incremental learning network (SSFE-Net) is proposed to solve this problem. Specifically, a self-supervised learning (SSL) and knowledge distillation (KD) framework is developed to enhance the feature extraction of the low-capacity backbone network for ultra-FGVC few-shot class incremental learning tasks. Besides, we for the first time create a series of benchmarks for FSCIL tasks on two public ultra-FGVC datasets and three normal finegrained datasets, which will facilitate the development of the Ultra-FGVC community. Extensive experimental results on public ultra-FGVC datasets and other state-of-the-art benchmarks consistently demonstrate the effectiveness of the proposed method. Zicheng Pan, Xiaohan Yu 0001, Miaohua Zhang, Yongsheng Gao 0001 |
WACV | 2 |
| 2023 | Learning consistent region features for lifelong person re-identification
Jinze Huang, Xiaohan Yu 0001, Dong An 0001, Yaoguang Wei, Xiao Bai 0001, Chen Wang 0026, Jun Zhou 0001 |
Pattern Recognit. | 2 |
| 2023 | A Lie algebra representation for efficient 2D shape classification
Xiaohan Yu 0001, Yongsheng Gao 0001, Mohammed Bennamoun, Shengwu Xiong 0001 |
Pattern Recognit. | 1 |
| 2023 | Mix-ViT: Mixing attentive vision transformer for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Jun Wang 0121, Yang Zhao 0019, Yongsheng Gao 0001 |
Pattern Recognit. | 1 |
| 2023 | Attribute subspaces for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Na Li 0014, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock |
Pattern Recognit. | 5 |
| 2023 | Gait-Assisted Video Person RetrievalabstractVideo person retrieval aims at matching video clips of the same person across non-overlapping camera views, where video sequences contain more comprehensive information, e.g., temporal cues. How to extract useful temporal cues is the key to the success of a video person retrieval system. Gait, as a unique biometric modality indicating the way people walk, contains informative temporal information. To date, it is not clear how to fully utilize gait to boost the performance of video person retrieval. In this paper, to validate whether gait could help retrieve person in videos, we build a two-stream architecture, named appearance-gait network (AGNet), to jointly learn the appearance features and gait features from RGB video clips and silhouette video clips. We further explore how to fully utilize gait features to enhance the video feature representation. Specifically, we propose an appearance-gait attention module (AGA) to fuse a discriminative feature representation for the person retrieval task. Furthermore, to eliminate the requirement of silhouette video clips during inference, we propose a simple yet effective appearance-gait distillation module (AGD) which transfers the gait knowledge to appearance stream. As such, we are able to perform the enhanced video person retrieval without silhouette video clips, which makes the inference more flexible and practical. To the best of our knowledge, our work is the first to successfully introduce such appearance-gait knowledge distillation design for video person retrieval. We verify the effectiveness of the proposed methods on two large-scale challenging benchmarks of MARS and DukeMTMC-VideoReID. Extensive experiments demonstrate superior or comparable performance compared to the state-of-the-art methods while being much simpler. Source code is publicly available athttps://github.com/yangyangkiki/Gait-Assisted-Video-Reid. Yang Zhao 0019, Xiaohan Yu 0001, Chunlei Liu 0001, Yongsheng Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock |
ECCV (24) | 6 |
| 2022 | PGTRNET: Two-Phase Weakly Supervised Object Detection with Pseudo Ground Truth RefinementabstractCurrent state-of-the-art weakly supervised object detection (WSOD) studies mainly follow a two-stage training strategy which integrates a fully supervised detector (FSD) with a pure WSOD model. There are two main problems hindering the performance of the two-phase WSOD approaches, i.e., insufficient learning problem and strict reliance between the FSD and the pseudo ground truth (PGT) generated by the WSOD model. This paper proposes pseudo ground truth refinement network (PGTRNet), a simple yet effective method with-out introducing any extra learnable parameters, to cope with these problems. PGTRNet utilizes multiple bounding boxes to establish the PGT, mitigating the insufficient learning problem. Besides, we propose a novel online PGT refinement approach to steadily improve the quality of PGT by fully taking advantage of the power of FSD during the second-phase training, decoupling the first and second-phase models. Elaborate experiments are conducted on the PASCAL VOC 2007 benchmark to verify the effectiveness of our methods. Experimental results demonstrate that PGTRNet boosts the backbone model by 2.1% mAP and achieves the state-of-the-art performance. Jun Wang 0121, Hefeng Zhou, Xiaohan Yu 0001 |
ICASSP | 3 |
| 2022 | SPARE: Self-supervised part erasing for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001 |
Pattern Recognit. | 1 |
| 2022 | Learning discriminative region representation for person retrieval
Yang Zhao 0019, Xiaohan Yu 0001, Yongsheng Gao 0001, Chunhua Shen |
Pattern Recognit. | 2 |
| 2021 | Feature Fusion Vision Transformer for Fine-Grained Visual Categorization
Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
BMVC | 2 |
| 2021 | Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human PerformanceabstractDeep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on the basis of such human visual observation. In contrast, few research has been done for visual categorization at the ultra-fine-grained level, i.e., a granularity where even human experts can hardly identify the visual differences or are not yet able to give affirmative labels by inferring observed pattern differences. This paper reports our efforts towards mitigating this research gap. We introduce the ultra-fine-grained (UFG) image dataset, a large collection of 47,114 images from 3,526 categories. All the images in the proposed UFG image dataset are grouped into categories with different confirmed cultivar names. In addition, we perform an extensive evaluation of state-of-the-art fine-grained classification methods on the proposed UFG image dataset as comparative baselines. The proposed UFG image dataset and evaluation protocols is intended to serve as a benchmark platform that can advance research of visual classification from approaching human performance to beyond human ability, via facilitating benchmark data of artificial intelligence (AI) not to be limited by the labels of human intelligence (HI). The dataset is available online at https://githuh.com/XiaohanYu-GU/Ultra-FGVC. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
ICCV | 1 |
| 2021 | Mask Guided Attention For Fine-Grained Patchy Image ClassificationabstractIn this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies in two folds, ultra-fine-grained inter-category variances among objects and very few data available for training. This motivates us to consider employing more useful supervision signal to train a discriminative model within limited training samples. Specifically, the proposed MGA integrates a pre-trained semantic segmentation model that produces auxiliary supervision signal, i.e., patchy attention mask, enabling a discriminative representation learning. The patchy attention mask drives the classifier to filter out the insignificant parts of images (e.g., common features between different categories), which enhances the robustness of MGA for the fine-grained patchy image classification. We verify the effectiveness of our method on three publicly available patchy image datasets. Experimental results demonstrate that our MGA method achieves superior performance on three datasets compared with the state-of-the-art methods. In addition, our ablation study shows that MGA improves the accuracy by 2.25% and 2% on the SoyCultivarVein and BtfPIS datasets, indicating its practicality towards solving the fine-grained patchy image classification. Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
ICIP | 2 |
| 2021 | MaskCOV: A random mask covariance network for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 1 |
| 2021 | Learning deep part-aware embedding for person retrieval
Yang Zhao 0019, Chunhua Shen, Xiaohan Yu 0001, Hao Chen 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 3 |
| 2020 | Patchy Image Structure Classification Using Multi-Orientation Region TransformabstractExterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal structure feature separately, and thus fail to function when classifying patchy image structures that have similar contours and flexible structures. To address above limitations, this paper proposes a novel Multi-Orientation Region Transform (MORT), which can effectively characterize both contour and structure features simultaneously, for patchy image structure classification. MORT is performed over multiple orientation regions at multiple scales to effectively integrate patchy features, and thus enables a better description of the shape in a coarse-to-fine manner. Moreover, the proposed MORT can be extended to combine with the deep convolutional neural network techniques, for further enhancement of classification accuracy. Very encouraging experimental results on the challenging ultra-fine-grained cultivar recognition task, insect wing recognition task, and large variation butterfly recognition task are obtained, which demonstrate the effectiveness and superiority of the proposed MORT over the state-of-the-art methods in classifying patchy image structures. Our code and three patchy image structure datasets are available at: https://github.com/XiaohanYu-GU/MReT2019. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
AAAI | 1 |
| 2019 | Contour Covariance: A Fast Descriptor for ClassificationabstractThis paper presents a novel shape descriptor to effectively and efficiently characterize the local image statistics. The proposed descriptor, termed contour covariance (CC), characterizes covariance features driven by a moving point on the shape contour at multiple scales. To calculate the covariance matrices, three basic features including texture, intensity and distance map, are extracted from the object image. Based on coefficients of the obtained covariance matrices, the proposed CC descriptor is compact yet informative, as well as invariant to rotation, translation and scale. The experimental results on two databases demonstrate the superiority and efficiency of the proposed method among the state-of-the-art methods for shape classification. Xiaohan Yu 0001, Shengwu Xiong 0001, Yongsheng Gao 0001 |
ICIP | 1 |
| 2017 | Bar charts detection and analysis in biomedical literature of PubMed Central
Xiaohan Yu 0001, Yangjing Gan, Tujin Zhu, Shengwu Xiong 0001, Lun Hu |
AMIA | 2 |
| 2016 | Research on campus traffic congestion detection using BP neural network and Markov model
Xiaohan Yu 0001, Shengwu Xiong 0001, W. Eric Wong, Yang Zhao 0019 |
J. Inf. Secur. Appl. | 1 |
| 2015 | PopGeV: a web-based large-scale population genome browserabstractMOTIVATION: The development of high-throughput sequencing technology has made it possible for more and more researchers to use population sequencing data to mine genes associated with specific traits. However, the massive amounts of sequencing data have also brought new challenges to the researchers. The question of how to browse population genomic data in an easy and intuitive manner must be addressed. Web-based genome browsers allow user to conveniently view the results of genomic analyses, but heavy usage can reduce the response speed of the webpage, which limits its usefulness in the display of large-scale genome data. IndexedDB technology is a good solution to this problem; it supports web browsers and so creates local databases. In this way, data can be read from the local storage, achieving a smooth display of population genomic data. RESULTS: PopGeV has the following characteristics. First, it uses a new encoding method for compression of population SNP and INDEL data. IndexedDB technology is used to download the results to local storage so that users can browse the results smoothly even when the network traffic is heavy. Second, PopGeV identify similar genomic regions between two individuals based on SNP data. Population diversity indexes are calculated when comparing two populations. Third, user defined annotation information can be integrated for user-friendly mining of gene functions. Simulation shows that PopGeV can smoothly display analysis results of population genome containing over 500 individuals with 2 millions SNP data. AVAILABILITY AND IMPLEMENTATION: PopGeV is available at www.soyomics.com/popgev/ CONTACT: [email protected]. Xinyi Shi, Xiaohan Yu 0001, Dongye Li, Baohui Liu, Fanjiang Kong |
Bioinform. | 3 |