EDBT 2026 Demo / reviewers in the wild / expert
Shuo Ye
dblp:318/8942
· DBLP profile ↗
18ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-7756-8233ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral InspectionabstractMinghui Jia, Qichao Zhang, Ali Luo, Linjing Li, Shuo Ye, Hailing Lu, Wen Hou, Dongbin Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Minghui Jia, A-Li Luo, Linjing Li, Shuo Ye, Hailing Lu, Wen Hou, Dongbin Zhao |
ACL (1) | 5 |
| 2026 | Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent PurificationabstractRecent Audio-Visual Question Answering (AVQA) methods have advanced significantly.However, most AVQA methods lack effective mechanisms for handling missing modalities, suffering from severe performance degradation in real-world scenarios with data interruptions.Furthermore, prevailing methods for handling missing modalities predominantly rely on generative imputation to synthesize missing features.While partially effective, these methods tend to capture inter-modal commonalities but struggle to acquire unique, modalityspecific knowledge within the missing data, leading to hallucinations and compromised reasoning accuracy.To tackle these challenges, we propose R 2 ScP, a novel framework that shifts the paradigm of missing modality handling from traditional generative imputation to retrieval-based recovery.Specifically, we leverage cross-modal retrieval via unified semantic embeddings to acquire missing domain-specific knowledge.To maximize semantic restoration, we introduce a context-aware adaptive purification mechanism that eliminates latent semantic noise within the retrieved data.Additionally, we employ a two-stage training strategy to explicitly model the semantic relationships between knowledge from different sources.Extensive experiments demonstrate that R 2 ScP significantly improves AVQA and enhances robustness in modal-incomplete scenarios.1 Jiayu Zhang 0002, Shuo Ye, Qilang Ye, Zihan Song 0009, Jiajian Huang, Zitong Yu |
ACL (1) | 2 |
| 2026 | FELT: Federated Ensemble Learning for Long-Tailed IoT Data via Communication-Efficient Private VotingabstractFederated learning (FL) enables collaborative model training on decentralized Internet of Things (IoT) data while keeping raw data local, thereby mitigating privacy risks. In practice, however, IoT deployments suffer from two coupled challenges: strict uplink bandwidth constraints and long-tailed label distributions where rare events are most critical. Existing work treats these issues separately—communication-efficient FL often neglects data imbalance, whereas federated long-tail methods typically incur heavy communication overheads and privacy risks. We propose FELT, short forFederated Ensemble Learning for Long-Tailed IoT Data, a framework that jointly addresses communication efficiency, long-tail robustness, and privacy through a communication-efficient private voting protocol. Instead of transmitting full model updates, FELT devices function as teacher models, transmitting only single-integer votes on public queries. This architecture reduces communication traffic by up to one order of magnitude (approx. 10×) compared to standard FL methods. On the server, FELT aggregates votes with calibrated noise to ensure (ε, δ)-differential privacy, followed by a post-hoc class-prior-based calibration that improves tail-class predictions with negligible extra computational overhead on the server side. Experiments on multiple real-world IoT datasets demonstrate that FELT substantially boosts tail-class F1 scores (up to 26% improvement) under the same privacy budget. These results highlight FELT as a practical solution for communication-constrained and privacy-sensitive IoT applications. Chaomeng Chen, Shuo Ye, Haochen Liang, Fei Luo 0003, Lu Wang 0002, Zitong Yu |
IEEE Internet Things J. | 2 |
| 2026 | Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and MethodabstractData is the foundation for the development of computer vision, and the establishment of datasets plays an important role in advancing the techniques of fine-grained visual categorization (FGVC). In the existing FGVC datasets used in computer vision, it is generally assumed that each collected instance has fixed characteristics and the distribution of different categories is relatively balanced. In contrast, the real world scenario reveals the fact that the characteristics of instances tend to vary with time and exhibit a long-tailed distribution. Hence, the collected datasets may mislead the optimization of the fine-grained classifiers, resulting in unpleasant performance in real applications. Starting from the real-world conditions and to promote the practical progress of fine-grained visual categorization, we present a Concept Drift and Long-Tailed Distribution (CDLT) dataset. Specifically, the dataset is collected by gathering 11195 images of 250 instances in different species for 47 consecutive months in their natural contexts. The collection process involves dozens of crowd workers for photographing and domain experts for labeling. Meanwhile, we propose a feature recombination framework to address the learning challenges associated with CDLT. Experimental results validate the efficacy of our method while also highlighting the limitations of popular large vision-language models (e.g., CLIP) in the context of long-tailed distributions. This emphasizes the significance of CDLT as a benchmark for investigating these challenges. Shuo Ye, Shiming Chen 0002, Ruxin Wang 0002, Tianxu Wu, Salman Khan 0001, Fahad Shahbaz Khan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | High-resolution underwater camouflaged object detection: GBU-UCOD dataset and topology-aware and frequency-decoupled networks
Wenji Wu, Shuo Ye, Yiyu Liu, Jiguang He, Zitong Yu |
Pattern Recognit. Lett. | 2 |
| 2026 | HRR: Hierarchical retrospection refinement for generated image detection
Peipei Yuan, Zijing Xie, Shuo Ye, Yanfang Tao, Xiangcheng Wu |
Pattern Recognit. Lett. | 3 |
| 2025 | GKA: Graph-guided knowledge association for fine-grained visual categorization
Yuetian Wang, Shuo Ye, Wenjin Hou, Duanquan Xu, Xinge You |
Neurocomputing | 2 |
| 2025 | VCE: Visual concept embedding for open-set fine-grained image retrieval
Yuetian Wang, Shuo Ye, Wenjin Hou, Shuhuang Chen |
Knowl. Based Syst. | 2 |
| 2025 | FN-NET: Adaptive data augmentation network for fine-grained visual categorization
Shuo Ye, Qinmu Peng, Yiu-Ming Cheung, Yu Wang 0106, Ziqian Zou, Xinge You |
Pattern Recognit. | 1 |
| 2025 | Toward Disentangled and Controllable Deep Metric Learning With Human-Like Concept DecompositionabstractDeep metric learning (DML) has shown significant advancements in learning discriminative embeddings for images, playing a crucial role in various vision tasks. However, existing methods typically rely on deep neural networks to extract holistic embeddings, which are challenging to disentangle and interpret. To address this issue, we take inspiration from human cognition, where objects are decomposed into distinct concepts for better understanding. Specifically, we propose the concept metrics network (CMNs) to achieve disentangled and controllable DML. CMN begins by initializing learnable concept vectors to represent various visual concepts. These vectors are then associated with regional visual features via cross-attention mechanism, ensuring each vector corresponds to specific visual properties. Finally, the concept values, determined by their presence in the image, form the output embedding. Comprehensive experiments demonstrate that CMN effectively disentangles visual concepts, with each embedding dimension corresponding to a specific concept. Our method not only outperforms existing state-of-the-art methods in conventional DML application (i.e., image retrieval), but also enables more flexible and controllable application. The code is available at https://github.com/shchen0001/CMN. Shuhuang Chen, Shiming Chen 0002, Shuo Ye, Yuetian Wang, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Filter Pruning Based on Information Capacity and IndependenceabstractFilter pruning has gained widespread adoption for the purpose of compressing and speeding up convolutional neural networks (CNNs). However, the existing approaches are still far from practical applications due to biased filter selection and heavy computation cost. This article introduces a new filter pruning method that selects filters in an interpretable, multiperspective, and lightweight manner. Specifically, we evaluate the contributions of filters from both individual and overall perspectives. For the amount of information contained in each filter, a new metric called information capacity is proposed. Inspired by the information theory, we utilize the interpretable entropy to measure the information capacity and develop a feature-guided approximation process. For correlations among filters, another metric called information independence is designed. Since the aforementioned metrics are evaluated in a simple but effective way, we can identify and prune the least important filters with less computation cost. We conduct comprehensive experiments on benchmark datasets employing various widely used CNN architectures to evaluate the performance of our method. For instance, on ILSVRC-2012, our method outperforms state-of-the-art methods by reducing floating-point operations (FLOPs) by 77.4% and parameters by 69.3% for ResNet-50 with only a minor decrease in an accuracy of 2.64%. Shuo Ye, Yufeng Shi 0003, Tianheng Hu, Qinmu Peng, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Causal Visual-semantic Correlation for Zero-shot Learning
Shuhuang Chen, Dingjie Fu, Shiming Chen 0002, Shuo Ye, Wenjin Hou, Xinge You |
ACM Multimedia | 4 |
| 2024 | R2-trans: Fine-grained visual categorization with redundancy reduction
Shuo Ye, Shujian Yu, Yu Wang 0106, Xinge You |
Image Vis. Comput. | 1 |
| 2024 | The Image Data and Backbone in Weakly Supervised Fine-Grained Visual Categorization: A Revisit and Further ThinkingabstractWeakly-supervised fine-grained visual categorization (FGVC) aims to achieve subclass classification within the same large class using only label information. Compared to general images, fine-grained images have similar appearances and features, and are often affected by disturbances such as viewpoint, lighting, and occlusion during data collection, resulting in significant intra-class variance and small inter-class variance. To achieve FGVC, carefully designed models are often needed to explore the locally discriminative regions of the image. This paper revisits high-quality FGVC publications based on deep learning and analyzes from two new perspective: fine-grained image data and backbone. We address two ignored but interesting problems in FGVC. First, we argue that the reasons for exacerbating intra-class variance are not the same in data of animal, plant, and commodity types, and it is necessary to consider the effects of posture, covariate shift, and structural changes. Additionally, the “soft boundary” between subclasses intensifies the difficulty of classification. Second, we highlight that convolutional networks and self-attention networks have different receptive fields and shape biases, leading to performance differences when processing different types of fine-grained data. Overall, our analysis provides new insights into recent advances, challenges, and future directions for FGVC based on deep learning, which can help researchers develop more effective models for FGVC. Shuo Ye, Yu Wang 0106, Qinmu Peng, Xinge You, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Discriminative Suprasphere Embedding for Fine-Grained Visual CategorizationabstractDespite the great success of the existing work in fine-grained visual categorization (FGVC), there are still several unsolved challenges, e.g., poor interpretation and vagueness contribution. To circumvent this drawback, motivated by the hypersphere embedding method, we propose a discriminative suprasphere embedding (DSE) framework, which can provide intuitive geometric interpretation and effectively extract discriminative features. Specifically, DSE consists of three modules. The first module is a suprasphere embedding (SE) block, which learns discriminative information by emphasizing weight and phase. The second module is a phase activation map (PAM) used to analyze the contribution of local descriptors to the suprasphere feature representation, which uniformly highlights the object region and exhibits remarkable object localization capability. The last module is a class contribution map (CCM), which quantitatively analyzes the network classification decision and provides insight into the domain knowledge about classified objects. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed method in comparison with state-of-the-art methods. Shuo Ye, Qinmu Peng, Wenju Sun, Jiamiao Xu, Yu Wang 0106, Xinge You, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Attention-Based Deep Convolutional Network for Speech Recognition Under Multi-scene Noise Environment
Chuanwu Yang, Shuo Ye, Zhishu Lin, Qinmu Peng, Jiamiao Xu, Peipei Yuan, Yuetian Wang, Xinge You |
ICONIP (9) | 2 |
| 2023 | Coping with change: Learning invariant and minimum sufficient representations for fine-grained visual categorization
Shuo Ye, Shujian Yu, Wenjin Hou, Yu Wang 0106, Xinge You |
Comput. Vis. Image Underst. | 1 |
| 2022 | Mask-Vit: an Object Mask Embedding in Vision Transformer for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) targets to accurately identify the subordinate categories from a target class. Convolutional neural network (CNN) based methods prove that the attention mechanism can enhance the representation of local regions and improve the recognition accuracy. Recently, vision transformer (ViT) has shown great application potential in image classification tasks by taking advantage of its inherent self-attention mechanism and early global information acquisition capability. However, this global information acquisition approach involves an irrelevant environment in the interaction process, which makes it difficult for fine-grained tasks that rely on local differences to quickly learn discriminant features. To this end, we propose a hybrid network termed Mask-ViT, which can effectively avoid environmental interference and express more robust features by focusing on the instance itself. Specifically, Contour Knowledge Embedding (CKE) is employed to transferred prior location information to ViT and guided the subsequent recognition. The experiments on three benchmarks demonstrate the effectiveness of the proposed method. Shuo Ye, Chengqun Song, Jun Cheng 0002 |
ICIP | 2 |