VLDB 2026 Research / reviewers in the wild / expert
Siya Mi
dblp:254/2399
· DBLP profile ↗
18ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-1751-7076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Effective In-context Demonstration Selection with CoresetabstractIn-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection. Zihua Wang, Jiarui Wang 0002, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002, Xu Yang 0021, Xiu-Shen Wei, Siya Mi, Yu Zhang 0004 |
AAAI | 8 |
| 2026 | Depth estimation based on RGB-event hybrid binocular cameras
Siya Mi |
Neurocomputing | 1 |
| 2025 | DVC2: Deep video cascade clustering from video structures
Zihua Wang, Siya Mi, Yu Zhang 0004 |
Neurocomputing | 2 |
| 2025 | Object Adaptive Self-Supervised Dense Visual Pre-TrainingabstractSelf-supervised visual pre-training models have achieved significant success without employing expensive annotations. Nevertheless, most of these models focus on iconic single-instance datasets (e.g. ImageNet), ignoring the insufficient discriminative representation for non-iconic multi-instance datasets (e.g. COCO). In this paper, we propose a novel Object Adaptive Dense Pre-training (OADP) method to learn the visual representation directly on the multi-instance datasets (e.g., PASCAL VOC and COCO) for dense prediction tasks (e.g., object detection and instance segmentation). We present a novel object-aware and learning-adaptive random view augmentation to focus the contrastive learning to enhance the discrimination of object presentations from large to small scale during different learning stages. Furthermore, the representations across different scale and resolutions are integrated so that the method can learn diverse representations. In the experiment, we evaluated OADP pre-trained on PASCAL VOC and COCO. Results show that our method has better performances than most existing state-of-the-art methods when transferring to various downstream tasks, including image classification, object detection, instance segmentation and semantic segmentation. Yu Zhang 0004, Hongyuan Zhu 0002, Siya Mi, Xi Peng 0001, Xin Geng 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Learning sample representativeness for class-imbalanced multi-label classification
Sichen Cao, Siya Mi, Yali Bian |
Pattern Anal. Appl. | 3 |
| 2024 | Temporal segment dropout for human action video recognition
Yu Zhang 0004, Zhengjie Chen, Siya Mi, Xin Geng 0001, Min-Ling Zhang |
Pattern Recognit. | 5 |
| 2023 | Person video alignment with human pose registration
Zihua Wang, Siya Mi |
Frontiers Comput. Sci. | 4 |
| 2023 | Asymmetric bi-encoder for image-text retrieval
Haoliang Liu, Siya Mi, Yu Zhang 0004 |
Multim. Syst. | 3 |
| 2023 | Self-label correction for image classification with noisy labels
Yu Zhang 0004, Fan Lin, Siya Mi, Yali Bian |
Pattern Anal. Appl. | 3 |
| 2023 | LPCL: Localized prominence contrastive learning for self-supervised dense visual pre-training
Hongyuan Zhu 0002, Siya Mi, Yu Zhang 0004, Xin Geng 0001 |
Pattern Recognit. | 4 |
| 2023 | Assisting Multimodal Named Entity Recognition by cross-modal auxiliary tasks
Zhengjie Chen, Yu Zhang 0004, Siya Mi |
Pattern Recognit. Lett. | 3 |
| 2023 | A Closer Look at Video Sampling for Sequential Action RecognitionabstractIn recent years, sequential action recognition has attracted increasingly attention as it requires long-term sequential and compositional reasoning of human actions and object interactions. Existing methods perform reasoning either by using snippets that cover very short consecutive frames or key frames sampled from segments, which take a bias process of local and global temporal information. We also find ad-hoc training and ensembling of two separate networks using existing sampling strategies can easily outperform complex state-of-the-art methods, which reveals the complementary nature of current sampling strategies. Motivated by this observation, we propose a simple yet efficient strategy named Dense Segmental Sampling (DSS) and a novel network architecture named Temporal Dense Segment Network (TDSN) to capture the complementary information from DSS. Our TDSN achieves excellent results on benchmark action recognition datasets, which not only validate the proposed strategy but also help highlight the importance along this direction for sequential video reasoning. Yu Zhang 0004, Zhengjie Chen, Siya Mi, Hongyuan Zhu 0002, Xin Geng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Pose-guided action recognition in static images using lie-group
Siya Mi, Yu Zhang 0004 |
Appl. Intell. | 1 |
| 2022 | Weakly supervised temporal action localization with proxy metric modeling
Yu Zhang 0004, Xin Geng 0001, Siya Mi, Zhihong Yang |
Frontiers Comput. Sci. | 5 |
| 2022 | Learning time-aware features for action quality assessment
Yu Zhang 0004, Siya Mi |
Pattern Recognit. Lett. | 3 |
| 2021 | Image captioning with transformer and knowledge graph
Yu Zhang 0004, Xinyu Shi 0003, Siya Mi, Xu Yang 0021 |
Pattern Recognit. Lett. | 3 |
| 2020 | Large-scale multi-label classification using unknown streaming images
Yu Zhang 0004, Xu-Ying Liu, Siya Mi, Min-Ling Zhang |
Pattern Recognit. | 4 |
| 2020 | Simultaneous 3D hand detection and pose estimation using single depth images
Yu Zhang 0004, Siya Mi, Jianxin Wu 0001, Xin Geng 0001 |
Pattern Recognit. Lett. | 2 |