VLDB 2026 Research / reviewers in the wild / expert
Xingdong Sheng
dblp:53/8066
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0003-9980-1392ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) extends traditional closed-set segmentation by enabling pixel-wise annotation for both seen and unseen categories using arbitrary textual descriptions. While existing methods leverage vision-language models (VLMs) like CLIP, their reliance on image-level pretraining often results in imprecise spatial alignment, leading to mismatched segmentations in ambiguous or cluttered scenes. However, most existing approaches lack strong object priors and region-level constraints, which can lead to object hallucination or missed detections, further degrading performance. To address these challenges, we propose LoGoSeg, an efficient single-stage framework that integrates three key innovations: (i) an object existence prior that dynamically weights relevant categories through global image-text similarity, effectively reducing hallucinations; (ii) a region-aware alignment module that establishes precise region-level visual-textual correspondences; and (iii) a dual-stream fusion mechanism that optimally combines local structural information with global semantic context. Unlike prior works, LoGoSeg eliminates the need for external mask proposals, additional backbones, or extra datasets, ensuring efficiency. Extensive experiments on six benchmarks (A-847, PC-459, A-150, PC-59, PAS-20, and PAS-20b) demonstrate its competitive performance and strong generalization in open-vocabulary settings. Xiangbo Lv, Zhiqiang Kou, Xingdong Sheng, Yiguo Qiao |
AAAI | 4 |
| 2026 | A Virtual Deployment System for Mobile Robot Inspection-Based IoT SensingabstractMobile robot inspection is a representative application of the Internet of Robotic Things (IoRT), where robots serve as mobile IoT sensing terminals for equipment monitoring and data acquisition. While embodied intelligence has greatly advanced locomotion, navigation, and perception, the critical deployment phase remains underexplored. Conventional inspection deployment relies on on-site manual configuration and repeated debugging, making it labor-intensive, costly, and potentially unsafe. To address these issues, this paper presents the design and implementation of a virtual deployment system for mobile robot inspection. The system constructs a high-fidelity digital twin of the physical site, migrating the entire deployment process into a virtual environment. The paper details the system architecture and workflow and analyzes key challenges with corresponding solutions. The proposed system enables scalable and cost-efficient IoRT inspection in industrial environments. Xingdong Sheng, Yunhui Liu 0006, Yuqiang Cao, Shijie Mao, Xiaokang Yang 0001 |
IEEE Internet Things J. | 1 |
| 2025 | General Compression Framework for Efficient Transformer Object TrackingabstractPrevious works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker. Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001 |
ICCV | 10 |
| 2025 | Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001 |
ICCV | 9 |
| 2025 | Enhancing Visual Localization with Cross-Domain Image GenerationabstractVisual localization aims to predict the absolute camera pose for a single query image. However, predominant methods focus on single-camera images and scenes with limited appearance variations, limiting their applicability to cross-domain scenes commonly encountered in real-world applications. Furthermore, the long-tail distribution of cross-domain datasets poses additional challenges for visual localization. In this work, we propose a novel cross-domain data generation method to enhance visual localization methods. To achieve this, we first construct a cross-domain 3DGS to accurately model photometric variations and mitigate the interference of dynamic objects in large-scale scenes. We introduce a text-guided image editing model to enhance data diversity for addressing the long-tail distribution problem and design an effective fine-tuning strategy for it. Then, we develop an anchor-based method to generate high-quality datasets for visual localization. Finally, we introduce positional attention to address data ambiguities in cross-camera images. Extensive experiments show that our method achieves state-of-the-art accuracy, outperforming existing cross-domain visual localization methods by an average of 59% across all domains. Project page: https://yzwang-sjtu.github.io/CDG-Loc. Yuanze Wang, Yichao Yan, Shiming Song 0003, Songchang Jin, Yilan Huang, Xingdong Sheng, Dian-xi Shi |
ICML | 6 |
| 2025 | IPAD: Industrial Process Anomaly Detection DatasetabstractVideo anomaly detection (VAD) is a challenging task aiming to recognize anomalies in video frames, and existing large-scale VAD researches primarily focus on road traffic and human activity scenes. In industrial scenes, there are often a variety of unpredictable anomalies, and the VAD method can play a significant role in these scenarios. However, there is a lack of applicable datasets and methods specifically tailored for industrial production scenarios due to concerns regarding privacy and security. To bridge this gap, we propose a new dataset, IPAD, specifically designed for VAD in industrial scenarios. The industrial processes in our dataset are chosen through on-site factory research and discussions with engineers. This dataset covers 16 different industrial devices and contains over 6 hours of both synthetic and real-world video footage. Moreover, we annotate the key feature of the industrial process, i.e., periodicity. Based on the proposed dataset, we introduce a period memory module and a sliding window inspection mechanism to effectively investigate the periodic information in a basic reconstruction model. Our framework leverages LoRA adapter to explore the effective migration of pretrained models, which are initially trained using synthetic data, into real-world scenarios. Our proposed dataset and method will fill the gap in the field of industrial video anomaly detection and drive the process of video understanding tasks as well as smart factory deployment. Project page:https://ljf1113.github.io/IPAD_VAD. Jinfan Liu, Yichao Yan, Weiming Zhao, Pengzhi Chu, Xingdong Sheng, Yunhui Liu 0006, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Inter-X: Towards Versatile Human-Human Interaction AnalysisabstractThe analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine- grained textual descriptions. To better perceive and generate human-human interactions, we propose Inter-X, a currently largest human-human interaction dataset with accurate body movements and diverse interaction patterns, together with detailed hand gestures. The dataset includes Liang Xu 0012, Xintao Lv, Yichao Yan, Xin Jin 0014, Shuwen Wu, Congsheng Xu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu 0006, Wenjun Zeng 0001, Xiaokang Yang 0001 |
CVPR | 10 |
| 2024 | E3Gen: Efficient, Expressive and Editable Avatars Generation
Weitian Zhang, Yichao Yan, Yunhui Liu 0006, Xingdong Sheng, Xiaokang Yang 0001 |
ACM Multimedia | 4 |
| 2009 | Visual Saliency Based Object Tracking
Zejian Yuan, Nanning Zheng 0001, Xingdong Sheng |
ACCV (2) | 4 |