EDBT 2026 Demo / reviewers in the wild / expert
Bowen Shi 0003
dblp:169/3160-3
· DBLP profile ↗
17ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-9169-7055ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
Yuchen Liu 0006, Bowen Shi 0003, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICCV | 3 |
| 2025 | Rethinking visual prompt learning as masked visual token modeling
Ning Liao, Bowen Shi 0003, Xiaopeng Zhang 0008, Min Cao 0005, Junchi Yan, Qi Tian 0001 |
Artif. Intell. | 2 |
| 2025 | MENSA: Multi-Dataset Harmonized Pretraining for Semantic SegmentationabstractExisting pretraining methods for semantic segmentation are hampered by the task gap between global image -level pretraining and local pixel-level finetuning. Joint dense-level pretraining is a promising alternative to exploit off-the-shelf annotations from diverse segmentation datasets but suffers from low-quality class embeddings and inconsistent data and supervision signals across multiple datasets by directly employing CLIP. To overcome these challenges, we propose a novelMulti-datasEt harmoNized pretraining framework forSemantic sEgmentation (MENSA). MENSA incorporates high-quality language embeddings and momentum-updated visual embeddings to effectively model the class relationships in the embedding space and thereby provide reliable supervision information for each category. To further adapt to multiple datasets, we achieve one-to-many pixel-embedding pairing with cross-dataset multi-label mapping through cross-modal information exchange to mitigate inconsistent supervision signals and introduce region-level and pixel-level cross-dataset mixing for varying data distribution. Experimental results demonstrate that MENSA is a powerful foundation segmentation model that consistently outperforms popular supervised or unsupervised ImageNet pretrained models for various benchmarks under standard fine-tuning. Furthermore, MENSA is shown to significantly benefit frozen-backbone fine-tuning and zero-shot learning by endowing pixel-level distinctiveness to learned representations. Bowen Shi 0003, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong |
IEEE Trans. Multim. | 1 |
| 2024 | UMG-CLIP: A Unified Multi-granularity Vision Generalist for Open-World Understanding
Bowen Shi 0003, Peisen Zhao, Yuhang Zhang 0012, Jin Li 0057, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001, Xiaopeng Zhang 0008 |
ECCV (38) | 1 |
| 2024 | Hybrid Distillation: Connecting Masked Autoencoders with Contrastive LearnersabstractAs two prominent strategies for representation learning, Contrastive Learning (CL) and Masked Image Modeling (MIM) have witnessed significant progress. Previous studies have demonstrated the advantages of each approach in specific scenarios. CL, resembling supervised pre-training, excels at capturing longer-range global patterns and enhancing feature discrimination, while MIM is adept at introducing local and diverse attention across transformer layers. Considering the respective strengths, previous studies utilize feature distillation to inherit both discrimination and diversity. In this paper, we thoroughly examine previous feature distillation methods and observe that the increase in diversity mainly stems from asymmetric designs, which may in turn compromise the discrimination ability. To strike a balance between the two properties, we propose a simple yet effective strategy termed Hybrid Distill, which leverages both the CL and MIM teachers to jointly guide the student model. Hybrid Distill emulates the token relations of the MIM teacher at intermediate layers for diversity, while simultaneously distilling the final features of the CL teacher to enhance discrimination. A progressive redundant token masking strategy is employed to reduce the expenses associated with distillation and aid in preventing the model from converging to local optima. Experimental results demonstrate that Hybrid Distill achieves superior performance on various benchmark datasets. Bowen Shi 0003, Xiaopeng Zhang 0008, Jin Li 0057, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
ICLR | 1 |
| 2024 | BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationabstractPre-training followed by full fine-tuning has gradually been substituted by Parameter-Efficient Tuning (PET) in the field of computer vision. PET has gained popularity, especially in the context of large-scale models, due to its ability to reduce transfer learning costs and conserve hardware resources. However, existing PET approaches primarily focus on recognition tasks and typically support uni-modal optimization, while neglecting dense prediction tasks and vision language interactions. To address this limitation, we propose a novel PET framework called **B**i-direction**a**l Inte**r**twined Vision **L**anguage Effici**e**nt Tuning for **R**eferring **I**mage Segment**a**tion (**BarLeRIa**), which leverages bi-directional intertwined vision language adapters to fully exploit the frozen pre-trained models' potential in cross-modal dense prediction tasks. In BarLeRIa, two different tuning modules are employed for efficient attention, one for global, and the other for local, along with an intertwined vision language tuning module for efficient modal fusion.
Extensive experiments conducted on RIS benchmarks demonstrate the superiority of BarLeRIa over prior PET methods with a significant margin, i.e., achieving an average improvement of 5.6\%. Remarkably, without requiring additional training datasets, BarLeRIa even surpasses SOTA full fine-tuning approaches. The code is available at https://github.com/NastrondAd/BarLeRIa. Jin Li 0057, Xiaopeng Zhang 0008, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICLR | 4 |
| 2024 | Bootstrap AutoEncoders With Contrastive Paradigm for Self-supervised Gaze EstimationabstractExisting self-supervised methods for gaze estimation using the dominant streams of contrastive and generative approaches are restricted to eye images and could fail in general full-face settings. In this paper, we reveal that contrastive methods are ineffective in data augmentation for self-supervised full-face gaze estimation, while generative methods are prone to trivial solutions due to the absence of explicit regularization on semantic representations. To address this challenge, we propose a novel approach called **B**ootstrap auto-**e**ncoders with **C**ontrastive p**a**radigm (**BeCa**), which combines the strengths of both generative and contrastive methods. Specifically, we revisit the Auto-Encoder used in generative approaches and incorporate the contrastive paradigm to introduce explicit regularization on gaze representation. Furthermore, we design the InfoMSE loss as an alternative to the vanilla MSE loss for Auto-Encoder to mitigate the inconsistency between reconstruction and representation learning. Experimental results demonstrate that the proposed approaches outperform state-of-the-art unsupervised gaze approaches on extensive datasets (including wild scenes) under both within-dataset and cross-dataset protocols. Jin Li 0057, Wenrui Dai, Bowen Shi 0003, Xiaopeng Zhang 0008, Hongkai Xiong |
ICML | 4 |
| 2023 | Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose EstimationabstractThere has been a recent surge of interest in introducing transformers to 3D human pose estimation (HPE) due to their powerful capabilities in modeling long-term dependencies. However, existing transformer-based methods treat body joints as equally important inputs and ignore the prior knowledge of human skeleton topology in the self-attention mechanism. To tackle this issue, in this paper, we propose a Pose-Oriented Transformer (POT) with uncertainty guided refinement for 3D HPE. Specifically, we first develop novel pose-oriented self-attention mechanism and distance-related position embedding for POT to explicitly exploit the human skeleton topology. The pose-oriented self-attention mechanism explicitly models the topological interactions between body joints, whereas the distance-related position embedding encodes the distance of joints to the root joint to distinguish groups of joints with different difficulties in regression. Furthermore, we present an Uncertainty-Guided Refinement Network (UGRN) to refine pose predictions from POT, especially for the difficult joints, by considering the estimated uncertainty of each joint with uncertainty-guided sampling strategy and self-attention mechanism. Extensive experiments demonstrate that our method significantly outperforms the state-of-the-art methods with reduced model parameters on 3D HPE benchmarks such as Human3.6M and MPI-INF-3DHP. Bowen Shi 0003, Wenrui Dai, Hongwei Zheng 0006, Junni Zou, Hongkai Xiong |
AAAI | 2 |
| 2023 | Adapting Shortcut with Normalizing Flow: An Efficient Tuning Framework for Visual RecognitionabstractPretraining followed by fine-tuning has proven to be effective in visual recognition tasks. However, fine-tuning all parameters can be computationally expensive, particularly for large-scale models. To mitigate the computational and storage demands, recent research has explored Parameter-Efficient Fine-Tuning (PEFT), which focuses on tuning a minimal number of parameters for efficient adaptation. Existing methods, however, fail to analyze the impact of the additional parameters on the model, resulting in an unclear and suboptimal tuning process. In this paper, we introduce a novel and effective PEFT paradigm, named SNF (Shortcut adaptation via Normalization Flow), which utilizes normalizing flows to adjust the shortcut layers. We highlight that layers without Lipschitz constraints can lead to error propagation when adapting to downstream datasets. Since modifying the over-parameterized residual connections in these layers is expensive, we focus on adjusting the cheap yet crucial shortcuts. Moreover, learning new information with few parameters in PEFT can be challenging, and information loss can result in label information degradation. To address this issue, we propose an information-preserving normalizing flow. Experimental results demonstrate the effectiveness of SNF. Specifically, with only 0.036M parameters, SNF surpasses previous approaches on both the FGVC and VTAB-1k benchmarks using ViT/B-16 as the backbone. The code is available at https://github.com/Wang-Yaoming/SNF Bowen Shi 0003, Xiaopeng Zhang 0008, Jin Li 0057, Yuchen Liu 0006, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
CVPR | 2 |
| 2023 | ActionPrompt: Action-Guided 3D Human Pose Estimation With Text and Pose PromptingabstractRecent 2D-to-3D human pose estimation (HPE) utilizes temporal consistency across sequences to alleviate the depth ambiguity problem but ignore the action related prior knowledge hidden in the pose sequence. In this paper, we propose a plug-and-play module named Action Prompt Module (APM) that effectively mines different kinds of action clues for 3D HPE. The highlight is that, the mining scheme of APM can be widely adapted to different frameworks and bring consistent benefits. Specifically, we first present a novel Action-related Text Prompt module (ATP) that directly embeds action labels and transfers the rich language information in the label to the pose sequence. Besides, we further introduce Action-specific Pose Prompt module (APP) to mine the position-aware pose pattern of each action, and exploit the correlation between the mined patterns and input pose sequence for further pose refinement. Experiments show that APM can improve the performance of most video-based 2D-to-3D HPE frameworks by a large margin. Hongwei Zheng 0006, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong |
ICME | 3 |
| 2023 | VioLET: Vision-Language Efficient Tuning with Collaborative Multi-modal GradientsabstractParameter-Efficient Tuning (PET) has emerged as a leading advancement in both Natural Language Processing and Computer Vision, enabling efficient accommodation of downstream tasks without costly fine-tuning. However, most existing PET approaches are limited to uni-modal tuning, even for vision-language models like CLIP. We investigate this limitation and demonstrate that simultaneous tuning of the two modalities in such models leads to multi-modal forgetting and catastrophic performance degradation, particularly when generalizing to new classes. To address this issue, we propose a novel PET approach called VioLET (Vision Language Efficient Tuning) that utilizes collaborative multi-modal gradients to unlock the full potential of both modalities. Specifically, we incorporate an additional visual encoder without learnable parameters and use these two visual encoders to compute the gradients of the context parameters separately. When conflicts arise, we replace the original gradient with an orthogonal gradient. Extensive experiments are conducted on few-shot recognition and unseen class generalization tasks using ResNet-50 or ViT/B-16 as the backbone. VioLET consistently outperforms several state-of-the-art methods on 11 datasets, showcasing its superiority over existing PET approaches. The code is available at https://github.com/Wang-Yaoming/VioLET. Yuchen Liu 0006, Xiaopeng Zhang 0008, Jin Li 0057, Bowen Shi 0003, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ACM Multimedia | 5 |
| 2023 | AiluRus: A Scalable ViT Framework for Dense PredictionabstractVision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, their complexity dramatically increases when handling long token sequences, particularly for dense prediction tasks that require high-resolution input. Notably, dense prediction tasks, such as semantic segmentation or object detection, emphasize more on the contours or shapes of objects, while the texture inside objects is less informative. Motivated by this observation, we propose to apply adaptive resolution for different regions in the image according to their importance. Specifically, at the intermediate layer of the ViT, we select anchors from the token sequence using the proposed spatial-aware density-based clustering algorithm. Tokens that are adjacent to anchors are merged to form low-resolution regions, while others are preserved independently as high-resolution. This strategy could significantly reduce the number of tokens, and the following layers only handle the reduced token sequence for acceleration. At the output end, the resolution of the feature map is recovered by unfolding merged tokens for task prediction. Consequently, we can considerably accelerate ViTs for dense prediction tasks. The proposed method is evaluated across three different datasets and demonstrates promising performance. For instance, "Segmenter ViT-L" can be accelerated by 48\% FPS without fine-tuning, while maintaining the performance. Moreover, our method can also be applied to accelerate fine-tuning. Experiments indicate that we can save 52\% training time while accelerating 2.46$\times$ FPS with only a 0.09\% performance drop. Jin Li 0057, Xiaopeng Zhang 0008, Bowen Shi 0003, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
NeurIPS | 4 |
| 2022 | A Transformer-Based Decoder for Semantic Segmentation with Multi-level Context Mining
Bowen Shi 0003, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001 |
ECCV (28) | 1 |
| 2021 | Hierarchical Graph Networks for 3D Human Pose Estimation
Bowen Shi 0003, Wenrui Dai, Yabo Chen, Junni Zou, Hongkai Xiong |
BMVC | 2 |
| 2020 | Tiny-Hourglassnet: An Efficient Design For 3d Human Pose EstimationabstractExisting deep learning based methods for 3D human pose estimation cannot be deployed on resource-constrained devices. In this paper, we propose a Tiny-HourglassNet for 3D human pose estimation that improves the efficiency of Stacked Hourglass Networks with a guarantee of estimation performance. To be concrete, we develop an efficient Tiny-Hourglass backbone with two different types of ShuffleNet V2 blocks. At the meantime, we present a Simple Estimation Head (SEH) to reduce the resource utility for depth prediction. Furthermore, we develop two feature enhancement modules, namely Feature Enhancement Module (FEM) and Intermediate Enhancement Module (IEM), to improve the feature representation ability of Tiny-HourglassNet without evident increment of resource usage and computational complexity. Experimental results demonstrate that Tiny-HourglassNet achieves equivalent accuracy with a reduction of 80% complexity (FLOPs) in comparison to the baselines. Bowen Shi 0003, Yuhui Xu 0002, Wenrui Dai, Shuai Zhang 0009, Junni Zou, Hongkai Xiong |
ICIP | 1 |
| 2019 | Deep Neural Network-Based Algorithm Approximation via Multivariate Polynomial RegressionabstractMany communication tasks have been formulated as optimization problems that can be solved by iterative algorithms. However, these algorithms are usually computationally intensive. To enable real-time processing of communication algorithms, in this paper, we propose a new deep neural network (DNN) architecture for algorithm approximation. Based on the idea of deep unfolding, we expand the iterative algorithm into a cascade network structure of basic blocks, with each network block corresponding to an iteration. Then, instead of directly approximating each iteration through network layers, we introduce the multivariate polynomial regression (MPR) as a bridge between the iterations and layers for constructing the unfolded network blocks. Specifically, we use a multivariate polynomial to approximate an iteration of the original iterative algorithm, and further construct a shallow neural network with one hidden layer to realize the polynomial, where the order of the polynomial provides theoretical guidance for determining the size of the network and controls the tradeoff between the approximation accuracy and computation speed. As empirical justifications, we apply the proposed DNN architecture in approximating the classic WMMSE algorithm for wireless interference management, showing that the proposed approach can achieve a good approximation accuracy with a much faster computation speed. Chunmiao Liu, Bowen Shi 0003, Junni Zou, Hongkai Xiong |
GLOBECOM | 2 |
| 2019 | ITS-Frame: A Framework for Multi-Aspect Analysis in the Field of Intelligent Transportation SystemsabstractIntelligent transportation systems (ITS) have been developed rapidly over the last few decades because of global urbanization and industrialization. ITS involve a wide range of different technologies and applications such as automatic road enforcement, dynamic traffic light sequence, and as a result, a significant number of scientific papers have been published in the field of ITS. In this paper, we present a useful insight into the development of ITS area by systematically analyzing the publications over the period of 20 years. First, we identify the most cited papers and most impactful authors in the field. Second, in the aspect of topic analysis, we identify some active keywords. To do so, we develop a keyword co-occurrence network to find topics in the ITS field. Finally, for the collaboration pattern analysis, we construct two networks to interpret collaboration patterns, including a co-authorship network, and an author co-keyword network to show the development and research tendency of ITS. Some most interesting findings from our investigation include the following: 1) Besides the USA, China and Europe have begun to play an increasingly significant role in this field and 2)GPS,traffic control, androad safetyshow an upward trend from the analysis of the evolution of ITS research topics, given the rise of new research areas such as autonomous vehicles. Our first-hand investigation and analysis of the literature provides a valuable reference to research activities in the development of ITS field and presents worthy insights on the current status and future technical trends. Xiujuan Xu, Yu Liu 0035, Wei Wang 0077, Xiaowei Zhao 0003, Quan Z. Sheng, Zhe Wang 0007, Bowen Shi 0003 |
IEEE Trans. Intell. Transp. Syst. | 7 |