VLDB 2026 Research / reviewers in the wild / expert
Xiuxiu Bai
dblp:77/8376
· DBLP profile ↗
18ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-8102-1596ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and PrunableabstractQuery-based models are extensively used in 3D object detection tasks, with a wide range of pre-trained checkpoints readily available online. However, despite their popularity, these models often require an excessive number of object queries, far surpassing the actual number of objects to detect. The redundant queries result in unnecessary computational and memory costs. In this paper, we find that not all queries contribute equally -- a significant portion of queries have a much smaller impact compared to others. Based on this observation, we propose an embarrassingly simple approach called Gradually Pruning Queries (GPQ), which prunes queries incrementally based on their classification scores. A key advantage of GPQ is that it requires no additional learnable parameters. It is straightforward to implement in any query-based method, as it can be seamlessly integrated as a fine-tuning step using an existing checkpoint after training. With GPQ, users can easily generate multiple models with fewer queries, starting from a checkpoint with an excessive number of queries. Experiments on various advanced 3D detectors show that GPQ effectively reduces redundant queries while maintaining performance. Using our method, model inference on desktop GPUs can be accelerated by up to 1.35x. Moreover, after deployment on edge devices, it achieves up to a 67.86% reduction in FLOPs and a 65.16% decrease in inference time. Lizhen Xu, Wenzhao Qiu, Shanmin Pang, Xiuxiu Bai, Jianru Xue |
AAAI | 5 |
| 2026 | GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
Ruifan Chu, Anbang Wang, Xiuxiu Bai, Xiaoshe Dong |
PAKDD (4) | 3 |
| 2026 | Mapo: Performance model driven GPU memory access code optimization
Xiaoshe Dong, Junkai Cao, Ruifan Chu, Ziheng Wang 0002, Qiang Wang 0062, Xiuxiu Bai |
Future Gener. Comput. Syst. | 7 |
| 2025 | Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningabstractQuery-based methods with dense features have demonstrated remarkable success in 3D object detection tasks. However, the computational demands of these models, particularly with large image sizes and multiple transformer layers, pose significant challenges for efficient running on edge devices. Existing pruning and distillation methods either need retraining or are designed for ViT models, which are hard to migrate to 3D detectors. To address this issue, we propose a zero-shot runtime pruning method for transformer decoders in 3D object detection models. The method, termed tgGBC (trim keys gradually Guided By Classification scores), systematically trims keys in transformer modules based on their importance. We expand the classification score to multiply it with the attention map to get the importance score of each key and then prune certain keys after each transformer layer according to their importance scores. Our method achieves a 1.99x speedup in the transformer decoder of the latest ToC3D model, with only a minimal performance loss of less than 1%. Interestingly, for certain models, our method even enhances their performance. Moreover, we deploy 3D detectors with tgGBC on an edge device, further validating the effectiveness of our method. The code can be found at https://github.com/iseri27/tg_gbc. Lizhen Xu, Xiuxiu Bai, Xiaojun Jia, Jianwu Fang, Shanmin Pang |
ICCV | 2 |
| 2025 | Refining CLIP's Spatial Awareness: A Visual-Centric PerspectiveabstractContrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches have introduced Region-Language Alignment (RLA) to enhance CLIP's performance in dense multimodal tasks by aligning regional visual representations with corresponding text inputs. However, we find that CLIP ViTs fine-tuned with RLA suffer from notable loss in spatial awareness, which is crucial for dense prediction tasks. To address this, we propose the Spatial Correlation Distillation (SCD) framework, which preserves CLIP's inherent spatial structure and mitigates above degradation. To further enhance spatial correlations, we introduce a lightweight Refiner that extracts refined correlations directly from CLIP before feeding them into SCD, based on an intriguring finding that CLIP naturally capture high-quality dense features. Together, these components form a robust distillation framework that enables CLIP ViTs to integrate both visual-language and visual-centric improvements, achieving state-of-the-art results across various open-vocabulary dense prediction benchmarks. Congpei Qiu, Yanhao Wu, Wei Ke 0003, Xiuxiu Bai, Tong Zhang 0023 |
ICLR | 4 |
| 2025 | Positional Prompt Tuning for Efficient 3D Representation Learning
Shaochen Zhang, Zekun Qi, Runpei Dong, Xiuxiu Bai, Xing Wei 0001 |
ACM Multimedia | 4 |
| 2024 | CONet: Crowd and occlusion-aware network for occluded human pose estimation
Xiuxiu Bai, Xing Wei 0001, Zengying Wang, Miao Zhang 0013 |
Neural Networks | 1 |
| 2023 | ProMask: Probability mask representation for skeleton detection
Xiuxiu Bai, Lele Ye |
Neural Networks | 1 |
| 2023 | Tensor-Based Incomplete Multi-View Clustering With Low-Rank Data Reconstruction and Consistency GuidanceabstractWe propose a new approach, called Tensor-based Incomplete Multi-view Clustering with Low-rank data Reconstruction and Consistency guidance (TIMC-RC), to perform clustering on multi-view data with missing views. Existing methods usually leverage original incomplete data to explore the partial correlations among multiple views, and do not make sufficient use of both consistent and complementary information across views. To explore the full information of missing and available views, TIMC-RC introduces low-rank data reconstruction and consistency view establishment. Specifically, 1) it adopts a low-rank constraint to reconstruct data representations so as to reduce the negative effect of missing data and obtain more reasonable data representations. 2) It builds a new consistency view by self-representation matrices and therefore explores the consistent correlation of different views. 3) It formalizes view-specific self-representation matrices and the consistent matrix as a tensor and utilizes the tensor singular value decomposition-based nuclear norm to enhance the consistency and complementarity of multi-view representations. Experiments conducted on eight benchmarks verify the effectiveness and advancement of the proposed TIMC-RC. Wenyu Hao, Shanmin Pang, Xiuxiu Bai, Jianru Xue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | SalS-GAN: Spatially-Adaptive Latent Space in StyleGAN for Real Image EmbeddingabstractMany GAN inversion methods have emerged to embed a given real image into the latent space of GAN for real image editing. These methods usually use a latent space composed of a series of one-dimensional vectors as an optimization space to reconstruct real images such as W+ latent space. However, the reconstructed image of these methods is usually difficult to maintain the rich detailed information in the real image. How to better preserve details in the real image is still a challenge. To solve this problem, we propose a spatially-adaptive latent space, called SA latent space, and adopt it as the optimization latent space in GAN inversion task. In particular, we use the affine transformation parameters of each convolutional layer in the generator to form the SA latent space and change affine transformation parameters from a one-dimensional vector to a spatially-adaptive three-dimensional tensor. With the more expressive latent space, we can better reconstruct the details of the real image. Extensive experiments suggest that the image reconstruction quality can be significantly improved while maintaining the semantic disentanglement ability of latent code. The code is available at https://github.com/zhang-lingyun/SalS-GAN. Xiuxiu Bai |
ACM Multimedia | 2 |
| 2021 | KEEP: Secure and Efficient Communication for Distributed IoT DevicesabstractSecurity over mobile Internet-of-Things (IoT) devices is critical due to the open nature of distributed wireless communication. To efficiently establish a secure connection between two communication parties, a fast mobile key extraction protocol, KEEP, is proposed. KEEP fastly generates similar bit sequences from two communication parties’ measurements of channel-state information (CSI) of different subcarriers. Then, a distributed “verification-recombination” mechanism is introduced to generate the same encryption key from bit sequences without the public-key authentication, digital signature, or key distribution center of the other party. We implemented real-world experiments using commercial off-the-shelf 802.11n devices to evaluate the performance of KEEP in various scenarios. Theoretical analysis and experimental verification show that KEEP is more secure, effective, and reliable than the state-of-the-art methods. Wei Xi 0003, Meichen Duan, Xiuxiu Bai, Kun Zhao 0002, Lufeng Mo, Jizhong Zhao |
IEEE Internet Things J. | 3 |
| 2021 | PFAN++: Bi-Directional Image-Text Retrieval With Position Focused Attention NetworkabstractBi-directional image-text retrieval and matching attract much attention recently. This cross-domain task demands a fine understanding of both modalities for learning a measure of different modality data. In this paper, we propose a novel position focused attention network to investigate the relation between the visual and the textual views. This work integrates the prior object position to enhance the visual-text joint-embedding learning. The image is first split into blocks, which are treated as the basic position cells, and the position of an image region is inferred. Then, we propose a position attention to model the relations between the image region and position cells. Finally, we generate a valuable position feature to further enhance the region expression and model a more reliable relationship between the visual image and the textual sentence. Experiments on the popular datasets Flickr30K and MS-COCO show the effectiveness of the proposed method. Besides the public datasets, we also conduct experiments on our collected practical large-scale news dataset (Tencent-News) to validate the practical application value of the proposed method. As far as we know, this is the first attempt to test the performance on the practical application. Our method achieves the competitive performance on all of these three datasets. Yaxiong Wang, Xiuxiu Bai, Xueming Qian, Lin Ma 0002 |
IEEE Trans. Multim. | 3 |
| 2020 | On the robustness of skeleton detection against adversarial attacks
Xiuxiu Bai |
Neural Networks | 1 |
| 2020 | Skeleton Filter: A Self-Symmetric Filter for Skeletonization in Noisy Text ImagesabstractRobustly computing the skeletons of objects in natural images is difficult due to the large variations in shape boundaries and the large amount of noise in the images. Inspired by recent findings in neuroscience, we propose the Skeleton Filter, which is a novel model for skeleton extraction from natural images. The Skeleton Filter consists of a pair of oppositely oriented Gabor-like filters; by applying the Skeleton Filter in various orientations to an image at multiple resolutions and fusing the results, our system can robustly extract the skeleton even under highly noisy conditions. We evaluate the performance of our approach using challenging noisy text datasets and demonstrate that our pipeline realizes state-of-the-art performance for extracting the text skeleton. Moreover, the presence of Gabor filters in the human visual system and the simple architecture of the Skeleton Filter can help explain the strong capabilities of humans in perceiving skeletons of objects, even under dramatically noisy conditions. Xiuxiu Bai, Lele Ye, Jihua Zhu, Li Zhu 0003, Taku Komura |
IEEE Trans. Image Process. | 1 |
| 2016 | Registration of Point Clouds Based on the Ratio of Bidirectional DistancesabstractDespite the fact that original Iterative Closest Point(ICP) algorithm has been widely used for registration, itcannot tackle the problem when two point clouds are par-tially overlapping. Accordingly, this paper proposes a ro-bust approach for the registration of partially overlappingpoint clouds. Given two initially posed clouds, it firstlybuilds up bilateral correspondence and computes bidirec-tional distances for each point in the data shape. Based onthe ratio of bidirectional distances, the exponential functionis selected and utilized to calculate the probability value,which can indicate whether the point pair belongs to theoverlapping part or not. Subsequently, the probability val-ue can be embedded into the least square function for reg-istration of partially overlapping point clouds and a novelvariant of ICP algorithm is presented to obtain the optimalrigid transformation. The proposed approach can achievegood registration of point clouds, even when their overlappercentage is low. Experimental results tested on public da-ta sets illustrate its superiority over previous approaches onrobustness. Jihua Zhu, Di Wang 0006, Xiuxiu Bai, Huimin Lu 0001, Congcong Jin, Zhongyu Li 0002 |
3DV | 3 |
| 2016 | Improving the Reliability of the Operating System Inside a VMabstractVirtualization technology can provide reusability and strong isolation between different virtual machines (VMs). However, there is no effective isolation mechanism inside a VM to solve an operating system's reliability problems, including driver faults. This paper describes Chariot, an architecture that provides effective and transparent driver isolation inside the VM, achieves fine-grained driver isolation and retains the reusability advantage of virtualization technology. First, Chariot transparently monitors an isolated driver with monitoring wrappers, and establishes an access control table (ACT) in a timely manner that records the driver write permissions. Secondly, Chariot protects the shadow page table of the VM (where the driver resides) in due time to capture its write operations. Next, the ACT examines the correctness of the write operations. Finally, if an illegal write operation is detected, Chariot recovers the faulty driver and prevents the spread of driver faults in the VM. The experimental results show that Chariot effectively isolates more than 90% of injected faults (with performance losses of |$<$|20% in most benchmarks) and effectively improves the reliability of the VM. In addition, Chariot can be easily extended to isolate new drivers and ported to other versions of OSs in the virtualization environment. Hao Zheng 0004, Xiaoshe Dong, Zhengdong Zhu, Baoke Chen, Xiuxiu Bai, Xingjun Zhang, Endong Wang |
Comput. J. | 5 |
| 2015 | Edge Propagation KD-Trees: Computing Approximate Nearest Neighbor FieldsabstractPropagation-assisted kd-tree is a state-of-the-art method for computing approximate nearest neighbor (ANN) fields. In this method, each query patch needs descending search in the kd-tree and propagation search in the nearby patches. We observed that the query patches in the edge region need descending search, while other query patches only need propagation search. This can be an opportunity to save plenty of search time. In this letter, we propose edge propagation kd-trees to quickly compute ANNs. Our method can distinguish between edge patches and propagation patches in choosing the proper search. Experiments on public data set VidPairs show that our search method is 2-3 times faster than the propagation-assisted kd-tree search method at nearly the same accuracy. Xiuxiu Bai, Xiaoshe Dong, Yuanqi Su |
IEEE Signal Process. Lett. | 1 |
| 2015 | A scalability prediction approach for multi-threaded applications on manycore processors
Xiuxiu Bai, Endong Wang, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 1 |