Rukai Wei

dblp:336/3910 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-1164-6360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Arbiter: Towards joint and fine-grained index and partition tuning in analytical databases
Rukai Wei, Hua Wang 0008, Zhaorui Ding, Zhongcong Mo, Ke Zhou 0001, Yu Liu 0040
Inf. Process. Manag.2
2026 PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002
Pattern Recognit.5
2025 TopTune: Tailored Optimization for Categorical and Continuous Knobs Towards Accelerated and Improved Database Performance Tuning
abstract
Using a machine learning (ML) model as a core component in database knob tuning has demonstrated remarkable advancements in recent years. However, a model that optimizes both categorical and continuous values in the same way may not guarantee efficiency and effectiveness in knob tuning. This is due to the fact that the usual assumption of a differentiable input space for efficient exploration of continuous spaces does not hold true in categorical spaces. Moreover, the inherent complexity of interdependences among knobs and the high-dimensionality of the configuration space compound the challenges of tuning. In this paper, we propose TopTune, which employs tailored optimization for continuous and categorical knobs, to achieve accelerated tuning efficiency and improved tuning performance. Specifically, we decompose the configuration space into two orthogonal subspaces: categorical and continuous spaces. Subsequently, we employ Bayesian optimization models, i.e., SMAC and GP to explore the categorical and continuous subspaces, respectively. These two models will alternately explore the two spaces with the proposed communication mechanism to ensure TopTune can capture the dependence between continuous and categorical knobs. Furthermore, to balance efficiency and accuracy, we utilize a knob-dimensional projection strategy to reduce the exploration domain by embedding the high-dimension configuration space into a lower-dimensional proxy space. In addition, we implement batch Bayesian optimization technology, which enables parallel knob evaluation while balancing exploration and exploitation. We evaluate TopTune under different benchmarks (SYSBENCH, TPC-C, and JOB), metrics (throughput and latency), and DBMSs (MySQL and Dameng). Extensive experiments demonstrate that TopTune identifies better configurations in up to approximately 12.2× less time while achieving a 10.7% improvement in throughput compared to state-of-the-art methods.
Rukai Wei, Yu Liu 0040, Yufeng Hou, Heng Cui, Ke Zhou 0001
ICDE1
2025 Graph Contrastive-and-Reconstructive Hashing for Unsupervised Cross-Modal Retrieval
abstract
Abstract Hashing-based unsupervised cross-modal retrieval has gained significant attention in the big data management community due to its low storage overhead and rapid retrieval speed. However, current methods often lack effective alignment strategies to reduce the modality gap. They also fail to explore the latent structural information of the training data for accurate relationship learning, resulting in sub-optimal cross-modal retrieval performance. To tackle these challenges, we propose a novel unsupervised cross-modal hashing method called G raph C ontrastive-and- R econstructive H ashing ( GCRH ). Specifically, GCRH first performs global graph contrastive learning , which involves both intra-modal and inter-modal pairs. This facilitates the learning of more discriminative hash codes through intra-modal discrimination and inter-modal alignment objectives. To further bridge the modality gap, GCRH conducts local graph reconstruction using GCN-based decoders to reconstruct the original features of one modality from the hash codes of another. The integration of contrastive-and-reconstructive learning with graph structural information enables GCRH to generate high-quality hash codes that are both well-aligned and discriminative. Extensive experiments on three benchmark datasets substantiate the superior cross-modal retrieval performance of GCRH .
Rukai Wei, Yu Liu 0040, Heng Cui, Yanzhao Xie, Ke Zhou 0001
Data Sci. Eng.1
2025 Angle Metric Learning for Discriminative Features on Vehicle Re-Identification
abstract
ABSTRACT Vehicle re‐identification (Re‐ID) facilitates the recognition and distinction of vehicles based on their visual characteristics in images or videos. However, accurately identifying a vehicle poses great challenges due to (i) the pronounced intra‐instance variations encountered under varying lighting conditions such as day and night and (ii) the subtle inter‐instance differences observed among similar vehicles. To address these challenges, the authors propose A ngle M etric learning for D iscriminative F eatures on vehicle Re‐ID (termed as AMDF), which aims to maximise the variance between visual features of different classes while minimising the variance within the same class. AMDF comprehensively measures the angle and distance discrepancies between features. First, to mitigate the impact of lighting conditions on intra‐class variation, the authors employ CycleGAN to generate images that simulate consistent lighting (either day or night), thereby standardising the conditions for distance measurement. Second, Swin Transformer was integrated to help generate more detailed features. At last, a novel angle metric loss based on cosine distance is proposed, which organically integrates angular metric and 2‐norm metric, effectively maximising the decision boundary in angular space. Extensive experimental evaluations on three public datasets including VERI‐776, VERI‐Wild, and VEHICLEID, indicate that the method achieves state‐of‐the‐art performance. The code of this project is released at https://github.com/ZnCu‐0906/AMDF .
Yutong Xie 0009, Shuoqi Zhang, Lide Guo, Rukai Wei, Yanzhao Xie, Yangtao Wang, Maobin Tang, Lisheng Fan
IET Comput. Vis.5
2024 Contrastive masked auto-encoders based self-supervised hashing for 2D image and 3D point cloud cross-modal retrieval
abstract
Implementing cross-modal hashing between 2D images and 3D point-cloud data is a growing concern in real-world retrieval systems. Simply applying existing cross-modal approaches to this new task fails to adequately capture latent multi-modal semantics and effectively bridge the modality gap between 2D and 3D. To address these issues without relying on hand-crafted labels, we propose contrastive masked autoencoders based self-supervised hashing (CMAH) for retrieval between images and point-cloud data. We start by contrasting 2D-3D pairs and explicitly constraining them into a joint Hamming space. This contrastive learning process ensures robust discriminability for the generated hash codes and effectively reduces the modality gap. Moreover, we utilize multi-modal auto-encoders to enhance the model’s understanding of multi-modal semantics. By completing the masked image/point-cloud data modeling task, the model is encouraged to capture more localized clues. In addition, the proposed multi-modal fusion block facilitates fine-grained interactions among different modalities. Extensive experiments on three public datasets demonstrate that the proposed CMAH significantly outperforms all baseline methods.
Rukai Wei, Heng Cui, Yu Liu 0040, Yanzhao Xie, Yufeng Hou, Ke Zhou 0001
ICME1
2024 EfficientMatting: Bilateral Matting Network for Real-Time Human Matting
Rongsheng Luo, Rukai Wei, Huaxin Zhang, Ming Tian, Changxin Gao, Nong Sang
PRCV (12)2
2024 Exploring Hierarchical Information in Hyperbolic Space for Self-Supervised Image Hashing
abstract
In real-world datasets, visually related images often form clusters, and these clusters can be further grouped into larger categories with more general semantics. These inherent hierarchical structures can help capture the underlying distribution of data, making it easier to learn robust hash codes that lead to better retrieval performance. However, existing methods fail to make use of this hierarchical information, which in turn prevents the accurate preservation of relationships between data points in the learned hash codes, resulting in suboptimal performance. In this paper, our focus is on applying visual hierarchical information to self-supervised hash learning and addressing three key challenges, including the construction, embedding, and exploitation of visual hierarchies. We propose a new self-supervised hashing method named Hierarchical Hyperbolic Contrastive Hashing (HHCH), making breakthroughs in three aspects. First, we propose to embed continuous hash codes into hyperbolic space for accurate semantic expression since embedding hierarchies in the hyperbolic space generates less distortion than in the hyper-sphere or Euclidean space. Second, we update the K-Means algorithm to make it run in the hyperbolic space. The proposed hierarchical hyperbolic K-Means algorithm can achieve the adaptive construction of hierarchical semantic structures. Last but not least, to exploit the hierarchical semantic structures in hyperbolic space, we propose the hierarchical contrastive learning algorithm, including hierarchical instance-wise and hierarchical prototype-wise contrastive learning. Extensive experiments on four benchmark datasets demonstrate that the proposed method outperforms state-of-the-art self-supervised hashing methods. Our codes are released at https://github.com/HUST-IDSM-AI/HHCH.git.
Rukai Wei, Yu Liu 0040, Jingkuan Song, Yanzhao Xie, Ke Zhou 0001
IEEE Trans. Image Process.1
2024 Supervised Hierarchical Online Hashing for Cross-modal Retrieval
abstract
Online cross-modal hashing has gained attention for its adaptability in processing streaming data. However, existing methods only define the hard similarity between data using labels. This results in poor retrieval performance, as they fail to exploit the semantic structure information of labels and miss the high-quality hash codes guided by the hierarchical relevance between labels. In addition, they ignore the bit-flipping problem, which leads to sub-optimal cross-modal retrieval performance. To address these issues, we propose Supervised Hierarchical Online Hashing (SHOH) for cross-modal retrieval. Our approach acquires hierarchical similarity via cross-layer affiliation of labels and explores its application to online hashing. We design a hierarchical similarity learning method in the online learning framework, which includes virtual center learning and hierarchical similarity embedding. Labels with soft similarity bridge the label hierarchy and cross-modal hash embedding. Furthermore, we propose a Weighted Retrieval Strategy (WRS) to mitigate the impact caused by bit-flipping errors. Extensive experiments and verification on hierarchical and non-hierarchical datasets demonstrate that SHOH preserves accurate inter-class distances and achieves performance improvements compared to state-of-the-art methods. The source code is available at https://github.com/HUST-IDSM-AI/SHOH .
Yu Liu 0040, Rukai Wei, Ke Zhou 0001, Kun Long
ACM Trans. Multim. Comput. Commun. Appl.3
2023 CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video Hashing
abstract
Compressing videos into binary codes can improve retrieval speed and reduce storage overhead. However, learning accurate hash codes for video retrieval can be challenging due to high local redundancy and complex global dependencies between video frames, especially in the absence of labels. Existing self-supervised video hashing methods have been effective in designing expressive temporal encoders, but have not fully utilized the temporal dynamics and spatial appearance of videos due to less challenging and unreliable learning tasks. To address these challenges, we begin by utilizing the contrastive learning task to capture global spatio-temporal information of videos for hashing. With the aid of our designed augmentation strategies, which focus on spatial and temporal variations to create positive pairs, the learning framework can generate hash codes that are invariant to motion, scale, and viewpoint. Furthermore, we incorporate two collaborative learning tasks, i.e., frame order verification and scene change regularization, to capture local spatio-temporal details within video frames, thereby enhancing the perception of temporal structure and the modeling of spatio-temporal relationships. Our proposed Contrastive Hashing with Global-Local Spatio-temporal Ibnformation (CHAIN) outperforms state-of-the-art self-supervised video hashing methods on four video benchmark datasets. Our codes will be released.
Rukai Wei, Yu Liu 0040, Jingkuan Song, Heng Cui, Yanzhao Xie, Ke Zhou 0001
ACM Multimedia1
2023 A hash centroid construction method with Swin transformer for multi-label image retrieval
Yanzhao Xie, Yangtao Wang, Rukai Wei, Yu Liu 0040, Ke Zhou 0001, Lisheng Fan
Neural Comput. Appl.3
2023 Deep debiased contrastive hashing
Rukai Wei, Yu Liu 0040, Jingkuan Song, Yanzhao Xie, Ke Zhou 0001
Pattern Recognit.1
2023 Label-Affinity Self-Adaptive Central Similarity Hashing for Image Retrieval
abstract
Due to the usage of global similarity, the hashing methods based on predefined hash centers have achieved more accurate retrieval results than the pairwise/triplet-based methods. Nevertheless, the fixed hash centers lack the perception of data distribution and are limited by the pre-determined Hadamard matrix, which consider neither the label semantic information nor the object scale size, resulting in sub-optimal retrieval performance and weak generalization ability. In this paper, we (1) adopt the label semantic information to generate self-adaptive hash centers and (2) propose the label-affinity coefficient (lac) that considers the scale size of each label/object appearing in the given image to calculate the real hash centroid for this image. Based on this, we proposeLabel-affinity Self-adaptive Central Similarity Hashing (LSCSH)for image retrieval. LSCSH consists of a hash code generator module and a hash center adapter module. First, we obtain the label word vector (i.e., the word vector representation of each class label) via the Word2Vector technique to generate and update the hash centers that adapt to the distribution of both label word vectors and generated hash codes. Second, we learnlacto indicate the dominance of different labels corresponding to objects in each given image, which considers the unequal scales of each object (corresponding to a label) to calculate a more accurate hash centroid for each image. Last but not least, we design an asynchronous learning mechanism to enable each hash code and its corresponding hash centroid to adapt to each other dynamically. We conduct extensive experiments on 5 image datasets including CIFAR-10, ImageNet, VOC2012, MS-COCO and NUS-WIDE. The experimental results demonstrate that LSCSH can achieve the state-of-the-art visual retrieval performance on both single-label and multi-label image datasets. The code of this work is released at:https://github.com/lzHZWZ/LSCSH_sourcecode.git.
Yanzhao Xie, Rukai Wei, Jingkuan Song, Yu Liu 0040, Yangtao Wang, Ke Zhou 0001
IEEE Trans. Multim.2