VLDB 2026 Research / reviewers in the wild / expert
Ruining Yang
dblp:256/3645
· DBLP profile ↗
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data ServerabstractWe explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usually exhibits distinct modes (i.e., semantic clusters representing data distribution), if the training set does not contain these target modes, the model performance would be compromised. While prior existing works improve algorithms iteratively, our research explores the often-overlooked potential of optimizing the structure of the data server. Inspired by the hierarchical nature of web search engines, we introduce a hierarchical data server, together with a bipartite mode matching algorithm (BMM) to align source and target modes. For each target mode, we look in the server data tree for the best mode match, which might be large or small in size. Through bipartite matching, we aim for all target modes to be optimally matched with source modes in a one-on-one fashion. Compared with existing training set search algorithms, we show that the matched server modes constitute training sets that have consistently smaller domain gaps with the target domain across object re-identification (re-ID) and detection tasks. Consequently, models trained on our searched training sets have higher accuracy than those trained otherwise. BMM allows data-centric unsupervised domain adaptation (UDA) orthogonal to existing model-centric UDA methods. By combining the BMM with existing UDA methods like pseudo-labeling, further improvement is observed. Yue Yao 0001, Ruining Yang, Tom Gedeon |
AAAI | 2 |
| 2026 | Infrastructure as Compromise: Abusing Residual Trust in Infrastructure as Code ToolsabstractInfrastructure as Code (IaC) has transformed the way developers deploy and manage their infrastructure, enabling automated, reproducible, and version-controlled builds. At the same time, developer errors in IaC configurations can result in thousands of identically-vulnerable servers. Ruining Yang, Narong Chaiwut, Nick Nikiforakis |
CODASPY | 1 |
| 2025 | ObfusLM: Privacy-preserving Language Model Service against Embedding Inversion AttacksabstractYu Lin, Ruining Yang, Yunlong Mao, Qizhi Zhang, Jue Hong, Quanwei Cai, Ye Wu, Huiqi Liu, Zhiyu Chen, Bing Duan, Sheng Zhong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruining Yang, Yunlong Mao, Qizhi Zhang 0007, Jue Hong, Quanwei Cai 0003, Huiqi Liu, Bing Duan, Sheng Zhong 0002 |
ACL (1) | 2 |
| 2025 | Unsupervised Search for Ethnic Minorities' Medical Segmentation Training SetabstractThis paper investigates the critical issue of dataset bias in medical imaging, with a particular emphasis on racial disparities caused by uneven population distribution in dataset collection. Our analysis reveals that medical segmentation datasets are significantly biased, primarily influenced by the demographic composition of their collection sites. For instance, Scanning Laser Ophthalmoscopy (SLO) fundus datasets collected in the United States predominantly feature images of White individuals, with minority racial groups underrepresented. This imbalance can result in biased model performance and inequitable clinical outcomes, particularly for minority populations. To address this challenge, we propose a novel training set search strategy aimed at reducing these biases by focusing on underrepresented racial groups. Our approach utilizes existing datasets and employs a simple greedy algorithm to identify source images that closely match the target domain distribution. By selecting training data that aligns more closely with the characteristics of minority populations, our strategy improves the accuracy of medical segmentation models on specific minorities, i.e., Black. Our experimental results demonstrate the effectiveness of this approach in mitigating bias. We also discuss the broader societal implications, highlighting how addressing these disparities can contribute to more equitable healthcare outcomes. Our code is available at https://github.com/yorkeyao/SnP. Yue Yao 0001, Ruining Yang, Ashu Gupta, Tom Gedeon |
ICASSP | 3 |
| 2025 | The Effect of Domain Terms on Password SecurityabstractThe predominant authentication method still relies on usernames and passwords. To enhance memorability, domain terms may have been opted to include as part of passwords. However, there is little analysis of the extent to which such practice affects password security, so there is a lack of guidance on how users use domain terms on websites with different domain characteristics. To address the problem, we propose a novel approach to analyze the security effect of using domain terms in passwords. The methodology primarily consists of three stages. First, we utilize Web crawlers to harvest domain vocabularies, subsequently leveraging the TextRank algorithm to rank their importance. Second, we propose an algorithm for constructing a simulated domain-specific password dataset by replacing password elements with domain terms. Third, password guessing experiments are done on the dataset using PCFG (Probabilistic Context-Free Grammar) and the Markov model to evaluate the impact of domain terms on password security. The experimental results indicate that, for systems without clear domain, 20% domain terms replacement in the test set can reduce the cracking rate by up to 5.45%. In contrast, for domain-specific systems, 20% domain terms replacement in the training set can increase the cracking rate by 6.45%. These findings provide practical guidance on the application of domain knowledge in password creation for different types of systems. In summary, this study offers a novel perspective for exploring the security implications of passwords influenced by specific domains. Yubing Bao, Jianping Zeng 0002, Jirui Yang, Ruining Yang, Zhihui Lu 0002 |
ACM Trans. Priv. Secur. | 4 |
| 2024 | SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D DetectionabstractLidar-based 3D Detection is one of the significant components of Autonomous Driving. However, current methods over-focus on improving the performance of 3D Lidar perception, which causes the architecture of networks becoming complicated and hard to deploy. Thus, the methods are difficult to apply in Autonomous Driving for real-time processing. In this paper, we propose a high-efficiency network, SwiftPillars, which includes Swift Pillar Encoder (SPE) and Multi-scale Aggregation Decoder (MAD). The SPE is constructed by a concise Dual-attention Module with lightweight operators. The Dual-attention Module utilizes feature pooling, matrix multiplication, etc. to speed up point-wise and channel-wise attention extraction and fusion. The MAD interconnects multiple scale features extracted by SPE with minimal computational cost to leverage performance. In our experiments, our proposal accomplishes 61.3% NDS and 53.2% mAP in nuScenes dataset. In addition, we evaluate inference time on several platforms (P4, T4, A2, MLU370, RTX3080), where SwiftPillars achieves up to 13.3ms (75FPS) on NVIDIA Tesla T4. Compared with PointPillars, SwiftPillars is on average 26.58% faster in inference speed with equivalent GPUs and a higher mAP of approximately 3.2% in the nuScenes dataset. Xin Jin 0014, Ruining Yang, Fei Hui, Wei Wu 0021 |
AAAI | 4 |
| 2024 | An End-to-End SoC for Brain-Inspired CNN-SNN Hybrid ApplicationsabstractInspired by the brain, Spiking Neural Network (SNN) applies temporally sparse spiking communication to gain more bio-mimetic and highly energy efficient computing. The current mainstream platforms for SNN applications are typically the combination of Host+FPGA+Chip Array, which requires an efficient host to preprocess and encode data. It’s not suitable for end-to-end tasks in edge due to its high system power consumption of host and non-negligible high latency of protocol conversion on FPGA. In addition, Convolutional Neural Network (CNN), exhibits strong feature extraction capabilities. Like the brain's visual system, a hierarchical CNN-SNN hybrid network, in which SNN can make use of CNN’s feature extraction capabilities during encoding, can achieve better performance. In this study, we design a 64Neural-Core Array and integrate it with a CNN encoder and a low-power RISC-V CPU within a System-on-Chip (SoC) to enable comprehensive end-to-end hybrid network application support. The proposed heterogeneous SoC is implemented on a Virtex UltraScale+ XCVU9P FPGA, featuring 32.8K neurons, 37.7M synapses and 578GOPS/s peak performance. It processes MNIST classification with a peak throughput of 2022 images per second at frequency of 250MHz. This design gains a balance between high throughput and recognition accuracy simultaneously. Zhaotong Zhang, Yi Zhong 0002, Yingying Cui, Yawei Ding, Yukun Xue, Qibin Li, Ruining Yang, Jian Cao 0002, Yuan Wang 0001 |
ISCAS | 7 |
| 2024 | Learning to disentangle and fuse for fine-grained multi-modality ship image retrieval
Pingliang Xu, Linzhou Huang, Ruining Yang |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | An Interpretable Fusion Siamese Network for Multi-Modality Remote Sensing Ship Image RetrievalabstractWith the increasing number of remote sensing ship images, it’s vitally important to search for the ship objects that users are interested in from the remote sensing image big data. The existing works are focused on single-modality remote sensing ship image processing. But there is no method of retrieving ship images from different remote sensing image modalities. Besides, the models in existing works only output a predicted result without offering a reasonable explanation. In this work, we propose an interpretable fusion siamese network (IFSN) for addressing the multi-modality remote sensing ship image retrieval (MRSSIR). 1) An interpretable attention feature representation module is proposed to generate multiple attention maps and aggregate the filters of the last convolutional layer, which can focus on the ship’s discriminative parts and make each divided convolutional filter group express specific visual information. 2) A multi-modality correlation learning module is proposed to overcome the intra-modality and inter-modality variations by designing some constraints. 3) A discriminative region mining module is proposed to exhaustively explore all the ship’s discriminative parts available for decision-making of the proposed network. We construct a multi-modality remote sensing ship images dataset (MRSSID) to evaluate the performance of the proposed IFSN. The experimental results exhibit that our IFSN outperforms the existing methods in retrieval accuracy and provides reasonable and intuitive interpretations for the retrieval results. Linzhou Huang, Ruining Yang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Development of Perceptual Training Software for Realizing High Variability Training Paradigm and Self Adaptive Training Paradigm
Ruining Yang, Hiroaki Nanjo, Masatake Dantsuji |
PACLIC | 1 |