VLDB 2026 Research / reviewers in the wild / expert
Ruoyi Zhang
dblp:226/1153
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly DetectionabstractThe increasing complexity of industrial anomaly detection (IAD) has positioned multimodal detection methods as a focal area of machine vision research. However, dedicated multimodal datasets specifically tailored for IAD remain limited. Pioneering datasets like MVTec 3D have laid essential groundwork in multimodal IAD by incorporating RGB+3D data, but still face challenges in bridging the gap with real industrial environments due to limitations in scale and resolution. To address these challenges, we introduce Real-IAD D3, a high-precision multimodal dataset that uniquely incorporates an additional pseudo-3D modality generated through photometric stereo, alongside high-resolution RGB images and micrometer-level 3D point clouds. Real-IAD D3features finer defects, diverse anomalies, and greater scale across 20 categories, providing a challenging benchmark for multimodal IAD Additionally, we introduce an effective approach that integrates RGB, point cloud, and pseudo-3D depth information to leverage the complementary strengths of each modality, enhancing detection performance. Our experiments highlight the importance of these modalities in boosting detection robustness and overall IAD performance. The dataset and code are publicly accessible for research purposes at https://realiad4ad.github.io/Real-IAD_D3. Wenbing Zhu, Ziqing Zhou, Chengjie Wang 0001, Yurui Pan, Ruoyi Zhang, Zhuhao Chen, Linjie Cheng, Bin-Bin Gao, Jiangning Zhang, Zhenye Gan, Yuxie Wang, Shuguang Qian, Mingmin Chi, Lizhuang Ma |
CVPR | 6 |
| 2025 | MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect LabelingabstractAcquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning. Ruoyi Zhang, Jiatong Shi |
INTERSPEECH | 2 |
| 2025 | Mamba-GIE: A visual state space models-based generalized image extrapolation method via dual-level adaptive feature fusion
Ruoyi Zhang, Shuyi Qu, Jun Wang 0078, Jinye Peng 0001 |
Expert Syst. Appl. | 1 |
| 2024 | FedVisual: Heterogeneity-Aware Model Aggregation for Federated Learning in Visual-Based Vehicular CrowdsensingabstractWith the advancement of assisted and autonomous driving technologies, vehicles are being outfitted with an ever-increasing number of sensors. Among these, visible light sensors, or dash-cameras, produce visual data rich in information. Analyzing this visual data through crowdsensing allows for low-cost and timely perception of urban road conditions, such as identifying dangerous driving behaviors and locating parking spaces. However, uploading such massive visual data to the cloud for centralized processing can lead to significant bandwidth challenges and also raise privacy concerns among vehicle owners. Federated learning (FL), in which vehicles serve as both data generators and computing nodes, presents a promising solution to address these challenges. Nevertheless, urban roads are complex and vehicles in different locations encounter completely different scenes, resulting in non-independently and identically distributed (non-i.i.d.) characteristics. Additionally, the diversity in dash-camera and onboard computation resources may lead to differences in the performance of locally trained models. Indiscriminate aggregating of local models from all vehicles can potentially degrade the global model’s performance. To overcome these challenges, we introduce FedVisual, a model aggregation approach for FL in vehicular visual crowdsensing. FedVisual leverages deep Q-network (DQN) to select appropriate local models, considering the heterogeneities in visual data contents and vehicles’ specifications. By leveraging the historical training experience, an effective model selection strategy can be obtained without complex mathematical modeling. Through the extensive simulations of our self-collected driving videos, FedVisual reduces model aggregation latency by up to 3.8% while improving the model’s performance by up to 3.2% compared to reference works. Wenjun Zhang 0013, Xiaoli Liu 0005, Ruoyi Zhang, Chao Zhu 0002, Sasu Tarkoma |
IEEE Internet Things J. | 3 |
| 2023 | Intra- and Inter-behavior Contrastive Learning for Multi-behavior Recommendation
Qingfeng Li 0001, Huifang Ma, Ruoyi Zhang, Wangyu Jin, Zhixin Li 0001 |
DASFAA (2) | 3 |
| 2023 | Dual-View Self-supervised Co-training for Knowledge Graph Recommendation
Ruoyi Zhang, Huifang Ma, Qingfeng Li 0001, Yike Wang 0001, Zhixin Li 0001 |
DASFAA (2) | 1 |
| 2023 | FloodSFCP: Quality and Latency Balanced Service Function Chain Placement for Remote Sensing in LEO Satellite NetworkabstractPrompted by the significant advancements in image processing technologies and their diverse range of applications, remote sensing satellites are poised for rapid expansion. Nonetheless, offloading the vast amount of remote sensing satellite images to the ground gateway station is inefficient due to the exorbitant costs induced by satellite links, while the limited resources of individual satellites hinder local task processing. With the advancement of the network function virtualization (NFV) technology, a new paradigm for service function chain (SFC) has emerged, which can significantly improve the flexibility and resource utilization of network services and alleviate resource conflicts by dividing large services into smaller ones organized in the form of SFCs. As mega-constellations (e.g., Starlink) developed, the number of low earth orbit (LEO) satellites is increasing. By dividing services into small sub-services and organizing them into SFCs throughout the LEO network, services that cannot be completed by a single satellite can be accomplished through multi-satellite cooperation. However, the quality of the remote sensing service is positively correlated with its latency, and the rapidly changing topology of LEO networks also adds complexity to the SFC placement. Hence, how to select appropriate satellites to place the SFC and modulate service levels, in order to obtain better remote sensing results within an acceptable latency, remains a question. To address these issues, this paper proposes the FloodSFCP, an SFC placement method that aims to increase service quality and decrease latency through offline training and online optimization via deep reinforcement learning, taking into account the variation in LEO network topology. By introducing NoisyNet, Dueling, and N-step learning, we improve the model’s generalization ability and reduce the state space, thus enhancing convergence speed while reducing decision and training time. Experimental results demonstrate that FloodSFCP significantly improves service quality while reducing total decision costs. Ruoyi Zhang, Chao Zhu 0002, Xiao Chen 0002, Qingyuan Gong, Xinlei Xie, Xiangyuan Bu |
SECON | 1 |
| 2023 | Dual-view co-contrastive learning for multi-behavior recommendation
Qingfeng Li 0001, Huifang Ma, Ruoyi Zhang, Wangyu Jin, Zhixin Li 0001 |
Appl. Intell. | 3 |
| 2023 | FIRE: knowledge-enhanced recommendation with feature interaction and intent-aware attention networks
Ruoyi Zhang, Huifang Ma, Qingfeng Li 0001, Yike Wang 0001, Zhixin Li 0001 |
Appl. Intell. | 1 |
| 2022 | Drug Side Effects Prediction via Heterogeneous Multi-Relational Graph Convolutional NetworksabstractNumerous clinical trials have revealed that a serious consequence of polypharmacy is that patients are at high risk of adverse side effects. However, designing clinical trials to determine the frequency of side effects from polypharmacy is both time-consuming and costly. Therefore, the computer-aided prediction of drug side effects is becoming an attractive proposition. Existing methods of drug side effects prediction introduce the target protein of a drug without screening. Although this alleviates the sparsity of the original data to some extent, the blind introduction of proteins as auxiliary information allows a large amount of noisy information to be added, which degrades the model efficiency and acheive sub-opitmal predicition results. To this end, we propose a novel method called DEP-GCN (Drug Side Effects Prediction via Heterogeneous Multi-Relational Graph Convolutional Networks). Specifically, we design two protein auxiliary pathways directly related to drugs and combine these two auxiliary pathways with a multi-relational graph of drug side effects, which both alleviate the sparsity of data and filter out noisy data. Then, to produce accurate drug representations, we distinguish the impact from different drug neighbors and introduce a query-aware attention mechanism to fine-grained determine how much messaging is delivered. Finally, in contrast to approaches limited to predicting the existence or associations of drug side effects, we output the exact frequency of drug side effects occurring via a tensor factorization decoder. Extensive experimental results demonstrate that DEP-GCN significantly outperforms all baseline methods. The further examination provides literature evidence for highly ranked predictions. Yike Wang 0001, Huifang Ma, Ruoyi Zhang, Zihao Gao 0001 |
ICTAI | 3 |
| 2022 | A Knowledge Graph Recommendation Model via High-order Feature Interaction and Intent DecompositionabstractKnowledge Graph(KG) contains structured attribute information which has been widely utilized for recommendations, as well as can effectively tackle the sparsity and cold start problems of collaborative filtering. In recent years, Graph Neural Networks (GNNs) serve as a novel deep learning technique that can significantly enhance recommendation performance. Unfortunately, existing KG-based GNN models are coarse-grained ignoring i)effective high-order feature interaction and fusion mechanism and ii)interpretable user latent intent decomposition. In this paper, we propose a new method named Knowledge Graph recommendation model via high-order feature Interaction and intent Decomposition(KGID), which explicitly models the fine-grained feature interaction and intent factors in KG-based GNN recommendation. Initially, high-order feature interactions are captured via the two-granularity convolutional neural networks on the item side. Next, the implicit intent factor behind the user decisions is modeled by two-level attention mechanisms. Ultimately, user representations and item representations are augmented simultaneously. We conduct experiments on three benchmark datasets to elucidate the superiority of the KGID to state-of-the-art baselines. Ruoyi Zhang, Huifang Ma, Qingfeng Li 0001, Zhixin Li 0001, Yike Wang 0001 |
IJCNN | 1 |
| 2022 | Co-contrastive Learning for Multi-behavior Recommendation
Qingfeng Li 0001, Huifang Ma, Ruoyi Zhang, Wangyu Jin, Zhixin Li 0001 |
PRICAI (3) | 3 |
| 2022 | Energy-Efficient Multi-Task Allocation for Antenna Array Empowered Vehicular Fog ComputingabstractWith the emergence of compute-intensive and latency-sensitive vehicular applications, vehicular fog computing (VFC) has been proposed for catering to the thriving demands for computing and communication resources close to vehicles. In VFC scenarios where multiple tasks need to be offloaded simultaneously, the data, often coming from multiple sources, must be transmitted at a high data-rate in parallel. An antenna array system, a set of multiple connected antennas which work together as a single antenna, could achieve a significantly higher data-rate than a traditional single antenna. However, data-rate of the antenna array system may decrease due to the presence of interference. On the other hand, an antenna array system consumes more energy than a single antenna, which is antagonistic to vehicles powered by limited electricity. To address these challenges, we propose EAAV, a multi-task allocation strategy that enables multiple tasks to be offloaded concurrently in antenna array empowered VFC. EAAV aims at reducing the transmission power consumption while maintaining a high transmission data-rate, taking into account the mobility of vehicles and communication interference. We transform the multi-task allocation problem into a convex solvable one and evaluate the effectiveness of EAAV based on real-world vehicle trajectories. Compared with the existing task allocation strategy, EAAV improves the average transmission data-rate by up to 8.2% and reduces the average power consumption by up to 38.3%. Xinlei Xie, Ruoyi Zhang, Chao Zhu 0002, Ruijin Li, Xiangyuan Bu, Yu Xiao 0001 |
VTC Spring | 2 |
| 2021 | Exploring Implicit Relationships in Social Network for Recommendation Systems
Yunhe Wei, Huifang Ma, Ruoyi Zhang, Zhixin Li 0001, Liang Chang 0003 |
PAKDD (2) | 3 |
| 2019 | Incorporating URL embedding into ensemble clustering to detect web anomalies
Bo Li 0005, Guiqin Yuan, Ruoyi Zhang, Yiyang Yao |
Future Gener. Comput. Syst. | 4 |
| 2018 | Robust eye detection using deeply-learned gaze shifting path
Ruoyi Zhang, Youqian Zhang |
J. Vis. Commun. Image Represent. | 3 |