VLDB 2026 Research / reviewers in the wild / expert
Junyao Yang
dblp:262/8668
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 62% Efficient and distributed learning · 22% Generative modeling · 10% | |
| Network and information security
2 papers |
Privacy and data protection · 67% Security and privacy of machine learning · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model merging |
2.0 | 2 | 2026 | ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging · ACL (1) 2026 RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior · AAAI 2026 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model
large reasoning model |
1.0 | 1 | 2026 | ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging · ACL (1) 2026 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
long chain-of-thought reasoning |
1.0 | 1 | 2026 | RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.9 | 1 | 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis · EMNLP 2025 |
Natural language and speech › Language models and text generation › trustworthy language model
privacy-preserving inference |
0.9 | 1 | 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model fine-tuning
private fine-tuning |
0.9 | 1 | 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis · EMNLP 2025 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis · EMNLP 2025 |
Medical and health informatics › medical imaging
medical image analysis |
0.9 | 1 | 2025 | Enhancing Bone Mineral Density Estimation from X-ray Images with Cross-Modal Knowledge Distillation · KDD (2) 2025 |
Privacy and data protection
differential privacy |
0.9 | 1 | 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis · EMNLP 2025 |
Security and privacy of machine learning
privacy-preserving inference |
0.9 | 1 | 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration · ACL (1) 2025 |
Privacy and data protection › differential privacy
synthetic data generation |
0.9 | 1 | 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis · EMNLP 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2026 | ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
model merging · 2.0reward model · 1.7privacy restoration · 1.7privacy removal · 1.7graph-based structural learning · 1.7differential privacy · 1.7data refinement · 1.7data filtering · 1.7reasoning capability indicator · 1.0contrastive gradient identification · 1.0multi-scale feature extraction · 0.9knowledge distillation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as PriorabstractLarge Language Models (LLMs) with long chain-of-thought (CoT) capability, termed Reasoning Models, demonstrate superior intricate problem-solving abilities through multi-step long CoT reasoning. To create a dual-capability model with long CoT capability and domain-specific knowledge without substantial computational and data costs, model merging emerges as a highly resource-efficient method. However, significant challenges lie in merging domain-specific LLMs with long CoT ones since nowadays merging methods suffer from reasoning capability degradation, even gibberish output and output collapse. To overcome this, we introduce RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior, a novel merging framework designed to integrate domain-specific LLMs with long CoT capability, meanwhile maintaining model performance in the original domain. Treating reasoning model weights as foundational prior, our method utilizes a reasoning capability indicator to preserve core long CoT capability model weights while selectively merging essential domain-specific weights. We conducted extensive experiments on Qwen2.5-7B, Llama3.1-8B, and Qwen2.5-1.5B models in BioMedicine and Finance domains. Our results show that RCP-Merging successfully merges a reasoning model with domain-specific ones, improving domain task performance by 9.5% and 9.2% over state-of-the-art methods, without significantly harming the original long CoT reasoning capability. Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng |
AAAI | 1 |
| 2026 | ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model MergingabstractLarge Reasoning Models (LRMs) with long chain-of-thought reasoning have recently achieved remarkable success.Yet, equipping domain-specialized models with such reasoning capabilities, referred to as "Reasoning + X", remains a significant challenge.While model merging offers a promising training-free solution, existing methods often suffer from a destructive performance collapse: existing methods tend to both weaken reasoning depth and compromise domain-specific utility.Interestingly, we identify a counter-intuitive phenomenon underlying this failure: reasoning ability predominantly resides in parameter regions with low gradient sensitivity, contrary to the common assumption that domain capabilities correspond to high-magnitude parameters.Motivated by this insight, we propose ReasonAny, a novel merging framework that resolves the reasoning-domain performance collapse through Contrastive Gradient Identification.Experiments across safety, biomedicine, and finance domains show that ReasonAny effectively synthesizes "Reasoning + X" capabilities, significantly outperforming state-of-theart baselines while retaining robust reasoning performance. Junyao Yang, Chen Qian 0010, Yong Liu 0007, Dongrui Liu |
ACL (1) | 1 |
| 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and RestorationabstractZiqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziqian Zeng, Junyao Yang, Zhengdong Lu, Huiping Zhuang, Cen Chen 0002 |
ACL (1) | 3 |
| 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data SynthesisabstractThe success of large language models (LLMs) has attracted many individuals to fine-tune them for domain-specific tasks by uploading their data.However, in sensitive areas like healthcare and finance, privacy concerns often arise.One promising solution is to generate synthetic data with Differential Privacy (DP) guarantees to replace private data.However, these synthetic data contain significant flawed data, which are considered as noise.Existing solutions typically rely on naive filtering by comparing ROUGE-L scores or embedding similarities, which are ineffective in addressing the noise.To address this issue, we propose RewardDS, a novel privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.Our RewardDS introduces two key modules, Reward Guided Filtering and Self-Optimizing Refinement, to both filter and refine the synthetic data, effectively mitigating the noise.Extensive experiments across medical, financial, and code generation domains demonstrate the effectiveness of our method. Chengming Shi, Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng |
EMNLP | 3 |
| 2025 | Enhancing Bone Mineral Density Estimation from X-ray Images with Cross-Modal Knowledge DistillationabstractDual-energy X-ray absorptiometry (DXA) enables accurate bone mineral density but requires specialized equipment and protocols. X-ray-based BMD screening offers opportunistic early detection, though prior methods struggle with X-ray intensity variations and demand large datasets. We introduce a cross-modal knowledge distillation BMD prediction framework (CMKD-BMD) fusing X-ray/CT data to enhance X-ray-only BMD prediction. Each single-modal student network employs a multi-scale visual extractor for hierarchical features, an unsupervised graph-based structural learner for anatomical relationships, and an adaptive fusion module to generate a unified representation. The teacher network integrates single-modal representations from students and transfers multimodal knowledge from the teacher to students. Our model outperforms existing methods on both the collected dataset (comprising 1620 X-rays and 280 CT cases) and the publicly available VerSe2019 dataset, demonstrating superior BMD estimation performance. Furthermore, we developed OrthoSim, an orthopedic surgical simulation platform with CMKD-BMD, which has shown promising clinical effectiveness in trial evaluations. Our code is available at https://github.com/KeyueShi/CMKD-BMD. Keyue Shi, Qianqian Shen, Zhongda Qi, Junyao Yang, Zhaoming Ye, Jiajun Bu, Haishuai Wang |
KDD (2) | 4 |
| 2023 | Multi-Tenant In-Memory Key-Value Cache Partitioning Using Efficient Random Sampling-Based LRU ModelabstractIn-memory key-value caches are widely used as a performance-critical layer in web applications, disk-based storage, and distributed systems. The Least Recently Used (LRU) replacement policy has become thede factostandard in those systems since it exploits workload locality well. However, the LRU implementation can be costly due to the rigid data structure in maintaining object priority, as well as the locks for object order updating. Redis as one of the most effective and prevalent deployed commercial systems adopts an approximated LRU policy, where the least recently used item from a small, randomly sampled set of items is chosen to evict. This random sampling-based policy is lightweight and shows its flexibility. We observe that there can exist a significant miss ratio gap between exact LRU and random sampling-based LRU under different sampling size$K$s. Therefore existing LRU miss ratio curve (MRC) construction techniques cannot be directly applied without loss of accuracy. In this article, we introduce a new probabilistic stack algorithm namedKRRto accurately model random sampling based-LRU, and extend it to handle both fixed and variable objects in key-value caches. We present an efficient stack update algorithm that reduces the expected running time of KRR significantly. To improve the performance of the in-memory multi-tenant key-value cache that utilizes random sampling-based replacement, we propose kRedis, a reference locality- and latency-aware memory partitioning scheme. kRedis guides the memory allocation among the tenants and dynamically customizes$K$to better exploit the locality of each individual tenant. Evaluation results over diverse workloads show that our model generates accurate miss ratio curves for both fixed and variable object size workloads, and enables practical, low-overhead online MRC prediction. Equipped with KRR, kRedis delivers up to a 50.2% average access latency reduction, and up to a 262.8% throughput improvement compared to Redis. Furthermore, by comparing with pRedis, a state-of-the-art design of memory allocation in Redis, kRedis shows up to 24.8% and 61.8% improvements in average access latency and throughput, respectively. Junyao Yang, Zhenlin Wang 0003 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Unsupervised Learning of Depth Estimation and Camera Pose With Multi-Scale GANsabstractUnsupervised learning methods have achieved remarkable performance in monocular depth estimation and camera pose, which mostly solve the multi-task learning problem by using their inner geometry consistency as the self-supervision signal. While most existing approaches mostly adopt the generative model to obtain the depth map prediction, so in the resolution of depth map there is room for improvement. To this end, we present our unsupervised learning architecture based on adversarial learning model, which is used for unsupervised learning of high-resolution single view depth and camera pose. Specifically, we present a multi-scale deep convolutional Generative Adversarial Network (GAN) based learning system, which consists of three networks (pose estimation network PCNN, Generator-D and Discriminator-D for depth map prediction). Furthermore, in order to generate high-resolution depth map, we propose a multi-scale GAN model (MSGAN) to decompose the hard high-quality image generation problem into more manageable sub-problems through a coarse-to-fine process. Then, we modify the overall generation architecture of GAN model by changing the down-sampling and up-sampling components to improve the quality and accuracy of the depth map prediction. Finally, in order to improve the rate of convergence, we use the Least Square Error to increase the penalty for outliers. Detailed quantitative and qualitative evaluations of the proposed framework on the KITTI dataset show that the proposed method provides better results for both pose estimation and depth recovery. Yu-Fan Xu, Yan Wang 0042, Rui Huang 0013, Zeyu Lei, Junyao Yang, Zijian Li 0005 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Efficient Modeling of Random Sampling-Based LRUabstractThe Miss Ratio Curve (MRC) is an important metric and effective tool for caching system performance prediction and optimization. Since the Least Recently Used (LRU) replacement policy is the de facto policy for many existing caching systems, most previous studies on efficient MRC construction are predominantly focused on the LRU replacement policy. Recently, the random sampling-based replacement mechanism, as opposed to replacement relying on the rigid LRU data structure, gains more popularity due to its lightweight and flexibility. To approximate LRU, at replacement times, the system randomly selects K objects and replaces the least recently used object among the sample. Redis implements this approximated LRU policy. We observe that there can exist a significant miss ratio gap between exact LRU and random sampling-based LRU under different sampling size K; therefore existing LRU MRC construction techniques cannot be directly applied to random sampling based LRU cache without loss of accuracy. Junyao Yang, Zhenlin Wang 0003 |
ICPP | 1 |
| 2021 | Attention based multilayer feature fusion convolutional neural network for unsupervised monocular depth estimation
Zeyu Lei, Yan Wang 0042, Zijian Li 0005, Junyao Yang |
Neurocomputing | 4 |