VLDB 2026 Research / reviewers in the wild / expert
Ruikai Liu
dblp:43/10207
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-8480-0154ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 48% Representation and self-supervised learning · 28% Language models and text generation · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
mathematical reasoning |
1.0 | 1 | 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning
policy optimization |
1.0 | 1 | 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.0 | 1 | 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning › reward design
reward shaping |
1.0 | 1 | 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › contrastive learning
multi-view contrastive learning |
0.9 | 1 | 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025 |
Medical and health informatics
clinical time series analysis |
0.9 | 1 | 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025 |
Machine learning › Time series and sequential data
anomaly detection |
0.3 | 1 | 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
reconstruction error · 1.7multi-head attention · 1.7contrastive learning · 1.7GAN-enhanced encoder-decoder · 1.7reward shaping · 1.0reinforcement learning · 1.0group relative policy optimization · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy OptimizationabstractReinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-correct or all-incorrect completions), leading to zero-variance rewards and ineffective gradient signals. Moreover, focusing solely on final answer correctness while ignoring the reasoning process, along with rigid length penalties, can hinder training stability and output quality. To address these issues, we introduce TAPO, a reinforcement learning framework that enhances optimization signals by modifying sampled completions within training groups. TAPO incorporates three core techniques: (1) Dynamic Teacher Injection (DTI), which selectively injects high-quality or adversarial examples to restore effective gradient signals in collapsed groups; (2) Perturbed Answer Injection (PAI), which makes partially correct completions to provide contrastive supervision separating reasoning correctness but wrong answer from the trajectories; and (3) InfoLen-Aware Reward Shaping, a fine-grained reward strategy that penalizes outputs based on both length and semantic redundancy, encouraging concise yet informative responses. Extensive experimental results demonstrate that TAPO significantly improves the mathematical reasoning capabilities of LLMs across multiple challenging benchmarks, outperforming the GRPO baseline by a substantial margin. Component-wise ablations further validate the contribution of each proposed technique. Maowei Jiang, Peter Bús, Moquan Chen, Quangao Liu, Ruiqi Li 0004, Pengyu Zeng, Ruikai Liu, Alan Liang, Yusong Hu, Zhiyong Dong |
AAAI | 10 |
| 2026 | Lightweight multi-classification pear fruit high-precision detection model in complex orchard scenes
Jinlin Xue, Tianxing Zhao, Ruikai Liu |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time seriesabstractMedical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi-view contrastive framework that integrates external normal samples and enhances feature learning through adaptive contrastive strategies. DAAC consists of two key modules: (1) a Discrepancy Estimator, built upon a GAN-enhanced encoder-decoder architecture, captures the distribution of normal data and computes reconstruction errors as indicators of abnormality. These discrepancy features augment the target dataset to mitigate overfitting. (2) an Adaptive Contrastive Learner uses multi-head attention to extract discriminative representations by contrasting embeddings across multiple views and data granularities (subject, trial, epoch, and temporal levels), eliminating the need for handcrafted positive-negative sample pairs. Extensive experiments on three clinical datasets—covering Alzheimer’s disease, Parkinson’s disease, and myocardial infarction—demonstrate that DAAC significantly outperforms existing methods, even when only 10\% of labeled data is available, showing strong generalization and diagnostic performance. Our code is available
at https://github.com/CUHKSZ-MED-BioE/DAAC. Hongfeng Ai, Ruiqi Li 0004, Maowei Jiang, Quangao Liu, Jiahua Dong 0001, Ruiyuan Kang, Alan Liang, Ruikai Liu, Chenzhong Li |
NeurIPS | 10 |
| 2022 | Flexible and Precision Snap-Fit Peg-in-Hole Assembly Based on Multiple Sensations and Damping IdentificationabstractSnap-fit peg-in-hole assembly widely exists in both industry and daily life, especially for consumer electronics. The buckle mechanism leads to a damping zone inside the port where insertion force needs to be increased. It is much difficult to automate this process by robots, for size and clearance of the components are always small, and the damping buckle should be perceived and distinguished from solid inner walls of the port. End-effector position control might be invalid, since grasping error will make it difficult to locate the plug accurately. In this article, we undertake this assembly challenge by taking advantage of fingertip tactile perception combined with visual images and force feedback. Raw sensor data is collected, processed, and fused together to be state input of a reinforcement learning network, generating continuous action vectors. We also propose a novel damping zone predictor through feature extraction and multimodal fusion, which is able to identify whether the plug has touched the buckle mechanism, so as to adjust the insertion force. The whole framework is implemented through a common USB Type-C insertion experiment on Franka Panda robot platform, reaching a success rate of 88%. Furthermore, system robustness is verified, and comparisons of different modalities are also conducted. Ruikai Liu, Xiansheng Yang, Ajian Li, Yunjiang Lou |
IROS | 1 |