Ruikai Liu

dblp:43/10207 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0002-8480-0154ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 48% Representation and self-supervised learning · 28% Language models and text generation · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
mathematical reasoning
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning › reward design
reward shaping
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Machine learning › Representation and self-supervised learning › contrastive learning
multi-view contrastive learning
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Medical and health informatics
clinical time series analysis
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Machine learning › Time series and sequential data
anomaly detection
0.312025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Machine learning › Generative modeling
generative adversarial network
0.312025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

reconstruction error · 1.7multi-head attention · 1.7contrastive learning · 1.7GAN-enhanced encoder-decoder · 1.7reward shaping · 1.0reinforcement learning · 1.0group relative policy optimization · 1.0
YearPublicationVenuePosition
2026 TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization
abstract
Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-correct or all-incorrect completions), leading to zero-variance rewards and ineffective gradient signals. Moreover, focusing solely on final answer correctness while ignoring the reasoning process, along with rigid length penalties, can hinder training stability and output quality. To address these issues, we introduce TAPO, a reinforcement learning framework that enhances optimization signals by modifying sampled completions within training groups. TAPO incorporates three core techniques: (1) Dynamic Teacher Injection (DTI), which selectively injects high-quality or adversarial examples to restore effective gradient signals in collapsed groups; (2) Perturbed Answer Injection (PAI), which makes partially correct completions to provide contrastive supervision separating reasoning correctness but wrong answer from the trajectories; and (3) InfoLen-Aware Reward Shaping, a fine-grained reward strategy that penalizes outputs based on both length and semantic redundancy, encouraging concise yet informative responses. Extensive experimental results demonstrate that TAPO significantly improves the mathematical reasoning capabilities of LLMs across multiple challenging benchmarks, outperforming the GRPO baseline by a substantial margin. Component-wise ablations further validate the contribution of each proposed technique.
Maowei Jiang, Peter Bús, Moquan Chen, Quangao Liu, Ruiqi Li 0004, Pengyu Zeng, Ruikai Liu, Alan Liang, Yusong Hu, Zhiyong Dong
AAAI10
2026 Lightweight multi-classification pear fruit high-precision detection model in complex orchard scenes
Jinlin Xue, Tianxing Zhao, Ruikai Liu
Eng. Appl. Artif. Intell.7
2025 DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series
abstract
Medical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi-view contrastive framework that integrates external normal samples and enhances feature learning through adaptive contrastive strategies. DAAC consists of two key modules: (1) a Discrepancy Estimator, built upon a GAN-enhanced encoder-decoder architecture, captures the distribution of normal data and computes reconstruction errors as indicators of abnormality. These discrepancy features augment the target dataset to mitigate overfitting. (2) an Adaptive Contrastive Learner uses multi-head attention to extract discriminative representations by contrasting embeddings across multiple views and data granularities (subject, trial, epoch, and temporal levels), eliminating the need for handcrafted positive-negative sample pairs. Extensive experiments on three clinical datasets—covering Alzheimer’s disease, Parkinson’s disease, and myocardial infarction—demonstrate that DAAC significantly outperforms existing methods, even when only 10\% of labeled data is available, showing strong generalization and diagnostic performance. Our code is available at https://github.com/CUHKSZ-MED-BioE/DAAC.
Hongfeng Ai, Ruiqi Li 0004, Maowei Jiang, Quangao Liu, Jiahua Dong 0001, Ruiyuan Kang, Alan Liang, Ruikai Liu, Chenzhong Li
NeurIPS10
2022 Flexible and Precision Snap-Fit Peg-in-Hole Assembly Based on Multiple Sensations and Damping Identification
abstract
Snap-fit peg-in-hole assembly widely exists in both industry and daily life, especially for consumer electronics. The buckle mechanism leads to a damping zone inside the port where insertion force needs to be increased. It is much difficult to automate this process by robots, for size and clearance of the components are always small, and the damping buckle should be perceived and distinguished from solid inner walls of the port. End-effector position control might be invalid, since grasping error will make it difficult to locate the plug accurately. In this article, we undertake this assembly challenge by taking advantage of fingertip tactile perception combined with visual images and force feedback. Raw sensor data is collected, processed, and fused together to be state input of a reinforcement learning network, generating continuous action vectors. We also propose a novel damping zone predictor through feature extraction and multimodal fusion, which is able to identify whether the plug has touched the buckle mechanism, so as to adjust the insertion force. The whole framework is implemented through a common USB Type-C insertion experiment on Franka Panda robot platform, reaching a success rate of 88%. Furthermore, system robustness is verified, and comparisons of different modalities are also conducted.
Ruikai Liu, Xiansheng Yang, Ajian Li, Yunjiang Lou
IROS1