VLDB 2026 Research / reviewers in the wild / expert
Huikang Su
dblp:415/0876
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 59% Language models and text generation · 41% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model reasoning › adaptive reasoning
adaptive reasoning depth |
1.0 | 1 | 2026 | Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language Models · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning |
1.0 | 1 | 2026 | Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language Models · AAAI 2026 |
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning |
0.9 | 1 | 2025 | Boundary-to-Region Supervision for Offline Safe Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.9 | 1 | 2025 | Boundary-to-Region Supervision for Offline Safe Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.3 | 1 | 2026 | Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language Models · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
reward management · 1.0chain-of-thought · 1.0sequence model · 0.9rotary positional embedding · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsabstractRecent advancements in large language models (LLMs) have greatly improved their ability to perform complex reasoning tasks through Long Chain-of-Thought (CoT). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. To improve efficiency, current methods often rely on human-defined difficulty priors, which do not align with the LLM's self-awared difficulty, leading to inefficiencies. In this paper, we introduce the Dynamic Reasoning-Boundary Self-Awareness Framework (DR. SAF), which enables LLMs to dynamically assess and adjust their reasoning depth in response to problem complexity. DR. SAF integrates three key components: Boundary Self-Awareness Alignment, Adaptive Reward Management, and a Boundary Preservation Mechanism. These components allow models to optimize their reasoning processes, balancing efficiency and accuracy without compromising performance. Our experimental results demonstrate that DR. SAF achieves a 49.27% reduction in total response tokens with minimal loss in accuracy. The framework also delivers a 6.59x gain in token efficiency and a 5x reduction in training time, making it well-suited to resource-limited settings. During extreme training, DR. SAF can even surpass traditional instruction-based models in token efficiency with more than 16% accuracy improvement. Qiguang Chen, Dengyun Peng, Huikang Su, Jiannan Guan, Libo Qin 0001, Wanxiang Che |
AAAI | 4 |
| 2025 | Boundary-to-Region Supervision for Offline Safe Reinforcement LearningabstractOffline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action generation on symmetric input tokens for return-to-go and cost-to-go, neglecting their intrinsic asymmetry: RTG serves as a flexible performance target, while CTG should represent a rigid safety boundary. This symmetric conditioning leads to unreliable constraint satisfaction, especially when encountering out-of-distribution cost trajectories. To address this, we propose Boundary-to-Region (B2R), a framework that enables asymmetric conditioning through cost signal realignment . B2R redefines CTG as a boundary constraint under a fixed safety budget, unifying the cost distribution of all feasible trajectories while preserving reward structures. Combined with rotary positional embeddings , it enhances exploration within the safe region. Experimental results show that B2R satisfies safety constraints in 35 out of 38 safety-critical tasks while achieving superior reward performance over baseline methods. This work highlights the limitations of symmetric token conditioning and establishes a new theoretical and practical approach for applying sequence models to safe RL. Huikang Su, Dengyun Peng, Zifeng Zhuang, Qiguang Chen, Qinghe Liu |
NeurIPS | 1 |