Chuhuai Yue

dblp:340/0731 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0004-4592-8109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 42% Image recognition and object detection · 37% Reinforcement learning · 21%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
Promoting Efficient Reasoning with Verifiable Stepwise Reward · AAAI 2026
Natural language and speech › Language models and text generation › large language model
large reasoning model
1.012026
Promoting Efficient Reasoning with Verifiable Stepwise Reward · AAAI 2026
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
1.012026
Promoting Efficient Reasoning with Verifiable Stepwise Reward · AAAI 2026
Computer vision › Image recognition and object detection
object detection
0.912025
UN-DETR: Promoting Objectness Learning via Joint Supervision for Unknown Object Detection · AAAI 2025
Computer vision › Image recognition and object detection › object detection › open-world object detection
unknown object detection
0.912025
UN-DETR: Promoting Objectness Learning via Joint Supervision for Unknown Object Detection · AAAI 2025

Methods — techniques the papers use, named apart from their topics

rule-based verifiable stepwise reward · 1.0REINFORCE++ · 1.0PPO · 1.0unsupervised pre-training · 0.9transformer · 0.9joint supervision · 0.9
YearPublicationVenuePosition
2026 Promoting Efficient Reasoning with Verifiable Stepwise Reward
abstract
Large reasoning models (LRMs) have recently achieved significant progress in complex reasoning tasks, aided by reinforcement learning with verifiable rewards. However, LRMs often suffer from overthinking, expending excessive computation on simple problems and reducing efficiency. Existing efficient reasoning methods typically require accurate task assessment to preset token budgets or select reasoning modes, which limits their flexibility and reliability. In this work, we revisit the essence of overthinking and identify that encouraging effective steps while penalizing ineffective ones is key to its solution. To this end, we propose a novel rule-based verifiable stepwise reward mechanism (VSRM), which assigns rewards based on the performance of intermediate states in the reasoning trajectory. This approach is intuitive and naturally fits the step-by-step nature of reasoning tasks. We conduct extensive experiments on standard mathematical reasoning benchmarks, including AIME24 and AIME25, by integrating VSRM with PPO and Reinforce++. Results show that our method achieves substantial output length reduction while maintaining original reasoning performance, striking an optimal balance between efficiency and accuracy. Further analysis of overthinking frequency and pass@k score before and after training demonstrates that our approach indeed effectively suppresses ineffective steps and encourages effective reasoning, fundamentally alleviating the overthinking problem.
Chuhuai Yue, Chengqi Dong, Yinan Gao, Hang He, Jiajun Chai, Guojun Yin
AAAI1
2025 UN-DETR: Promoting Objectness Learning via Joint Supervision for Unknown Object Detection
abstract
Unknown Object Detection (UOD) aims to identify objects of unseen categories, differing from the traditional detection paradigm limited by the closed-world assumption. A key component of UOD is learning a generalized representation, i.e. objectness for both known and unknown categories to distinguish and localize objects from the background in a class-agnostic manner. However, previous methods obtain supervision signals for learning objectness in isolation from either localization or classification information, leading to poor performance for UOD. To address this issue, we propose a transformer-based UOD framework, UN-DETR. Based on this, we craft Instance Presence Score (IPS) to represent the probability of an object's presence. For the purpose of information complementarity, IPS employs a strategy of joint supervised learning, integrating attributes representing general objectness from the positional and the categorical latent space as supervision signals. To enhance IPS learning, we introduce a one-to-many assignment strategy to incorporate more supervision. Then, we propose Unbiased Query Selection to provide premium initial query vectors for the decoder. Additionally, we propose an IPS-guided post process strategy to filter redundant boxes and correct classification predictions for known and unknown objects. Finally, we pretrain the entire UN-DETR in an unsupervised manner, in order to obtain objectness prior. Our UN-DETR is comprehensively evaluated on multiple UOD and known detection benchmarks, demonstrating its effectiveness and achieving state-of-the-art performance.
Haomiao Liu, Hao Xu 0050, Chuhuai Yue, Bo Ma 0012
AAAI3
2025 Adaptive objectness learning for enhanced unknown object detection
Haomiao Liu, Hao Xu 0050, Chuhuai Yue, Bo Ma 0012
Vis. Comput.3
2024 UOA-RCNN: Detect Anything with Unknown Object Aware RCNN
Haomiao Liu, Hao Xu 0050, Chuhuai Yue, Bo Ma 0012
ICONIP (8)3