EDBT 2026 Demo / reviewers in the wild / expert
Changxin Huang
dblp:207/8606
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 36% Robot manipulation · 16% Efficient and distributed learning · 16% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
reward design |
1.7 | 2 | 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025 Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025 |
Robotics › Motion planning and robot control › robot learning
robot skill learning |
1.7 | 2 | 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025 Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025 |
Computer vision › Video understanding and tracking
action recognition |
1.0 | 1 | 2026 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models · AAAI 2026 |
Machine learning › Efficient and distributed learning › distillation
cross-architecture knowledge distillation |
1.0 | 1 | 2026 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models · AAAI 2026 |
Robotics › Robot manipulation › grasping › grasp control
dynamic grasping |
1.0 | 1 | 2026 | Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators · AAAI 2026 |
Robotics › Robot manipulation
grasping |
1.0 | 1 | 2026 | Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models · AAAI 2026 |
Robotics › Robot manipulation › mobile manipulation
legged manipulators |
1.0 | 1 | 2026 | Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators · AAAI 2026 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.0 | 1 | 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model · AAAI 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models · AAAI 2026 |
Computer vision › Vision and language
vision-and-language navigation |
1.0 | 1 | 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model · AAAI 2026 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
1.0 | 1 | 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model · AAAI 2026 |
Machine learning › Reinforcement learning › reward learning
LLM-based reward generation |
0.9 | 1 | 2025 | Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.9 | 1 | 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025 |
Robotics › Legged, aerial and field robots › bipedal robot
bipedal locomotion control |
0.7 | 1 | 2023 | Reward-Adaptive Reinforcement Learning: Dynamic Policy Gradient Optimization for Bipedal Locomotion · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion |
0.7 | 1 | 2023 | Reward-Adaptive Reinforcement Learning: Dynamic Policy Gradient Optimization for Bipedal Locomotion · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.7 | 1 | 2023 | Reward-Adaptive Reinforcement Learning: Dynamic Policy Gradient Optimization for Bipedal Locomotion · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Vision and language
multimodal reasoning |
0.3 | 1 | 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model · AAAI 2026 |
Robotics › Motion planning and robot control
whole-body control |
0.3 | 1 | 2026 | Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators · AAAI 2026 |
Machine learning › Reinforcement learning
policy optimization |
0.3 | 1 | 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025 |
Machine learning › Reinforcement learning
value function |
0.3 | 1 | 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.7temporal modeling · 1.0teacher-student framework · 1.0multimodal world model · 1.0hierarchical prediction-feedback · 1.0dual-teacher distillation · 1.0discrepancy-aware weighting · 1.0policy hybridization · 0.9multi-branch value network · 0.9bayesian optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelabstractVision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such methods is limited to the linguistic modality, lacking visual reasoning capabilities. Moreover, existing reasoning modules are optimized separately from navigation policies, leading to incompatibility and potential conflicts in optimization objectives. To tackle these challenges, we introduce UNeMo, a novel framework designed for the collaborative optimization of visual state reasoning and navigational decision-making. It introduces a Multimodal World Model (MWM) that takes visual features, language instructions, and navigational actions as inputs to jointly predict subsequent visual states, enabling cross-modal reasoning. Via a Hierarchical Prediction-Feedback (HPN) mechanism, MWM collaborates with navigation policies: the first layer generates actions using current vision-and-language features; MWM then infers post-action visual states to guide the second layer’s fine-grained decisions. This forms a dynamic bidirectional promotion mechanism where MWM reasoning optimizes navigation policies, while policy decisions feedback to improve MWM’s reasoning accuracy. Experiments on R2R and REVERIE datasets show UNeMo outperforms state-of-the-art methods by 2.1% and 0.7% in navigation accuracy for unseen scenes, validating its effectiveness. Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu, Runhao Zeng, Zun Liu |
AAAI | 1 |
| 2026 | Whole-Body Coordination for Dynamic Object Grasping with Legged ManipulatorsabstractQuadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, existing research has predominantly focused on static-object grasping, neglecting the challenges posed by dynamic targets and thus limiting applicability in dynamic scenarios such as logistics sorting and human–robot collaboration. To address this, we introduce DQ-Bench, a new benchmark that systematically evaluates dynamic grasping across varying object motions, velocities, heights, object types, and terrain complexities, along with comprehensive evaluation metrics. Building upon this benchmark, we propose DQ-Net, a compact teacher–student framework designed to infer grasp configurations from limited perceptual cues. During training, the teacher network leverages privileged information to holistically model both the static geometric properties and dynamic motion characteristics of the target, and integrates a grasp fusion module to deliver robust guidance for motion planning. Concurrently, we design a lightweight student network that performs dual-viewpoint temporal modeling using only the target mask, depth map, and proprioceptive state, enabling closed-loop action outputs without reliance on privileged data. Extensive experiments on DQ-Bench demonstrate that DQ-Net achieves robust dynamic objects grasping across multiple task settings, substantially outperforming baseline methods in both success rate and responsiveness. We will release our codebase and benchmark publicly. Qiwei Liang, Boyang Cai, Rongyi He, Tao Teng, Haihan Duan, Changxin Huang, Runhao Zeng |
AAAI | 7 |
| 2026 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video ModelsabstractVision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more efficient but suffer from accuracy gaps. Cross-Architecture Knowledge Distillation (CAKD) addresses this by transferring knowledge from ViTs to CNNs, yet existing methods often struggle with architectural mismatch and overlook the value of stronger homogeneous CNN teachers. To tackle these challenges, we propose a Dual-Teacher Knowledge Distillation framework that leverages both a heterogeneous ViT teacher and a homogeneous CNN teacher to collaboratively guide a lightweight CNN student. We introduce two key components: (1) Discrepancy-Aware Teacher Weighting, which dynamically fuses the predictions from ViT and CNN teachers by assigning adaptive weights based on teacher confidence and prediction discrepancy with the student, enabling more informative and effective supervision; and (2) a Structure Discrepancy-Aware Distillation strategy, where the student learns the residual features between ViT and CNN teachers via a lightweight auxiliary branch, focusing on transferable architectural differences without mimicking all of ViT’s high-dimensional patterns. Extensive experiments on benchmarks including HMDB51, EPIC-KITCHENS-100, and Kinetics-400, demonstrate that our method consistently outperforms state-of-the-art distillation approaches, achieving notable performance improvements with a maximum accuracy gain of 5.95% on HMDB51. Hongsen Ye, Changxin Huang, Xiping Hu, Jian Chen 0011, Runhao Zeng |
AAAI | 3 |
| 2025 | Efficient Language-instructed Skill Acquisition via Reward-Policy Co-EvolutionabstractThe ability to autonomously explore and resolve tasks with minimal human guidance is crucial for the self-development of embodied intelligence. Although reinforcement learning methods can largely ease human effort, it's challenging to design reward functions for real-world tasks, especially for high-dimensional robotic control, due to complex relationships among joints and tasks. Recent advancements large language models (LLMs) enable automatic reward function design. However, approaches evaluate reward functions by re-training policies from scratch placing an undue burden on the reward function, expecting it to be effective throughout the whole policy improvement process. We argue for a more practical strategy in robotic autonomy, focusing on refining existing policies with policy-dependent reward functions rather than a universal one. To this end, we propose a novel reward-policy co-evolution framework where the reward function and the learned policy benefit from each other's progressive on-the-fly improvements, resulting in more efficient and higher-performing skill acquisition. Specifically, the reward evolution process translates the robot's previous best reward function, descriptions of tasks and environment into text inputs. These inputs are used to query LLMs to generate a dynamic amount of reward function candidates, ensuring continuous improvement at each round of evolution. For policy evolution, our method generates new policy populations by hybridizing historically optimal and random policies. Through an improved Bayesian optimization, our approach efficiently and robustly identifies the most capable and plastic reward-policy combination, which then proceeds to the next round of co-evolution. Despite using less data, our approach demonstrates an average normalized improvement of 95.3\% across various high-dimensional robotic skill learning tasks. Changxin Huang, Yanbin Chang, Junfan Lin, Junyang Liang, Runhao Zeng, Jianqiang Li 0001 |
AAAI | 1 |
| 2025 | Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill LearningabstractEnabling a high-degree-of-freedom robot to learn specific skills is a challenging task due to the complexity of robotic dynamics. Reinforcement learning (RL) has emerged as a promising solution; however, addressing such problems requires the design of multiple reward functions to account for various constraints in robotic motion. Existing approaches typically sum all reward components indiscriminately to optimize the RL value function and policy. We argue that this uniform inclusion of all reward components in policy optimization is inefficient and limits the robot's learning performance. To address this, we propose an Automated Hybrid Reward Scheduling (AHRS) framework based on Large Language Models (LLMs). This paradigm dynamically adjusts the learning intensity of each reward component throughout the policy optimization process, enabling robots to acquire skills in a gradual and structured manner. Specifically, we design a multi-branch value network, where each branch corresponds to a distinct reward component. During policy optimization, each branch is assigned a weight that reflects its importance, and these weights are automatically computed based on rules designed by LLMs. The LLM generates a rule set in advance, derived from the task description, and during training, it selects a weight calculation rule from the library based on language prompts that evaluate the performance of each branch. Experimental results demonstrate that the AHRS method achieves an average$\mathbf{6. 4 8 \%}$performance improvement across multiple high-degree-of-freedom robotic tasks. Changxin Huang, Junyang Liang, Yanbin Chang, Jingzhao Xu, Jianqiang Li 0001 |
ICRA | 1 |
| 2024 | Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior LearningabstractLearning behavior in legged robots presents a significant challenge due to its inherent instability and complex constraints. Recent research has proposed the use of a large language model (LLM) to generate reward functions in reinforcement learning, thereby replacing the need for manually designed rewards by experts. However, this approach, which relies on textual descriptions to define learning objectives, fails to achieve controllable and precise behavior learning with clear directionality. In this paper, we introduce a new video2reward method, which directly generates reward functions from videos depicting the behaviors to be mimicked and learned. Specifically, we first process videos containing the target behaviors, converting the motion information of individuals in the videos into keypoint trajectories represented as coordinates through a video2text transforming module. These trajectories are then fed into an LLM to generate the reward function, which in turn is used to train the policy. To enhance the quality of the reward function, we develop a video-assisted iterative reward refinement scheme that visually assesses the learned behaviors and provides textual feedback to the LLM. This feedback guides the LLM to continually refine the reward function, ultimately facilitating more efficient behavior learning. Experimental results on tasks involving bipedal and quadrupedal robot motion control demonstrate that our method surpasses the performance of state-of-the-art LLM-based reward generation methods by over 37.6% in terms of human normalized score. More importantly, by switching video inputs, we find our method can rapidly learn diverse motion behaviors such as walking and running. Runhao Zeng, Dingjie Zhou, Qiwei Liang, Changxin Huang, Jianqiang Li 0001, Xiping Hu |
ECAI | 6 |
| 2023 | Reward-Adaptive Reinforcement Learning: Dynamic Policy Gradient Optimization for Bipedal LocomotionabstractControlling a non-statically bipedal robot is challenging due to the complex dynamics and multi-criterion optimization involved. Recent works have demonstrated the effectiveness of deep reinforcement learning (DRL) for simulation and physical robots. In these methods, the rewards from different criteria are normally summed to learn a scalar function. However, a scalar is less informative and may be insufficient to derive effective information for each reward channel from the complex hybrid rewards. In this work, we propose a novel reward-adaptive reinforcement learning method for biped locomotion, allowing the control policy to be simultaneously optimized by multiple criteria using a dynamic mechanism. The proposed method applies a multi-head critic to learn a separate value function for each reward component, leading to hybrid policy gradients. We further propose dynamic weight, allowing each component to optimize the policy with different priorities. This hybrid and dynamic policy gradient (HDPG) design makes the agent learn more efficiently. We show that the proposed method outperforms summed-up-reward approaches and is able to transfer to physical robots. The MuJoCo results further demonstrate the effectiveness and generalization of HDPG. Changxin Huang, Guangrun Wang, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Deductive Reinforcement Learning for Visual Autonomous Urban Driving NavigationabstractExisting deep reinforcement learning (RL) are devoted to research applications on video games, e.g., The Open Racing Car Simulator (TORCS) and Atari games. However, it remains under-explored for vision-based autonomous urban driving navigation (VB-AUDN). VB-AUDN requires a sophisticated agent working safely in structured, changing, and unpredictable environments; otherwise, inappropriate operations may lead to irreversible or catastrophic damages. In this work, we propose a deductive RL (DeRL) to address this challenge. A deduction reasoner (DR) is introduced to endow the agent with ability to foresee the future and to promote policy learning. Specifically, DR first predicts future transitions through a parameterized environment model. Then, DR conducts self-assessment at the predicted trajectory to perceive the consequences of current policy resulting in a more reliable decision-making process. Additionally, a semantic encoder module (SEM) is designed to extract compact driving representation from the raw images, which is robust to the changes of the environment. Extensive experimental results demonstrate that DeRL outperforms the state-of-the-art model-free RL approaches on the public CAR Learning to Act (CARLA) benchmark and presents a superior performance on success rate and driving safety for goal-directed navigation. Changxin Huang, Meizi Ouyang, Pengxu Wei, Junfan Lin, Jiang Su, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | A heading adjustment method in wireless directional sensor networks
Wei Li 0075, Changxin Huang, Songchen Han |
Comput. Networks | 2 |