Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chenyang Gu

dblp:239/5631 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Robot manipulation · 17% 3D vision · 15% Motion planning and robot control · 12%

Topics — the 24 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.012026
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO · ACL (1) 2026
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
1.012026
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO · ACL (1) 2026
Machine learning › Generative modeling › diffusion model
human motion generation
1.012026
CrowdMoGen: Event-Driven Collective Human Motion Generation · Int. J. Comput. Vis. 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO · ACL (1) 2026
Natural language and speech › Language models and text generation › text generation
scientific hypothesis generation
1.012026
MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › debiasing
selection bias mitigation
1.012026
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO · ACL (1) 2026
Robotics › Robot manipulation › dexterous manipulation
3d object manipulation
0.912025
Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation · CVPR 2025
Computer vision › 3D vision
3d scene understanding
0.912025
SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation · ICRA 2025
Robotics › Robot manipulation
diffusion policy
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › Vision and language › multimodal reasoning
embodied reasoning
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Computer vision › Segmentation and scene understanding › scene understanding
indoor scene understanding
0.912025
SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation · ICRA 2025
Robotics › Robot manipulation
mobile manipulation
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud representation learning
0.912025
Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation · CVPR 2025
Robotics › Motion planning and robot control
robot control
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.912025
SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation · ICRA 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Robotics › Motion planning and robot control
whole-body control
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Machine learning › Generative modeling
motion generation
0.812024
Large Motion Model for Unified Multi-modal Motion Generation · ECCV (13) 2024
Natural language and speech › Language models and text generation
large language model training
0.312026
MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models · ACL (1) 2026
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.312025
Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation · CVPR 2025
Computer vision › 3D vision
multimodal perception
0.312025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › 3D vision › point cloud processing
point cloud and image fusion
0.312025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.0supervised fine-tuning · 1.0permutation-aware training · 1.0event-driven generation · 1.0diffusion model · 1.0contrastive learning · 1.0self-supervised fine-tuning · 0.9masked autoencoder · 0.9cross-attention · 0.92d foundation model lifting · 0.9
YearPublicationVenuePosition
2026 MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models
abstract
Scientific ideation aims to propose novel solutions within a given scientific context.Existing LLM-based agentic approaches emulate human research workflows, yet inadequately model scientific reasoning, resulting in surface-level conceptual recombinations that lack technical depth and scientific grounding.To address this issue, we propose MoRI (Motivation-grounded Reasoning for Scientific Ideation), a framework that enables LLMs to explicitly learn the reasoning process from research motivations to methodologies.The base LLM is initialized via supervised fine-tuning to generate a research motivation from a given context, and is subsequently trained under a composite reinforcement learning reward that approximates scientific rigor: (1) entropy-aware information gain encourages the model to uncover and elaborate high-complexity technical details grounded in ground-truth methodologies, and (2) contrastive semantic gain constrains the reasoning trajectory to remain conceptually aligned with scientifically valid solutions.Empirical results show that MoRI consistently outperforms strong commercial LLMs and complex agentic baselines across multiple dimensions, including novelty, technical rigor, and feasibility.The code is available on GitHub.
Chenyang Gu, Meicong Zhang, Pujun Zheng, Jinquan Zheng, Guoxiu He
ACL (1)1
2026 Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
abstract
Large language models (LLMs) used for multiple-choice and pairwise evaluation tasks often exhibit selection bias due to non-semantic factors like option positions and label symbols.Existing inference-time debiasing is costly and may harm reasoning, while pointwise training ignores that the same question should yield consistent answers across permutations.To address this issue, we propose Permutation-Aware Group Relative Policy Optimization (PA-GRPO), which mitigates selection bias by enforcing permutation-consistent semantic reasoning.PA-GRPO constructs a permutation group for each instance by generating multiple candidate permutations, and optimizes the model using two complementary mechanisms: (1) cross-permutation advantage, which computes advantages relative to the mean reward over all permutations of the same instance, and (2) consistency-aware reward, which encourages the model to produce consistent decisions across different permutations.Experimental results demonstrate that PA-GRPO outperforms strong baselines across seven benchmarks, substantially reducing selection bias while maintaining high overall performance.The code is available on GitHub.
Jinquan Zheng, Jia Yuan, Jiacheng Yao, Chenyang Gu, Pujun Zheng, Guoxiu He
ACL (1)4
2026 CrowdMoGen: Event-Driven Collective Human Motion Generation
Haozhe Xie, Chenyang Gu, Ziwei Liu 0002
Int. J. Comput. Vis.5
2025 Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation
abstract
3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on the explicit extraction of 3D features, while still facing challenges such as the lack of large-scale robotic 3D data and the potential loss of spatial geometry. To address these limitations, we propose the Lift3D framework, which progressively enhances 2D foundation models with implicit and explicit 3D robotic representations to construct a robust 3D manipulation policy. Specifically, we first design a task-aware masked autoencoder that masks task-relevant affordance patches and reconstructs depth information, enhancing the 2D foundation model’s implicit 3D robotic representation. After self-supervised fine-tuning, we introduce a 2D model-lifting strategy that establishes a positional mapping between the input 3D points and the positional embeddings of the 2D model. Based on the mapping, Lift3D utilizes the 2D foundation model to directly encode point cloud data, leveraging large-scale pretrained knowledge to construct explicit 3D robotic representations while minimizing spatial information loss. In experiments, Lift3D consistently outperforms previous state-of-the-art methods across several simulation benchmarks and real-world scenarios.
Yueru Jia, Jiaming Liu 0003, Sixiang Chen, Chenyang Gu, Zhilue Wang, Longzan Luo, Xiaoqi Li 0009, Pengwei Wang 0004, Zhongyuan Wang 0006, Renrui Zhang, Shanghang Zhang
CVPR4
2025 SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation
abstract
3D semantic occupancy prediction is a crucial task in visual perception, as it requires the simultaneous comprehension of both scene geometry and semantics. It plays a crucial role in understanding 3D scenes and has great potential for various applications, such as robotic vision perception and autonomous driving. Many existing works utilize planar-based representations such as Bird's Eye View (BEV) and Tri-Perspective View (TPV). These representations aim to simplify the complexity of 3D scenes while preserving essential object information, thereby facilitating efficient scene representation. However, in dense indoor environments with prevalent occlusions, directly applying these planar-based methods often leads to difficulties in capturing global semantic occupancy, ultimately degrading model performance. In this paper, we present a new vertical slice representation that divides the scene along the vertical axis and projects spatial point features onto the nearest pair of parallel planes. To utilize these slice features, we propose SliceOcc, an RGB camera-based model specifically tailored for indoor 3D semantic occupancy prediction. SliceOcc utilizes pairs of slice queries and cross-attention mechanisms to extract planar features from input images. These local planar features are then fused to form a global scene representation, which is employed for indoor occupancy prediction. Experimental results on the EmbodiedScan dataset demonstrate that SliceOcc achieves a mIoU of 15.45 % across 81 indoor categories, setting a new state-of-the-art performance among RGB camera-based models for indoor 3D semantic occupancy prediction.
Jianing Li 0001, Ming Lu 0002, Hao Wang 0073, Chenyang Gu, Wenzhao Zheng, Shanghang Zhang
ICRA5
2025 RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
abstract
Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others directly use human hand demonstrations to obtain training data. However, the former requires both a robotic system and a skilled operator, limiting scalability, while the latter faces challenges in aligning the visual gap between human hand demonstrations and the deployed robot observations. To address this, we propose a human hand data collection system combined with our hand-to-gripper generative model, which translates human hand demonstrations into robot gripper demonstrations, effectively bridging the observation gap. Specifically, a GoPro fisheye camera is mounted on the human wrist to capture human hand demonstrations. We then train a generative model on a self-collected dataset of paired human hand and UMI gripper demonstrations, which have been processed using a tailored data pre-processing strategy to ensure alignment in both timestamps and observations. Therefore, given only human hand demonstrations, we are able to automatically extract the corresponding SE(3) actions and integrate them with high-quality generated robot demonstrations through our generation pipeline for training robotic policy model. In experiments, the robust manipulation performance demonstrates not only the quality of the generated robot demonstrations but also the efficiency and practicality of our data collection method. More demonstrations can be found at: https://rwor.github.io/.
Liang Heng, Xiaoqi Li 0020, Shangqing Mao, Jiaming Liu 0003, Ruolin Liu, Jingli Wei, Yu-Kai Wang, Yueru Jia, Chenyang Gu, Rui Zhao 0010, Shanghang Zhang, Hao Dong 0003
IROS9
2025 Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
abstract
Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency. To mitigate this dilemma, dual-system approaches have been proposed to leverage a VLM-based System 2 module for handling high-level decision-making, and a separate System 1 action module for ensuring real-time control. However, existing designs maintain both systems as separate models, limiting System 1 from fully leveraging the rich pretrained knowledge from the VLM-based System 2. In this work, we propose Fast-in-Slow (FiS), a unified dual-system vision-language-action (VLA) model that embeds the System 1 execution module within the VLM-based System 2 by partially sharing parameters. This innovative paradigm not only enables high-frequency execution in System 1, but also facilitates coordination between multimodal reasoning and execution components within a single foundation model of System 2. Given their fundamentally distinct roles within FiS-VLA, we design the two systems to incorporate heterogeneous modality inputs alongside asynchronous operating frequencies, enabling both fast and precise manipulation. To enable coordination between the two systems, a dual-aware co-training strategy is proposed that equips System 1 with action generation capabilities while preserving System 2’s contextual understanding to provide stable latent conditions for System 1. For evaluation, FiS-VLA outperforms previous state-of-the-art methods by 8% in simulation and 11% in real-world tasks in terms of average success rate, while achieving a 117.7 Hz control frequency with action chunk set to eight. Project web page: https://fast-in-slow.github.io.
Hao Chen 0193, Jiaming Liu 0003, Chenyang Gu, Zhuoyang Liu, Renrui Zhang, Xiaoqi Li 0020, Yandong Guo, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng
NeurIPS3
2025 AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
abstract
Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicitly model the influence of the mobile base on manipulator control, which easily leads to error accumulation under high degrees of freedom. On the other hand, they treat the entire mobile manipulation process with the same visual observation modality (e.g., either all 2D or all 3D), overlooking the distinct multimodal perception requirements at different stages during mobile manipulation. To address this, we propose the Adaptive Coordination Diffusion Transformer (AC-DiT), which enhances mobile base and manipulator coordination for end-to-end mobile manipulation. First, since the motion of the mobile base directly influences the manipulator's actions, we introduce a mobility-to-body conditioning mechanism that guides the model to first extract base motion representations, which are then used as context prior for predicting whole-body actions. This enables whole-body control that accounts for the potential impact of the mobile base’s motion. Second, to meet the perception requirements at different stages of mobile manipulation, we design a perception-aware multimodal conditioning strategy that dynamically adjusts the fusion weights between various 2D visual images and 3D point clouds, yielding visual features tailored to the current perceptual needs. This allows the model to, for example, adaptively rely more on 2D inputs when semantic information is crucial for action prediction, while placing greater emphasis on 3D geometric information when precise spatial understanding is required. We empirically validate AC-DiT through extensive experiments on both simulated and real-world mobile manipulation tasks, demonstrating superior performance compared to existing methods.
Sixiang Chen, Jiaming Liu 0003, Siyuan Qian, Han Jiang 0003, Zhuoyang Liu, Chenyang Gu, Xiaoqi Li 0009, Chengkai Hou, Pengwei Wang 0004, Zhongyuan Wang 0006, Renrui Zhang, Shanghang Zhang
NeurIPS6
2025 Molecular dynamics simulations of human cohesin subunits identify DNA binding sites and their potential roles in DNA loop extrusion
abstract
The SMC complex cohesin mediates interphase chromatin structural formation in eukaryotic cells through DNA loop extrusion. Here, we sought to investigate its mechanism using molecular dynamics simulations. To achieve this, we first constructed the amino-acid-residue-resolution structural models of the cohesin subunits, SMC1, SMC3, STAG1, and NIPBL. By simulating these subunits with double-stranded DNA molecules, we predicted DNA binding patches on each subunit and quantified the affinities of these patches to DNA using their dissociation rate constants as a proxy. Then, we constructed the structural model of the whole cohesin complex and mapped the predicted high-affinity DNA binding patches on the structure. From the spatial relations of the predicted patches, we identified that multiple patches on the SMC1, SMC3, STAG1, and NIPBL subunits form a DNA clamping patch group. The simulations of the whole complex with double-stranded DNA molecules suggest that this patch group facilitates DNA bending and helps capture a DNA segment in the cohesin ring formed by the SMC1 and SMC3 subunits. In previous studies, these have been identified as critical steps in DNA loop extrusion. Therefore, this study provides experimentally testable predictions of DNA binding sites implicated in previously proposed DNA loop extrusion mechanisms and highlights the essential roles of the accessory subunits STAG1 and NIPBL in the mechanism.
Chenyang Gu, Shoji Takada, Giovanni B. Brandani, Tsuyoshi Terakawa
PLoS Comput. Biol.1
2024 Large Motion Model for Unified Multi-modal Motion Generation
Daisheng Jin, Chenyang Gu, Fangzhou Hong, Zhongang Cai, Jingfang Huang, Chongzhi Zhang, Lei Yang 0045, Ying He 0001, Ziwei Liu 0002
ECCV (13)3
2023 Mask-guided image person removal with data synthesis
abstract
Abstract As a special case of common object removal, image person removal is playing an increasingly important role in social media and criminal investigation domains. Due to the integrity of person area and the complexity of human posture, person removal has its own dilemmas. In this paper, a novel idea is proposed to tackle these problems from the perspective of data synthesis. Concerning the lack of a dedicated dataset for image person removal, two dataset production methods are proposed to automatically generate images, masks and ground truths, respectively. Then, a learning framework similar to local image degradation is proposed so that the masks can be used to guide the feature extraction process and more texture information can be gathered for final prediction. A coarse‐to‐fine training strategy is further applied to refine the details. The data synthesis and learning framework combine well with each other. Experimental results verify the effectiveness of the method quantitatively and qualitatively, and the trained network proves to have good generalization ability either on real or synthetic images.
Yunliang Jiang, Chenyang Gu, Zhenfeng Xue, Xiongtao Zhang, Yong Liu 0007
IET Image Process.2