Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Liang Heng

dblp:129/2783 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0002-8205-6747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 36% 3D vision · 27% Efficient and distributed learning · 24%
Computer networks
1 paper
Wireless sensing and localization · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.922026
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026
Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025
Machine learning › Efficient and distributed learning › adaptive computation
layer skipping
1.012026
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
0.912025
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025
Robotics › Robot manipulation
grasping
0.912025
Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025
Computer vision › Vision and language
visual grounding
0.912025
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025
Computer vision › 3D vision › 3d scene understanding › 3d visual grounding
weakly supervised 3d visual grounding
0.912025
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.312026
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026
Computer vision › 3D vision › 3d object detection
3d object localization
0.312025
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025
Robotics › Robot manipulation
long-horizon task execution
0.312025
Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025
Computer vision › 3D vision
point cloud analysis
0.312025
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025
Wireless sensing and localization
range-based localization
0.212013
Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013
Wireless sensing and localization
sensor network localization
0.212013
Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013
Wireless sensing and localization › localization performance analysis
geometric dilution of precision
0.012013
Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013

Methods — techniques the papers use, named apart from their topics

routing · 1.0mixture of experts · 1.0knowledge distillation · 1.0pre-trained detector · 0.9multimodal prompting · 0.9instance-level alignment · 0.9category-level alignment · 0.9SE(3) pose prediction · 0.9effective degree analysis · 0.2GDOP lower bound · 0.2
YearPublicationVenuePosition
2026 MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
abstract
Vision-Language-Action (VLA) models enable robotic systems to perform embodied tasks but face deployment challenges due to the high computational demands of the dense Large Language Models (LLMs), with existing early-exit-based sparsification methods often overlooking the critical semantic role of final layers in downstream tasks. Aligning with the recent breakthrough of the Shallow Brain Hypothesis (SBH) in neuroscience and the mixture of experts in model sparsification, we conceptualize each LLM layer as an expert and propose a Mixture-of-LayEr Vision Language Action model (MoLe-VLA or simply MoLe) architecture for dynamic LLM layer activation. Specifically, we introduce a Spatial-Temporal Aware Router (STAR) for MoLe to selectively activate only parts of the layers based on the robot’s current state, mimicking the brain's distinct signal pathways specialized for cognition and causal reasoning. Additionally, to compensate for the cognition ability of LLM lost during the layer-skipping, we devise a Cognitive self-Knowledge Distillation (CogKD) to enhance the understanding of task demands and generate task-relevant action sequences by leveraging cognition features. Extensive experiments in RLBench simulations and real-world environments demonstrate the superiority of MoLe-VLA in both efficiency and performance, improving the mean success rate by 9.7% across ten simulation tasks while accelerating inference by 36.8% over OpenVLA.
Rongyu Zhang, Menghang Dong, Yuan Zhang 0020, Liang Heng, Xiaowei Chi, Gaole Dai, Dan Wang 0002, Yuan Du, Shanghang Zhang
AAAI4
2025 Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
abstract
In robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To tackle these challenges, we introduce CrayonRobo that leverages comprehensive multi-modal prompts that explicitly convey both low-level actions and high-level planning in a simple manner. Specifically, for each key-frame in the task sequence, our method allows for manual or automatic generation of simple and expressive 2D visual prompts overlaid on RGB images. These prompts represent the required task goals, such as the end-effector pose and the desired movement direction after contact. We develop a training strategy that enables the model to interpret these visual-language prompts and predict the corresponding contact poses and movement directions in SE(3) space. Furthermore, by sequentially executing all key-frame steps, the model can complete long-horizon tasks. This approach not only helps the model explicitly understand the task objectives but also enhances its robustness on unseen tasks by providing easily interpretable prompts. We evaluate our method in both simulated and real-world environments, demonstrating its robust manipulation capabilities.
Xiaoqi Li 0009, Mingxu Zhang, Jiaming Liu 0003, Yan Shen 0035, Iaroslav Ponomarenko, Liang Heng, Siyuan Huang 0004, Shanghang Zhang, Hao Dong 0003
CVPR8
2025 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
abstract
The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Category-level ambiguity arises from representing objects of fine-grained categories in a highly sparse point cloud format, making category distinction challenging. Instance-level complexity stems from multiple instances of the same category coexisting in a scene, leading to distractions during grounding. To address these challenges, we propose a novel weaklysupervised grounding approach that explicitly differentiates between categories and instances. In the category-level branch, we utilize extensive category knowledge from a pre-trained external detector to align object proposal features with sentencelevel category features, thereby enhancing category awareness. In the instance-level branch, we utilize spatial relationship descriptions from language queries to refine object proposal features, ensuring clear differentiation among objects. These designs enable our model to accurately identify target-category objects while distinguishing instances within the same category. Compared to previous methods, our approach achieves state-of-the-art performance on three widely used benchmarks: Nr3D, Sr3D, and ScanRef.
Xiaoqi Li 0020, Jiaming Liu 0003, Nuowei Han, Liang Heng, Yandong Guo, Hao Dong 0003, Yang Liu 0105
ICRA4
2025 RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
abstract
Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others directly use human hand demonstrations to obtain training data. However, the former requires both a robotic system and a skilled operator, limiting scalability, while the latter faces challenges in aligning the visual gap between human hand demonstrations and the deployed robot observations. To address this, we propose a human hand data collection system combined with our hand-to-gripper generative model, which translates human hand demonstrations into robot gripper demonstrations, effectively bridging the observation gap. Specifically, a GoPro fisheye camera is mounted on the human wrist to capture human hand demonstrations. We then train a generative model on a self-collected dataset of paired human hand and UMI gripper demonstrations, which have been processed using a tailored data pre-processing strategy to ensure alignment in both timestamps and observations. Therefore, given only human hand demonstrations, we are able to automatically extract the corresponding SE(3) actions and integrate them with high-quality generated robot demonstrations through our generation pipeline for training robotic policy model. In experiments, the robust manipulation performance demonstrates not only the quality of the generated robot demonstrations but also the efficiency and practicality of our data collection method. More demonstrations can be found at: https://rwor.github.io/.
Liang Heng, Xiaoqi Li 0020, Shangqing Mao, Jiaming Liu 0003, Ruolin Liu, Jingli Wei, Yu-Kai Wang, Yueru Jia, Chenyang Gu, Rui Zhao 0010, Shanghang Zhang, Hao Dong 0003
IROS1
2024 Lidar-Assisted Hitch Angle Estimation System for Self-Driving Truck
abstract
Level-4 autonomous trucks are essential for operating in challenging real-world environments. Accurate estimation of the hitch angle in the tractor-trailer system is crucial for safe maneuvering. Traditional sensor-based measurement techniques for estimating the hitch angle can be complex and expensive. To address this challenge, we propose a lidar-assisted hitch angle estimation (HAE) approach, leveraging existing two-sided lidars installed on the tractor. Overcoming challenges related to accurate lidar extrinsic calibration, time synchronization, online calibration of the trailer’s hitch point, and robustness to environmental changes, our system achieves highly accurate and robust HAE, even in adverse conditions. Our contributions include introducing a novel lidar-assisted HAE system, successfully implementing it in real-time level4 self-driving trucks, and curating a challenging real-world dataset for comprehensive evaluation and benchmarking of HAE in autonomous driving systems.
Cansen Jiang, Zhichen Pan, Liang Heng
IV5
2015 GNSS Multipath and Jamming Mitigation Using High-Mask-Angle Antennas and Multiple Constellations
abstract
Multipath and jamming interference affects the accuracy, availability, and continuity of global navigation satellite systems. The U.S. Global Positioning System and the Russian Global'naya Navigatsionnaya Sputnikovaya Sistema (GLONASS) are being joined by the European Galileo and the Chinese BeiDou. An increasing number of satellites in multiple constellations enable users to use high-mask-angle antennas (HMAAs) to mitigate interference signals coming from a low-elevation angle. This paper studies the optimal antenna mask angle that maximizes the suppression of interference but still maintains the performance of a single constellation with a low-mask-angle antenna. This paper first proves a novel lower bound on the expectation of dilution of precision (DOP) and derives closed-form formulas that relate the lower bound to the antenna mask angle and the number of satellites. Then, through extensive simulations, a variety of optimal mask angles are obtained with respect to different constellation settings, different DOP metrics, and different assumptions of range accuracy. The numerical results highly agree with our theory. Both of them show that two constellations can match the performance of one constellation with a 5°-14° higher mask, and three constellations can match the performance of one constellation with an 11°-23° higher mask, depending on the DOP metric and the range error model used. The numerical results also show that using HMAAs is more beneficial to users interested in positioning accuracy than to users interested in time transfer accuracy.
Liang Heng, Todd Walter, Per K. Enge, Grace Xingxin Gao
IEEE Trans. Intell. Transp. Syst.1
2015 GPS Signal Authentication From Cooperative Peers
abstract
Secure reliable position information is indispensable for many transportation systems and services, such as traffic monitoring, fleet management, electronic toll collection, route guidance, vehicle telematics, and emergency response. Unfortunately, civil Global Positioning System (GPS) signals are vulnerable to spoofing attacks. This paper introduces a signal authentication architecture based on a network of cooperative GPS receivers. A receiver in the network correlates its received military P(Y) signal with those received by other receivers (hereinafter referred to as cross-check receivers) to detect spoofing attacks. This paper describes three candidate structures to implement this architecture and evaluates spoofing detection performance through theoretical analyses and field experiments. We show that the spoofing detection performance improves exponentially with increasing number of cross-check receivers. Even if the cross-check receivers are low cost, unreliable, and in challenging environment, cooperative authentication can match, if not outperform, a single high-quality reliable reference receiver in terms of spoofing detection performance.
Liang Heng, Daniel B. Work, Grace Xingxin Gao
IEEE Trans. Intell. Transp. Syst.1
2013 Poster abstract: Range-based localization in sensor networks: localizability and accuracy
abstract
Localizability and accuracy are two fundamental problems in many range-based localization schemes for sensor networks. This poster addresses the two problems theoretically by introducing two new concepts: the effective degree and the lower bound of geometric dilution of precision (GDOP). We prove that the network is not localizable unless the average effective degree is greater than or equal to the dimension of the location space. We further show that the average localization accuracy is approximately inversely proportional to the average degree.
Liang Heng, Grace Xingxin Gao
IPSN1
2013 Accuracy of range-based localization schemes in random sensor networks: A lower bound analysis
abstract
Accuracy is a fundamental performance requirement in network localization. This paper studies the accuracy of range-based localization schemes for random sensor networks with respect to network connectivity and scale. We show that the variance of localization errors is proportional to the average geometric dilution of precision (AGDOP). The paper proves a novel lower bound of expectation of AGDOP (LB-E-AGDOP). Our analysis based on LB-E-AGDOP shows that localization accuracy is approximately inversely proportional to the average degree of network. A further analysis shows that when network connectivity merely guarantees localizability, increasing sensor nodes leads to bounded monotonic increase in AGDOP; when a network is densely connected, increasing sensor nodes leads to bounded monotonic decrease in AGDOP. Finally, these conclusions are validated by numerical simulations.
Liang Heng, Grace Xingxin Gao
IROS1