VLDB 2026 Research / reviewers in the wild / expert
Liang Heng
dblp:129/2783
· DBLP profile ↗
9ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0002-8205-6747ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Robot manipulation · 36% 3D vision · 27% Efficient and distributed learning · 24% | |
| Computer networks
1 paper |
Wireless sensing and localization · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.9 | 2 | 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026 Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025 |
Machine learning › Efficient and distributed learning › adaptive computation
layer skipping |
1.0 | 1 | 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026 |
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
0.9 | 1 | 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025 |
Robotics › Robot manipulation
grasping |
0.9 | 1 | 2025 | Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025 |
Computer vision › Vision and language
visual grounding |
0.9 | 1 | 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025 |
Computer vision › 3D vision › 3d scene understanding › 3d visual grounding
weakly supervised 3d visual grounding |
0.9 | 1 | 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation · AAAI 2026 |
Computer vision › 3D vision › 3d object detection
3d object localization |
0.3 | 1 | 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025 |
Robotics › Robot manipulation
long-horizon task execution |
0.3 | 1 | 2025 | Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation · CVPR 2025 |
Computer vision › 3D vision
point cloud analysis |
0.3 | 1 | 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment · ICRA 2025 |
Wireless sensing and localization
range-based localization |
0.2 | 1 | 2013 | Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013 |
Wireless sensing and localization
sensor network localization |
0.2 | 1 | 2013 | Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013 |
Wireless sensing and localization › localization performance analysis
geometric dilution of precision |
0.0 | 1 | 2013 | Poster abstract: Range-based localization in sensor networks: localizability and accuracy · IPSN 2013 |
Methods — techniques the papers use, named apart from their topics
routing · 1.0mixture of experts · 1.0knowledge distillation · 1.0pre-trained detector · 0.9multimodal prompting · 0.9instance-level alignment · 0.9category-level alignment · 0.9SE(3) pose prediction · 0.9effective degree analysis · 0.2GDOP lower bound · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationabstractVision-Language-Action (VLA) models enable robotic systems to perform embodied tasks but face deployment challenges due to the high computational demands of the dense Large Language Models (LLMs), with existing early-exit-based sparsification methods often overlooking the critical semantic role of final layers in downstream tasks. Aligning with the recent breakthrough of the Shallow Brain Hypothesis (SBH) in neuroscience and the mixture of experts in model sparsification, we conceptualize each LLM layer as an expert and propose a Mixture-of-LayEr Vision Language Action model (MoLe-VLA or simply MoLe) architecture for dynamic LLM layer activation. Specifically, we introduce a Spatial-Temporal Aware Router (STAR) for MoLe to selectively activate only parts of the layers based on the robot’s current state, mimicking the brain's distinct signal pathways specialized for cognition and causal reasoning. Additionally, to compensate for the cognition ability of LLM lost during the layer-skipping, we devise a Cognitive self-Knowledge Distillation (CogKD) to enhance the understanding of task demands and generate task-relevant action sequences by leveraging cognition features. Extensive experiments in RLBench simulations and real-world environments demonstrate the superiority of MoLe-VLA in both efficiency and performance, improving the mean success rate by 9.7% across ten simulation tasks while accelerating inference by 36.8% over OpenVLA. Rongyu Zhang, Menghang Dong, Yuan Zhang 0020, Liang Heng, Xiaowei Chi, Gaole Dai, Dan Wang 0002, Yuan Du, Shanghang Zhang |
AAAI | 4 |
| 2025 | Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic ManipulationabstractIn robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To tackle these challenges, we introduce CrayonRobo that leverages comprehensive multi-modal prompts that explicitly convey both low-level actions and high-level planning in a simple manner. Specifically, for each key-frame in the task sequence, our method allows for manual or automatic generation of simple and expressive 2D visual prompts overlaid on RGB images. These prompts represent the required task goals, such as the end-effector pose and the desired movement direction after contact. We develop a training strategy that enables the model to interpret these visual-language prompts and predict the corresponding contact poses and movement directions in SE(3) space. Furthermore, by sequentially executing all key-frame steps, the model can complete long-horizon tasks. This approach not only helps the model explicitly understand the task objectives but also enhances its robustness on unseen tasks by providing easily interpretable prompts. We evaluate our method in both simulated and real-world environments, demonstrating its robust manipulation capabilities. Xiaoqi Li 0009, Mingxu Zhang, Jiaming Liu 0003, Yan Shen 0035, Iaroslav Ponomarenko, Liang Heng, Siyuan Huang 0004, Shanghang Zhang, Hao Dong 0003 |
CVPR | 8 |
| 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level AlignmentabstractThe 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Category-level ambiguity arises from representing objects of fine-grained categories in a highly sparse point cloud format, making category distinction challenging. Instance-level complexity stems from multiple instances of the same category coexisting in a scene, leading to distractions during grounding. To address these challenges, we propose a novel weaklysupervised grounding approach that explicitly differentiates between categories and instances. In the category-level branch, we utilize extensive category knowledge from a pre-trained external detector to align object proposal features with sentencelevel category features, thereby enhancing category awareness. In the instance-level branch, we utilize spatial relationship descriptions from language queries to refine object proposal features, ensuring clear differentiation among objects. These designs enable our model to accurately identify target-category objects while distinguishing instances within the same category. Compared to previous methods, our approach achieves state-of-the-art performance on three widely used benchmarks: Nr3D, Sr3D, and ScanRef. Xiaoqi Li 0020, Jiaming Liu 0003, Nuowei Han, Liang Heng, Yandong Guo, Hao Dong 0003, Yang Liu 0105 |
ICRA | 4 |
| 2025 | RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without RobotabstractRecent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others directly use human hand demonstrations to obtain training data. However, the former requires both a robotic system and a skilled operator, limiting scalability, while the latter faces challenges in aligning the visual gap between human hand demonstrations and the deployed robot observations. To address this, we propose a human hand data collection system combined with our hand-to-gripper generative model, which translates human hand demonstrations into robot gripper demonstrations, effectively bridging the observation gap. Specifically, a GoPro fisheye camera is mounted on the human wrist to capture human hand demonstrations. We then train a generative model on a self-collected dataset of paired human hand and UMI gripper demonstrations, which have been processed using a tailored data pre-processing strategy to ensure alignment in both timestamps and observations. Therefore, given only human hand demonstrations, we are able to automatically extract the corresponding SE(3) actions and integrate them with high-quality generated robot demonstrations through our generation pipeline for training robotic policy model. In experiments, the robust manipulation performance demonstrates not only the quality of the generated robot demonstrations but also the efficiency and practicality of our data collection method. More demonstrations can be found at: https://rwor.github.io/. Liang Heng, Xiaoqi Li 0020, Shangqing Mao, Jiaming Liu 0003, Ruolin Liu, Jingli Wei, Yu-Kai Wang, Yueru Jia, Chenyang Gu, Rui Zhao 0010, Shanghang Zhang, Hao Dong 0003 |
IROS | 1 |
| 2024 | Lidar-Assisted Hitch Angle Estimation System for Self-Driving TruckabstractLevel-4 autonomous trucks are essential for operating in challenging real-world environments. Accurate estimation of the hitch angle in the tractor-trailer system is crucial for safe maneuvering. Traditional sensor-based measurement techniques for estimating the hitch angle can be complex and expensive. To address this challenge, we propose a lidar-assisted hitch angle estimation (HAE) approach, leveraging existing two-sided lidars installed on the tractor. Overcoming challenges related to accurate lidar extrinsic calibration, time synchronization, online calibration of the trailer’s hitch point, and robustness to environmental changes, our system achieves highly accurate and robust HAE, even in adverse conditions. Our contributions include introducing a novel lidar-assisted HAE system, successfully implementing it in real-time level4 self-driving trucks, and curating a challenging real-world dataset for comprehensive evaluation and benchmarking of HAE in autonomous driving systems. Cansen Jiang, Zhichen Pan, Liang Heng |
IV | 5 |
| 2015 | GNSS Multipath and Jamming Mitigation Using High-Mask-Angle Antennas and Multiple ConstellationsabstractMultipath and jamming interference affects the accuracy, availability, and continuity of global navigation satellite systems. The U.S. Global Positioning System and the Russian Global'naya Navigatsionnaya Sputnikovaya Sistema (GLONASS) are being joined by the European Galileo and the Chinese BeiDou. An increasing number of satellites in multiple constellations enable users to use high-mask-angle antennas (HMAAs) to mitigate interference signals coming from a low-elevation angle. This paper studies the optimal antenna mask angle that maximizes the suppression of interference but still maintains the performance of a single constellation with a low-mask-angle antenna. This paper first proves a novel lower bound on the expectation of dilution of precision (DOP) and derives closed-form formulas that relate the lower bound to the antenna mask angle and the number of satellites. Then, through extensive simulations, a variety of optimal mask angles are obtained with respect to different constellation settings, different DOP metrics, and different assumptions of range accuracy. The numerical results highly agree with our theory. Both of them show that two constellations can match the performance of one constellation with a 5°-14° higher mask, and three constellations can match the performance of one constellation with an 11°-23° higher mask, depending on the DOP metric and the range error model used. The numerical results also show that using HMAAs is more beneficial to users interested in positioning accuracy than to users interested in time transfer accuracy. Liang Heng, Todd Walter, Per K. Enge, Grace Xingxin Gao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | GPS Signal Authentication From Cooperative PeersabstractSecure reliable position information is indispensable for many transportation systems and services, such as traffic monitoring, fleet management, electronic toll collection, route guidance, vehicle telematics, and emergency response. Unfortunately, civil Global Positioning System (GPS) signals are vulnerable to spoofing attacks. This paper introduces a signal authentication architecture based on a network of cooperative GPS receivers. A receiver in the network correlates its received military P(Y) signal with those received by other receivers (hereinafter referred to as cross-check receivers) to detect spoofing attacks. This paper describes three candidate structures to implement this architecture and evaluates spoofing detection performance through theoretical analyses and field experiments. We show that the spoofing detection performance improves exponentially with increasing number of cross-check receivers. Even if the cross-check receivers are low cost, unreliable, and in challenging environment, cooperative authentication can match, if not outperform, a single high-quality reliable reference receiver in terms of spoofing detection performance. Liang Heng, Daniel B. Work, Grace Xingxin Gao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Poster abstract: Range-based localization in sensor networks: localizability and accuracyabstractLocalizability and accuracy are two fundamental problems in many range-based localization schemes for sensor networks. This poster addresses the two problems theoretically by introducing two new concepts: the effective degree and the lower bound of geometric dilution of precision (GDOP). We prove that the network is not localizable unless the average effective degree is greater than or equal to the dimension of the location space. We further show that the average localization accuracy is approximately inversely proportional to the average degree. Liang Heng, Grace Xingxin Gao |
IPSN | 1 |
| 2013 | Accuracy of range-based localization schemes in random sensor networks: A lower bound analysisabstractAccuracy is a fundamental performance requirement in network localization. This paper studies the accuracy of range-based localization schemes for random sensor networks with respect to network connectivity and scale. We show that the variance of localization errors is proportional to the average geometric dilution of precision (AGDOP). The paper proves a novel lower bound of expectation of AGDOP (LB-E-AGDOP). Our analysis based on LB-E-AGDOP shows that localization accuracy is approximately inversely proportional to the average degree of network. A further analysis shows that when network connectivity merely guarantees localizability, increasing sensor nodes leads to bounded monotonic increase in AGDOP; when a network is densely connected, increasing sensor nodes leads to bounded monotonic decrease in AGDOP. Finally, these conclusions are validated by numerical simulations. Liang Heng, Grace Xingxin Gao |
IROS | 1 |