Zhan Ling

dblp:254/1980 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 20% Language models and text generation · 19% Vision and language · 11%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
context compression
1.012026
Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression · ACL (1) 2026
Machine learning › Reinforcement learning
LLM agent training
1.012026
Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression · ACL (1) 2026
Natural language and speech › Language models and text generation
in-context learning
0.912025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › procedural reasoning
multimodal procedural planning
0.812024
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities · EMNLP 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › task planning
procedure planning
0.812024
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities · EMNLP 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities · EMNLP 2024
Computer vision › 3D vision › 3d shape analysis › 3d shape segmentation
3d part segmentation
0.712023
PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models · CVPR 2023
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023
Computer vision › 3D vision › range sensing
depth sensing
0.712023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023
Robotics › Motion planning and robot control › robot learning
manipulation skill learning
0.712023
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills · ICLR 2023
Machine learning › Reinforcement learning
model-based reinforcement learning
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.712023
Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023
Machine learning › Reinforcement learning
policy learning
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Machine learning › Trustworthy machine learning › verification
self-verification
0.712023
Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.712023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
Robotics › Motion planning and robot control
trajectory optimization
0.712023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Computer vision › Vision and language
vision-language model distillation
0.712023
Distilling Large Vision-Language Model with Out-of-Distribution Generalizability · ICCV 2023
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.612022
Improving Policy Optimization with Generalist-Specialist Learning · ICML 2022
Machine learning › Reinforcement learning
policy optimization
0.612022
Improving Policy Optimization with Generalist-Specialist Learning · ICML 2022
Geometric modeling and processing › shape decomposition
approximate convex decomposition
0.612022
Approximate convex decomposition for 3D meshes with collision-aware concavity and tree search · ACM Trans. Graph. 2022
Geometric modeling and processing
shape decomposition
0.612022
Approximate convex decomposition for 3D meshes with collision-aware concavity and tree search · ACM Trans. Graph. 2022
Machine learning › Reinforcement learning
imitation learning
0.412020
State Alignment-based Imitation Learning · ICLR 2020
Robotics › Motion planning and robot control
robot learning
0.422023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills · ICLR 2023
Natural language and speech › Language models and text generation › large language model reasoning
long-context reasoning
0.312025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.212023
Reparameterized Policy Learning for Multimodal Trajectory Optimization · ICML 2023
Computer vision › 3D vision
point cloud analysis
0.212023
PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models · CVPR 2023
Automated reasoning and model checking › program verification
deductive verification
0.212023
Deductive Verification of Chain-of-Thought Reasoning · NeurIPS 2023
Geometric modeling and processing
collision detection
0.212022
Approximate convex decomposition for 3D meshes with collision-aware concavity and tree search · ACM Trans. Graph. 2022

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.4summarization · 1.0policy gradient · 1.0retrieval-augmented generation · 0.9many-shot ICL · 0.9large language model · 0.8fine-tuning · 0.8prompt tuning · 0.7natural program · 0.7multi-view rendering · 0.7knowledge distillation · 0.7chain-of-thought prompting · 0.7tree search · 0.6plane cutting · 0.6collision-aware concavity · 0.6
YearPublicationVenuePosition
2026 Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression
abstract
We study reinforcement learning (RL) finetuning of large language model (LLM) agents for long-horizon multi-turn tool use, where context length quickly becomes a fundamental bottleneck.Existing multi-turn RL pipelines suffer from degraded instruction following, excessive rollout costs, and most importantly, strict context limits.In this work, to address these challenges, we introduce summarization-based context management to training.In specific, it periodically compresses the tool using history by LLM-generated summaries that retain taskrelevant information to keep a compact context while enabling the agent to scale beyond the fixed context window.Building on this formulation, we derive a policy gradient representation that seamlessly enables standard LLM RL infrastructures to optimize both tool-use behaviors as well as summarization strategies in an end-to-end fashion.We instantiate this framework with SUmmarization augmented Policy Optimization (SUPO), an LLM RL algorithm that enables long-horizon training beyond a fixed context limit.Experiments on interactive function calling and searching tasks demonstrate that SUPO significantly improves the success rate while maintaining the same or even lower working context length compared to baselines.We also demonstrate that for complex searching tasks SUPO can further improve the evaluation performance when scaling test-time maximum round of summarization beyond that of training time.
Miao Lu, Weihua Du, Zhan Ling, Xuesong Yao, Jiecao Chen
ACL (1)4
2025 MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
abstract
The ability to recognize patterns from examples and apply them to new ones is a primal ability for general intelligence, and is widely studied by psychology and AI researchers. Many benchmarks have been proposed to measure such ability for Large Language Models (LLMs); however, they focus on few-shot (usually <10) setting and lack evaluation for aggregating many pieces of information from long contexts. On the other hand, the ever-growing context length of LLMs have brought forth the novel paradigm of many-shot In-Context Learning (ICL), which addresses new tasks with hundreds to thousands of examples without expensive and inefficient fine-tuning. However, many-shot evaluations often focus on classification, and popular long-context LLM tasks such as Needle-In-A-Haystack (NIAH) seldom require complicated intelligence for integrating many pieces of information. To fix the issues from both worlds, we propose MIR-Bench, the first many-shot in-context reasoning benchmark for pattern recognition that asks LLM to predict output via input-output examples from underlying functions with diverse data format. Based on MIR-Bench, we study many novel problems for many-shot in-context reasoning, and acquired many insightful findings including scaling effect, robustness, inductive vs. transductive reasoning, retrieval Augmented Generation (RAG), coding for inductive reasoning, cross-domain generalizability, etc. Our dataset is available at https://huggingface.co/datasets/kaiyan289/MIR-Bench.
Zhan Ling, Ting-Han Fan, Lingfeng Shen, Zhengyin Du, Jiecao Chen
NeurIPS2
2024 ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
abstract
Large language models (LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability.However, there is still lack of study on how vision language models (VLMs) behave when multi-modal task inputs are considered.Counterfactual planning that evaluates the model's reasoning ability over alternative task situations are also under exploited.In order to evaluate the planning ability of both multimodal and counterfactual aspects, we propose ActPlan-1K.ActPlan-1K is a multi-modal planning benchmark constructed based on ChatGPT and household activity simulator iGibson2.The benchmark consists of 153 activities and 1,187 instances.Each instance describing one activity has a natural language task description and multiple environment images from the simulator.The gold plan of each instance is action sequences over the objects in provided scenes.Both the correctness and commonsense satisfaction are evaluated on typical VLMs.It turns out that current VLMs are still struggling at generating human-level procedural plans for both normal activities and counterfactual activities.We further provide automatic evaluation metrics by finetuning over BLEURT model to facilitate future research on our benchmark.
Zhan Ling, Cheng Jiayang, Yauwai Yim, Yangqiu Song
EMNLP2
2023 PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models
abstract
Generalizable 3D part segmentation is important but challenging in vision and robotics. Training deep models via conventional supervised methods requires large-scale 3D datasets with fine-grained part annotations, which are costly to collect. This paper explores an alternative way for low-shot part segmentation of 3D point clouds by leveraging a pretrained image-language model, GLIP. which achieves superior performance on open-vocabulary 2D detection. We transfer the rich knowledge from 2D to 3D through GLIP-based part detection on point cloud rendering and a novel 2D-to-3D label lifting algorithm. We also utilize multi-view 3D priors and few-shot prompt tuning to boost performance significantly. Extensive evaluation on PartNet and PartNet-Mobility datasets shows that our method enables excellent zero-shot 3D part segmentation. Our few-shot version not only outperforms existing few-shot approaches by a large margin but also achieves highly competitive results compared to the fully supervised counterpart. Furthermore, we demonstrate that our method can be directly applied to iPhone-scanned point clouds without significant domain gaps.
Minghua Liu, Yinhao Zhu, Shizhong Han, Zhan Ling, Fatih Porikli, Hao Su 0001
CVPR5
2023 Distilling Large Vision-Language Model with Out-of-Distribution Generalizability
abstract
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distillation, the process of creating smaller, faster models that maintain the performance of larger models, is a promising direction towards the solution. This paper investigates the distillation of visual representations in large teacher vision-language models into lightweight student models using a small- or mid-scale dataset. Notably, this study focuses on open-vocabulary out-of-distribution (OOD) generalization, a challenging problem that has been overlooked in previous model distillation literature. We propose two principles from vision and language modality perspectives to enhance student’s OOD generalization: (1) by better imitating teacher’s visual representation space, and carefully promoting better coherence in vision-language alignment with the teacher; (2) by enriching the teacher’s language representations with informative and fine-grained semantic attributes to effectively distinguish between different labels. We propose several metrics and conduct extensive experiments to investigate their techniques. The results demonstrate significant improvements in zero-shot and few-shot student performance on open-vocabulary out-of-distribution classification, highlighting the effectiveness of our proposed approaches. Code released at this link.
Yunhao Fang, Minghua Liu, Zhan Ling, Zhuowen Tu, Hao Su 0001
ICCV4
2023 ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Jiayuan Gu, Fanbo Xiang, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen 0019, Hao Su 0001
ICLR4
2023 Reparameterized Policy Learning for Multimodal Trajectory Optimization
abstract
We investigate the challenge of parametrizing policies for reinforcement learning (RL) in high-dimensional continuous action spaces. Our objective is to develop a multimodal policy that overcomes limitations inherent in the commonly-used Gaussian parameterization. To achieve this, we propose a principled framework that models the continuous RL policy as a generative model of optimal trajectories. By conditioning the policy on a latent variable, we derive a novel variational bound as the optimization objective, which promotes exploration of the environment. We then present a practical model-based RL method, called Reparameterized Policy Gradient (RPG), which leverages the multimodal policy parameterization and learned world model to achieve strong exploration capabilities and high data efficiency. Empirical results demonstrate that our method can help agents evade local optima in tasks with dense rewards and solve challenging sparse-reward environments by incorporating an object-centric intrinsic reward. Our method consistently outperforms previous approaches across a range of tasks. Code and supplementary materials are available on the project page https://haosulab.github.io/RPG/
Zhiao Huang, Litian Liang, Zhan Ling, Chuang Gan 0001, Hao Su 0001
ICML3
2023 Deductive Verification of Chain-of-Thought Reasoning
abstract
Large Language Models (LLMs) significantly benefit from Chain-of-thought (CoT) prompting in performing various reasoning tasks. While CoT allows models to produce more comprehensive reasoning processes, its emphasis on intermediate reasoning steps can inadvertently introduce hallucinations and accumulated errors, thereby limiting models’ ability to solve complex reasoning tasks. Inspired by how humans engage in careful and meticulous deductive logical reasoning processes to solve tasks, we seek to enable language models to perform explicit and rigorous deductive reasoning, and also ensure the trustworthiness of their reasoning process through self-verification. However, directly verifying the validity of an entire deductive reasoning process is challenging, even with advanced models like ChatGPT. In light of this, we propose to decompose a reasoning verification process into a series of step-by-step subprocesses, each only receiving their necessary context and premises. To facilitate this procedure, we propose Natural Program, a natural language-based deductive reasoning format. Our approach enables models to generate precise reasoning steps where subsequent steps are more rigorously grounded on prior steps. It also empowers language models to carry out reasoning self-verification in a step-by-step manner. By integrating this verification process into each deductive reasoning stage, we significantly enhance the rigor and trustfulness of generated reasoning steps. Along this process, we also improve the answer correctness on complex reasoning tasks.
Zhan Ling, Yunhao Fang, Zhiao Huang, Mingu Lee, Roland Memisevic, Hao Su 0001
NeurIPS1
2023 Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation
abstract
In this article, we focus on the simulation of active stereovision depth sensors, which are popular in both academic and industry communities. Inspired by the underlying mechanism of the sensors, we designed a fully physics-grounded simulation pipeline that includes material acquisition, ray-tracing-based infrared (IR) image rendering, IR noise simulation, and depth estimation. The pipeline is able to generate depth maps with material-dependent error patterns similar to a real depth sensor in real time. We conduct real experiments to show that perception algorithms and reinforcement learning policies trained in our simulation platform could transfer well to the real-world test cases without any fine-tuning. Furthermore, due to the high degree of realism of this simulation, our depth sensor simulator can be used as a convenient testbed to evaluate the algorithm performance in the real world, which will largely reduce the human effort in developing robotic algorithms. The entire pipeline has been integrated into the SAPIEN simulator and is open-sourced to promote the research of vision and robotics communities.
Xiaoshuai Zhang, Rui Chen 0019, Ang Li 0010, Fanbo Xiang, Yuzhe Qin, Jiayuan Gu, Zhan Ling, Minghua Liu, Peiyu Zeng, Songfang Han, Zhiao Huang, Tongzhou Mu, Jing Xu 0011, Hao Su 0001
IEEE Trans. Robotics7
2022 Improving Policy Optimization with Generalist-Specialist Learning
abstract
Generalization in deep reinforcement learning over unseen environment variations usually requires policy learning over a large set of diverse training variations. We empirically observe that an agent trained on many variations (a generalist) tends to learn faster at the beginning, yet its performance plateaus at a less optimal level for a long time. In contrast, an agent trained only on a few variations (a specialist) can often achieve high returns under a limited computational budget. To have the best of both worlds, we propose a novel generalist-specialist training framework. Specifically, we first train a generalist on all environment variations; when it fails to improve, we launch a large population of specialists with weights cloned from the generalist, each trained to master a selected small subset of variations. We finally resume the training of the generalist with auxiliary rewards induced by demonstrations of all specialists. In particular, we investigate the timing to start specialist training and compare strategies to learn generalists with assistance from specialists. We show that this framework pushes the envelope of policy learning on several challenging and popular benchmarks including Procgen, Meta-World and ManiSkill.
Zhiwei Jia, Zhan Ling, Yiran Wu, Hao Su 0001
ICML3
2022 Approximate convex decomposition for 3D meshes with collision-aware concavity and tree search
abstract
Approximate convex decomposition aims to decompose a 3D shape into a set of almost convex components, whose convex hulls can then be used to represent the input shape. It thus enables efficient geometry processing algorithms specifically designed for convex shapes and has been widely used in game engines, physics simulations, and animation. While prior works can capture the global structure of input shapes, they may fail to preserve fine-grained details (e.g., filling a toaster's slots), which are critical for retaining the functionality of objects in interactive environments. In this paper, we propose a novel method that addresses the limitations of existing approaches from three perspectives: (a) We introduce a novel collision-aware concavity metric that examines the distance between a shape and its convex hull from both the boundary and the interior. The proposed concavity preserves collision conditions and is more robust to detect various approximation errors. (b) We decompose shapes by directly cutting meshes with 3D planes. It ensures generated convex hulls are intersection-free and avoids voxelization errors. (c) Instead of using a one-step greedy strategy, we propose employing a multi-step tree search to determine the cutting planes, which leads to a globally better solution and avoids unnecessary cuttings. Through extensive evaluation on a large-scale articulated object dataset, we show that our method generates decompositions closer to the original shape with fewer components. It thus supports delicate and efficient object interaction in downstream applications.
Xinyue Wei, Minghua Liu, Zhan Ling, Hao Su 0001
ACM Trans. Graph.3
2020 State Alignment-based Imitation Learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, Hao Su 0001
ICLR2