EDBT 2026 Demo / reviewers in the wild / expert
Chilam Cheang
dblp:295/9397
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Robot manipulation · 27% 3D vision · 26% Generative modeling · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Interaction techniques and input · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
grasping |
1.3 | 3 | 2022 | I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches · ICRA 2022 Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language Instructions · ICRA 2022 SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Computer vision › 3D vision › object pose estimation
6d object pose estimation |
1.1 | 2 | 2022 | Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language Instructions · ICRA 2022 SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Computer vision › Video understanding and tracking › video prediction
action-conditioned video prediction |
0.9 | 1 | 2025 | IRASim: A Fine-Grained World Model for Robot Manipulation · ICCV 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | IRASim: A Fine-Grained World Model for Robot Manipulation · ICCV 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | IRASim: A Fine-Grained World Model for Robot Manipulation · ICCV 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.9 | 1 | 2025 | IRASim: A Fine-Grained World Model for Robot Manipulation · ICCV 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | Vision-Language Foundation Models as Effective Robot Imitators · ICLR 2024 |
Robotics › Motion planning and robot control
robot learning |
0.8 | 1 | 2024 | Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation · ICLR 2024 |
Computer vision › Vision and language › vision-language model › vision-language model adaptation
vision-language model fine-tuning |
0.8 | 1 | 2024 | Vision-Language Foundation Models as Effective Robot Imitators · ICLR 2024 |
Computer vision › 3D vision
3d shape reconstruction |
0.6 | 1 | 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
category-level object pose estimation |
0.6 | 1 | 2022 | Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language Instructions · ICRA 2022 |
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
category-level pose and size estimation |
0.6 | 1 | 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Computer vision › 3D vision › 3d shape reconstruction
category-level shape reconstruction |
0.6 | 1 | 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Robotics › Robot manipulation › grasping
grasp detection |
0.6 | 1 | 2022 | I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches · ICRA 2022 |
Machine learning › Reinforcement learning
policy evaluation |
0.3 | 1 | 2025 | IRASim: A Fine-Grained World Model for Robot Manipulation · ICCV 2025 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.2 | 1 | 2024 | Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation · ICLR 2024 |
Robotics › Robot manipulation › grasping › unknown object grasping
category-level grasping |
0.2 | 1 | 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation · CVPR 2022 |
Interaction techniques and input › sketch-based interaction
sketch recognition |
0.2 | 1 | 2022 | I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches · ICRA 2022 |
Methods — techniques the papers use, named apart from their topics
imitation learning · 1.5diffusion transformer · 0.9action conditioning · 0.9vision-language model · 0.8transformer · 0.8policy head · 0.8generative pre-training · 0.8symmetric correspondence prediction · 0.6point cloud deformation · 0.6language-conditioned grounding · 0.6graph neural network · 0.6end-to-end learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IRASim: A Fine-Grained World Model for Robot ManipulationabstractWorld models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the visual space using existing methods which overlooks precise alignment between each action and the corresponding frame. In this paper, we present IRASim, a novel world model capable of generating videos with fine-grained robotobject interaction details, conditioned on historical observations and robot action trajectories. We train a diffusion transformer and introduce a novel frame-level actionconditioning module within each transformer block to explicitly model and strengthen the action-frame alignment. Extensive experiments show that: (1) the quality of the videos generated by our method surpasses all the baseline methods and scales effectively with increased model size and computation; (2) policy evaluations using IRASim exhibit a strong correlation with those using the ground-truth simulator, highlighting its potential to accelerate real-world policy evaluation; (3) testing-time scaling through model-based planning with IRASim significantly enhances policy performance, as evidenced by an improvement in the IoU metric on the Push-T benchmark from 0.637 to 0.961; (4) IRASim provides flexible action controllability, allowing virtual robotic arms in datasets to be controlled via a keyboard or VR controller. Video and code are available at https://gen-irasim.github.io/. Fangqi Zhu, Song Guo 0001, Chilam Cheang, Tao Kong |
ICCV | 5 |
| 2024 | Vision-Language Foundation Models as Effective Robot ImitatorsabstractRecent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data.
To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control.
Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy. Our code will be made public upon acceptance. Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Chilam Cheang, Ya Jing, Weinan Zhang 0001, Huaping Liu 0001, Tao Kong |
ICLR | 7 |
| 2024 | Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationabstractGenerative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training. We introduce GR-1, a GPT-style model designed for multi-task language-conditioned visual robot manipulation. GR-1 takes as inputs a language instruction, a sequence of observation images, and a sequence of robot states. It predicts robot actions as well as future images in an end-to-end manner. Thanks to a flexible design, GR-1 can be seamlessly finetuned on robot data after pre-trained on a large-scale video dataset. We perform extensive experiments on the challenging CALVIN benchmark and a real robot. On CALVIN benchmark, our method outperforms state-of-the-art baseline methods and improves the success rate from 88.9% to 94.9%. In the setting of zero-shot unseen scene generalization, GR-1 improves the success rate from 53.3% to 85.4%. In real robot experiments, GR-1 also outperforms baseline methods and shows strong potentials in generalization to unseen scenes and objects. We provide inaugural evidence that a unified GPT-style transformer, augmented with large-scale video generative pre-training, exhibits remarkable generalization to multi-task visual robot manipulation. Project page: https://GR1-Manipulation.github.io Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Tao Kong |
ICLR | 3 |
| 2023 | Multi-view Shape Generation for a 3D Human-like BodyabstractThree-dimensional (3D) human-like body reconstruction via a single RGB image has attracted significant research attention recently. Most of the existing methods rely on the Skinned Multi-Person Linear model and thus can only predict unified human bodies. Moreover, meshes reconstructed by current methods sometimes perform well from a canonical view but not from other views, as the reconstruction process is commonly supervised by only a single view. To address these limitations, this article proposes a multi-view shape generation network for a 3D human-like body. Particularly, we propose a coarse-to-fine learning model that gradually deforms a template body toward the ground truth body. Our model utilizes the information of multi-view renderings and corresponding 3D vertex transformation as supervision. Such supervision will help to generate 3D bodies well aligned to all views. To accurately operate mesh deformation, a graph convolutional network structure is introduced to support the shape generation from 3D vertex representation. Additionally, a graph up-pooling operation is designed over the intermediate representations of the graph convolutional network, and thus our model can generate 3D shapes with higher resolution. Novel loss functions are employed to help optimize the whole multi-view generation model, resulting in smoother surfaces. In addition, two multi-view human body datasets are produced and contributed to the community. Extensive experiments conducted on the benchmark datasets demonstrate the efficacy of our model over the competitors. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size EstimationabstractGiven a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the visual cues in RGB images, we rely on the shape information predominately from the depth (D) channel. The key idea is to explore the shape alignment of each instance against its corresponding category-level template shape, and the symmetric correspondence of each object category for estimating a coarse 3D object shape. Our framework deforms the point cloud of the category-level template shape to align the observed instance point cloud for implicitly representing its 3D rotation. Then we model the symmetric correspondence by predicting symmetric point cloud from the partially observed point cloud. The concatenation of the observed point cloud and symmetric one reconstructs a coarse object shape, thus facilitating object center (3D translation) and 3D size estimation. Extensive experiments on the category-level NOCS benchmark demonstrate that our lightweight model still competes with state-of-the-art approaches that require labeled real-world images. We also deploy our approach to a physical Baxter robot to perform grasping tasks on unseen but category-known instances, and the results further validate the efficacy of our proposed model. Code and pre-trained models are available on the project webpage11Project webpage. https://hetolin.github.io/SAR-Net. Zichang Liu, Chilam Cheang, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001 |
CVPR | 3 |
| 2022 | Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language InstructionsabstractThis paper studies the task of any objects grasping from the known categories by free-form language instructions. This task demands the technique in computer vision, natural language processing, and robotics. We bring these disciplines together on this open challenge, which is essential to human-robot interaction. Critically, the key challenge lies in inferring the category of objects from linguistic instructions and accurately estimating the 6-DoF information of unseen objects from the known classes. In contrast, previous works focus on inferring the pose of object candidates at the instance level. This significantly limits its applications in real-world scenarios. In this paper, we propose a language-guided 6-DoF category-level object localization model to achieve robotic grasping by comprehending human intention. To this end, we propose a novel two-stage method. Particularly, the first stage grounds the target in the RGB image through language description of names, attributes, and spatial relations of objects. The second stage extracts and segments point clouds from the cropped depth image and estimates the full 6-DoF object pose at category-level. Under such a manner, our approach can locate the specific object by following human instructions, and estimate the full 6-DoF pose of a category-known but unseen instance which is not utilized for training the model. Extensive experimental results show that our method is competitive with the state-of-the-art language-conditioned grasp method. Importantly, we deploy our approach on a physical robot to validate the usability of our framework in real-world applications. Please refer to the supplementary for the demo videos of our robot experiments. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICRA | 1 |
| 2022 | I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand SketchesabstractIn this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language and the cases where a textual description is not available on the fly. However, very few works are aware of the usability of this novel interactive way between humans and robots. To this end, we propose a method to generate a potential grasp configuration relevant to the sketch -depicted objects. Due to the inherent ambiguity of sketches with abstract details, we take the advantage of the graph by incorporating the structure of the sketch to enhance the representation ability. This graph-represented sketch is further validated to improve the generalization of the network, capable of learning the sketch-queried grasp detection by using a small collection (around 100 samples) of hand-drawn sketches. Additionally, our model is trained and tested in an end-to-end manner which is easy to be implemented in real-world applications. Experiments on the multi-object VMRD and GraspNet-1Billion datasets demonstrate the good generalization of the proposed method. The physical robot experiments confirm the utility of our method in object-cluttered scenes. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICRA | 2 |
| 2022 | HandO: a hybrid 3D hand-object reconstruction model for unknown objects
Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
Multim. Syst. | 2 |