EDBT 2026 Demo / reviewers in the wild / expert
Te Cui
dblp:362/9470
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0002-3265-3148ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic ManipulationabstractLearning from demonstration is a powerful method for robotic skill acquisition. Nevertheless, a critical limitation lies in the substantial costs associated with gathering demonstration datasets, typically action-labeled robot data, which creates a fundamental constraint in the field. Video data offer a compelling solution as an alternative rich data source, containing diverse behavioral and physical knowledge. This study introduces G3M, an innovative framework that exploits video data viaGraph-to-GraphsGenerativeModeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. The proposed G3M abstracts video frame into graph representations by identifying object and visual action vertices for capturing state information. It then effectively models internal structures and spatial relationships present in these graph constructions, with the objective of predicting forthcoming graphs. The generated graphs function as conditional inputs that guide the control policy in determining robotic behaviors. This concise method effectively encodes critical spatial relationships while facilitating accurate prediction of subsequent graph sequences, thus allowing the development of resilient control policy despite constraints in action-annotated training samples. Furthermore, these transferable graph representations enable the effective extraction of manipulation knowledge through human videos as well as recordings from robots with different embodiments. The experimental results demonstrate that G3M attains superior performance using merely 20% action-labeled data relative to comparable approaches. Moreover, our method outperforms the state-of-the-art method, showing performance gains exceeding 19% in simulated environments and 23% in real-world experiments, while delivering improvements of over 35% in cross-embodiment transfer experiments and exhibiting strong performance on long-horizon tasks. Our project page is available athttps://g3m-project.github.io/. Guangyan Chen, Meiling Wang 0002, Te Cui, Chengcai Yang, Mengxiao Hu, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
IEEE Trans. Robotics | 3 |
| 2025 | GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningabstractLearning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alternative. In this paper, we present GraphMimic, a novel paradigm that leverages video data via graph-to-graphs generative modeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. Specifically, GraphMimic abstracts video frames into object and visual action vertices, and constructs graphs for state representations. The graph generative modeling network then effectively models internal structures and spatial relationships within the constructed graphs, aiming to generate future graphs. The generated graphs serve as conditions for the control policy, mapping to robot actions. Our concise approach captures important spatial relations and enhances future graph generation accuracy, enabling the acquisition of robust policies from limited action-labeled data. Furthermore, the transferable graph representations facilitate the effective learning of manipulation skills from cross-embodiment videos. Our experiments exhibit that GraphMimic achieves superior performance using merely 20% action-labeled data. Moreover, our method outperforms the state-of-the-art method by over 17% and 23% in simulation and real-world experiments, and delivers improvements of over 33% in cross-embodiment transfer experiments. Guangyan Chen, Te Cui, Meiling Wang 0002, Chengcai Yang, Mengxiao Hu, Yao Mu 0001, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue |
CVPR | 2 |
| 2025 | High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic GraspingabstractIn various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. However, previous methods for pose estimation typically over-look dynamic situations, such as slippage of grasped objects or movement of contacted objects during interactions with the environment, thus increasing the complexity of pose estimation. To address these challenges, we propose an efficient method that utilizes visual and tactile sensing to estimate object poses through particle filtering. We leverage visual information to track the pose of the contacted object in real-time and estimate the pose changes of the grasped object using displacement data obtained from tactile sensors. Our experimental evaluation on 13 objects with diverse geometric shapes demonstrated the ability to estimate high-precision poses, which revealed the robot's powerful ability to cope with dynamic scenes for compelled motion of objects, proving our framework's adaptability in practical scenarios with uncertainty. Zicai Peng, Te Cui, Guangyan Chen, Yi Yang 0009, Yufeng Yue |
ICRA | 2 |
| 2025 | ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region AlignmentabstractImage feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, characterized by rotation and scale changes. To alleviate this limitation, we introduce a novel oriented Overlapping Region Alignment method, named ORA-NET, which presents a concise and efficient approach to enhance the performance of image feature matching methods. We introduce the Multidirectional Cross-scale Feature Aggregation module to aggregate rotation-equivariant features across multiple scales and model long-range dependencies. Additionally, the Oriented Overlap Alignment module estimates scale and rotation differences within overlapping regions using a coarse-to-fine rotation correction approach. Importantly, our method serves as a plug-and-play module that can be seamlessly integrated into other correspondence matching pipelines. Experimental results demonstrate that ORA-NET significantly enhances the matching performance of existing local feature matching methods, particularly in scenarios involving substantial perspective differences. Te Cui, Meiling Wang 0002, Guangyan Chen, Yufeng Yue |
IROS | 1 |
| 2025 | Human Demonstrations are Generalizable Knowledge for RobotsabstractLearning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object instances. In this paper, we propose a different perspective, considering human demonstration videos not as mere instructions, but as a source of knowledge for robots. Motivated by this perspective and the remarkable comprehension and generalization capabilities exhibited by large language models (LLMs), we propose DigKnow, a method that DIstills Generalizable KNOWledge with a hierarchical structure. Specifically, DigKnow begins by converting human demonstration video frames into observation knowledge. This knowledge is then subjected to analysis to extract human action knowledge and further distilled into pattern knowledge that comprises task and object instances, resulting in the acquisition of generalizable knowledge with a hierarchical structure. In settings with different tasks or object instances, DigKnow retrieves relevant knowledge for the current task and object instances. Subsequently, the LLM-based planner conducts planning based on the retrieved knowledge, and the policy executes actions in line with the plan to achieve the designated task. Utilizing the retrieved knowledge, we validate and rectify planning and execution outcomes, resulting in a substantial enhancement of the success rate. Experimental results across a range of tasks and scenes demonstrate the effectiveness of this approach in facilitating real-world robots to accomplish tasks with the knowledge derived from human demonstrations. Te Cui, Tianxing Zhou, Mengxiao Hu, Zicai Peng, Haizhou Li 0004, Guangyan Chen, Meiling Wang 0002, Yufeng Yue |
IROS | 1 |
| 2025 | TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation ModelsabstractReliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when categories share similar visual characteristics. 2) While SAM excels in instance-level segmentation, integrating it with thermal images and text is hindered by modality heterogeneity and computational inefficiency. To address these, we propose TASeg, a text-aware RGB-T segmentation framework by using Low-Rank Adaptation (LoRA) fine-tuning technology to adapt vision foundation models. Specifically, we propose a Dynamic Feature Fusion Module (DFFM) in the image encoder, which effectively merges features from multiple visual modalities while freezing SAM’s original transformer blocks. Additionally, we incorporate CLIP-generated text embeddings in the mask decoder to enable semantic alignment, which further rectifies the classification error and improves the semantic understanding accuracy. Experimental results across diverse datasets demonstrate that our method achieves superior performance in challenging scenarios with fewer trainable parameters. Te Cui, Qitong Chu, Wenjie Song 0001, Yi Yang 0009, Yufeng Yue |
IROS | 2 |
| 2025 | Unifying Latent Action and Latent State Pre-training for Policy Learning from VideosabstractVideo data provides an accessible and rich source beyond expensive action-labeled robot data for advancing robotic learning paradigms. Motivated by this potential, researchers investigate methods to exploit video data in robotic learning. Recent approaches can be primarily divided into two categories: Action-based approaches tokenize latent actions from videos for policy pre-training. State-based approaches pre-train models to predict subsequent states. The former establishes rich motion priors, while the latter empowers the robot to anticipate future events. These complementary capabilities suggest significant potential for integration into a unified framework. In this paper, we propose UniMimic, a novel approach unifying latent action and latent state pre-training from videos. We first train a unified tokenizer to learn latent states from video frames while deriving latent actions between state tokens. Subsequently, the policy is pre-trained on videos to predict these latent actions and subsequent latent states. Finally, the policy is fine-tuned on an action-labeled robot dataset to transfer the learned priors to precise robot execution. Experiments exhibit that our pre-training stage enhances the performance by 19% in the Libero benchmark and improves the average number of tasks completed in a row of 5 from 2.50 and 2.35 to 3.89 and 3.73 in the CALVIN benchmark. In the real-world experiments, our method still delivers improvements exceeding 36%. Guangyan Chen, Meiling Wang 0002, Te Cui, Luojie Yang, Lin Zhao 0016, Yi Yang 0009, Yufeng Yue |
SIGGRAPH Asia | 3 |
| 2024 | Robust Collaborative Perception against Temporal Information DisturbanceabstractCollaborative perception facilitates a more comprehensive representation of the environment by leveraging complementary information shared among various agents and sensors. However, practical applications often encounter information disturbance which includes perception packet loss and time delays, and a comprehensive framework that can simultaneously address such issues is absent. In addition, the feature extraction process prior to fusion is not sufficient, as it lacks exploration of the local semantics and context dependencies of individual features. To enhance both accuracy and robustness, this paper introduces a novel framework named Robust Collaborative Perception against Temporal Information Disturbance, which predicts perception information when disturbance occurs. Specifically, the Historical Frame Prediction (HFP) module is introduced to make compensation for information loss with temporal association excavation of historical features. Based on the predicted features generated by the HFP module, the Pyramid Attention Integration (PAI) module is introduced to augment local semantics and incorporate global long-range dependencies through multi-scale window attention. Compared with existing methods on the publicly available dataset OPV2V, our approach exhibits superior performance and expanded robustness in the 3D object detection task. The code will be publicly available at https://github.com/hexunjie/Ro-temd. Xunjie He, Te Cui, Yufeng Yue |
ICRA | 3 |
| 2024 | VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsabstractVisual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage. Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
NeurIPS | 3 |
| 2024 | VIFNet: An end-to-end visible-infrared fusion network for image dehazing
Te Cui, Yufeng Yue |
Neurocomputing | 2 |