Yi-Lin Wei

dblp:376/2518 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0004-9210-7370ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WaveC2R: Wavelet-Driven Coarse-to-Refined Hierarchical Learning for Radar Retrieval
abstract
Satellite-based radar retrieval methods are widely employed to fill coverage gaps in ground-based radar systems, especially in remote areas affected by terrain blockage and limited detection range. Existing methods predominantly rely on overly simplistic spatial-domain architectures constructed from a single data source, limiting their ability to accurately capture complex precipitation patterns and sharply defined meteorological boundaries. To address these limitations, we propose WaveC2R, a novel wavelet-driven coarse-to-refined framework for radar retrieval. WaveC2R integrates complementary multi-source data and leverages frequency-domain decomposition to separately model low-frequency components for capturing precipitation patterns and high-frequency components for delineating sharply defined meteorological boundaries. Specifically, WaveC2R consists of two stages (i) Intensity-Boundary Decoupled Learning, which leverages wavelet decomposition and frequency-specific loss functions to separately optimize low-frequency intensity and high-frequency boundaries; and (ii) Detail-Enhanced Diffusion Refinement, which employs frequency-aware conditional priors and multi-source data to progressively enhance fine-scale precipitation structures while preserving coarse-scale meteorological consistency. Experimental results on the publicly available SEVIR dataset demonstrate that WaveC2R achieves state-of-the-art performance in satellite-based radar retrieval, particularly excelling at preserving high-intensity precipitation features and sharply defined meteorological boundaries.
Yi-Lin Wei, Yongchao Feng, Yecheng Zhang, Dan Niu
AAAI4
2026 CastDiffuser: Cascaded latent diffusion framework for high-resolution precipitation nowcasting via multi source fusion
abstract
Precipitation nowcasting is critical for meteorological disaster warning, water resource management, and severe rainfall prediction, directly impacting public safety and daily operations. Existing methods based on discriminative modeling tend to produce ambiguous extrapolation maps. While generative models have improved perceptual metrics, they still suffer from limited prediction accuracy (e.g., skills scores such as Critical Success Index (CSI) and Fractions Skill Score (FSS)), high computational costs, and inadequate convective initiation prediction-with the latter stemming from constraints of single-source data. To address these challenges, we propose CastDiffuser, a cascaded latent diffusion framework for high-resolution and high-precision precipitation nowcasting. The framework performs cascaded modeling in the latent space and leverages complementary satellite and radar data, enabling both enhanced predictive accuracy and reduced computational cost. Specifically, CastDiffuser downscales and reconstructs high-resolution radar by variational autocoders. Then we utilize a spatio-temporal translator (ST-Translator) to model the deterministic components of precipitation evolution. Subsequently, a satellite-guided diffusion model is introduced to refine these deterministic features, which are then used to generate high-resolution radar predictions. Experiments on Jiangsu provincial meteorological datasets show CastDiffuser outperforms state-of-the-art methods in prediction accuracy and fine-detail preservation, particularly for heavy rainfall and convective initiation events. • We propose CastDiffuser, a Cascaded Latent Diffusion Framework thatintegrates radar and satellite data in latent space for precipitation forecasting.By modeling in a low-dimensional latent space with a cascaded design, CastDiffuser achieves superior spatio-temporal prediction accuracy while significantly reducing the computational cost of high-resolution forecasting. • We design the Spatio-Temporal Translator composed of hierarchical ST-Inception modules for robust multi-scale feature extraction. This module provides stable and structured representations for subsequent diffusion learning, thereby addressing the issue of unstable input features and improving forecast accuracy. • We introduce FsrFormer, a multi-source fusion denoising network acting as a spatial refinement module. FsrFormer adaptively modulates the influence of satellite-derived features during the diffusion process across different time steps, facilitating efficient and dynamic fusion of multi-source information. • Experimental results on real-world datasets show that the proposed CastDiffuser significantly improves the prediction performance of heavy precipitation and maintains this accuracy over a longer forecast period, especially in the convective incipient prediction task.
Dan Niu, Daben Niu, Yi-Lin Wei, Zengliang Zang, Jun Yang 0011
Neurocomputing4
2025 ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
abstract
We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly modeling joint-level interactions is more natural and effective for generating realistic HOIs, as it directly captures the geometric and semantic relationships between joints, rather than modeling interactions in the latent pose space. To this end, ChainHOI introduces a novel joint graph to capture potential interactions with objects, and a Generative Spatiotemporal Graph Convolution Network to explicitly model interactions at the joint level. Furthermore, we propose a Kinematics-based Interaction Module that explicitly models interactions at the kinetic chain level, ensuring more realistic and biomechanically coherent motions. Evaluations on two public datasets demonstrate that ChainHOI significantly outperforms previous methods, generating more realistic, and semantically consistent HOIs. Code is available here.
Ling-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu, Yu-Ming Tang, Jingke Meng, Wei-Shi Zheng 0001
CVPR3
2025 Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction Framework
abstract
Bimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their actions. However, we think bimanual manipulation involves not only coordinated tasks but also various uncoordinated tasks that do not require explicit cooperation during execution, such as grasping objects with the closest hand, which integrated control frameworks ignore to consider due to their enforced cooperation in the early inputs. In this paper, we propose a novel decoupled interaction framework that considers the characteristics of different tasks in bimanual manipulation. The key insight of our framework is to assign an independent model to each arm to enhance the learning of uncoordinated tasks, while introducing a selective interaction module that adaptively learns weights from its own arm to improve the learning of coordinated tasks. Extensive experiments on seven tasks in the RoboTwin dataset demonstrate that: (1) Our framework achieves outstanding performance, with a 23.5% boost over the SOTA method. (2) Our framework is flexible and can be seamlessly integrated into existing methods. (3) Our framework can be effectively extended to multi-agent manipulation tasks, achieving a 28% boost over the integrated control SOTA. (4) The performance boost stems from the decoupled design itself, surpassing the SOTA by 16.5% in success rate with only 1/6 of the model size.
Jian-Jian Jiang, Xiao-Ming Wu 0002, Yi-Xiang He, Ling-An Zeng, Yi-Lin Wei, Wei-Shi Zheng 0001
ICCV5
2025 AffordDexGrasp: Open-Set Language-Guided Dexterous Grasp With Generalizable-Instructive Affordance
Yi-Lin Wei, Mu Lin, Jian-Jian Jiang, Xiao-Ming Wu 0002, Ling-An Zeng, Wei-Shi Zheng 0001
ICCV1
2025 iManip: Skill-Incremental Learning for Robotic Manipulation
Zexin Zheng, Jia-Feng Cai, Xiao-Ming Wu 0002, Yi-Lin Wei, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001
ICCV4
2025 TacCap: A Wearable FBG-Based Tactile Sensor for Efficient Human-to-Robot Skill Transfer
abstract
Tactile sensing is essential for dexterous manipulation, yet large-scale human demonstration datasets lack tactile feedback, limiting their effectiveness in skill transfer to robots. To address this, we introduce TacCap, a wearable Fiber Bragg Grating (FBG)-based tactile sensor designed for seamless human-to-robot transfer. TacCap is lightweight, durable, and immune to electromagnetic interference, making it ideal for real-world data collection. We detail its design and fabrication, evaluate its sensitivity, repeatability, and cross-sensor consistency, and assess its effectiveness through grasp stability prediction and ablation studies. Our results demonstrate that TacCap enables transferable tactile data collection, bridging the gap between human demonstrations and robotic execution, with broad implications for fine-motor disciplines such as surgical training and musical performance. To support further research and development, we open-source our hardware design and software.
Chengyi Xing, Hao Li 0076, Yi-Lin Wei, Tian-Ao Ren, Tianyu Tu, Elizabeth Schumann, Wei-Shi Zheng 0001, Mark R. Cutkosky
IROS3
2024 Single-View Scene Point Cloud Human Grasp Generation
abstract
In this work, we explore a novel task of generating human grasps based on single-view scene point clouds, which more accurately mirrors the typical real-world situation of observing objects from a single viewpoint. Due to the incompleteness of object point clouds and the presence of numerous scene points, the generated hand is prone to penetrating into the invisible parts of the object and the model is easily affected by scene points. Thus, we introduce S2HGrasp, a framework composed of two key modules: the Global Perception module that globally perceives partial object point clouds, and the DiffuGrasp module designed to generate high-quality human grasps based on complex inputs that include scene points. Additionally, we introduce S2HGD dataset, which comprises approximately 99,000 single-object single-view scene point clouds of 1,668 unique objects, each annotated with one human grasp. Our extensive experiments demonstrate that S2HGrasp can not only generate natural human grasps regardless of scene points, but also effectively prevent penetration between the hand and invisible parts of the object. Moreover, our model showcases strong generalization capability when applied to unseen objects. Our code and dataset are available at https://github.com/iSEE-Laboratory/S2HGrasp.
Yan-Kang Wang, Chengyi Xing, Yi-Lin Wei, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001
CVPR3
2024 Dexterous Grasp Transformer
abstract
In this work, we propose a novel discriminative frame-work for dexterous grasp generation, named Dexterous Grasp TRansformer (DGTR), capable of predicting a di-verse set of feasible grasp poses by processing the object point cloud with only one forward pass. We formulate dex-terous grasp generation as a set prediction task and design a transformer-based grasping model for it. However, we identify that this set prediction paradigm encounters sev-eral optimization challenges in the field of dexterous grasping and results in restricted performance. To address these issues, we propose progressive strategies for both the training and testing phases. First, the dynamic-static matching training (DSMT) strategy is presented to enhance the opti-mization stability during the training phase. Second, we in-troduce the adversarial-balanced test-time adaptation (AB-TTA) with a pair of adversarial losses to improve grasping quality during the testing phase. Experimental results on the DexGraspNet dataset demonstrate the capability of DGTR to predict dexterous grasp poses with both high quality and diversity. Notably, while keeping high qual-ity, the diversity of grasp poses predicted by DGTR sig-nificantly outperforms previous works in multiple metrics without any data pre-processing. Codes are available at https://github.com/iSEE-Laboratory/DGTR.
Guo-Hao Xu, Yi-Lin Wei, Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001
CVPR2
2024 An Economic Framework for 6-DoF Grasp Detection
Xiao-Ming Wu 0002, Jia-Feng Cai, Jian-Jian Jiang, Dian Zheng, Yi-Lin Wei, Wei-Shi Zheng 0001
ECCV (27)5
2024 iGrasp: An Interactive 2D-3D Framework for 6-DoF Grasp Detection
Jian-Jian Jiang, Xiao-Ming Wu 0002, Zibo Chen, Yi-Lin Wei, Wei-Shi Zheng 0001
ICPR (30)4
2024 Grasp as You Say: Language-guided Dexterous Grasp Generation
abstract
This paper explores a novel task "Dexterous Grasp as You Say'' (DexGYS), enabling robots to perform dexterous grasping based on human commands expressed in natural language. However, the development of this field is hindered by the lack of datasets with natural human guidance; thus, we propose a language-guided dexterous grasp dataset, named DexGYSNet, offering high-quality dexterous grasp annotations along with flexible and fine-grained human language guidance. Our dataset construction is cost-efficient, with the carefully-design hand-object interaction retargeting strategy, and the LLM-assisted language guidance annotation system. Equipped with this dataset, we introduce the DexGYSGrasp framework for generating dexterous grasps based on human language instructions, with the capability of producing grasps that are intent-aligned, high quality and diversity. To achieve this capability, our framework decomposes the complex learning process into two manageable progressive objectives and introduce two components to realize them. The first component learns the grasp distribution focusing on intention alignment and generation diversity. And the second component refines the grasp quality while maintaining intention consistency. Extensive experiments are conducted on DexGYSNet and real world environments for validation.
Yi-Lin Wei, Jian-Jian Jiang, Chengyi Xing, Xiantuo Tan, Xiao-Ming Wu 0002, Hao Li 0076, Mark R. Cutkosky, Wei-Shi Zheng 0001
NeurIPS1