EDBT 2026 Demo / reviewers in the wild / expert
Yunzhi Lin
dblp:118/4321
· DBLP profile ↗
15ranked-venue papers
7as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Systems, architecture and hardware · 11 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGBabstractTo address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive comparison of pose tracking algorithms. We propose a pipeline featuring an uncertainty-aware keypoint refinement network, employing probabilistic modeling to refine pose estimation. Comparative evaluations demonstrate that our approach achieves performance superior to existing baselines on real datasets, underscoring the effectiveness of our synthetic dataset and refinement technique in enhancing tracking precision in dynamic contexts. Our contributions set a new precedent for the development and assessment of object pose tracking methodologies in complex scenes. Yunzhi Lin, Yipu Zhao, Fu-Jen Chu, Weiyao Wang 0001, Patricio A. Vela, Matt Feiszli, Kevin J. Liang |
IROS | 1 |
| 2023 | WDiscOOD: Out-of-Distribution Detection via Whitened Linear Discriminant AnalysisabstractDeep neural networks are susceptible to generating overconfident yet erroneous predictions when presented with data beyond known concepts. This challenge underscores the importance of detecting out-of-distribution (OOD) samples in the open world. In this work, we propose a novel feature-space OOD detection score based on class-specific and class-agnostic information. Specifically, the approach utilizes Whitened Linear Discriminant Analysis to project features into two subspaces the discriminative and residual subspaces - for which the in-distribution (ID) classes are maximally separated and closely clustered, respectively. The OOD score is then determined by combining the deviation from the input data to the ID pattern in both subspaces. The efficacy of our method, named WDiscOOD, is verified on the large-scale ImageNet-1k benchmark, with six OOD datasets that cover a variety of distribution shifts. WDiscOOD demonstrates superior performance on deep classifiers with diverse backbone architectures, including CNN and vision transformer. Furthermore, we also show that WDiscOOD more effectively detects novel concepts in representation spaces trained with contrastive objectives, including supervised contrastive loss and multi-modality contrastive loss. Yiye Chen, Yunzhi Lin, Ruinian Xu, Patricio A. Vela |
ICCV | 2 |
| 2023 | Keypoint-GraspNet: Keypoint-based 6-DoF Grasp Generation from the Monocular RGB-D inputabstractThe success of 6-DoF grasp learning with point cloud input is tempered by the computational costs resulting from their unordered nature and pre-processing needs for reducing the point cloud to a manageable size. These properties lead to failure on small objects with low point cloud cardinality. Instead of point clouds, this manuscript explores grasp generation directly from the RGB-D image input. The approach, called Keypoint-GraspNet (KGN), operates in perception space by detecting projected gripper keypoints in the image, then recovering their SE(3) poses with a$\mathrm{P}n\mathrm{P}$algorithm. Training of the network involves a synthetic dataset derived from primitive shape objects with known continuous grasp families. Trained with only single-object synthetic data, Keypoint-GraspNet achieves superior result on our single-object dataset, comparable performance with state-of-art baselines on a multi-object test set, and outperforms the most competitive baseline on small objects. Keypoint-GraspNet is more than 3x faster than tested point cloud methods. Robot experiments show high success rate, demonstrating KGN's practical potential. Yiye Chen, Yunzhi Lin, Ruinian Xu, Patricio A. Vela |
ICRA | 2 |
| 2023 | Parallel Inversion of Neural Radiance Fields for Robust Pose EstimationabstractWe present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual between pixels rendered from a fast NeRF model and pixels in the observed image. We integrate a momentum-based camera extrinsic optimization procedure into Instant Neural Graphics Primitives, a recent exceptionally fast NeRF implementation. By introducing parallel Monte Carlo sampling into the pose estimation task, our method overcomes local minima and improves efficiency in a more extensive search space. We also show the importance of adopting a more robust pixel-based loss function to reduce error. Experiments demonstrate that our method can achieve improved generalization and robustness on both synthetic and real-world benchmarks. Yunzhi Lin, Thomas Müller 0013, Jonathan Tremblay, Bowen Wen, Stephen Tyree, Alex Evans, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 1 |
| 2023 | KGNv2: Separating Scale and Pose Prediction for Keypoint-Based 6-DoF Grasp Synthesis on RGB-D InputabstractWe propose an improved keypoint approach for 6-DoF grasp pose synthesis from RGB-D input. Keypoint-based grasp detection from image input demonstrated promising results in a previous study, where the visual information provided by color imagery compensates for noisy or imprecise depth measurements. However, it relies heavily on accurate keypoint prediction in image space. We devise a new grasp generation network that reduces the dependency on precise keypoint estimation. Given an RGB-D input, the network estimates both the grasp pose and the camera-grasp length scale. Re-design of the keypoint output space mitigates the impact of keypoint prediction noise on Perspective-n-Point (PnP) algorithm solutions. Experiments show that the proposed method outperforms the baseline by a large margin, validating its design. Though trained only on simple synthetic objects, our method demonstrates sim-to-real capacity through competitive results in real-world robot experiments. Yiye Chen, Ruinian Xu, Yunzhi Lin, Patricio A. Vela |
IROS | 3 |
| 2022 | Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty EstimationabstractWe propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predict the bounding cuboid and 6- DoF pose (up to scale). Internally, a deep network predicts distributions over object keypoints (vertices of the bounding cuboid) in image coordinates, after which a novel probabilistic filtering process integrates across estimates before computing the final pose using PnP. Our framework allows the system to take previous uncertainties into consideration when predicting the current frame, resulting in predictions that are more accurate and stable than single frame methods. Extensive experiments show that our method outperforms existing approaches on the challenging Objectron benchmark of annotated object videos. We also demonstrate the usability of our work in an augmented reality setting. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 1 |
| 2022 | Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB ImageabstractPrior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstructured, real-world scenarios. In this work, we propose a single-stage, keypoint-based approach for category-level object pose estimation that operates on unknown object instances within a known category using a single RGB image as input. The proposed network performs 2D object detection, detects 2D keypoints, estimates 6- DoF pose, and regresses relative bounding cuboid dimensions. These quantities are estimated in a sequential fashion, leveraging the recent idea of convGRU for propagating information from easier tasks to those that are more difficult. We favor simplicity in our design choices: generic cuboid vertex coordinates, single-stage network, and monocular RGB input. We conduct extensive experiments on the challenging Objectron benchmark, outperforming state-of-the-art methods on the 3D IoU metric (27.6% higher than the MobilePose single-stage approach and 7.1 % higher than the related two-stage approach). Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 1 |
| 2021 | Poster: Accelerate Cross-Device Federated Learning With Semi-Reliable Model Multicast Over The AirabstractTo achieve efficient model multicast for cross-device Federated Learning (FL) over shared wireless channels, we propose SRMP, a transport protocol that performs semi-reliable model multicast over the air by leveraging existing PHY-aided wireless multicast techniques. The preliminary study shows that, with novel designs, SRMP could reduce the communication time involved in each round of training significantly. Yunzhi Lin, Shouxi Luo |
ICNP | 1 |
| 2021 | A Joint Network for Grasp Detection Conditioned on Natural Language CommandsabstractWe consider the task of grasping a target object based on a natural language command query. Previous work primarily focused on localizing the object given the query, which requires a separate grasp detection module to grasp it. The cascaded application of two pipelines incurs errors in overlapping multi-object cases due to ambiguity in the individal outputs. This work proposes a model named Command Grasping Network (CGNet) to directly output command satisficing grasps from RGB image and textual command inputs. A dataset with ground truth (image, command, grasps) tuple is generated based on the VMRD dataset to train the proposed network. Experimental results on the generated test set show that CGNet outperforms a cascaded object-retrieval and grasp detection baseline by a large margin. Three physical experiments demonstrate the functionality and performance of CGNet. Yiye Chen, Ruinian Xu, Yunzhi Lin, Patricio A. Vela |
ICRA | 3 |
| 2021 | Multi-view Fusion for Multi-level Robotic Scene UnderstandingabstractWe present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of unknown objects from categories corresponding to primitive shapes (e.g., cuboids and cylinders), and 3) full 6-DoF pose of known objects. By developing and fusing recent techniques in these domains, we provide a rich scene representation for robot awareness. We demonstrate the importance of each of these modules, their complementary nature, and the potential benefits of the system in the context of robotic manipulation. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
IROS | 1 |
| 2020 | Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object GraspingabstractA segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is trained on synthetic data with 6 classes of primitive shapes generated by a simulation engine. Each primitive shape is designed with parametrized grasp families, permitting the pipeline to identify multiple grasp candidates per shape primitive region. The grasps are priority ordered via proposed ranking algorithm, with the first feasible one chosen for execution. On task-free grasping of individual objects, the method achieves a 94% success rate. On task-oriented grasping, it achieves a 76% success rate. Overall, the method supports the hypothesis that shape primitives can support task-free and task-relevant grasp prediction. Yunzhi Lin, Chao Tang 0001, Fu-Jen Chu, Patricio A. Vela |
ICRA | 1 |
| 2020 | Joint Optimization of Spectrum and Energy Efficiency Considering the C-V2X Security: A Deep Reinforcement Learning ApproachabstractCellular vehicle-to-everything (C-V2X) communication, as a part of 5G wireless communications, has been considered one of the most significant techniques for Smart City. Vehicles platooning is an application of Smart City that improves traffic capacity and safety by C-V2X. However, different from vehicles platooning travelling on highways, C-V2X could be more easily eavesdropped and the spectrum resource could be limited when vehicles converge at an intersection. Satisfying the secrecy rate of C-V2X, how to increase the spectrum efficiency (SE) and energy efficiency (EE) in the platooning network is a big challenge. In this paper, to solve this problem, a Security-Aware Approach to Enhancing SE and EE Based on Deep Reinforcement Learning is proposed, named SEED. The SEED formulates an objective optimization function considering both SE and EE, and the secrecy rate of C-V2X is treated as a critical constraint of this function. The optimization problem is transformed into the spectrum and transmission power selections of V2X links using deep Q network (DQN). The heuristic result of SE and EE is obtained by the DQN based on rewards mechanism. Finally, the traffic and communication environments are simulated by Python 3. The evaluation results demonstrate that the SEED outperforms the DQN-wopa algorithm and the baseline algorithm by 31.83% and 68.40% in efficiency, respectively. Yinhui Han, Jianwei Fan, Yunzhi Lin |
INDIN | 5 |
| 2020 | Redundant Forward Gateway in Heterogeneous WLAN and LTE-M Networks for Reliable Metro CommunicationabstractWith the advance of urban railway transportation and wireless communication technology, an increasing concern is placed on the reliability and low latency of urban mass transit communication system. Nowadays, both of WLAN and LTEM systems, known as two mainstream subway communication systems, are deployed redundantly with single communication network in case of unexpected failure. However, using a single communication system has irresistibility to frequency interference and traditional redundant cold backup manual has an intolerable delay in switching. In this paper, we adopt a multipath transmission network architecture with heterogeneous communication systems to solve the problems mentioned above. To realize the hot backup networks, the Redundant Forward Gateway (RFG) device is designed to connect terminals to the two networks. Information forwarded by source RFGs can be transmitted redundantly and simultaneously via both WLAN and LTE-M communication systems. At the other side, the redundant data packets will be removed before being forwarded to the destination terminals. Besides, tests were performed in various communication scenarios to evaluate the availability and effectiveness of the RFG and the network architecture. The results indicate that the RFG has a negligible processing delay and improves the robustness and performance compared to the traditional LTE-M network according to the transmission delay and packet loss tests. In conclusion, it is considered that the RFG and the heterogeneous LTE-M and WLAN networks are of vital importance in enhancing reliability in the metro communication system. Jianwei Fan, Yunzhi Lin |
INDIN | 5 |
| 2020 | Gradient-based discriminative modeling for blind image deblurring
Wenze Shao, Yunzhi Lin, Li-Qian Wang, Qi Ge, Bing-Kun Bao, Haibo Li 0001 |
Neurocomputing | 2 |
| 2018 | Blind Deblurring Using Discriminative Image Smoothing
Wenze Shao, Yunzhi Lin, Bing-Kun Bao, Liqian Wang, Qi Ge, Haibo Li 0001 |
PRCV (1) | 2 |