Lei Yang 0048

dblp:50/2484-48 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0002-3284-4019ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Evidential Multimodal Fusion Network for Trusted Pedestrian Crossing Intent Prediction
abstract
Accurate prediction of pedestrians’ behavior poses formidable challenges for autonomous vehicles in urban environments. Multimodal data, such as pedestrians’ motion data, context images, and ego vehicle speed, offer complementary and comprehensive information that can significantly enhance prediction performance. However, the previous methods, despite yielding promising results, integrate different modalities to form a uniform representation, which falls short of fully exploiting the heterogeneity and complementarity of all modalities. Besides, the uncertainty inherent in predictions is also a major concern for safety-critical systems such as autonomous vehicles. In light of the above concerns, this study puts forward a novel evidential multimodal fusion network called EMFNet, which leverages multimodal data and evidence fusion techniques for trusted pedestrian crossing intention prediction. Specifically, modal-specific embedding is first developed to project raw multimodal data into high-dimensional space by considering the unique properties of each modality. Then, we devise intra-modal and cross-modal feature learning to capture the temporal correlation within each modality and the interaction across different modalities, respectively. By doing so, modal-invariant features and modal-specific features can be effectively extracted. Subsequently, we introduce the Dempster–Shafer’s evidence theory (DST) to amalgamate evidence associated with different modalities, thus allowing the model to estimate the uncertainty and achieve trusted crossing prediction. Finally, we design an adaptive multiloss function that can effectively supervise the learning of modal-invariant and model-specific and facilitate the evidence fusion process. Experimental evaluation on real-world benchmark datasets demonstrates the improved prediction performance and enhanced reliability of the proposed method.
Xiaobo Chen 0001, Wei Xu 0052, Lei Yang 0048, Jian Yang 0003
IEEE Trans. Comput. Soc. Syst.4
2026 Internal State Estimation in Crowds via Active Information Gathering
abstract
Accurately estimating human internal states, such as personality traits or behavioral patterns, is critical for enhancing the effectiveness of human–robot interaction, particularly in multi-agent settings. These insights are key in applications ranging from social navigation to autism diagnosis. However, prior methods are limited by scalability and passive observation, making real-time estimation in complex, multi-human settings difficult. In this work, we propose a practical method for active human personality estimation in crowds, with a focus on applications related to Autism Spectrum Disorder (ASD). Our method combines a personality-conditioned behavior model, based on the Eysenck 3-Factor theory, with an active robot information-gathering policy that triggers human behaviors through a receding-horizon planner. The robot’s belief about human personality is then updated via Bayesian inference. We demonstrate the effectiveness of our approach through proof-of-concept studies in simulation, user studies with typical adults, and preliminary experiments involving participants with ASD. Our results show that our method can scale to tens of humans and reduce personality estimation error by 29.2% and uncertainty by 79.9% in simulation compared to the passive baseline. User studies with typical adults confirm the method’s ability to generalize across complex personality distributions. Additionally, we explore its application in autism-related scenarios, demonstrating that the method can identify the difference between neurotypical and autistic behavior. The results suggest that our framework could serve as a foundation for future ASD-specific applications.
Xuebo Ji, Zherong Pan, Xifeng Gao, Lei Yang 0048, Xinxin Du, Kaiyun Li, Yong-Jin Liu 0001, Wenping Wang 0001, Changhe Tu, Jia Pan 0001
ACM Trans. Hum. Robot Interact.4
2026 NeuBase: Spline Surfaces with Neural Basis Functions
abstract
We introduce NeuBase , a neural parametric surface representation that both accurately fits target surfaces with fine geometric detail and supports intuitive real time surface deformation. NeuBase consists of a Catmull-Clark subdivision base surface and an offset field defined by a set of neural basis functions encoded via a neural map. By construction, NeuBase surfaces exhibit four fundamental geometric properties, i.e., linearity, locality, smoothness, and affine equivariance, enabling real-time, direct manipulation without retraining the neural network. In addition, we propose a scalable neural map that maintains memory efficiency even for complex shapes with dense control meshes. Experiments on a large-scale dataset demonstrate that our method achieves better fitting accuracy than state-of-the-art neural parametric surface representations.
Anshul Mendiratta, Lei Yang 0048, Xin Li 0003, John Keyser, Scott Schaefer, Wenping Wang 0001
ACM Trans. Graph.2
2026 NeuPPS: Neural Piecewise Parametric Surfaces
abstract
Piecewise parametric surfaces have long been established as prevalent geometric representations; however, they often require surface refinement or sophisticated quadrangulation to accurately represent complex geometries. Geometric deep learning has shown that neural networks can provide greater representational power than conventional methods. Nevertheless, approaches using a single parametric surface for shape fitting struggle to capture fine-grained geometric details, while multi-patch methods fail to ensure seamless connections between adjacent patches. We present Neural Piecewise Parametric Surfaces ( NeuPPS ), the first piecewise neural surface representation that allows for coarse patch layouts composed of arbitrary n -sided surface patches to model complex surface geometries with high precision, offering enhanced flexibility compared with traditional parametric surfaces. This new surface representation guarantees, by construction, the continuity between adjacent patches, a property that other neural patch-based approaches cannot ensure. Two novel components are introduced: a learnable feature complex and a continuous mapping function approximated by multi-layer perceptrons (MLPs). We apply the proposed NeuPPS to surface fitting and shape space learning tasks. Extensive experiments demonstrate the advantages of NeuPPS over traditional parametric representations and existing patch-based learning approaches.
Lei Yang 0048, Yongqing Liang 0001, Xin Li 0003, Congyi Zhang 0001, Guying Lin, Cheng Lin 0001, Alla Sheffer, Scott Schaefer, John Keyser, Wenping Wang 0001
ACM Trans. Graph.1
2025 A Coarse-to-Fine Robotic Fabric Alignment System Integrating Visual Servoing and Admittance Control
abstract
Fabric alignment is essential to key production processes such as cutting, sewing, and fusing in garment manufacturing. Traditionally, this task has relied heavily on the dexterity and expertise of skilled human workers. Although automated systems have been introduced, they often lack the flexibility required for complex alignment tasks. In this paper, we present a novel robotic fabric alignment framework that fully automates the process with high precision and adaptability. First, we propose a coarse-to-fine alignment strategy, where an initial imprecise target position is roughly computed based on a basic perception module and eye-to-hand calibration. This is followed by a sliding mode control (SMC)-based visual servoing approach (in an eye-in-hand configuration) to ensure a close-up view of feedback features for the fine alignment process. Additionally, we consider system disturbances estimated by a fuzzy logic system (FLS) and combine it with the controller to further enhance the system’s robustness. Finally, we developed an advanced end-effector equipped with force/torque (F/T) sensors and air-powered needle grippers for gentle fabric manipulation using admittance control. We validate our framework through a series of experiments that demonstrate its effectiveness in fabric alignment tasks.
Jiaming Qi, Liang Lu 0005, Lei Yang 0048, Yan Ding 0002, Pai Zheng, David Navarro-Alarcon, Jia Pan 0001, Peng Zhou 0018
IEEE Trans Autom. Sci. Eng.3
2025 Patch-Grid: An Efficient and Feature-Preserving Neural Implicit Surface Representation
abstract
Neural implicit representations are increasingly used to depict three-dimensional (3D) shapes owing to their inherent smoothness and compactness, contrasting with traditional discrete representations. Yet, the multilayer perceptron–based neural representation, because of its smooth nature, rounds sharp corners or edges, rendering it unsuitable for representing objects with sharp features like computer-aided design (CAD) models. Moreover, neural implicit representations need long training times to fit 3D shapes. While previous works address these issues separately, we present a unified neural implicit representation called Patch-Grid , which efficiently fits complex shapes, preserves sharp features delineating different patches, and can also represent surfaces with open boundaries and thin geometric features. Patch-Grid learns a signed distance field (SDF) to approximate an encompassing surface patch of the shape with a learnable patch feature volume. To form sharp edges and corners in a CAD model, Patch-Grid merges the learned SDFs via the constructive solid geometry (CSG) approach. Core to the merging process is a novel merge grid design that organizes different patch feature volumes in a common octree structure. This design choice ensures robust merging of multiple learned SDFs by confining the CSG operations to localized regions. Additionally, it drastically reduces the complexity of the CSG operations in each merging cell, allowing the proposed method to be trained in seconds to fit a complex shape at high fidelity. Experimental results demonstrate that the proposed Patch-Grid representation is capable of accurately reconstructing shapes with complex sharp features, open boundaries, and thin geometric elements, achieving state-of-the-art reconstruction quality with high computational efficiency within seconds.
Guying Lin, Lei Yang 0048, Congyi Zhang 0001, Hao Pan 0001, Yuhan Ping, Guodong Wei, Taku Komura, John Keyser, Wenping Wang 0001
ACM Trans. Graph.2
2025 NeuVAS: Neural Implicit Surfaces for Variational Shape Modeling
abstract
Neural implicit shape representation has drawn significant attention in recent years due to its smoothness, differentiability, and topological flexibility. However, directly modeling the shape of a neural implicit surface, especially as the zero-level set of a neural signed distance function (SDF), with sparse geometric control is still a challenging task. Sparse input shape control typically includes 3D curve networks or, more generally, 3D curve sketches, which are unstructured and cannot be connected to form a curve network, and therefore more difficult to deal with. While 3D curve networks or curve sketches provide intuitive shape control, their sparsity and varied topology pose challenges in generating high-quality surfaces to meet such curve constraints. In this paper, we propose NeuVAS, a variational approach to shape modeling using neural implicit surfaces constrained under sparse input shape control, including unstructured 3D curve sketches as well as connected 3D curve networks. Specifically, we introduce a smoothness term based on a functional of surface curvatures to minimize shape variation of the zero-level set surface of a neural SDF. We also develop a new technique to faithfully model G 0 sharp feature curves as specified in the input curve sketches. Comprehensive comparisons with the state-of-the-art methods demonstrate the significant advantages of our method.
Qiujie Dong, Fangtian Liang, Hao Pan 0001, Lei Yang 0048, Congyi Zhang 0001, Guying Lin, Caiming Zhang 0001, Yuanfeng Zhou, Changhe Tu, Shi-Qing Xin, Alla Sheffer, Xin Li 0003, Wenping Wang 0001
ACM Trans. Graph.5
2025 On Optimal Sampling for Learning SDF Using MLPs Equipped With Positional Encoding
abstract
Neural implicit fields, such as the neural signed distance field (SDF) of a shape, have emerged as a powerful representation for many applications, e.g., encoding a 3D shape and performing collision detection. Typically, implicit fields are encoded by Multi-layer Perceptrons (MLP) with positional encoding (PE) to capture high-frequency geometric details. However, a notable side effect of such PE-equipped MLPs is the noisy artifacts present in the learned implicit fields. While increasing the sampling rate could in general mitigate these artifacts, in this paper we aim to explain this adverse phenomenon through the lens of Fourier analysis. We devise a tool to determine the appropriate sampling rate for learning an accurate neural implicit field without undesirable side effects. Specifically, we propose a simple yet effective method to estimate the intrinsic frequency of a given network with randomized weights based on the Fourier analysis of the network's responses. It is observed that a PE-equipped MLP has an intrinsic frequency much higher than the highest frequency component in the PE layer. Sampling against this intrinsic frequency following the Nyquist-Sannon sampling theorem allows us to determine an appropriate training sampling rate. We empirically show in the setting of SDF fitting that this recommended sampling rate is sufficient to secure accurate fitting results, while further increasing the sampling rate would not further noticeably reduce the fitting error. Training PE-equipped MLPs simply with our sampling strategy leads to performances superior to the existing methods.
Guying Lin, Lei Yang 0048, Yuan Liu 0025, Congyi Zhang 0001, Junhui Hou, Xiaogang Jin 0001, Taku Komura, John Keyser, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.2
2025 A Potential Field Method for Tooth Motion Planning in Orthodontic Treatment
abstract
Invisible orthodontics, commonly known as clear alignment treatment, offers a more comfortable and aesthetically pleasing alternative in orthodontic care, attracting considerable attention in the dental community in recent years. It replaces conventional metal braces with a series of removable, and transparent aligners. Each aligner is crafted to facilitate a gradual adjustment of the teeth, ensuring progressive stages of dental correction. This necessitates the design for teeth motion. Here we present an automatic method and a system for generating collision-free teeth motion planning while avoiding gaps between adjacent teeth, which is unacceptable in clinical practice. To tackle this task, we formulate it as a constrained optimization problem and utilize the interior point method for its solution. We also developed an interactive system that enables dentists to easily visualize and edit the paths. Our method significantly speeds up the clear aligner planning process, creating the desired motion paths for a full set of teeth in under five minutes-a task that typically requires several hours of manual work. Our experiments and user studies confirm the effectiveness of this method in planning teeth movement, showcasing its potential to streamline orthodontic procedures.
Yuexin Ma, Lei Yang 0048, Congyi Zhang 0001, Guangshun Wei, Runnan Chen, Min Gu 0003, Jia Pan 0001, Zhengbao Yang, Taku Komura, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Generated realistic noise and rotation-equivariant models for data-driven mesh denoising
Sipeng Yang, Wenhui Ren, Xiwen Zeng, Qingchuan Zhu, Hongbo Fu 0001, Kaijun Fan, Lei Yang 0048, Jingping Yu, Qilong Kou, Xiaogang Jin 0001
Comput. Aided Geom. Des.7
2024 Learning Autonomous Viewpoint Adjustment from Human Demonstrations for Telemanipulation
abstract
Teleoperation systems find many applications from earlier search-and-rescue to more recent daily tasks. It is widely acknowledged that using external sensors can decouple the view of the remote scene from the motion of the robot arm during manipulation, facilitating the control task. However, this design requires the coordination of multiple operators or may exhaust a single operator as s/he needs to control both the manipulator arm and the external sensors. To address this challenge, our work introduces a viewpoint prediction model, the first data-driven approach that autonomously adjusts the viewpoint of a dynamic camera to assist in telemanipulation tasks. This model is parameterized by a deep neural network and trained on a set of human demonstrations. We propose a contrastive learning scheme that leverages viewpoints in a camera trajectory as contrastive data for network training. We demonstrated the effectiveness of the proposed viewpoint prediction model by integrating it into a real-world robotic system for telemanipulation. User studies reveal that our model outperforms several camera control methods in terms of control experience and reduces the perceived task load compared to manual camera control. As an assistive module of a telemanipulation system, our method significantly reduces task completion time for users who choose to adopt its recommendation.
Ruixing Jia, Lei Yang 0048, Ying Cao 0001, Calvin K. L. Or, Wenping Wang 0001, Jia Pan 0001
ACM Trans. Hum. Robot Interact.2
2024 Neuromorphic Synergy for Video Binarization
abstract
Bimodal objects, such as the checkerboard pattern used in camera calibration, markers for object tracking, and text on road signs, to name a few, are prevalent in our daily lives and serve as a visual form to embed information that can be easily recognized by vision systems. While binarization from intensity images is crucial for extracting the embedded information in the bimodal objects, few previous works consider the task of binarization of blurry images due to the relative motion between the vision sensor and the environment. The blurry images can result in a loss in the binarization quality and thus degrade the downstream applications where the vision system is in motion. Recently, neuromorphic cameras offer new capabilities for alleviating motion blur, but it is non-trivial to first deblur and then binarize the images in a real-time manner. In this work, we propose an event-based binary reconstruction method that leverages the prior knowledge of the bimodal target's properties to perform inference independently in both event space and image space and merge the results from both domains to generate a sharp binary image. We also develop an efficient integration method to propagate this binary image to high frame rate binary video. Finally, we develop a novel method to naturally fuse events and images for unsupervised threshold identification. The proposed method is evaluated in publicly available and our collected data sequence, and shows the proposed method can outperform the SOTA methods to generate high frame rate binary video in real-time on CPU-only devices.
Shijie Lin, Xiang Zhang 0022, Lei Yang 0048, Lei Yu 0006, Wenping Wang 0001, Jia Pan 0001
IEEE Trans. Image Process.3
2024 DreamMat: High-quality PBR Material Generation with Geometry- and Light-aware Diffusion Models
abstract
Recent advancements in 2D diffusion models allow appearance generation on untextured raw meshes. These methods create RGB textures by distilling a 2D diffusion model, which often contains unwanted baked-in shading effects and results in unrealistic rendering effects in the downstream applications. Generating Physically Based Rendering (PBR) materials instead of just RGB textures would be a promising solution. However, directly distilling the PBR material parameters from 2D diffusion models still suffers from incorrect material decomposition, such as baked-in shading effects in albedo. We introduce DreamMat , an innovative approach to resolve the aforementioned problem, to generate high-quality PBR materials from text descriptions. We find out that the main reason for the incorrect material distillation is that large-scale 2D diffusion models are only trained to generate final shading colors, resulting in insufficient constraints on material decomposition during distillation. To tackle this problem, we first finetune a new light-aware 2D diffusion model to condition on a given lighting environment and generate the shading results on this specific lighting condition. Then, by applying the same environment lights in the material distillation, DreamMat can generate high-quality PBR materials that are not only consistent with the given geometry but also free from any baked-in shading effects in albedo. Extensive experiments demonstrate that the materials produced through our methods exhibit greater visual appeal to users and achieve significantly superior rendering quality compared to baseline methods, which are preferable for downstream tasks such as game and film production.
Yuqing Zhang 0005, Yuan Liu 0025, Zhiyu Xie 0004, Lei Yang 0048, Zhongyuan Liu, Mengzhou Yang, Qilong Kou, Cheng Lin 0001, Wenping Wang 0001, Xiaogang Jin 0001
ACM Trans. Graph.4
2023 Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos
abstract
Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices.
Yilin Wen 0001, Hao Pan 0001, Lei Yang 0048, Jia Pan 0001, Taku Komura, Wenping Wang 0001
CVPR3
2023 Evaluating Explanation Methods for Vision-and-Language Navigation
abstract
The ability to navigate robots with natural language instructions in an unknown environment is a crucial step for achieving embodied artificial intelligence (AI). With the improving performance of deep neural models proposed in the field of vision-and-language navigation (VLN), it is equally interesting to know what information the models utilize for their decision-making in the navigation tasks. To understand the inner workings of deep neural models, various explanation methods have been developed for promoting explainable AI (XAI). But they are mostly applied to deep neural models for image or text classification tasks and little work has been done in explaining deep neural models for VLN tasks. In this paper, we address these problems by building quantitative benchmarks to evaluate explanation methods for VLN models in terms of faithfulness. We propose a new erasure-based evaluation pipeline to measure the step-wise textual explanation in the sequential decision-making setting. We evaluate several explanation methods for two representative VLN models on two popular VLN datasets and reveal valuable findings through our experiments.
Guanqi Chen, Lei Yang 0048, Guanhua Chen 0001, Jia Pan 0001
ECAI2
2023 Surface Extraction from Neural Unsigned Distance Fields
abstract
We propose a method, named DualMesh-UDF, to extract a surface from unsigned distance functions (UDFs), encoded by neural networks, or neural UDFs. Neural UDFs are becoming increasingly popular for surface representation because of their versatility in presenting surfaces with arbitrary topologies, as opposed to the signed distance function that is limited to representing a closed surface. However, the applications of neural UDFs are hindered by the notorious difficulty in extracting the target surfaces they represent. Recent methods for surface extraction from a neural UDF suffer from significant geometric errors or topological artifacts due to two main difficulties: (1) A UDF does not exhibit sign changes; and (2) A neural UDF typically has substantial approximation errors.DualMesh-UDF addresses these two difficulties. Specifically, given a neural UDF encoding a target surface $\bar S$ to be recovered, we first estimate the tangent planes of $\bar S$ at a set of sample points close to $\bar S$. Next, we organize these sample points into local clusters, and for each local cluster, solve a linear least squares problem to determine a final surface point. These surface points are then connected to create the output mesh surface, which approximates the target surface. The robust estimation of the tangent planes of the target surface and the subsequent minimization problem constitute our core strategy, which contributes to the favorable performance of DualMesh-UDF over other competing methods. To efficiently implement this strategy, we employ an adaptive Octree. Within this framework, we estimate the location of a surface point in each of the octree cells identified as containing part of the target surface. Extensive experiments show that our method outperforms existing methods in terms of surface reconstruction quality while maintaining comparable computational efficiency.
Congyi Zhang 0001, Guying Lin, Lei Yang 0048, Xin Li 0003, Taku Komura, Scott Schaefer, John Keyser, Wenping Wang 0001
ICCV3
2023 Polymer-Based Self-Calibrated Optical Fiber Tactile Sensor
abstract
Human skin can accurately sense the self-decoupled normal and shear forces when in contact with objects of different sizes. Although there exist many soft and conformable tactile sensors on robotic applications able to decouple the normal force and shear forces, the impact of the size of object in contact on the force calibration model has been commonly ignored. Here, using the principle that contact force can be derived from the light power loss in the soft optical fiber core, we present a soft tactile sensor that decouples normal and shear forces and calibrates the measurement results based on the object size, by designing a two-layered weaved polymer-based optical fiber anisotropic structure embedded in a soft elastomer. Based on the anisotropic response of optical fibers, we developed a linear calibration algorithm to simultaneously measure the size of the contact object and the decoupled normal and shear forces calibrated the object size. By calibrating the sensor at the robotic arm tip, we show that robots can reconstruct the force vector at an average accuracy of 0.15N for normal forces, 0.17N for shear forces in X-axis, and 0.18N for shear forces in Y-axis, within the sensing range of 0-2N in all directions, and the average accuracy of object size measurement of 0.4mm, within the test indenter diameter range of 5-12mm.
Youcan Yan, Zeqing Zhang, Lei Yang 0048, Jia Pan 0001
IROS4
2023 CreatureShop: Interactive 3D Character Modeling and Texturing From a Single Color Drawing
abstract
Creating 3D shapes from 2D drawings is an important problem with applications in content creation for computer animation and virtual reality. We introduce a new sketch-based system, CreatureShop, that enables amateurs to create high-quality textured 3D character models from 2D drawings with ease and efficiency. CreatureShop takes an input bitmap drawing of a character (such as an animal or other creature), depicted from an arbitrary descriptive pose and viewpoint, and creates a 3D shape with plausible geometric details and textures from a small number of user annotations on the 2D drawing. Our key contributions are a novel oblique view modeling method, a set of systematic approaches for producing plausible textures on the invisible or occluded parts of the 3D character (as viewed from the direction of the input drawing), and a user-friendly interactive system. We validate our system and methods by creating numerous 3D characters from various drawings, and compare our results with related works to show the advantages of our method. We perform a user study to evaluate the usability of our system, which demonstrates that our system is a practical and efficient approach to create fully-textured 3D character models for novice users.
Congyi Zhang 0001, Lei Yang 0048, Nenglun Chen, Nicholas Vining, Alla Sheffer, Francis C. M. Lau 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.2
2022 DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation
Yilin Wen 0001, Hao Pan 0001, Lei Yang 0048, Zheng Wang 0002, Taku Komura, Wenping Wang 0001
ECCV (9)4
2022 Visual-tactile Sensing for Real-time Liquid Volume Estimation in Grasping
abstract
We propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrations. The robotic system is well controlled and adjusted based on the estimation model in real time. The main contributions and novelties of our work are listed as follows: 1) Explore a proprioceptive way for liquid volume estimation by developing an end-to-end predictive model with multi-modal convolutional networks, which achieve a high precision with an error of ~ 2 ml in the experimental validation. 2) Propose a multi-task learning architecture which comprehensively considers the losses from both classification and regression tasks, and comparatively evaluate the performance of each variant on the collected data and actual robotic platform. 3) Utilize the proprioceptive robotic system to accurately serve and control the requested volume of liquid, which is continuously flowing into a deformable container in real time. 4) Adaptively adjust the grasping plan to achieve more stable grasping and manipulation according to the real-time liquid volume prediction.
Ruixing Jia, Lei Yang 0048, Youcan Yan, Zheng Wang 0002, Jia Pan 0001, Wenping Wang 0001
IROS3
2022 Dense representative tooth landmark/axis detection network on 3D model
Guangshun Wei, Zhiming Cui 0001, Lei Yang 0048, Yuanfeng Zhou, Pradeep Singh 0003, Min Gu 0003, Wenping Wang 0001
Comput. Aided Geom. Des.4
2022 Coverage Axis: Inner Point Selection for 3D Shape Skeletonization
abstract
Abstract In this paper, we present a simple yet effective formulation called Coverage Axis for 3D shape skeletonization. Inspired by the set cover problem, our key idea is to cover all the surface points using as few inside medial balls as possible. This formulation inherently induces a compact and expressive approximation of the Medial Axis Transform (MAT) of a given shape. Different from previous methods that rely on local approximation error, our method allows a global consideration of the overall shape structure, leading to an efficient high‐level abstraction and superior robustness to noise. Another appealing aspect of our method is its capability to handle more generalized input such as point clouds and poor‐quality meshes. Extensive comparisons and evaluations demonstrate the remarkable effectiveness of our method for generating compact and expressive skeletal representation to approximate the MAT.
Zhiyang Dou, Cheng Lin 0001, Rui Xu 0016, Lei Yang 0048, Shi-Qing Xin, Taku Komura, Wenping Wang 0001
Comput. Graph. Forum4
2022 RPC: a large-scale and fine-grained retail product checkout dataset
Xiu-Shen Wei, Quan Cui, Lei Yang 0048, Peng Wang 0023, Lingqiao Liu, Jian Yang 0003
Sci. China Inf. Sci.3
2021 VertNet: Accurate Vertebra Localization and Identification Network from CT Images
Zhiming Cui 0001, Changjian Li 0001, Lei Yang 0048, Chunfeng Lian, Feng Shi 0001, Wenping Wang 0001, Dijia Wu, Dinggang Shen
MICCAI (5)3
2021 Distributed Attention for Grounded Image Captioning
abstract
We study the problem of weakly supervised grounded image captioning. That is, given an image, the goal is to automatically generate a sentence describing the context of the image with each noun word grounded to the corresponding region in the image. This task is challenging due to the lack of explicit fine-grained region word alignments as supervision. Previous weakly supervised methods mainly explore various kinds of regularization schemes to improve attention accuracy. However, their performances are still far from the fully supervised ones. One main issue that has been ignored is that the attention for generating visually groundable words may only focus on the most discriminate parts and can not cover the whole object. To this end, we propose a simple yet effective method to alleviate the issue, termed as partial grounding problem in our paper. Specifically, we design a distributed attention mechanism to enforce the network to aggregate information from multiple spatially different regions with consistent semantics while generating the words. Therefore, the union of the focused region proposals should form a visual region that encloses the object of interest completely. Extensive experiments have demonstrated the superiority of our proposed method compared with the state-of-the-arts.
Nenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang 0048, Zhiwen Lin, Yuqiang Ren, Haolei Yuan, Feiyue Huang, Wenping Wang 0001
ACM Multimedia4
2021 ScaffoldGAN: Synthesis of Scaffold Materials based on Generative Adversarial Networks
Hui Zhang 0027, Lei Yang 0048, Changjian Li 0001, Bojian Wu, Wenping Wang 0001
Comput. Aided Des.2
2021 Self-attention implicit function networks for 3D dental data completion
Yuhan Ping, Guodong Wei, Lei Yang 0048, Zhiming Cui 0001, Wenping Wang 0001
Comput. Aided Geom. Des.3
2021 Structure-Driven Unsupervised Domain Adaptation for Cross-Modality Cardiac Segmentation
abstract
Performance degradation due to domain shift remains a major challenge in medical image analysis. Unsupervised domain adaptation that transfers knowledge learned from the source domain with ground truth labels to the target domain without any annotation is the mainstream solution to resolve this issue. In this paper, we present a novel unsupervised domain adaptation framework for cross-modality cardiac segmentation, by explicitly capturing a common cardiac structure embedded across different modalities to guide cardiac segmentation. In particular, we first extract a set of 3D landmarks, in a self-supervised manner, to represent the cardiac structure of different modalities. The high-level structure information is then combined with another complementary feature, the Canny edges, to produce accurate cardiac segmentation results both in the source and target domains. We extensively evaluate our method on the MICCAI 2017 MM-WHS dataset for cardiac segmentation. The evaluation, comparison and comprehensive ablation studies demonstrate that our approach achieves satisfactory segmentation results and outperforms state-of-the-art unsupervised domain adaptation methods by a significant margin.
Zhiming Cui 0001, Changjian Li 0001, Zhixu Du, Nenglun Chen, Guodong Wei, Runnan Chen, Lei Yang 0048, Dinggang Shen, Wenping Wang 0001
IEEE Trans. Medical Imaging7
2020 Mapping in a Cycle: Sinkhorn Regularized Unsupervised Learning for Point Cloud Shapes
Lei Yang 0048, Wenxi Liu, Zhiming Cui 0001, Nenglun Chen, Wenping Wang 0001
ECCV (10)1
2019 Real-time editing of man-made mesh models under geometric constraints
Congyi Zhang 0001, Lei Yang 0048, Liyou Xu, Wenping Wang 0001
Comput. Graph.2