EDBT 2026 Demo / reviewers in the wild / expert
Yangming Li
dblp:62/8367
· DBLP profile ↗
46ranked-venue papers
23as first author
14since 2021 · last 2025
0000-0002-9794-8054ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 21 first-author · 13 since 2021Systems, architecture and hardware · 14 · 8 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy SamplesabstractDiffusion models are mainly studied on image data. However, non-image data (e.g., tabular data) are also prevalent in real applications and tend to be noisy due to some inevitable factors in the stage of data collection, degrading the generation quality of diffusion models. In this paper, we consider a novel problem setting where every collected sample is paired with a vector indicating the data quality: risk vector. This setting applies to many scenarios involving noisy data and we propose risk-sensitive SDE, a type of stochastic differential equation (SDE) parameterized by the risk vector, to address it. With some proper coefficients, risk-sensitive SDE can minimize the negative effect of noisy samples on the optimization of diffusion models. We conduct systematic studies for both Gaussian and non-Gaussian noise distributions, providing analytical forms of risk-sensitive SDE. To verify the effectiveness of our method, we have conducted extensive experiments on multiple tabular and time-series datasets, showing that risk-sensitive SDE permits a robust optimization of diffusion models with noisy samples and significantly outperforms previous baselines. Yangming Li, Max Ruiz Luyten, Mihaela van der Schaar |
ICLR | 1 |
| 2025 | Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite AttacksabstractText watermarking aims to subtly embeds statistical signals into text by controlling the Large Language Model (LLM)’s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The robustness of these watermarking algorithms has become a key factor in evaluating their effectiveness. Current text watermarking algorithms embed watermarks in high-entropy tokens to ensure text quality. In this paper, we reveal that this seemingly benign design can be exploited by attackers, posing a significant risk to the robustness of the watermark. We introduce a generic efficient paraphrasing attack, the Self-Information Rewrite Attack (SIRA), which leverages the vulnerability by calculating the self-information of each token to identify potential pattern tokens and perform targeted attack. Our work exposes a widely prevalent vulnerability in current watermarking algorithms. The experimental results show SIRA achieves nearly 100% attack success rates on seven recent watermarking methods with only $0.88 per million tokens cost. Our approach does not require any access to the watermark algorithms or the watermarked LLM and can seamlessly transfer to any LLM as the attack model even mobile-level models. Our findings highlight the urgent need for more robust watermarking. Yixin Cheng, Hongcheng Guo, Yangming Li, Leonid Sigal |
ICML | 3 |
| 2024 | Soft Mixture Denoising: Beyond the Expressive Bottleneck of Diffusion ModelsabstractBecause diffusion models have shown impressive performances in a number of tasks, such as image synthesis, there is a trend in recent works to prove (with certain assumptions) that these models have strong approximation capabilities. In this paper, we show that current diffusion models actually have an expressive bottleneck in backward denoising and some assumption made by existing theoretical guarantees is too strong. Based on this finding, we prove that diffusion models have unbounded errors in both local and global denoising. In light of our theoretical studies, we introduce soft mixture denoising (SMD), an expressive and efficient model for backward denoising. SMD not only permits diffusion models to well approximate any Gaussian mixture distributions in theory, but also is simple and efficient for implementation. Our experiments on multiple image datasets show that SMD significantly improves different types of diffusion models (e.g., DDPM), espeically in the situation of few backward iterations. Yangming Li, Boris van Breugel, Mihaela van der Schaar |
ICLR | 1 |
| 2024 | On Error Propagation of Diffusion ModelsabstractAlthough diffusion models (DMs) have shown promising performances in a number of tasks (e.g., speech synthesis and image generation), they might suffer from error propagation because of their sequential structure. However, this is not certain because some sequential models, such as Conditional Random Field (CRF), are free from this problem. To address this issue, we develop a theoretical framework to mathematically formulate error propagation in the architecture of DMs, The framework contains three elements, including modular error, cumulative error, and propagation equation. The modular and cumulative errors are related by the equation, which interprets that DMs are indeed affected by error propagation. Our theoretical study also suggests that the cumulative error is closely related to the generation quality of DMs. Based on this finding, we apply the cumulative error as a regularization term to reduce error propagation. Because the term is computationally intractable, we derive its upper bound and design a bootstrap algorithm to efficiently estimate the bound for optimization. We have conducted extensive experiments on multiple image datasets, showing that our proposed regularization reduces error propagation, significantly improves vanilla DMs, and outperforms previous baselines. Yangming Li, Mihaela van der Schaar |
ICLR | 1 |
| 2023 | A Simple Yet Effective Approach to Structured Knowledge DistillationabstractStructured prediction models aim at solving tasks where the output is a complex structure, rather than a single variable. Performing knowledge distillation for such problems is non- trivial due to their exponentially large output space. Previous works address this problem by developing particular distillation strategies (e.g., dynamic programming) that are both complicated and of low run-time efficiency. In this work, we propose an approach that is much simpler in its formulation, far more efficient for training than existing methods, and even performs better than our baselines. Specifically, we transfer the knowledge from a teacher model to its student by locally matching their computations on all internal structures rather than the final outputs. In this manner, we avoid time-consuming techniques like Monte Carlo Sampling for decoding output structures, permitting parallel computation and efficient training. Besides, we show that it encourages the student model to better mimic the internal behavior of the teacher model. Experiments on two structured prediction tasks demonstrate that our approach not only halves the time cost, but also outperforms previous methods on two widely adopted benchmark datasets.1 2 Wenye Lin, Yangming Li, Lemao Liu, Shuming Shi 0001, Hai-Tao Zheng 0002 |
ICASSP | 2 |
| 2022 | Rethinking Negative Sampling for Handling Missing Entity AnnotationsabstractNegative sampling is highly effective in handling missing annotations for named entity recognition (NER).One of our contributions is an analysis on how it makes sense through introducing two insightful concepts: missampling and uncertainty.Empirical studies show low missampling rate and high uncertainty are both essential for achieving promising performances with negative sampling.Based on the sparsity of named entities, we also theoretically derive a lower bound for the probability of zero missampling rate, which is only relevant to sentence length.The other contribution is an adaptive and weighted sampling distribution that further improves negative sampling via our former analysis.Experiments on synthetic datasets and well-annotated datasets (e.g., CoNLL-2003) show that our proposed approach benefits negative sampling in terms of F1 score and loss convergence.Besides, models with improved negative sampling have achieved new state-of-the-art results on realworld datasets (e.g., EC). Yangming Li, Lemao Liu, Shuming Shi 0001 |
ACL (1) | 1 |
| 2022 | Multi-domain Spoken Language Understanding Using Domain- and Task-aware ParameterizationabstractSpoken language understanding (SLU) has been addressed as a supervised learning problem, where a set of training data is available for each domain. However, annotating data for a new domain can be both financially costly and non-scalable. One existing approach solves the problem by conducting multi-domain learning where parameters are shared for joint training across domains, which is domain-agnostic and task-agnostic . In the article, we propose to improve the parameterization of this method by using domain-specific and task-specific model parameters for fine-grained knowledge representation and transfer. Experiments on five domains show that our model is more effective for multi-domain SLU and obtain the best results. In addition, we show its transferability when adapting to a new domain with little data, outperforming the prior best model by 12.4%. Finally, we explore the strong pre-trained model in our framework and find that the contributions from our framework do not fully overlap with contextualized word representations (RoBERTa). Libo Qin 0001, Fuxuan Wei, Minheng Ni, Yue Zhang 0004, Wanxiang Che, Yangming Li, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2021 | Interpretable NLG for Task-oriented Dialogue Systems with Heterogeneous Rendering MachinesabstractEnd-to-end neural networks have achieved promising performances in natural language generation (NLG). However, they are treated as black boxes and lack interpretability. To address this problem, we propose a novel framework, heterogeneous rendering machines (HRM), that interprets how neural generators render an input dialogue act (DA) into an utterance. HRM consists of a renderer set and a mode switcher. The renderer set contains multiple decoders that vary in both structure and functionality. For every generation step, the mode switcher selects an appropriate decoder from the renderer set to generate an item (a word or a phrase). To verify the effectiveness of our method, we have conducted extensive experiments on 5 benchmark datasets. In terms of automatic metrics (e.g., BLEU), our model is competitive with the current state-of-the-art method. The qualitative analysis shows that our model can interpret the rendering process of neural generators well. Human evaluation also confirms the interpretability of our proposed approach. Yangming Li, Kaisheng Yao |
AAAI | 1 |
| 2021 | Rewriter-Evaluator Architecture for Neural Machine TranslationabstractYangming Li, Kaisheng Yao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yangming Li, Kaisheng Yao |
ACL/IJCNLP (1) | 1 |
| 2021 | Fine-grained Entity Typing without Knowledge BaseabstractExisting work on Fine-grained Entity Typing (FET) typically trains automatic models on the datasets obtained by using Knowledge Bases (KB) as distant supervision.However, the reliance on KB means this training setting can be hampered by the lack of or the incompleteness of the KB.To alleviate this limitation, we propose a novel setting for training FET models: FET without accessing any knowledge base.Under this setting, we propose a two-step framework to train FET models.In the first step, we automatically create pseudo data with fine-grained labels from a large unlabeled dataset.Then a neural network model is trained based on the pseudo data, either in an unsupervised way or using self-training under the weak guidance from a coarse-grained Named Entity Recognition (NER) model.Experimental results show that our method achieves competitive performance with respect to the models trained on the original KB-supervised datasets.* The first two authors (Jing and Yibin) contributed equally to this work during the internships at Tencent AI Lab. Lemao Liu, Yangming Li, Haiyun Jiang, Haisong Zhang, Shuming Shi 0001 |
EMNLP (1) | 4 |
| 2021 | Empirical Analysis of Unlabeled Entity Problem in Named Entity Recognition
Yangming Li, Lemao Liu, Shuming Shi 0001 |
ICLR | 1 |
| 2021 | Learning Surgical Motion Pattern from Small Data in Endoscopic Sinus and Skull Base SurgeriesabstractExisting studies demonstrated that surgical motion patterns are strongly correlated with surgical outcomes. Real surgeries are complicated and it is expensive to harvest surgical data. Consequently, existing researches on surgical motion patterns focus on specific concise surgical tasks or simple surgical procedures. The paper presents a surgical motion pattern modeling technique that uses small data but can be applied to virtually any Endoscopic Sinus and Skull Base Surgeries (ESSBSs). The proposed method decreases the dimensionalities of the feature space through projecting surgical instrument motions into the endoscope coordinate, based on human expert domain knowledge. Furthermore, the method uses kinematic features and learns the motion pattern with Gaussian Process learning techniques. Comparing with existing surgical motion pattern modeling methods, the proposed method: 1, learns the motion model from small data; 2, can be generally applied to ESSBSs because it neither assumes nor depends on specific surgical tasks; 3, provides informative results in a real-time manner for optimizing surgical motions for improving surgical outcomes. The proposed method was verified by predicting surgical skill levels on cadaver surgeries. The results show the real-time prediction precision is higher than 81% and the offline accumulated precision reach 100%. Yangming Li, Randall A. Bly, Sarah Akkina, Fangbo Qin, Rajeev C. Saxena, Ian Humphreys, Mark Whipple, Kris S. Moe, Blake Hannaford |
ICRA | 1 |
| 2021 | Neural Sequence Segmentation as Determining the Leftmost SegmentsabstractPrior methods to text segmentation are mostly at token level.Despite the adequacy, this nature limits their full potential to capture the long-term dependencies among segments.In this work, we propose a novel framework that incrementally segments natural language sentences at segment level.For every step in segmentation, it recognizes the leftmost segment of the remaining sequence.Implementations involve LSTM-minus technique to construct the phrase representations and recurrent neural networks (RNN) to model the iterations of determining the leftmost segments.We have conducted extensive experiments on syntactic chunking and Chinese part-of-speech (POS) tagging across 3 datasets, demonstrating that our methods have significantly outperformed previous all baselines and achieved new stateof-the-art results.Moreover, qualitative analysis and the study on segmenting long-length sentences verify its effectiveness in modeling long-term dependencies. Yangming Li, Lemao Liu, Kaisheng Yao |
NAACL-HLT | 1 |
| 2021 | Knowing Where to Leverage: Context-Aware Graph Convolutional Network With an Adaptive Fusion Layer for Contextual Spoken Language UnderstandingabstractSpoken language understanding (SLU) systems aim to understand users’ utterance, which is a key component of task-oriented dialogue systems. In this paper, we focus on improving the contextual SLU. The contextual SLU systems mainly focus on how to effectively incorporate dialog context information (contextual information). The existing approaches all use the same contextual information to guide slot filling at all tokens, which may inject the irrelevant information and result in ambiguity. To tackle this problem, we propose a context-aware graph convolutional network (GCN) with an adaptive fusion layer for contextual SLU. The context-aware GCN is proposed to automatically aggregate the contextual information, which frees our model from the manually designed heuristic aggregation function. Meanwhile, an adaptive fusion layer is applied at each token to dynamically incorporate relevant contextual information, which achieves a fine-grained contextual information transfer to guide the token-level slot filling. Experiments on the Simulated Dialog Dataset show that our model achieves state-of-the-art performance and outperforms other previous methods by a large margin (+3.67% on Sim-R, +4.18% on Sim-M and +3.75% on Overall dataset). In addition, we explore and analyze the pre-trained model (i.e., BERT) in our framework. We show that incorporating BERT brings a large improvement in low-resource setting. Libo Qin 0001, Wanxiang Che, Minheng Ni, Yangming Li, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Span-Based Neural Buffer: Towards Efficient and Effective Utilization of Long-Distance Context for Neural Sequence ModelsabstractNeural sequence model, though widely used for modeling sequential data such as the language model, has sequential recency bias (Kuncoro et al. 2018) to the local context, limiting its full potential to capture long-distance context. To address this problem, this paper proposes augmenting sequence models with a span-based neural buffer that efficiently represents long-distance context, allowing a gate policy network to make interpolated predictions from both the neural buffer and the underlying sequence model. Training this policy network to utilize long-distance context is however challenging due to the simple sentence dominance problem (Marvin and Linzen 2018). To alleviate this problem, we propose a novel training algorithm that combines an annealed maximum likelihood estimation with an intrinsic reward-driven reinforcement learning. Sequence models with the proposed span-based neural buffer significantly improve the state-of-the-art perplexities on the benchmark Penn Treebank and WikiText-2 datasets to 43.9 and 35.2 respectively. We conduct extensive analysis and confirm that the proposed architecture and the training algorithm both contribute to the improvements. Yangming Li, Kaisheng Yao, Libo Qin 0001, Shuang Peng 0009, Xiaolong Li 0005 |
AAAI | 1 |
| 2020 | DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers' intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately (Kim and Kim 2018). Most of the existing systems either treat them as separate tasks or just jointly model the two tasks by sharing parameters in an implicit way without explicitly modeling mutual interaction and relation. To address this problem, we propose a Deep Co-Interactive Relation Network (DCR-Net) to explicitly consider the cross-impact and model the interaction between the two tasks by introducing a co-interactive relation layer. In addition, the proposed relation layer can be stacked to gradually capture mutual knowledge with multiple steps of interaction. Especially, we thoroughly study different relation layers and their effects. Experimental results on two public datasets (Mastodon and Dailydialog) show that our model outperforms the state-of-the-art joint model by 4.3% and 3.4% in terms of F1 score on dialog act recognition task, 5.7% and 12.4% on sentiment classification respectively. Comprehensive analysis empirically verifies the effectiveness of explicitly modeling the relation between the two tasks and the multi-steps interaction mechanism. Finally, we employ the Bidirectional Encoder Representation from Transformer (BERT) in our framework, which can further boost our performance in both tasks. Libo Qin 0001, Wanxiang Che, Yangming Li, Minheng Ni, Ting Liu 0001 |
AAAI | 3 |
| 2020 | Handling Rare Entities for Neural Sequence LabelingabstractOne great challenge in neural sequence labeling is the data sparsity problem for rare entity words and phrases.Most of test set entities appear only few times and are even unseen in training corpus, yielding large number of out-of-vocabulary (OOV) and low-frequency (LF) entities during evaluation.In this work, we propose approaches to address this problem.For OOV entities, we introduce local context reconstruction to implicitly incorporate contextual information into their representations.For LF entities, we present delexicalized entity identification to explicitly extract their frequency-agnostic and entity-typespecific representations.Extensive experiments on multiple benchmark datasets show that our model has significantly outperformed all previous methods and achieved new startof-the-art results.Notably, our methods surpass the model fine-tuned on pre-trained language models without external resource. Yangming Li, Kaisheng Yao |
ACL | 1 |
| 2020 | Slot-consistent NLG for Task-oriented Dialogue Systems with Iterative Rectification NetworkabstractData-driven approaches using neural networks have achieved promising performances in natural language generation (NLG).However, neural generators are prone to make mistakes, e.g., neglecting an input slot value and generating a redundant slot value.Prior works refer this to hallucination phenomenon.In this paper, we study slot consistency for building reliable NLG systems with all slot values of input dialogue act (DA) properly generated in output sentences.We propose Iterative Rectification Network (IRN) for improving general NLG systems to produce both correct and fluent responses.It applies a bootstrapping algorithm to sample training candidates and uses reinforcement learning to incorporate discrete reward related to slot inconsistency into training.Comprehensive studies have been conducted on multiple benchmark datasets, showing that the proposed methods have significantly reduced the slot error rate (ERR) for all strong baselines.Human evaluations also have confirmed its effectiveness. Yangming Li, Kaisheng Yao, Libo Qin 0001, Wanxiang Che, Xiaolong Li 0005, Ting Liu 0001 |
ACL | 1 |
| 2020 | LC-GAN: Image-to-image Translation Based on Generative Adversarial Network for Endoscopic ImagesabstractIntelligent vision is appealing in computer-assisted and robotic surgeries. Vision-based analysis with deep learning usually requires large labeled datasets, but manual data labeling is expensive and time-consuming in medical problems. We investigate a novel cross-domain strategy to reduce the need for manual data labeling by proposing an image-to-image translation model live-cadaver GAN (LC-GAN) based on generative adversarial networks (GANs). We consider a situation when a labeled cadaveric surgery dataset is available while the task is instrument segmentation on an unlabeled live surgery dataset. We train LC-GAN to learn the mappings between the cadaveric and live images. For live image segmentation, we first translate the live images to fake-cadaveric images with LC-GAN and then perform segmentation on the fake-cadaveric images with models trained on the real cadaveric dataset. The proposed method fully makes use of the labeled cadaveric dataset for live image segmentation without the need to label the live dataset. LC-GAN has two generators with different architectures that leverage the deep feature representation learned from the cadaveric image based segmentation task. Moreover, we propose the structural similarity loss and segmentation consistency loss to improve the semantic consistency during translation. Our model achieves better image-to-image translation and leads to improved segmentation performance in the proposed cross-domain segmentation task. Fangbo Qin, Yangming Li, Randall A. Bly, Kris S. Moe, Blake Hannaford |
IROS | 3 |
| 2020 | Discrete Computational Neural Dynamics Models for Solving Time-Dependent Sylvester Equation With Applications to Robotics and MIMO SystemsabstractIn this article, a neural dynamics model is constructed and investigated for solving time-dependent Sylvester equation with matrix inversion involved in the solving process. Besides, to eliminate the matrix inversion in the model, the quasi-Newton Broyden-Fletcher-Goldfarb-Shanno method is leveraged to construct a new model. Moreover, the global convergence performance and the effectiveness of the two discrete computational models are testified by providing theoretical analyses and numerical experiments with comparisons to the existing solutions, respectively. Two applications to robotics and the multiple-input multiple-output system are given to elucidate the feasibility of the proposed models for solving time-dependent Sylvester equation. Yimeng Qi, Long Jin 0001, Yangming Li |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | A Stack-Propagation Framework with Token-Level Intent Detection for Spoken Language UnderstandingabstractLibo Qin, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Libo Qin 0001, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Entity-Consistent End-to-end Task-Oriented Dialogue System with KB RetrieverabstractLibo Qin, Yijia Liu, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Libo Qin 0001, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Surgical Instrument Segmentation for Endoscopic Vision with Data Fusion of rediction and Kinematic PoseabstractThe real-time and robust surgical instrument segmentation is an important issue for endoscopic vision. We propose an instrument segmentation method fusing the convolutional neural networks (CNN) prediction and the kinematic pose information. First, the CNN model ToolNet-C is designed, which cascades a convolutional feature extractor trained over numerous unlabeled images and a pixel-wise segmentor trained on few labeled images. Second, the silhouette projection of the instrument body onto the endoscopic image is implemented based on the measured kinematic pose. Third, the particle filter with the shape matching likelihood and the weight suppression is proposed for data fusion, whose estimate refines the kinematic pose. The refined pose determines an accurate silhouette mask, which is the final segmentation output. The experiments are conducted with a surgical navigation system, several animal-tissue backgrounds, and a debrider instrument. Fangbo Qin, Yangming Li, Yun-Hsuan Su, De Xu, Blake Hannaford |
ICRA | 2 |
| 2019 | A Model-Based Recurrent Neural Network With Randomness for Efficient Control With ApplicationsabstractRecently, Recurrent Neural Network (RNN) control schemes for redundant manipulators have been extensively studied. These control schemes demonstrate superior computational efficiency, control precision, and control robustness. However, they lack planning completeness. This paper explains why RNN control schemes suffer from the problem. Based on the analysis, this work presents a new random RNN control scheme, which 1) introduces randomness into RNN to address the planning completeness problem, 2) improves control precision with a new optimization target, 3) improves planning efficiency through learning from exploration. Theoretical analyses are used to prove the global stability, the planning completeness, and the computational complexity of the proposed method. Software simulation is provided to demonstrate the improved robustness against noise, the planning completeness and the improved planning efficiency of the proposed method over benchmark RNN control schemes. Real-world experiments are presented to demonstrate the application of the proposed method. Yangming Li, Shuai Li 0002, Blake Hannaford |
IEEE Trans. Ind. Informatics | 1 |
| 2018 | A Novel Recurrent Neural Network for Improving Redundant Manipulator Motion Planning CompletenessabstractRecurrent Neural Networks (RNNs) demonstrated advantages on control precision, system robustness and computational efficiency, and have been widely applied to redundant manipulator control optimization. Existing RNN control schemes locally optimize trajectories and are efficient and reliable on obstacle avoidance. However, for motion planning, they suffer from local minimum and do not have planning completeness. This work explained the cause of the planning incompleteness and addressed the problem with a novel RNN control scheme. The paper presented the proposed method in detail and analyzed the global stability and the planning completeness in theory. The proposed method was compared with other three control schemes on the precision, the robustness and the planning completeness in software simulation and the results shows the proposed method has improved precision and robustness, and planning completeness. Yangming Li, Shuai Li 0002, Blake Hannaford |
ICRA | 1 |
| 2018 | Soft-obstacle Avoidance for Redundant Manipulators with Recurrent Neural NetworkabstractCompressing soft-obstacles secondary to a controlled motion task is common for human beings. While these tasks are nearly trivial for teleoperated robots, they remain a challenging problem in robotic autonomy. Addressing the problem is significant. For example, in Minimally Invasive Surgeries (MISs), safely compressing soft tissues ensures the surgical safety and decreases tissue removal, thus dramatically decreases surgical trauma and operating room time, and leads to improved surgical outcomes. In this work, we define the problem of soft-obstacle avoidance and project the safety motion constraints into the task space and the velocity space. We illustrate the significance of addressing this problem in the robotic surgery scenario. We present a Recurrent Neural Networks (RNNs) based solution, which formulates the problem as an inequality constrained optimization problem and solves it in its dual space. The application of the proposed method was demonstrated in the Raven II surgical robot. Experimental results demonstrated that the proposed method is effective in addressing the soft-obstacle avoidance problem. Yangming Li, Blake Hannaford |
IROS | 1 |
| 2018 | STMVO: biologically inspired monocular visual odometry
Yangming Li, Shuai Li 0002 |
Neural Comput. Appl. | 1 |
| 2018 | Neural Dynamics for Cooperative Control of Redundant Robot ManipulatorsabstractIn this paper, a neural-dynamic distributed scheme is proposed for the cooperative control of multiple redundant manipulators with limited communications. It is guaranteed that, with the communication network being connected, all manipulators can jointly reach the same desired motion. The proposed distributed scheme is rearranged as a time-varying quadratic program and solved online by a Zhang neural network. Then, theoretical analyses show that, without noise, the proposed distributed scheme is able to execute a given task with exponentially convergent position errors. Moreover, an explicit bound relationship between the control input noise and the end-effector position error is analytically derived. Furthermore, numerical comparisons substantiate the superiority, effectiveness, and accuracy of the proposed distributed scheme. Long Jin 0001, Shuai Li 0002, Xin Luo 0001, Yangming Li |
IEEE Trans. Ind. Informatics | 4 |
| 2017 | Roboscope: A flexible and bendable surgical robot for single portal Minimally Invasive SurgeryabstractMinimally Invasive Surgery (MIS) can reduce iatrogenic injury and decrease the possibility of surgical complications. This paper presents a novel flexible and bendable endoscopic device, “Roboscope”, which delivers two instruments, two miniature scanning fiber endoscopes, and a suction/irrigation port to the operation site through a single portal. Compared with existing bendable and steerable robotic surgical systems, Roboscope provides two bending degrees of freedom for its outer sheath and two insertion degrees of freedom, while simultaneously delivering two instruments and two endoscopes to the surgical site. Each bending axis and insertion freedom of Roboscope is independently controllable via an external actuation pack. Surgical tools can be changed without retracting the robot arm. This paper presents the design of the Roboscope mechanical system, electrical system, and control and software systems, design requirements and prototyping validation as well as analysis of Roboscope workspece. Jacob Rosen 0001, Laligam N. Sekhar, Daniel Glozman, Muneaki Miyasaka, Jesse Dosher, Brian Dellon, Kris S. Moe, Aylin Kim, Louis J. Kim, Thomas S. Lendvay, Yangming Li, Blake Hannaford |
ICRA | 11 |
| 2017 | Improving control precision and motion adaptiveness for surgical robot with recurrent neural networkabstractSurgical robot research is driven by the desire of improving surgical outcomes. This paper proposed a Recurrent Neural Network based controller to address two problems: 1) improving control precision, 2) increasing adaptiveness for robot motion (explained in Section I). RNN was adopted in this work mainly because 1) the problem formulation naturally matches RNN structure, 2) RNN has advantages as an biologically inspired method. The proposed method was explained in detail and analysis shows that the proposed method is able to dynamically regulate outputs to increase the adaptiveness and the control precision. This paper uses Raven II surgical robot as an example to show the application of the proposed method, and the numeral simulation results from the proposed method and three other controllers show that the proposed method has improved precision, improved high robustness against noise and increased movement smoothness, and it keeps the manipulator links as far away as possible from physical boundaries, which potentially increases surgical safety and leads to improved surgical outcomes. Yangming Li, Shuai Li 0002, David E. Caballero, Muneaki Miyasaka, Andrew Lewis 0001, Blake Hannaford |
IROS | 1 |
| 2017 | Distributed Recurrent Neural Networks for Cooperative Control of Manipulators: A Game-Theoretic PerspectiveabstractThis paper considers cooperative kinematic control of multiple manipulators using distributed recurrent neural networks and provides a tractable way to extend existing results on individual manipulator control using recurrent neural networks to the scenario with the coordination of multiple manipulators. The problem is formulated as a constrained game, where energy consumptions for each manipulator, saturations of control input, and the topological constraints imposed by the communication graph are considered. An implicit form of the Nash equilibrium for the game is obtained by converting the problem into its dual space. Then, a distributed dynamic controller based on recurrent neural networks is devised to drive the system toward the desired Nash equilibrium to seek the optimal solution of the cooperative control. Global stability and solution optimality of the proposed neural networks are proved in the theory. Simulations demonstrate the effectiveness of the proposed method. Shuai Li 0002, Jinbo He, Yangming Li, Muhammad Usman Rafique |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Unscented Kalman Filter and 3D vision to improve cable driven surgical robot joint angle estimationabstractCable driven manipulators are popular in surgical robots due to compact design, low inertia, and remote actuation. In these manipulators, encoders are usually mounted on the motor, and joint angles are estimated based on transmission kinematics. However, due to non-linear properties of cables such as cable stretch, lower stiffness, and uncertainties in kinematic model parameters, the precision of joint angle estimation is limited with transmission kinematics approach. To improve the positioning of these manipulators, we use a pair of low cost stereo camera as the observation for joint angles and we input these noisy measurements into an Unscented Kalman Filter (UKF) for state estimation. We use the dual UKF to estimate cable parameters and states offline. We evaluated the effectiveness of the proposed method on a Raven-II experimental surgical research platform. Additional encoders at the joint output were employed as a reference system. From the experiments, the UKF improved the accuracy of joint angle estimation by 33- 72%. Also, we tested the reliability of state estimation under camera occlusion. We found that when the system dynamics is tuned with offline UKF parameter estimation, the camera occlusion has no effect on the online state estimation. Mohammad Haghighipanah, Muneaki Miyasaka, Yangming Li, Blake Hannaford |
ICRA | 3 |
| 2016 | Dynamic modeling of cable driven elongated surgical instruments for sensorless grip force estimationabstractHaptic feedback plays a key role in surgeries, but it is still a missing component in robotic Minimally Invasive Surgeries. This paper proposes a dynamic model-based sensorless grip force estimation method to address the haptic perception problem for commonly used elongated cable-driven surgical instruments. Cable and cable-pulley properties are studied for dynamic modeling; grip forces, along with driven motor and gripper jaw positions and velocities are jointly estimated with Unscented Kalman Filter and only motor encoder readings and motor output torques are assumed to be known. A bounding filter is used to compensate for model inaccuracy and to improve method robustness. The proposed method was validated on a 10mm gripper which is driven by a Raven-II surgical robot. The gripper was equipped with 1-dimensional force sensors which served as ground truth data. The experimental results showed that the proposed method provides sufficiently good grip force estimation, while only motor encoder and the motor torques are used as observations. Yangming Li, Muneaki Miyasaka, Mohammad Haghighipanah, Blake Hannaford |
ICRA | 1 |
| 2016 | Hysteresis model of longitudinally loaded cable for cable driven robots and identification of the parametersabstractIn this paper, we propose model of longitudinally loaded cable based on the Bouc-Wen hysteresis model and within the framework of the Duhem operator. By optimizing the 9 hysteresis model parameters with a genetic algorithm, the proposed model is shown to be capable of representing quasi-static response of two different diameter cables, 0.61 mm (thin) and 1.19 mm (thick), used for the RAVEN II surgical robotic surgery platform. The construction of the cable is 7 strands with 19 individual wires per strand. Furthermore, it is shown that the dynamic response of the cables are captured by adding a linear damping term. The hysteresis model and linear damper with the optimized parameters accurately models a longitudinal vibration test result in terms of frequency, steady state stretch, and logarithmic decrement. Energy dissipation due solely to the hysteresis term is approximately calculated to be 57 and 71% of the total energy loss for the thin and thick cables respectively. The proposed model may be used for cables with different contraction and diameter and can be applied for control of cable driven robots in which cables are stretched longitudinally without large excitation of other modes. Muneaki Miyasaka, Mohammad Haghighipanah, Yangming Li, Blake Hannaford |
ICRA | 3 |
| 2015 | Improving position precision of a servo-controlled elastic cable driven surgical robot using Unscented Kalman FilterabstractCable driven power transmission is popular in many manipulator applications including medical arms. In spite of advantages obtained by removing motors from the mechanism, cable transmission introduces higher non-linearity and more uncertainties such as cable stretch and cable coupling. In order to improve the control precision and robustness of the Raven-II surgical robot, particularly for automation applications, the Unscented Kalman Filter (UKF) was adopted for state estimation. The UKF estimated state variables of the Raven-II dynamic model from sensor data. The dual UKF was used offline to estimate cable coupling parameters. The experimental results showed that the proposed method improved joint position estimation precision and the estimation consistency, especially on the more elastic links. The improvements for links 2 and 3 of the Raven were 36.76%, and 62.99%, respectively. For link 1 the improvement was 1.43% because the transmission is very stiff. Mohammad Haghighipanah, Yangming Li, Muneaki Miyasaka, Blake Hannaford |
IROS | 2 |
| 2014 | Nonlinearly Activated Neural Network for Solving Time-Varying Complex Sylvester EquationabstractThe Sylvester equation is often encountered in mathematics and control theory. For the general time-invariant Sylvester equation problem, which is defined in the domain of complex numbers, the Bartels-Stewart algorithm and its extensions are effective and widely used with an O(n³) time complexity. When applied to solving the time-varying Sylvester equation, the computation burden increases intensively with the decrease of sampling period and cannot satisfy continuous realtime calculation requirements. For the special case of the general Sylvester equation problem defined in the domain of real numbers, gradient-based recurrent neural networks are able to solve the time-varying Sylvester equation in real time, but there always exists an estimation error while a recently proposed recurrent neural network by Zhang et al [this type of neural network is called Zhang neural network (ZNN)] converges to the solution ideally. The advancements in complex-valued neural networks cast light to extend the existing real-valued ZNN for solving the time-varying real-valued Sylvester equation to its counterpart in the domain of complex numbers. In this paper, a complex-valued ZNN for solving the complex-valued Sylvester equation problem is investigated and the global convergence of the neural network is proven with the proposed nonlinear complex-valued activation functions. Moreover, a special type of activation function with a core function, called sign-bi-power function, is proven to enable the ZNN to converge in finite time, which further enhances its advantage in online processing. In this case, the upper bound of the convergence time is also derived analytically. Simulations are performed to evaluate and compare the performance of the neural network with different parameters and activation functions. Both theoretical analysis and numerical simulations validate the effectiveness of the proposed method. Shuai Li 0002, Yangming Li |
IEEE Trans. Cybern. | 2 |
| 2014 | Fast and Robust Data Association Using Posterior Based Approximate Joint Compatibility TestabstractData association is a fundamental problem in multisensor fusion, tracking, and localization. The joint compatibility test is commonly regarded as the true solution to the problem. However, traditional joint compatibility tests are computationally expensive, are sensitive to linearization errors, and require the knowledge of the full covariance matrix of state variables. The paper proposes a posterior-based joint compatibility test scheme to conquer the three problems mentioned above. The posterior-based test naturally separates the test of state variables from the test of observations. Therefore, through the introduction of the robot movement and proper approximation, the joint test process is sequentialized to the sum of individual tests; therefore, the test has$O(n)$complexity (compared with$O(n^{2})$for traditional tests), where$n$denotes the total number of related observations. At the same time, the sequentialized test neither requires the knowledge to the full covariance matrix of state variables nor is sensitive to linearization errors caused by poor pose estimates. The paper also shows how to apply the proposed method to various simultaneous localization and mapping (SLAM) algorithms. Theoretical analysis and experiments on both simulated data and popular datasets show the proposed method outperforms some classical algorithms, including sequential compatibility nearest neighbor (SCNN), random sample consensus (RANSAC), and joint compatibility branch and bound (JCBB), on precision, efficiency, and robustness. Yangming Li, Shuai Li 0002, Quanjun Song, Max Q.-H. Meng |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | A biologically inspired solution to simultaneous localization and consistent mapping in dynamic environments
Yangming Li, Shuai Li 0002, Yunjian Ge |
Neurocomputing | 1 |
| 2013 | Decentralized control of collaborative redundant manipulators with partial command coverage via locally connected recurrent neural networks
Shuai Li 0002, Hongzhu Cui, Yangming Li, Bo Liu 0006, Yuesheng Lou |
Neural Comput. Appl. | 3 |
| 2013 | A class of finite-time dual neural networks for solving quadratic programming problems and its k-winners-take-all application
Shuai Li 0002, Yangming Li, Zheng Wang 0009 |
Neural Networks | 2 |
| 2013 | Using Laplacian Eigenmap as Heuristic Information to Solve Nonlinear Constraints Defined on a Graph and Its Application in Distributed Range-Free Localization of Wireless Sensor Networks
Shuai Li 0002, Zheng Wang 0009, Yangming Li |
Neural Process. Lett. | 3 |
| 2013 | Selective Positive-Negative Feedback Produces the Winner-Take-All Competition in Recurrent Neural NetworksabstractThe winner-take-all (WTA) competition is widely observed in both inanimate and biological media and society. Many mathematical models are proposed to describe the phenomena discovered in different fields. These models are capable of demonstrating the WTA competition. However, they are often very complicated due to the compromise with experimental realities in the particular fields; it is often difficult to explain the underlying mechanism of such a competition from the perspective of feedback based on those sophisticate models. In this paper, we make steps in that direction and present a simple model, which produces the WTA competition by taking advantage of selective positive-negative feedback through the interaction of neurons via p-norm. Compared to existing models, this model has an explicit explanation of the competition mechanism. The ultimate convergence behavior of this model is proven analytically. The convergence rate is discussed and simulations are conducted in both static and dynamic competition scenarios. Both theoretical and numerical results validate the effectiveness of the dynamic equation in describing the nonlinear phenomena of WTA competition. Shuai Li 0002, Bo Liu 0006, Yangming Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2012 | IPJC: The Incremental Posterior Joint Compatibility test for fast feature cloud matchingabstractOne of the fundamental challenges in robotics is data-association: determining which sensor observations correspond to the same physical object. A common approach is to consider groups of observations simultaneously: a constellation of observations can be significantly less ambiguous than the observations considered individually. The Joint Compatibility Branch and Bound (JCBB) test is the gold standard method for these data association problems. But its computational complexity and its sensitivity to non-linearities limit its practical usefulness. We propose the Incremental Posterior Joint Compatibility (IPJC) test. While equivalent to JCBB on linear problems, it is significantly more accurate on non-linear problems. When used for feature-cloud matching (an important special case), IPJC is also dramatically faster than JCBB. We demonstrate the advantages of IPJC over JCBB and other commonly-used methods on both synthetic and real-world datasets. Yangming Li, Edwin Olson |
IROS | 1 |
| 2012 | Decentralized kinematic control of a class of collaborative redundant manipulators via recurrent neural networks
Shuai Li 0002, Sanfeng Chen, Bo Liu 0006, Yangming Li, Yongsheng Liang 0001 |
Neurocomputing | 4 |
| 2011 | Structure tensors for general purpose LIDAR feature extractionabstractThe detection of features from Light Detection and Ranging (LIDAR) data is a fundamental component of feature-based mapping and SLAM systems. Classical approaches are often tied to specific environments, computationally expensive, or do not extract precise features. We describe a general purpose feature detector that is not only efficient, but also applicable to virtually any environment. Our method shares its mathematical foundation with feature detectors from the computer vision community, where structure tensor based methods have been successful. Our resulting method is capable of identifying stable and repeatable features at a variety of spatial scales, and produces uncertainty estimates for use in a state estimation algorithm. We verify the proposed method on standard datasets, including the Victoria Park dataset and the Intel Research Center dataset. Yangming Li, Edwin Olson |
ICRA | 1 |
| 2010 | Extracting general-purpose features from LIDAR dataabstractThe detection of features from Light Detection and Ranging (LIDAR) data is a fundamental component of feature-based mapping and SLAM systems. Existing detectors tend to exploit characteristics of specific environments: corners and lines from indoor (rectilinear) environments, and trees from outdoor environments. While these detectors work well in their intended environments, their performance in different environments can be very poor. We describe a general purpose feature detector for LIDAR data that is applicable to virtually any environment. Our methods adapt classic feature detection methods from the image processing literature, specifically the multi-scale Kanade-Tomasi corner detector. Our resulting method is capable of identifying stable features at a variety of spatial scales and produces uncertainty estimates for use in a state estimation algorithm. We present results on standard datasets, including Victoria Park and Intel Research Center (both 2D), and the MIT DARPA Urban Challenge dataset (3D). Yangming Li, Edwin Olson |
ICRA | 1 |