VLDB 2026 Research / reviewers in the wild / expert
Yu Sun 0004
dblp:62/3689-4
· DBLP profile ↗
57ranked-venue papers
9as first author
20since 2021 · last 2025
0000-0001-9953-291XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 8 first-author · 17 since 2021Systems, architecture and hardware · 28 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Deep Learning Approach for Post-Operative Neonatal Pain Detection and Prediction Through Physiological SignalsabstractIt is well-known that severe pain and powerful pain medications cause short- and long-term damage to the developing nervous system of newborns. Caregivers routinely use physiological vital signs [Heart Rate (HR), Respiration Rate (RR), Oxygen Saturation (SR)] to monitor post-surgical pain in the Neonatal Intensive Care Unit (NICU). Here we present a novel approach that combines continuous, non-invasive monitoring of these vital signs and Computer Vision/Deep Learning to make automatic neonate pain detection with an accuracy of 74% AUC, 67.59% mAP. Further, we report for the first time our Early Pain Detection (EPD) approach that explores prediction of the time to onset of post-surgical pain in neonates. Our EPD can alert NICU workers to postoperative neonatal pain about 5 to 10 minutes prior to pain onset. In addition to alleviating the need for intermittent pain assessments by busy NICU nurses via long-term observation, our EPD approach creates a time window prior to pain onset for the use of less harmful pain mitigation strategies. Through effective pain mitigation prior to spinal sensitization, EPD could minimize or eliminate severe post-surgical pain and the consequential need for powerful analgesics in post-surgical neonates. Jacqueline Hausmann, Marcia Kneusel, Stephanie Prescott, Peter R. Mouton, Yu Sun 0004, Dmitry B. Goldgof |
CBMS | 6 |
| 2025 | Few-Shot Prompting with Vision Language Model for Pain Classification in Infant Cry SoundsabstractAccurately detecting pain in infants remains a complex challenge. Conventional deep neural networks used for analyzing infant cry sounds typically demand large labeled datasets, substantial computational power, and often lack interpretability. In this work, we introduce a novel approach that leverages OpenAI's vision-language model, GPT-4(V), combined with mel spectrogram-based representations of infant cries through prompting. This prompting strategy significantly reduces the dependence on large training datasets while enhancing transparency and interpretability. Using the USF-MNPAD-II dataset, our method achieves an accuracy of 83.33% with only 16 training samples, in contrast to the 4,914 samples required in the baseline model. To our knowledge, this represents the first application of few-shot prompting with vision-language models such as GPT-4o for infant pain classification. Anthony McCofie, Abhiram Kandiyana, Peter R. Mouton, Yu Sun 0004, Dmitry B. Goldgof |
CBMS | 4 |
| 2025 | Only Pick Once: Algorithms for Efficiently Picking an Exact Number of Multiple Identical ObjectsabstractPicking up multiple objects at once is a grasping skill that makes a human worker efficient in many domains. This work tackles the problem of getting a requested number of identical objects in a shallow bin by only pick once (OPO) using a simple parallel gripper. The proposed system contains several graph-based algorithms that convert the layout of objects into a graph, cluster vertices in the graph, rank and select candidate clusters based on their topology. Our algorithm also has a multi-object picking predictor based on a convolutional neural network for estimating how many objects would be picked up with a given gripper location and orientation. This paper presents four evaluation metrics and three protocols to evaluate the proposed system. The results show that our proposed system has very high success rates for two and three objects when only picking once. Utilizing our approach can significantly outperform single object picking two to three times in terms of efficiency. The results also show our algorithm can be applied to one unseen shape (hexagon) and unseen sizes cube and cylinder during training without fine-tuning to achieve decent accuracy. Zihe Ye, Ricardo Frumento, Yu Sun 0004 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Enhancing Concept-Based Explanation with Vision-Language ModelsabstractAlthough concept-based approaches are widely used to explain a model's behavior and assess the contributions of different concepts in decision-making, identifying relevant concepts can be challenging for non-experts. This paper introduces a novel method that simplifies concept selection by leveraging the capabilities of a state-of-the-art large Vision-Language Model (VLM). Our method employs a VLM to select textual concepts that describe the classes in the target dataset. We then transform these influential textual concepts into human-readable image concepts using a text-to-image model. This process allows us to explain the targeted network in a post-hoc manner. Further, we use directional derivatives and concept activation vectors to quantify the importance of the generated concepts. We evaluate our method on a neonatal pain classification task, analyzing the sensitivity of the model's output for the generated concepts. The results demonstrate that the VLM not only generates coherent and meaningful concepts that are easily understandable by non-experts but also achieves performance comparable to that of natural image concepts without the need for additional annotation costs. Md Imran Hossain, Ghada Zamzmi, Peter R. Mouton, Yu Sun 0004, Dmitry B. Goldgof |
CBMS | 4 |
| 2024 | On Training Data Influence of GPT ModelsabstractAmidst the rapid advancements in generative language models, the investigation of how training data shapes the performance of GPT models is still emerging.This paper presents GPTfluence, a novel approach that leverages a featurized simulation to assess the impact of training examples on the training dynamics of GPT models.Our approach not only traces the influence of individual training instances on performance trajectories, such as loss and other key metrics, on targeted test points but also enables a comprehensive comparison with existing methods across various training scenarios in GPT models, ranging from 14 million to 2.8 billion parameters, across a range of downstream tasks.Contrary to earlier methods that struggle with generalization to new data, GPTfluence introduces a parameterized simulation of training dynamics, demonstrating robust generalization capabilities to unseen training data.This adaptability is evident across both fine-tuning and instruction-tuning scenarios, spanning tasks in natural language understanding and generation.We make our Yekun Chai, Qingyi Liu, Shuohuan Wang, Yu Sun 0004, Qiwei Peng 0002, Hua Wu 0003 |
EMNLP | 4 |
| 2024 | Tool-Augmented Reward ModelingabstractReward modeling (*a.k.a.*, preference modeling) is instrumental for aligning large language models with human preferences, particularly within the context of reinforcement learning from human feedback (RLHF). While conventional reward models (RMs) have exhibited remarkable scalability, they oft struggle with fundamental functionality such as arithmetic computation, code execution, and factual lookup. In this paper, we propose a tool-augmented preference modeling approach, named Themis, to address these limitations by empowering RMs with access to external environments, including calculators and search engines. This approach not only fosters synergy between tool utilization and reward grading but also enhances interpretive capacity and scoring reliability. Our study delves into the integration of external tools into RMs, enabling them to interact with diverse external sources and construct task-specific tool engagement and reasoning traces in an autoregressive manner. We validate our approach across a wide range of domains, incorporating seven distinct external tools. Our experimental results demonstrate a noteworthy overall improvement of 17.7% across eight tasks in preference ranking. Furthermore, our approach outperforms Gopher 280B by 7.3% on TruthfulQA task in zero-shot evaluation. In human evaluations, RLHF trained with Themis attains an average win rate of 32% when compared to baselines across four distinct tasks. Additionally, we provide a comprehensive collection of tool-related RM datasets, incorporating data from seven distinct tool APIs, totaling 15,000 instances. We have made the code, data, and model checkpoints publicly available to facilitate and inspire further research advancements (https://github.com/ernie-research/Tool-Augmented-Reward-Model). Lei Li 0040, Yekun Chai, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Ningyu Zhang 0001, Hua Wu 0003 |
ICLR | 4 |
| 2024 | From Cooking Recipes to Robot Task Trees - Improving Planning Correctness and Task Efficiency by Leveraging LLMs with a Knowledge NetworkabstractTask planning for robotic cooking involves generating a sequence of actions for a robot to prepare a meal successfully. This paper introduces a novel task tree generation pipeline producing correct planning and efficient execution for cooking tasks. Our method first uses a large language model (LLM) to retrieve recipe instructions and then utilizes a fine-tuned GPT-3 to convert them into a task tree, capturing sequential and parallel dependencies among subtasks. The pipeline then mitigates the uncertainty and unreliable features of LLM outputs using task tree retrieval. We combine multiple LLM task tree outputs into a graph and perform a task tree retrieval to avoid questionable nodes and high-cost nodes to improve planning correctness and execution efficiency. Our evaluation results show its superior performance in task planning accuracy and efficiency compared to previous works. Md Sadman Sakib, Yu Sun 0004 |
ICRA | 2 |
| 2023 | Uncertainty-aware Unsupervised Video HashingabstractLearning to hash has become popular for video retrieval due to its fast speed and low storage consumption. Previous efforts formulate video hashing as training a binary auto-encoder, for which noncontinuous latent representations are optimized by the biased straight-through (ST) back-propagation heuristic. We propose to formulate video hashing as learning a discrete variational auto-encoder with the factorized Bernoulli latent distribution, termed as Bernoulli variational auto-encoder (BerVAE). The corresponding evidence lower bound (ELBO) in our BerVAE implementation leads to closed-form gradient expression, which can be applied to achieve principled training along with some other unbiased gradient estimators. BerVAE enables uncertainty-aware video hashing by predicting the probability distribution of video hash code-words, thus providing reliable uncertainty quantification. Experiments on both simulated and real-world large-scale video data demonstrate that our BerVAE trained with unbiased gradient estimators can achieve the state-of-the-art retrieval performance. Furthermore, we show that quantified uncertainty is highly correlated to video retrieval performance, which can be leveraged to further improve the retrieval accuracy. Our code is available at https://github.com/wangyucheng1234/BerVAE Mingyuan Zhou, Yu Sun 0004, Xiaoning Qian |
AISTATS | 3 |
| 2023 | Enhancing Neonatal Pain Assessment Transparency via Explanatory Training Examples IdentificationabstractDeep Learning (DL)-based solutions have shown promising performance in assessing neonatal pain. However, the occlusion of the visual modality (face and body) is common in clinical settings due to several factors, including a prone sleeping position, low light, or swaddling. In such scenarios, other pain signals, such as audio signals, can be used as the major behavioral signs of pain. Although DL-based methods are proposed to assess pain from audio, these methods lack transparency and explainability (black box), which can decrease the user's trust in the automated decision. In this work, we visualize the neonate's audio signal as a spectrogram image to classify it as pain or no pain and present an instance-based approach for explaining the decision of the black-box model. Further, this work provides an analysis of the most helpful and harmful training instances using an influence score followed by assessing their impact on pain prediction. Experimental results demonstrate that the proposed approach can detect and remove harmful instances, eventually leading to a compressed dataset. Our results also show that the proposed work can add explainability to the current DL-based pain detection methods, which can enhance users' trust and provide a viable approach toward pain assessment in clinical settings. Md Imran Hossain, Ghada Zamzmi, Peter R. Mouton, Yu Sun 0004, Dmitry B. Goldgof |
CBMS | 4 |
| 2023 | ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-ExpertsabstractRecent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of image fidelity and text relevancy. In this paper, we propose ERNIE-ViLG 2.0, a large-scale Chinese text-to-image diffusion model, to progressively upgrade the quality of generated images by: (1) incorporating fine-grained textual and visual knowledge of key elements in the scene, and (2) utilizing different denoising experts at different denoising stages. With the proposed mechanisms, ERNIE-ViLG 2.01not only achieves a new state-of-the-art on MS-COCO with zero-shot FID-30k score of 6.75, but also significantly outperforms recent models in terms of image fidelity and image-text alignment, with side-by-side human evaluation on the bilingual prompt set ViLG-300. Zhida Feng, Zhenyu Zhang 0006, Yewei Fang, Lanxin Li, Xuyi Chen, Jiaxiang Liu 0004, Weichong Yin, Shikun Feng, Yu Sun 0004, Li Chen 0011, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
CVPR | 11 |
| 2022 | Multi-Object Grasping - Types and TaxonomyabstractThis paper proposes 12 multi-object grasps (MOGs) types from a human and robot grasping data set. The grasp types are then analyzed and organized into a MOG taxonomy. This paper first presents three MOG data collection setups: a human finger tracking setup for multi-object grasping demonstrations, a real system with Barretthand, UR5e arm, and a MOG algorithm, a simulation system with the same settings as the real system. Then the paper describes a novel stochastic grasping routine designed based on a biased random walk to explore the robotic hand's configuration space for feasible MOGs. Based on obser-vations in both the human demonstrations and robotic MOG solutions, this paper proposes 12 MOG types in two groups: shape-based types and function-based types. The new MOG types are compared using six characteristics and then compiled into a taxonomy. This paper then introduces the observed MOG type combinations and shows examples of 16 different combinations. Yu Sun 0004, Eliza Amatova, Tianze Chen |
ICRA | 1 |
| 2022 | Multi-Object Grasping - Efficient Robotic Picking and Transferring Policy for Batch PickingabstractIn a typical fulfillment center, the order fulfilling process is managed by a warehouse management system (WMS). For efficiency, WMS usually applies batch picking, also called multi-order picking, to collect the same items for multiple orders. Suppose an item appears in multiple orders, instead of repeatedly revisiting the exact picking location multiple times, a picker will be instructed to pick up multiple same items at once and bring them to a sorting station, also called a re-bin station. It is at the re-bin station, where the workers sort the picked items into separate orders. We have seen many robotic technologies being developed for sorting. However, we have not seen any feasible robotic technology for batch picking. Transferring multiple objects between bins is a common task. In robotics, a standard approach is to transfer a single object at a time. However, grasping multiple objects and transferring them at once is more efficient. This paper presents a set of novel strategies for efficiently grasping and transferring multiple objects. The grasping strategies enable a robotic hand to grasp multiple objects by identifying an optimal ready hand configuration (pre-grasp), calculating a flexion synergy based on the desired quantity of objects to be grasped, and utilizing a deep learning model to signal the completion of a grasp. The transferring strategies demonstrate an approach that models the problem as a Markov decision process (MDP) and defines specific grasping actions to efficiently transfer objects when the required quantity is larger than the capability of a single grasp. Using the MDP model, the approach can generate an optimal pick-transfer policy that minimizes the number of transfers. The complete proposed approach has been evaluated in both a simulation environment and on a real robotic system. The proposed approach reduces the number of transfers by 59% and the number of lifts by 58% compared to an optimal single object pick-transfer solution. Adheesh Shenoy, Tianze Chen, Yu Sun 0004 |
IROS | 3 |
| 2022 | Attentional Generative Multimodal Network for Neonatal Postoperative Pain Estimation
Md Sirajus Salekin, Ghada Zamzmi, Dmitry B. Goldgof, Peter R. Mouton, Kanwaljeet J. S. Anand, Terri Ashmeade, Stephanie Prescott, Yangxin Huang, Yu Sun 0004 |
MICCAI (3) | 9 |
| 2022 | A Comprehensive and Context-Sensitive Neonatal Pain Assessment Using Computer VisionabstractInfants receiving care in the Neonatal Intensive Care Unit (NICU) experience several painful procedures during their hospitalization. Assessing neonatal pain is difficult because the current standard for assessment is subjective, inconsistent, and discontinuous. The intermittent and inconsistent assessment can induce poor treatment and, therefore, cause serious long-term outcomes. In this paper, we present a comprehensive pain assessment system that utilizes facial expressions along with crying sounds, body movement, and vital sign changes. The proposed automatic system generates a standardized pain assessment comparable to those obtained by conventional nurse-derived pain scores. The system achieved 95.56 percent accuracy using decision fusion of different pain responses that were recorded in a challenging clinical environment. In addition to the decision fusion, we present the performance of multimodal assessment using other fusion schemes as well as a unimodal assessment approach. We also discuss the impact of different factors (e.g., gestational age) on pain, propose several group-specific models for pain assessment (e.g., pre-term and full-term models), and compare the performance of these models with the performance of general models. While further research is needed, our results show that the automatic assessment of neonatal pain is a viable and more efficient alternative to the manual assessment. Ghada Zamzmi, Chih-Yun Pai, Dmitry B. Goldgof, Rangachar Kasturi, Terri Ashmeade, Yu Sun 0004 |
IEEE Trans. Affect. Comput. | 6 |
| 2021 | ERNIE-Doc: A Retrospective Long-Document Modeling TransformerabstractSiYu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Task Planning with a Weighted Functional Object-Oriented NetworkabstractIn reality, there is still much to be done for robots to be able to perform manipulation actions with full autonomy. Complicated manipulation tasks, such as cooking, may still require a person to perform some actions that are very risky for a robot to perform. On the other hand, some other actions may be very risky for a human with physical disabilities to perform. Therefore, it is necessary to balance the workload of a robot and a human based on their limitations while minimizing the effort needed from a human in a collaborative robot (cobot) set-up. This paper proposes a new version of our functional object-oriented network (FOON) that integrates weights in its functional units to reflect a robot’s chance of successfully executing an action of that functional unit. The paper also presents a task planning algorithm for the weighted FOON to allocate manipulation action load to the robot and human to achieve optimal performance while minimizing human effort. Through a number of experiments, this paper shows several successful cases in which using the proposed weighted FOON and the task planning algorithm allow a robot and a human to successfully complete complicated tasks together with higher success rates than a robot doing them alone. David Paulius, Kelvin Sheng Pei Dong, Yu Sun 0004 |
ICRA | 3 |
| 2021 | Multi-Object Grasping - Estimating the Number of Objects in a Robotic GraspabstractA human hand can grasp a desired number of objects at once from a pile based solely on tactile sensing. To do so, a robot needs to make a grasp in a pile, sense the number of objects in the grasp before lifting, and predict how many will remain in the grasp after lifting. It is a very challenging problem because when making the prediction, the robotic hand is still in the pile and the objects in the grasp are not observable to vision systems. Moreover, some objects in the hand before lifting may fall out the grasp when the lifting starts because they were supported by other objects in the pile instead of the fingers. A robotic hand should sense how many objects are in a grasp using its tactile sensors before lifting. This paper presents novel multi-object grasping analyzing methods to solve this problem. They include a grasp volume calculation, tactile force analysis, and a data-driven deep learning approach. The methods have been implemented on a Barrett hand and then evaluated in simulations and a real setup with a robotic system. The evaluation results conclude that once the Barrett hand grasps multiple objects in the pile, the data-driven models can make a good prediction before lifting on how many objects will remain in the hand after lifting. The root-mean-square errors are 0.74 for balls and 0.58 for cubes in simulations, and 1.06 for balls and 1.45 for cubes in the real system. Tianze Chen, Adheesh Shenoy, Anzhelika Kolinko, Syed Shah, Yu Sun 0004 |
IROS | 5 |
| 2021 | Learning State-Dependent Sensor Measurement Models with Limited Sensor MeasurementsabstractWe present a two-stage transfer learning method for training state-dependent sensor measurement models (SDSMMs) with limited sensor data. This method can alleviate collecting sizeable sensor and ground truth data to learn accurate sensor models, especially when we must learn many sensor models (for example, a fleet of autonomous cars, drones, or warehouse robots). In the first stage, we use prior knowledge of the sensor (such as a physical model) to generate a sizeable artificial dataset. Then the artificial dataset is used to pre-train an SDSMM. The second stage fine-tunes the pre-trained SDSMM using a "small" number of data collected by our target real sensor. To our knowledge, we are the first to learn measurement distributions using data generated from a physical model and data from a real sensor. We evaluated our proposed method using the Extended Kalman Particle Filter and a real-world localization dataset collected by several robots. Compared to the prior method, the proposed method achieved comparable performance with as little as ~19% of the real training data. Troi Williams, Yu Sun 0004 |
IROS | 2 |
| 2021 | ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language UnderstandingabstractDongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dongling Xiao, Yu-Kun Li, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
NAACL-HLT | 4 |
| 2021 | Pattern Recognition in Vital Signs Using SpectrogramsabstractSpectrograms visualize the frequency components of a given signal which may be an audio signal or even a time-series signal. Audio signals have higher sampling rate and high variability of frequency with time. Spectrograms can capture such variations well. But, vital signs which are time-series signals have less sampling frequency and low-frequency variability due to which, spectrograms fail to express variations and patterns. In this paper, we propose a novel solution to introduce frequency variability using frequency modulation on vital signs. Then we apply spectrograms on frequency modulated signals to capture the patterns. The proposed approach has been evaluated on 4 different medical datasets across both prediction and classification tasks. Significant results are found showing the efficacy of the approach for vital sign signals. The results from the proposed approach are promising with an accuracy of 91.55% and 91.67% in prediction and classification tasks respectively. Sidharth Srivatsav Sribhashyam, Md Sirajus Salekin, Dmitry B. Goldgof, Ghada Zamzmi, Mark Last, Yu Sun 0004 |
SMC | 6 |
| 2020 | ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingabstractRecently pre-trained models have achieved state-of-the-art results in various language understanding tasks. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides co-occurring information, there exists other valuable lexical, syntactic and semantic information in training corpora, such as named entities, semantic closeness and discourse relations. In order to extract the lexical, syntactic and semantic information from training corpora, we propose a continual pre-training framework named ERNIE 2.0 which incrementally builds pre-training tasks and then learn pre-trained models on these constructed tasks via continual multi-task learning. Based on this framework, we construct several tasks and train the ERNIE 2.0 model to capture lexical, syntactic and semantic aspects of information in the training data. Experimental results demonstrate that ERNIE 2.0 model outperforms BERT and XLNet on 16 tasks including English tasks on GLUE benchmarks and several similar tasks in Chinese. The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE. Yu Sun 0004, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 1 |
| 2020 | First Investigation into the Use of Deep Learning for Continuous Assessment of Neonatal Postoperative PainabstractThis paper presents the first investigation into the use of fully automated deep learning framework for assessing neonatal postoperative pain. It specifically investigates the use of Bilinear Convolutional Neural Network (B-CNN) to extract facial features during different levels of postoperative pain followed by modeling the temporal pattern using Recurrent Neural Network (RNN). Although acute and postoperative pain have some common characteristics (e.g., visual action units), postoperative pain has a different dynamic, and it evolves in a unique pattern over time. Our experimental results indicate a clear difference between the pattern of acute and postoperative pain. They also suggest the efficiency of using a combination of bilinear CNN with RNN model for the continuous assessment of postoperative pain intensity. Md Sirajus Salekin, Ghada Zamzmi, Dmitry B. Goldgof, Rangachar Kasturi, Thao Ho, Yu Sun 0004 |
FG | 6 |
| 2020 | Developing Motion Code Embedding for Action Recognition in VideosabstractIn this work, we propose a motion embedding strategy known as motion codes, which is a vectorized representation of motions based on a manipulation's salient mechanical attributes. These motion codes provide a robust motion representation, and they are obtained using a hierarchy of features called the motion taxonomy. We developed and trained a deep neural network model that combines visual and semantic features to identify the features found in our motion taxonomy to embed or annotate videos with motion codes. To demonstrate the potential of motion codes as features for machine learning tasks, we integrated the extracted features from the motion embedding model into the current state-of-the-art action recognition model. The obtained model achieved higher accuracy than the baseline model for the verb classification task on egocentric videos from the EPIC-KITCHENS dataset. Maxat Alibayev, David Paulius, Yu Sun 0004 |
ICPR | 3 |
| 2020 | ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language GenerationabstractCurrent pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks. To address this issue, we propose an enhanced multi-flow sequence to sequence pre-training and fine-tuning framework named ERNIE-GEN, which bridges the discrepancy between training and inference with an infilling generation mechanism and a noise-aware generation method. To make generation closer to human writing patterns, this framework introduces a span-by-span generation flow that trains the model to predict semantically-complete spans consecutively rather than predicting word by word. Unlike existing pre-training methods, ERNIE-GEN incorporates multi-granularity target sampling to construct pre-training data, which enhances the correlation between encoder and decoder. Experimental results demonstrate that ERNIE-GEN achieves state-of-the-art results with a much smaller amount of pre-training data and parameters on a range of language generation tasks, including abstractive summarization (Gigaword and CNN/DailyMail), question generation (SQuAD), dialogue generation (Persona-Chat) and generative question answering (CoQA). The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE/ernie-gen. Dongling Xiao, Yu-Kun Li, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
IJCAI | 4 |
| 2020 | Estimating Motion Codes from Demonstration VideosabstractA motion taxonomy can encode manipulations as a binary-encoded representation, which we refer to as motion codes. These motion codes innately represent a manipulation action in an embedded space that describes the motion's mechanical features, including contact and trajectory type. The key advantage of using motion codes for embedding is that motions can be more appropriately defined with robotic-relevant features, and their distances can be more reasonably measured using these motion features. In this paper, we develop a deep learning pipeline to extract motion codes from demonstration videos in an unsupervised manner so that knowledge from these videos can be properly represented and used for robots. Our evaluations show that motion codes can be extracted from demonstrations of action in the EPIC-KITCHENS dataset. Maxat Alibayev, David Paulius, Yu Sun 0004 |
IROS | 3 |
| 2020 | Generalizing Learned Manipulation Skills in PracticeabstractRobots should be able to learn and perform a manipulation task across different settings. This paper presents an approach that learns an RNN-based manipulation skill model from demonstrations and then generalizes the learned skill in new settings. The manipulation skill model learned from demonstrations in an initial set of setting performs well in those settings and similar ones. However, the model may perform poorly in a novel setting that is significantly different from the learned settings. Therefore a novel approach called generalization in practice (GiP) is developed to tackle this critical problem. In this approach, the robot practices in the new setting to obtain new training data and refine the learned skill using the new data to gradually improve the learned skill model. The proposed approach has been implemented for one type of manipulation task - pouring that is the most performed manipulation in cooking applications. The presented approach enables a pouring robot to pour gracefully like a person in terms of speed and accuracy in learned setups and gradually improve the pouring performance in novel setups after several practices. Juan Wilches, Yongqiang Huang 0001, Yu Sun 0004 |
IROS | 3 |
| 2019 | Joint Object and State Recognition Using Language KnowledgeabstractThe state of an object is an important piece of knowledge in robotics applications. States and objects are intertwined together, meaning that object information can help recognize the state of an image and vice versa. This paper addresses the state identification problem in cooking related images and uses state and object predictions together to improve the classification accuracy of objects and their states from a single image. The pipeline presented in this paper includes a CNN with a double classification layer and the Concept-Net language knowledge graph on top. The language knowledge creates a semantic likelihood between objects and states. The resulting object and state confidences from the deep architecture are used together with object and state relatedness estimates from a language knowledge graph to produce marginal probabilities for objects and states. The marginal probabilities and confidences of objects (or states) are fused together to improve the final object (or state) classification results. Experiments on a dataset of cooking objects show that using a language knowledge graph on top of a deep neural network effectively enhances object and state classification. Ahmad Babaeian Jelodar, Yu Sun 0004 |
ICIP | 2 |
| 2019 | Pain Assessment From Facial Expression: Neonatal Convolutional Neural Network (N-CNN)abstractThe current standard for assessing neonatal pain is discontinuous and suffers from inter-observer variations, which can result in delayed intervention and inconsistent treatment of pain. Therefore, it is critical to address the shortcomings of the current standard and develop continuous and less subjective pain assessment tools. Convolutional Neural Networks have gained much popularity in the last decades due to the wide range of its successful applications in medical image analysis, object recognition, and emotion recognition. In this paper, we propose a Neonatal Convolutional Neural Network, designed and trained end-to-end to detect neonatal pain. We evaluated the proposed network in two data sets of neonates and compared its performance to the performance of ResNet architecture in the same data sets. Our proposed method outperformed ResNet in recognizing neonates' pain and achieved around 91.00% accuracy. While further research is needed, our preliminary results suggest that the presented network can be used for automatic pain assessment, and possibly similar applications. It also suggests that the automatic recognition of neonatal pain provides a viable and more efficient alternative to the current standard of pain assessment. Ghada Zamzmi, Rahul Paul, Dmitry B. Goldgof, Rangachar Kasturi, Yu Sun 0004 |
IJCNN | 5 |
| 2019 | Accurate Pouring using Model Predictive Control Enabled by Recurrent Neural NetworkabstractHumans perform the task of pouring often and in which exhibit consistent accuracy regardless of the complicated dynamics of the liquid. Model predictive control (MPC) appears to be a natural candidate solution for the task of accurate pouring considering its wide use in industrial applications. However, MPC requires the model of the system in question. Since an accurate model of the liquid dynamics is difficult to obtain, the usefulness of MPC for the pouring task is uncertain. In this work, we model the dynamics of water using a recurrent neural network (RNN), which enables the use of MPC for pouring control. We evaluated our RNN-enabled MPC controller using a physical system we made ourselves and averaged a pouring error of 16.4 mL over 5 different source containers. We also compared our controller with a baseline switch controller and showed that our controller achieved a much higher accuracy than the baseline controller. Tianze Chen, Yongqiang Huang 0001, Yu Sun 0004 |
IROS | 3 |
| 2019 | Manipulation Motion Taxonomy and Coding for RobotsabstractThis paper introduces a taxonomy of manipulations as seen especially in cooking for 1) grouping manipulations from the robotics point of view, 2) consolidating aliases and removing ambiguity for motion types, and 3) provide a path to transferring learned manipulations to new unlearned manipulations. Using instructional videos as a reference, we selected a list of common manipulation motions seen in cooking activities grouped into similar motions based on several trajectory and contact attributes. Manipulation codes are then developed based on the taxonomy attributes to represent the manipulation motions. The manipulation taxonomy is then used for comparing motion data in the Daily Interactive Manipulation (DIM) data set to reveal their motion similarities. David Paulius, Yongqiang Huang 0001, Jason Meloncon, Yu Sun 0004 |
IROS | 4 |
| 2019 | Learning State-Dependent, Sensor Measurement Models for LocalizationabstractA robot typically relies on sensor measurements to infer its state and the state of its environment. Unfortunately, sensor measurements are noisy, and the amount of noise can vary with state. The literature provides a collection of methods that estimate and adapt measurement noise over time. However, many methods do not assume that measurement noise is stochastic, or they do not estimate sensor measurement bias and noise based on state. In this paper, we propose a novel method called state-dependent, sensor measurement models(SDSMMs). This method: 1) learns to estimate measurement probability density functions directly from sensor measurements and 2) stochastically estimates an expected measurement (which includes measurement bias) and a measurement noise, both of which are conditioned upon the states of a robot and its environment. Throughout this paper, we discuss how to learn an SDSMM and use it with the Extended Kalman Filter (EKF). We then apply our method to solve an EKF localization problem using a real robot dataset. Our localization results showed that at least one of our proposed methods outperformed a standard EKF in all 15 cases for 2D position error and 10 of 15 cases for 1D orientation error. Our methods had a mean improvement of 39% for position and 15% for orientation. Troi Williams, Yu Sun 0004 |
IROS | 2 |
| 2019 | Multi-Channel Neural Network for Assessing Neonatal Pain from VideosabstractNeonates do not have the ability to either articulate pain or communicate it non-verbally by pointing. The current clinical standard for assessing neonatal pain is intermittent and highly subjective. This discontinuity and subjectivity can lead to inconsistent assessment, and therefore, inadequate treatment. In this paper, we propose a multi-channel deep learning framework for assessing neonatal pain from videos. The proposed framework integrates information from two pain indicators or channels, namely facial expression and body movement, using convolutional neural network (CNN). It also integrates temporal information using a recurrent neural network (LSTM). The experimental results prove the efficiency and superiority of the proposed temporal and multi-channel framework as compared to existing similar methods. Md Sirajus Salekin, Ghada Zamzmi, Dmitry B. Goldgof, Rangachar Kasturi, Thao Ho, Yu Sun 0004 |
SMC | 6 |
| 2019 | Long Activity Video Understanding Using Functional Object-Oriented NetworkabstractVideo understanding is one of the most challenging topics in computer vision. In this paper, a four-stage video understanding pipeline is presented to simultaneously recognize all atomic actions and the single ongoing activity in a video. This pipeline uses objects and motions from the video and a graph-based knowledge representation network as prior reference. Two deep networks are trained to identify objects and motions in each video sequence associated with an action and low level image features are used to identify objects of interest in the video sequence. Confidence scores are assigned to objects of interest to represent their involvement in the action and to motion classes based on results from a deep neural network that classifies an ongoing action in video into motion classes. Confidence scores are computed for each candidate functional unit to associate them with an action using a knowledge representation network, object confidences, and motion confidences. Each action, therefore, is associated with a functional unit, and the sequence of actions is evaluated to identify the sole activity occurring in the video. The knowledge representation used in the pipeline is called the functional object-oriented network, which is a graph-based network useful for encoding knowledge about manipulation tasks. Experiments are performed on a dataset of cooking videos to test the proposed algorithm with action inference and activity classification. Experiments show that using a functional object-oriented network improves video understanding significantly. Ahmad Babaeian Jelodar, David Paulius, Yu Sun 0004 |
IEEE Trans. Multim. | 3 |
| 2018 | Functional Object-Oriented Network: Construction & ExpansionabstractWe build upon the functional object-oriented network (FOON), a structured knowledge representation which is constructed from observations of human activities and manipulations. A FOON can be used for representing object-motion affordances. Knowledge retrieval through graph search allows us to obtain novel manipulation sequences using knowledge spanning across many video sources, hence the novelty in our approach. However, we are limited to the sources collected. To further improve the performance of knowledge retrieval as a follow up to our previous work, we discuss generalizing knowledge to be applied to objects which are similar to what we have in FOON without manually annotating new sources of knowledge. We discuss two means of generalization: 1) expanding our network through the use of object similarity to create new functional units from those we already have, and 2) compressing the functional units by object categories rather than specific objects. We discuss experiments which compare the performance of our knowledge retrieval algorithm with both expansion and compression by categories. David Paulius, Ahmad Babaeian Jelodar, Yu Sun 0004 |
ICRA | 3 |
| 2017 | Learning to pourabstractPouring is a simple task people perform daily. It is the second most frequently executed motion in cooking scenarios, after pick-and-place. We present a pouring trajectory generation approach, which uses force feedback from the cup to determine the future velocity of pouring. The approach uses recurrent neural networks as its building blocks. We collected the pouring demonstrations which we used for training. To test our approach in simulation, we also created and trained a force estimation system. The simulated experiments show that the system is able to generalize to single unseen element of the pouring characteristics. Yongqiang Huang 0001, Yu Sun 0004 |
IROS | 2 |
| 2016 | An approach for automated multimodal analysis of infants' painabstractCurrent practices of assessing infants' pain depends on the observer's subjective and potentially inconsistent judgment and requires continuous monitoring by care providers. Therefore, pain may be misinterpreted or totally missed leading to misdiagnosis and over/under treatment. To address these shortcomings, current practices can be augmented with a machine-based assessment system that monitors various pain cues and provides an objective and continuous assessment of pain. Although several machine-based pain assessment approaches have been introduced, the majority of these approaches assess pain based on analysis of a single pain indicator (i.e., unimodal). In this paper, we propose an automated multimodal approach that utilizes a combination of both behavioral and physiological pain indicators to assess infants' pain. We also present a unimodal approach that depends on a single pain indicator for assessment. Recogsnizing pain using a single indicator yielded 88%, 85%, and 82% overall accuracies for facial expression, body movement, and vital signs, respectively. Combining facial expression, body movement, and changes in vital signs (i.e., the multimodal approach) for assessment achieved 95% overall accuracy. These preliminarily results indicate that utilizing both behavioral and physiological pain indicators could provide a better and more reliable assessment of infants' pain. Ghada Zamzmi, Chih-Yun Pai, Dmitry B. Goldgof, Rangachar Kasturi, Terri Ashmeade, Yu Sun 0004 |
ICPR | 6 |
| 2016 | Functional object-oriented network for manipulation learningabstractThis paper presents a novel structured knowledge representation called the functional object-oriented network (FOON) to model the connectivity of the functional-related objects and their motions in manipulation tasks. The graphical model FOON is learned by observing object state change and human manipulations with the objects. Using a well-trained FOON, robots can decipher a task goal, seek the correct objects at the desired states on which to operate, and generate a sequence of proper manipulation motions. The paper describes FOON's structure and an approach to form a universal FOON with extracted knowledge from online instructional videos. A graph retrieval approach is presented to generate manipulation motion sequences from the FOON to achieve a desired goal, demonstrating the flexibility of FOON in creating a novel and adaptive means of solving a problem using knowledge gathered from multiple sources. The results are demonstrated in a simulated environment to illustrate the motion sequences generated from the FOON to carry out the desired tasks. David Paulius, Yongqiang Huang 0001, Roger Milton, William D. Buchanan, Jeanine Sam, Yu Sun 0004 |
IROS | 6 |
| 2015 | Generating manipulation trajectory using motion harmonicsabstractThis paper presents a novel manipulation trajectory generating algorithm that constructs trajectories from learned motion harmonics and user defined constraints. The algorithm uses functional eigenanalysis to learn motion harmonics from demonstrated motions and then use the motion harmonics to compute the optimal trajectory that resembles the demonstrated motions and also satisfies the constraints. The algorithm has been tested on five real human motion data sets to obtain motion harmonics and then generate motions of each task for a NAO robot. The generated trajectories were compared with the trajectories generated using linear segment with parabolic blend approach and with the Open Motion Planning Library. The approach can also work with motion planners. Yongqiang Huang 0001, Yu Sun 0004 |
IROS | 2 |
| 2015 | Task-based grasp quality measures for grasp synthesisabstractTo facilitate manipulation tasks, grasp should be selected intelligently to fulfill different stability properties and manipulative requirements in the tasks. In this paper, two task-dependent grasp quality measures are introduced: task wrench coverage measure and the manipulator efficiency measure. The first one measures the ability of a grasp to provide required interactive wrench during a task, while the second measures the effort that the manipulator takes for the whole manipulation process in facilitating the required instrument motion, which is determined by the grasp when the motion of the instrument is defined. The proposed measures are then used in selecting grasps for three typical manipulation tasks in simulations and using a real robotic system and produced successful grasp synthesis outcomes that satisfy manipulative requirements. Yun Lin 0003, Yu Sun 0004 |
IROS | 2 |
| 2014 | Grasp planning based on strategy extracted from demonstrationabstractIn this paper, we discuss information that is beneficial to robotic grasp planning and can be extracted from human demonstration. We present a method that integrates grasp intention: grasp type, and the relative thumb positions and orientations on the grasped object to the force-closure-based grasp planning procedure. Instead of completely mimicking the human grasp, grasp type and the relative thumb position are partially extracted from the demonstration to represent the task properties and grasp strategies, and avoid the challenging kinematic correspondence problem. Instead of mapping the demonstrated motion, the grasp type and thumb position provide meaningful constraints on hand posture and wrist position. Both the feasible workspace of a robotic hand and the search space of grasp planning are thereby highly reduced by the constraints. This approach has been evaluated in a simulation with a Barrett hand and a Shadow hand on eight daily objects. Yun Lin 0003, Yu Sun 0004 |
IROS | 2 |
| 2013 | Grasp mapping using locality preserving projections and kNN regressionabstractIn this paper, we propose a novel mapping approach to map a human grasp to a robotic grasp based on human grasp motion trajectories rather than grasp poses, since the grasp trajectories of a human grasp provide more information to disambiguate between different grasp types than grasp poses. Human grasp motions usually contain complex and nonlinear patterns in a high-dimensional space. In this paper, we reduced the high-dimensionality of motion trajectories by using locality preserving projections (LPP). Then, a Hausdorff distance was performed to find the k-nearest neighbor trajectories in the reduced low-dimensional subspace, and k-nearest neighbor (kNN) regression was used to map a demonstrated grasp motion by a human hand to a robotic hand. Several experiments were designed and carried out to compare the robotic grasping trajectory generated with and without the trajectory-based mapping approach. The regression errors of the mapping results show that our approach generates more robust grasps than using only grasp poses. In addition, our approach has the ability to successfully map a grasp motion of a new grasp demonstration that has not been trained before to a robotic hand. Yun Lin 0003, Yu Sun 0004 |
ICRA | 2 |
| 2013 | Functional analysis of grasping motionabstractThis paper presents a novel grasping motion analysis technique based on functional principal component analysis (fPCA). The functional analysis of grasping motion provides an effective representation of grasping motion and emphasizes motion dynamic features that are omitted by classic PCA-based approaches. The proposed approach represents, processes, and compares grasping motion trajectories in a low-dimensional space. An experiment was conducted to record grasping motion trajectories of 15 different grasp types in Cutkosky grasp taxonomy. We implemented our method for the analysis of collected grasping motion in the PCA+fPCA space, which generated a new data-driven taxonomy of the grasp types, and naturally clustered grasping motion into 5 consistent groups across 5 different subjects. The robustness of the grouping was evaluated and confirmed using a tenfold cross validation approach. Yu Sun 0004, Xiaoning Qian |
IROS | 2 |
| 2013 | Task-Oriented Grasp Planning Based on Disturbance Distribution
Yun Lin 0003, Yu Sun 0004 |
ISRR | 2 |
| 2013 | Determining the benefit of human input in human-in-the-loop robotic systemsabstractIn this work, we analyze the pick and place task for a human-in-the-loop robotic system to determine where human input can be most beneficial to a collaborative task. This is accomplished by implementing a pick and place task on a commercial robotic arm system and determining which segments of the task, when replaced by human guidance, provide the most improvement to overall task performance and require the least cognitive effort. The pick and place task can be segmented into two main areas: coarse approach towards goal object and fine pick motion. For the fine picking phase, we look at the importance of user guidance in terms of position and orientation of the end effector. Results from our experiment show that the most successful strategy for our human-in-the-loop system is the one in which the human specifies a general region for grasping, and the robotic system completes the remaining elements of the task. Our experimental setup and procedures could be generalized and used to guide similar analysis of human impact in other human-in-the-loop systems performing other tasks. Christine Bringes, Yun Lin 0003, Yu Sun 0004, Redwan Alqasemi |
RO-MAN | 3 |
| 2013 | Exploration of spatial augmented reality on personabstractSpatial Augmented Reality (SAR) allows users to collaborate without need for see-through screens or head-mounted displays. We explore natural on-person interfaces using SAR. Spatial Augmented Reality on Person (SARP) leverages self-based psychological effects such as Self-Referential Encoding (SRE) and ownership by intertwining augmented body interactions with the self. Applications based on SARP could provide powerful tools in education, health awareness, and medical visualization. The goal of this paper is to explore benefits and limitations of generating ownership and SRE using the SARP technique. We implement a hardware platform which provides a Spatial Augmented Game Environment to allow SARP experimentation. We test a STEM educational game entitled `Augmented Anatomy' designed for our proposed platform with experts and a student population in US and China. Results indicate that learning of anatomy on-self does appear correlated with increased interest in STEM and is rated more engaging, effective and fun than textbook-only teaching of anatomical structures. Adrian S. Johnson, Yu Sun 0004 |
VR | 2 |
| 2012 | Exploration of intention expression for robotsabstractThis paper presents a novel exploration on how to enable a robot to express its intention so that the humans and robot can form a synergic relationship. A systematic design approach is proposed to obtain a set of possible intentions for a given robot from three levels of intentions. A visual intention expression system approach is developed to visualize the intentions and implemented on a mobile robot and a manipulator to demonstrate the intention expression concept. Ivan Shindev, Yu Sun 0004, Michael D. Coovert, Jenny Pavlova, Tiffany Lee 0001 |
HRI | 2 |
| 2012 | MARVEL: A wireless Miniature Anchored Robotic Videoscope for Expedited LaparoscopyabstractThis paper describes the design and implementation of a Miniature Anchored Robotic Videoscope for Expedited Laparoscopy (MARVEL) and Camera Module (CM) that features wireless communications and control. The CM decreases the surgical-tool bottleneck experienced by surgeons in state-of-the art Laparoscopic Endoscopic Single-Site (LESS) procedures for minimally invasive abdominal surgery. The system includes: (1) a near-zero latency video wireless communications link, (2) a pan/tilt camera platform, actuated by two motors that provides surgeons a full hemisphere field of view inside the abdominal cavity, (3) a small wireless camera, (4) a wireless illumination control system, and (5) a wireless human-machine interface (HMI) to control the CM. An in-vivo experiment on a porcine subject was carried out to test the performance of the system. The robotic design is a Research Platform for a broad range of experiments in a range of domains for faculty and students in the Colleges of Engineering and Medicine and at Tampa General Hospital. This research is the first step in developing semi-autonomous wirelessly controlled and networked laparoscopic devices to enable a paradigm shift in minimally invasive surgery and other domains such as Wireless Body Area Networks. Cristian A. Castro, Sara Smith, Adham Alqassis, Thomas Ketterl, Yu Sun 0004, Sharona Ross, Alexander Rosemurgy, Peter P. Savage, Richard D. Gitlin |
ICRA | 5 |
| 2012 | Learning grasping force from demonstrationabstractThis paper presents a novel force learning framework to learn fingertip force for a grasping and manipulation process from a human teacher with a force imaging approach. A demonstration station is designed to measure fingertip force without attaching force sensor on fingertips or objects so that this approach can be used with daily living objects. A Gaussian Mixture Model (GMM) based machine learning approach is applied on the fingertip force and position to obtain the motion and force model. Then a force and motion trajectory is generated with Gaussian Mixture Regression (GMR) from the learning result. The force and motion trajectory is applied to a robotic arm and hand to carry out a grasping and manipulation task. An experiment was designed and carried out to verify the learning framework by teaching a Fanuc robotic arm and a BarrettHand a pick-and-place task with demonstration. Experimental results show that the robot applied proper motions and forces in the pick-and-place task from the learned model. Yun Lin 0003, Shaogang Ren, Matthew Clevenger, Yu Sun 0004 |
ICRA | 4 |
| 2012 | Visual servoing control of a 9-DoF WMRA to perform ADL tasksabstractThe wheelchair-mounted robotic arm (WMRA) is a mobile manipulator that consists of a 7-DoF robotic arm and a 2-DoF power wheelchair platform. Previous works combined mobility and manipulation control using weighted optimization for dual-trajectory tracking [7]. In this work, we present an image-based visual servoing (IBVS) approach with scale-invariant feature transform (SIFT) using an eye-in-hand monocular camera for combined control of mobility and manipulation for the 9-DoF WMRA system to execute activities of daily living (ADL) autonomously. We also present results of the physical implementation with a simple “Go to and Pick Up” task and the “Go to and Open the Door” task previously published in simulation, using IBVS to aid the task performance. William G. Pence, Fabian Farelo, Redwan Alqasemi, Yu Sun 0004, Rajiv V. Dubey |
ICRA | 4 |
| 2011 | 5-D force control system for fingernail imaging calibrationabstractThis paper presents a low-cost automated system that is able to apply a 5-degree-of-freedom (DOF) force on a human fingertip with high precision. It is designed to be used as a calibration platform for the previous proposed fingernail imaging system, and as a haptic system. The system is composed of two Novint Falcon devices linked by two universal joints and a rigid bar to provide 5-DOF motion and force, with feedback from a 6-DOF force sensor. A force controller is designed with an inner position control to meet the calibration goal and requirement. Experiment result and analysis showed that the system was capable of controlling the force with a settling time of less than 0.25 seconds. Two force trajectories are designed for fast and sufficient calibrations. A calibration experiments demonstrated that the system tracked the trajectories with an interval of 0.3 seconds, and step sizes of 0.1 N and 1 N·mm with root-mean-squared errors of 0.02 – 0.04 N for forces and 0.39 N·mm for torque. Yun Lin 0003, Yu Sun 0004 |
ICRA | 2 |
| 2011 | Fingertip force and contact position and orientation sensorabstractThis paper presents a novel integrated system that is composed of a fingerprint sensor and a force sensor to measure contact position and orientation of the fingertip along with the contact force. The system uses fingerprints from the fingerprint sensor to identify the contact position and orientation with fingerprint features such as core point and ridge orientations. The contact position and orientation are represented in a fingerpad coordinate system for grasping studies. An experiment has been designed to evaluate the proposed system in terms of accuracy and resolution with three subjects. The proposed system can be used in human grasping studies to characterize the fingerpad contact. Yu Sun 0004 |
ICRA | 1 |
| 2009 | Estimation of Fingertip Force Direction With Computer VisionabstractThis paper presents a method of imaging the coloration pattern in the fingernail and surrounding skin to infer fingertip force direction (which includes four major shear-force directions plus normal force) during planar contact. Nail images from 15 subjects were registered to reference images with random sample consensus (RANSAC) and then warped to an atlas with elastic registration. With linear discriminant analysis, common linear features corresponding to force directions, but irrelevant to subjects, are automatically extracted. The common feature regions in the fingernail and surrounding skin are consistent with observation and previous studies. Without any individual calibration, the overall recognition accuracy on test images of 15 subjects was 90%. With individual training, the overall recognition accuracy on test images of 15 subjects was 94%. The lowest imaging resolution, without sacrificing classification accuracy, was found to be between 10-by-10 and 20-by-20 pixels. Yu Sun 0004, John M. Hollerbach, Stephen A. Mascaro |
IEEE Trans. Robotics | 1 |
| 2008 | Observability index selection for robot calibrationabstractThis paper relates 5 observability indexes for robot calibration to the "alphabet optimalities" from the experimental design literature. These 5 observability indexes are shown to be the upper and lower bounds of one another. All observability indexes are proved to be equivalent when the design is optimal after a perfect column scaling. It is shown that when the goal is to minimize the variance of the parameters, D-optimality is the best criterion. When the goal is to minimize the uncertainty of the end-effector position, E-optimality is the best criterion. It is proved that G-optimality is equivalent to E-optimality for exact design. Yu Sun 0004, John M. Hollerbach |
ICRA | 1 |
| 2008 | Active robot calibration algorithmabstractThis paper presents a new updating algorithm to reduce the complexity of computing an observability index for kinematic calibration of robots. An active calibration algorithm is developed to include an updating algorithm in the pose selection process. Simulations on a 6-DOF PUMA robot with 27 unknown parameters shows that the proposed algorithm performs more than 50,000 times better than exhaustive search based on randomly generated designs. Yu Sun 0004, John M. Hollerbach |
ICRA | 1 |
| 2007 | Imaging the Finger Force DirectionabstractThis paper presents a method of imaging the coloration pattern in the fingernail and surrounding skin to infer fingertip force direction during planar contact. Nail images from 7 subjects were registered to reference images with RANSAC and then warped to an atlas with elastic registration. Recognition of fingertip force direction, based on linear discriminant analysis, shows that there are common color pattern features in the fingernail and surrounding skin for different subjects. Based on the common features, the overall recognition accuracy is 92%. Yu Sun 0004, John M. Hollerbach, Stephen A. Mascaro |
CVPR | 1 |
| 2007 | EigenNail for Finger Force Direction RecognitionabstractThis paper presents a technique termed eigennails to classify fingertip force during contact based on the coloration patterns in the fingernail and surrounding skin. Fingertip force is classified into six directions: no force, normal force only, two directions (left/right) of lateral shear force, and two directions (forward/backward) of longitudinal shear forces. Based on the face recognition technique eigenfaces, a small number of eigennails are sufficient to express the color pattern features for shear force direction classification. Results show that 98% of 960 fingernail images of 8 different subjects are correctly classified. The lowest imaging resolution without sacrificing classification accuracy is found to be 10-by-10. Yu Sun 0004, John M. Hollerbach, Stephen A. Mascaro |
ICRA | 1 |
| 2006 | Dynamic Features and Prediction Model for Imaging the Fingernail to Measure Fingertip ForcesabstractAs an extension of our previous work on estimating fingertip forces by imaging the fingernail (Y. Sun, et al., 2006), the dynamic features of the coloration response of different parts of the fingernail and surrounding skin to different force levels are studied. The effect of the cardiovascular state on measurable coloration is also characterized. The accuracy of normal force estimated by generalized least squares is presented. A time compensation method for fast force estimation is presented, based on a knowledge of the time constants from the dynamic response of individual fingernail regions Yu Sun 0004, John M. Hollerbach, Stephen A. Mascaro |
ICRA | 1 |