Yuxiang Gao

dblp:203/5807 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GTA: Generating high-performance tensorized program with dual-task scheduling
Anxing Xie, Yonghua Hu, Yuxiang Gao, Zenghua Cheng
J. Syst. Archit.5
2024 An Empirical Study on Model-Agnostic Techniques for Source Code-Based Defect Prediction
abstract
Interpretation is important for adopting software defect prediction in practice. Model-agnostic techniques such as Local Interpretable Model-agnostic Explanation (LIME) can help practitioners understand the factors which contribute to the prediction. They are effective and useful for models constructed on tabular data with traditional features. However, when they are applied on source code-based models, they cannot differentiate the contribution of code tokens in different locations for deep learning-based models with Bag-of-Word features. Besides, only using limited features as explanation may result in information loss about actual riskiness. Such limitations may lead to inaccurate explanation for source code-based models, and make model-agnostic techniques not useful and helpful as expected. Thus, we apply a perturbation-based approach Randomized Input Sampling Explanation (RISE) for source code-based defect prediction. Besides, to fill the gap that there lacks a systematical evaluation on model-agnostic techniques on source code-based defect models, we also conduct an extensive case study on the model-agnostic techniques on both token frequency-based and deep learning-based models. We find that (1) model-agnostic techniques are effective to identify the most important code tokens for an individual prediction and predict defective lines based on the importance scores, (2) using limited features (code tokens) for explanation may result in information loss about actual riskiness, and (3) RISE is more effective than others as it can generate more accurate explanation, achieve better cost-effectiveness for line-level prediction, and result in less information loss about actual riskiness. Based on such findings, we suggest that model-agnostic techniques can be a supplement to file-level source code-based defect models, while such explanations should be used with caution as actual risky tokens may be ignored. Also, compared with LIME, we would recommend RISE for a more effective explanation.
Yi Zhu 0008, Yuxiang Gao, Qiao Yu 0001
Int. J. Softw. Eng. Knowl. Eng.2
2024 Forging Productive Human-Robot Partnerships Through Task Training
abstract
Productive human-robot partnerships are vital to successful integration of assistive robots into everyday life. Although prior research has explored techniques to facilitate collaboration during human-robot interaction, the work described here aims to forge productive partnerships prior to human-robot interaction, drawing upon team-building activities’ aid in establishing effective human teams. Through a 2 (group membership: ingroup and outgroup) ×3 (robot error: main task errors, side task errors, and no errors) online study ( N=62 ), we demonstrate that (1) a non-social pre-task exercise can help form ingroup relationships; (2) an ingroup robot is perceived as a better, more committed teammate than an outgroup robot (despite the two behaving identically); and (3) participants are more tolerant of negative outcomes when working with an ingroup robot. We discuss how pre-task exercises may serve as an active task failure mitigation strategy.
Maia Stiber, Yuxiang Gao, Russell H. Taylor, Chien-Ming Huang 0001
ACM Trans. Hum. Robot Interact.2
2023 Typical stochastic resonance models and their applications in steady-state visual evoked potential detection technology
Ruiquan Chen, Guanghua Xu 0001, Jinju Pei, Yuxiang Gao, Sicong Zhang, Chengcheng Han 0001
Expert Syst. Appl.4
2023 Driving Style Feature Extraction and Recognition Based on Hyperdimensional Computing and Semi-Supervised Twin Projection Vector Machine
abstract
Driving style recognition is one of the most crucial requirements for human-centric autonomous and assistive driving systems. Existing studies either require a large amount of labeled data or suffer from hand-crafted features, thus hindering the practical application of these methods. To address the above challenges, this paper proposes a driving style recognition approach from the perspective of brain-inspired hyperdimensional computing and semi-supervised learning. Specifically, raw sensor signals are treated as multivariate time series which are encoded by hyperdimensional computing, thus generating a large holistic feature representation. Then, considering the high dimension of the resulting representation, we put forward a semi-supervised twin projection vector machine (SSTPVM) model that can take full advantage of unlabeled data and jointly optimize multiple projection directions tailored for dimension reduction. In addition, a heuristic method based on particle swarm optimization is developed for the parameter selection of SSTPVM. Finally, extensive experiment comparisons with other related methods are performed on naturalistic driving style data. The results show that hyperdimensional computing is rather suitable for semi-supervised learning while our proposed SSTPVM can significantly improve recognition performance even with a small portion of labeled data.
Xiaobo Chen 0001, Yuxiang Gao, Haoze Yu, Hai Wang 0003, Yingfeng Cai
IEEE Trans. Intell. Transp. Syst.2
2022 Learning a Group-Aware Policy for Robot Navigation
abstract
Human-aware robot navigation promises a range of applications in which mobile robots bring versatile assistance to people in common human environments. While prior research has mostly focused on modeling pedestrians as independent, intentional individuals, people move in groups; consequently, it is imperative for mobile robots to respect human groups when navigating around people. This paper explores learning group-aware navigation policies based on dynamic group formation using deep reinforcement learning. Through simulation experiments, we show that group-aware policies, compared to baseline policies that neglect human groups, achieve greater robot navigation performance (e.g., fewer collisions), minimize violation of social norms and discomfort, and reduce the robot's movement impact on pedestrians. Our results contribute to the development of social navigation and the integration of mobile robots into human environments.
Kapil D. Katyal, Yuxiang Gao, Jared Markowitz, Sara Pohland, Corban G. Rivera, I-Jeng Wang, Chien-Ming Huang 0001
IROS2
2022 Evaluating the effectiveness of local explanation methods on source code-based defect prediction models
abstract
Interpretation has been considered as one of key factors for applying defect prediction in practice. As one way for interpretation, local explanation methods has been widely used for certain predictions on datasets of traditional features. There are also attempts to use local explanation methods on source code-based defect prediction models, but unfortunately, it will get poor results. Since it is unclear how effective those local explanation methods are, we evaluate such methods with automatic metrics which focus on local faithfulness and explanation precision. Based on the results of experiments, we find that the effectiveness of local explanation methods depends on the adopted defect prediction models. They are effective on token frequency-based models, while they may not be effective enough to explain all predictions of deep learning-based models. Besides, we also find that the hyperparameter of local explanation methods should be carefully optimized to get more precise and meaningful explanation.
Yuxiang Gao, Yi Zhu 0008, Qiao Yu 0001
MSR1
2022 Dealing with imbalanced data for interpretable defect prediction
abstract
Context Interpretation has been considered as a key factor to apply defect prediction in practice. As interpretation from rule-based interpretable models can provide insights about past defects with high quality, many prior studies attempt to construct interpretable models for both accurate prediction and comprehensible interpretation. However, class imbalance is usually ignored, which may bring huge negative impact on interpretation. Objective In this paper, we are going to investigate resampling techniques, a popular solution to deal with imbalanced data , on interpretation for interpretable models. We also investigate the feasibility to construct interpretable defect prediction models directly on original data. Further, we are going to propose a rule-based interpretable model which can deal with imbalanced data directly. Method We conduct an empirical study on 47 publicly available datasets to investigate the impact of resampling techniques on rule-based interpretable models and the feasibility to construct such models directly on original data. We also improve gain function and tolerate lower confidence based on rule induction algorithms to deal with imbalanced data. Results We find that (1) resampling techniques impact on interpretable models heavily from both feature importance and model complexity, (2) it is not feasible to construct meaningful interpretable models on original but imbalanced data due to low coverage of defects and poor performance, and (3) our proposed approach is effective to deal with imbalanced data compared with other rule-based models. Conclusion Imbalanced data heavily impacts on the interpretable defect prediction models. Resampling techniques tend to shift the learned concept, while constructing rule-based interpretable models on original data may also be infeasible. Thus, it is necessary to construct rule-based models which can deal with imbalanced data well in further studies.
Yuxiang Gao, Yi Zhu 0008
Inf. Softw. Technol.1
2019 PATI: a projection-based augmented table-top interface for robot programming
abstract
As robots begin to provide daily assistance to individuals in human environments, their end-users, who do not necessarily have substantial technical training or backgrounds in robotics or programming, will ultimately need to program and "re-task" their robots to perform a variety of custom tasks. In this work, we present PATI---a Projection-based Augmented Table-top Interface for robot programming---through which users are able to use simple, common gestures (e.g., pinch gestures) and tools (e.g., shape tools) to specify table-top manipulation tasks (e.g., pick-and-place) for a robot manipulator. PATI allows users to interact with the environment directly when providing task specifications; for example, users can utilize gestures and tools to annotate the environment with task-relevant information, such as specifying target landmarks and selecting objects of interest. We conducted a user study to compare PATI with a state-of-the-art, standard industrial method for end-user robot programming. Our results show that participants needed significantly less training time before they felt confident in using our system than they did for the industrial method. Moreover, participants were able to program a robot manipulator to complete a pick-and-place task significantly faster with PATI. This work indicates a new direction for end-user robot programming.
Yuxiang Gao, Chien-Ming Huang 0001
IUI1
2018 QuantCloud: A Software with Automated Parallel Python for Quantitative Finance Applications
abstract
Quantitative Finance is a field that replies on data analysis and big data enabling software to discover market signals. In this, a decisive factor is the speed that concerns execution speed and software development speed. So, an efficient software plays a key role in helping trading firms. Inspired by this, we present a novel software: QuantCloud to integrate a parallel Python system with a C++-coded Big Data system. C++ is used to implement this big data system and Python is used to code the user methods. The automated parallel execution of Python codes is built upon a coprocess-based parallel strategy. We test our software using two popular algorithms: moving-window and autoregressive moving-average (ARMA). We conduct an extensive comparative study between Intel Xeon E5 and Xeon Phi processors. The results show that our method achieved a nearly linear speedup for executing Python codes in parallel, prefect for today's multicore processors.
Yuxiang Gao
QRS2
2017 Family Relationship Inference Using Knights Landing Platform
abstract
Using genetic data to infer relatedness has been crucial for genetics studies for decades. In a previously published paper together with the KING software, we demonstrated that the kinship coefficient, a measure of relatedness between a pair of individuals, can be accurately estimated using their genome-wide SNP data, without estimating the allele frequencies at each SNP in the whole dataset. The computational efficiency of this algorithm has been substantially improved in the second generation of KING. Three levels of computational speed-up are implemented in KING 2.0, including: 1) bit-level parallelism; 2) multiple-core parallelism using OpenMP; and 3) a multi-stage procedure to eliminate unrelated or distantly related pairs of individuals. The efficient implementation in KING 2.0 allows instant relationship inference in a matter of seconds in a typical dataset (with 10,000s individuals). To demonstrate the computational performance and scalability of KING 2.0, we use the Knights Landing platform to infer relatedness in a dataset consisting of 303,750 individuals each typed at 168,749 autosome SNPs. The computational time to identify all first-degree relatives by scanning 46 billion pairs of individuals is ∼10 minutes using 256 threads, a noticeable speed-up comparing to the general-purpose CPU. Algorithm improvement in the second generation of KING and the use of the latest computing system such as the Knights Landing platform makes it feasible for researchers to infer relatedness in their genetic datasets in the largest size up-to-date on a single computer.
Yuxiang Gao
CSCloud1
2017 Finding the Best Box-Cox Transformation in Big Data with Meta-Model Learning: A Case Study on QCT Developer Cloud
abstract
Finding the best model to reveal potential relationships of a given set of data is not an easy job and often requires many iterations of trial and errors for model sections, feature selections and parameters tuning. This problem is greatly complicated in the big data era where the I/O bottlenecks significantly slowed down the time needed to finding the best model. In this article, we examine the case of Box-Cox transformation when assumptions of a regression model are violated. Specifically, we construct and compute a set of summary statistics and transformed the maximum likelihood computation into a per-role operational fashion. The innovative algorithms reduced the big data machine learning problem into a stream based small data learning problem. Once the Box-Cox information array is obtained, the optimal power transformation as well as the corresponding estimates of model parameters can be quickly computed. To evaluate the performance, we implemented the proposed Box-Cox algorithms on QCT developer cloud. Our results showed that by leveraging both the algorithms and the QCT cloud technology, find the fittest model from 101 potential parameters is much faster than the conventional approach.
Yuxiang Gao, Tonglin Zhang, Baijian Yang 0001
CSCloud1
2017 Evaluation of Combining Bootstrap with Multiple Imputation Using R on Knights Landing Platform
abstract
Cloud computing and big data technologies are converging to offer a cost-effective delivery model for cloud-based big data analytics. Though impacts of size and scaling of big data on cloud have been extensively studied, the effects of complexity of underlying analytic methods on cloud performance have received less attention. This paper will develop and evaluate a computationally intensive statistical methodology to perform inference in the presence of both non-Gaussian data and missing data. Two well-established statistical approaches, bootstrap and multiple imputations (MI), will be combined to form the methodology. Bootstrap is a computer-based nonparametric resampling procedure that involves randomly selecting data many thousands of times to construct an empirical distribution, which is then used to construct confidence intervals for significance tests. This statistical technique enables scientists who conduct studies on data with known non-normality to obtain higher quality significance tests than is possible with a traditional asymptotic, normal-theory based significance test. However, the bootstrapping procedure only works when no data are missing or the data are missing completely at random (MCAR). Missing data can lead to biased estimates when the MCAR assumption is violated. It is unclear how to best implement a bootstrapping procedure in the presence of missing data. The proposed methods will provide guidelines and procedures that will enable researchers to use the technique in all areas of health, behavior and developmental science in which a study has missing data and cannot rely on parametric inference. Either bootstrapping or MI can be computationally expensive, and combining these two can lead to further computation costs in the cloud. Using carefully constructed simulation examples, we demonstrate that it is feasible to implement the proposed methodology in a high performance Knights Landing platform. However, the computation costs are substantial even with small data size. Further studies are needed to study the effects of optimizing the implementation and its performance with big data.
Chuan Zhou 0005, Yuxiang Gao, Waylon Howard
CSCloud2