Ying Yin 0001

dblp:96/3627-1 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0002-6798-9293ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 GPU-Powered Evolutionary Auxiliary Multitasking for Fast SNP Interaction Detection
abstract
Identifying complex interactions among millions of single nucleotide polymorphisms (SNPs) is a key challenge in Genome-Wide Association Studies (GWAS), offering crucial insights into the genetic architecture of complex diseases. Evolutionary algorithm (EA)-based methods have gained significant attention for their global search capabilities, controllable runtime, and multi-objective optimization potential. However, when applied to high-dimensional GWAS datasets, many existing EA-based methods encounter challenges such as getting trapped in local optima and facing high computational demands. To address these issues, the evolutionary multitasking (EMT) paradigm presents a promising solution, enhancing population diversity and convergence speed through collaborative, cross-task knowledge sharing. Furthermore, the multi-tasking framework and EA can be seamlessly deployed across multiple Graphics Processing Units (GPUs), leveraging their high parallelism and aggregated memory bandwidth. Therefore, we introduce a GPU-powered evolutionary auxiliary multitasking algorithm (GEAMT) for fast SNP interaction detection. GEAMT first constructs a main task along with several low-dimensional auxiliary tasks to redefine the original task. The main task explores the entire search space, while the auxiliary tasks search distinct subspaces to enhance local optimization capabilities. In each iteration, the auxiliary tasks transfer high-quality information to the main task via an information transfer mechanism. Subsequently, an auxiliary task update strategy based on feature regrouping is employed to switch the search subspaces of the auxiliary tasks. The final results are derived from the Pareto-optimal solutions of the main task. Implemented across multiple GPUs, GEAMT achieves notable scalability and efficiency. Comprehensive experiments on both synthetic and real-world datasets demonstrate that GEAMT can significantly enhance search accuracy and speed up the search process.
Ying Yin 0001, Xin Wang 0124, Changyong Yu, Yuhai Zhao
IEEE Trans. Comput. Biol. Bioinform.2
2024 Multi-graph learning-based software defect location
abstract
Abstract Software quality is key to the success of software systems. Modern software systems are growing in their worth based on industry needs and becoming more complex, which inevitably increases the possibility of more defects in software systems. Software repairing is time‐consuming, especially locating the source files related to specific software defect reports. To locate defective source files more quickly and accurately, automated software defect location technology is generated and has a huge application value. The existing deep learning‐based software defect location method focuses on extracting the semantic correlation between the source file and the corresponding defect reports. However, the extensive code structure information contained in the source files is ignored. To this end, we propose a software defect location method, namely, multi‐graph learning‐based software defect location (MGSDL). By extracting the program dependency graphs for functions, each source file is converted into a graph bag containing multiple graphs (i.e., multi‐graph). Further, a multi‐graph learning method is proposed, which learns code structure information from multi‐graph to establish the semantic association between source files and software defect reports. Experiments' results on four publicly available datasets, AspectJ, Tomcat, Eclipse UI, and SWT, show that MGSDL improves on average 3.88%, 5.66%, 13.23%, 9.47%, and 3.26% over the competitive methods in five evaluation metrics, rank@10, rank@5, MRR, MAP, and AUC, respectively.
Ying Yin 0001, Yucen Shi, Yuhai Zhao, Fazal Wahab
J. Softw. Evol. Process.1
2022 How to better utilize code graphs in semantic code search?
abstract
Semantic code search greatly facilitates software reuse, which enables users to find code snippets highly matching user-specified natural language queries. Due to the rich expressive power of code graphs (e.g., control-flow graph and program dependency graph), both of the two mainstream research works (i.e., multi-modal models and pre-trained models) have attempted to incorporate code graphs for code modelling. However, they still have some limitations: First, there is still much room for improvement in terms of search effectiveness. Second, they have not fully considered the unique features of code graphs.
Yucen Shi, Ying Yin 0001, Zhengkui Wang, David Lo 0001, Tao Zhang 0001, Xin Xia 0001, Yuhai Zhao
ESEC/SIGSOFT FSE2
2022 Detecting Disease-Associated SNP-SNP Interactions Using Progressive Screening Memetic Algorithm
abstract
Hundreds of thousands of single nucleotide polymorphisms (SNPs)are currently available for genome-wide association study (GWAS). Detecting disease-associated SNP-SNP interactions is considered an important way to capture the underlying genetic causes of complex diseases. In the combinatorially explosive search space, evolutionary algorithms are promising in solving this difficult problem because of their controllable time complexity. However, in existing evolutionary algorithms, some possible SNP-SNP interactions are evaluated multiple times by the fitness function. Such reevaluations not only waste computing resources but also make these algorithms easy to fall into local optima. To tackle this drawback, a progressive screening memetic algorithm (PSMA)is proposed in the paper. PSMA first represents all possible SNP-SNP interactions in a constructed graph. Then, the proposed algorithm uses the progressive screening strategy to guarantee that every possible SNP-SNP interaction can only be evaluated once by reducing the constructed graph. Furthermore, two types of local search algorithms are introduced to enhance the detecting power of PSMA. For detecting disease-associated SNP-SNP interactions, experimental results show that our proposed method outperforms other existing state-of-the-art methods in terms of accuracy and time.
Boxin Guan, Yuhai Zhao, Ying Yin 0001, Yuan Li 0008
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 A differential evolution based feature combination selection algorithm for high-dimensional data
Boxin Guan, Yuhai Zhao, Ying Yin 0001, Yuan Li 0008
Inf. Sci.3
2021 Multi-objective evolutionary clustering for large-scale dynamic community detection
Ying Yin 0001, Yuhai Zhao, Xiangjun Dong 0001
Inf. Sci.1
2020 Effective and Efficient Dense Subgraph Query in Large-Scale Social Internet of Things
abstract
Social Internet of Things (SIoT) is the integration of social network (SN) and the Internet of Things (IOT). Community search in SIOT is an important problem beneficial to the resource/service discovery. In this article, we address the problem from the perspective of a dense subgraph query. Specifically, we propose a core-based static dense subgraph query and a graph kernel based dynamic dense subgraph query. The two algorithms consider the large scale and the time-varying nature of the SIoT, respectively. Unlike the existing works, the static method is inspired by the first-connection-last-expansion idea. Top-k neighbors of each query node are first found by a random walk. Then, all the query nodes and their top-k neighbors are connected as the core using Steiner tree expansion. The rank constraint random sampling is utilized to extend the core to a dense subgraph. Further, by leveraging a graph kernel index and identifying the updates that may affect the results, we conduct the dynamic query in an incremental update way instead of executing it from scratch. The experiments on synthetic and real datasets show that the proposed algorithms are both effective and efficient.
Yuhai Zhao, Xiangjun Dong 0001, Ying Yin 0001
IEEE Trans. Ind. Informatics3
2019 Ant Colony Optimization with Self-Evolving Parameter for Detecting Epistatic Interactions
abstract
The epistatic interactions of single nucleotide poly-morphisms (SNPs) are fundamentally important for understanding the genetic causes of complex diseases. Due to the intensive computational burden and the diversity of disease models, existing methods suffer from low detection power, high computational cost, and preferences for some types of disease models. To tackle these drawbacks, an ant colony optimization with self-evolving parameter (SEPACO) is proposed in the paper. In the proposed algorithm, the self-evolving parameter control (SEPC) strategy is used to select the best parameters of the algorithm during the running process. In this way, SEPACO can set different optimal parameters for different disease models, which leads to the enhancement of the detection ability of the algorithm. Furthermore, the probability distribution function and the pheromone evaporation formula are improved to adapt SEPACO to the detection of epistatic interactions. SEPACO is compared with other recent algorithms on a variety of simulated datasets and a real biological dataset. The experimental results show that our algorithm, compared to the other test algorithms, can improve the average detection power from no more than 32% up to 68%. Moreover, our algorithm uses less running time.
Boxin Guan, Yuzhai Zhao, Yuan Li 0008, Ying Yin 0001
BIBM4
2017 Enhancing ELM by Markov Boundary based feature selection
Ying Yin 0001, Yuhai Zhao, Bin Zhang 0001
Neurocomputing1
2016 An Efficient and Effective Overlapping Communities Discovery Based on Agglomerative Graph
abstract
Community discovery is a popular way to solve the personal service recommendation problem and has recently attracted more and more attentions of the researchers. The communities are often practically overlapping with each other, thus more and more research focus on the problem of overlapping communities detection. A common drawback of the existing algorithms to this problem is the low efficiency when dealing the large scale network. In this paper, we propose a graph compression based overlapping communities discovery algorithm, which greatly enhances the power of handling large networks even using a single computer. First, a graph compression based social network model, namely agglomerative graph, is introduced, which is a lossless compression to the original network. Then, inspired by the idea of iteration based on the selected seeds, the algorithm expands the selected seeds to the communities by optimizing the proposed community fitness function iteratively. Finally, it merges the communities of high similarity with each other to get the final results. Since the network is lossless compressed, and massive redundant computations are avoided, the results can be exactly obtained in an efficient and effective way. The experiments based on both real and synthetic datasets demonstrate efficiency and effectiveness of the proposal method in detecting overlapping communities over large scale networks.
Ying Yin 0001, Yuhai Zhao, Bin Zhang 0001, Yongming Yan
ICWS1
2016 Correctness Verification of Outsourced Inner Product of Vectors with Error Localization
abstract
Security challenges are the vital issues to be addressed in the research arena of cloud computing, among which the data security plays an important role. The technique of correctness verification enables the cloud user to check whether or not the returned results by the cloud service provider are correct. Inner product of vectors, i.e. weighted sum, is studied and applied in wide area, such as similar document detection, face recognition, etc. In this paper, we aim at the goal of result correctness verification with error localization of the outsourced inner products of vectors in the cloud. Our proposed scheme can help the cloud user to verify the correctness of the results. Further more, the cloud user can localize which result(s) is(are) incorrect if the results are checked to be incorrect. The efficiency analysis show the running efficiency of the proposed scheme.
Gang Sheng, Hongyan Han, Ying Yin 0001
ISPDC4
2016 MD-VCMatrix: An Efficient Scheme for Publicly Verifiable Computation of Outsourced Matrix Multiplication
Gang Sheng, Chunming Tang 0003, Wei Gao 0007, Ying Yin 0001
NSS4
2016 Improving ELM-based microarray data classification by diversified sequence features selection
Yuhai Zhao, Guoren Wang, Ying Yin 0001, Yuan Li 0008, Zhanghui Wang
Neural Comput. Appl.3
2014 An Incremental Updating Based Fast Phenotype Structure Learning Algorithm
Yuhai Zhao, Ying Yin 0001
ICIC (3)3
2013 An Active Service Reselection Triggering Mechanism
Ying Yin 0001, Tiancheng Zhang 0001, Bin Zhang 0001, Gang Sheng, Yuhai Zhao
APWeb1
2010 Reliable Web Service Selection based on Transactional Risk
Ying Yin 0001, Bin Zhang 0001
SEKE1
2009 A self-healing composite Web service model
abstract
Composite Web services are often long-running, loosely coupled and cross-organizational applications. They always run in a highly dynamic environment. For the applications and environment, advanced transaction support is required to ensure the quality of reliable execution. Towards composite service adaptive mechanism unavailable for lacking transaction support, this paper proposes a self-healing model for Web service reliable execution, which is an integration of flexible compensation service in selection and reselecting in execution. In order to make the composite service healing itself as quickly as possible and minimize the number of reselections, away of mining cascading scope of replacement in advance by considering full multi-relation among transaction Web services is proposed in this paper. Further more, a new comprehensive, objective QoS-driven service replacement model with compensation support is presented, and the self-healing algorithm is proposed. Experiments show that the model guarantees business process reliability.
Ying Yin 0001, Bin Zhang 0001, Yuhai Zhao
APSCC1
2009 A composite web services discovery technique based on community mining
abstract
Community structure has been recognized as an important statistical feature of network systems over the past decade. The web service in SOA system naturally forms into some service community during execution process, within which the links between nodes are very dense, but between which they are quite sparse. These service communities were generated by repeatedly interaction between composite services which accomplish the same task. Mining and analysis web service community will help design SOA system and predict service behavior. This article addresses the problem of how to discovering and quantifying web services community formed by closely interactive web services and gives the composite web service discovery technique. We consider the case where the details usage record is logging by execution engine. We proposed a novel approach which construct web service interactive network (WSIN) from usage log and get community structure by spectrum clustering. Generally, the web services belong to same cluster have strong relative to same task object and we call it web service community. The approach has been implemented in an experience system for web services dynamic composition and discovery, and the experimental results demonstrated the efficiency and effectiveness of the proposed algorithm.
Ying Yin 0001, Mingwei Zhang 0001, Bin Zhang 0001
APSCC2
2007 Identifying Synchronous and Asynchronous Co-regulations from Time Series Gene Expression Data
Ying Yin 0001, Yuhai Zhao, Bin Zhang 0001
PAKDD1
2007 A Novel Approach to Revealing Positive and Negative Co-Regulated Genes
Yuhai Zhao, Guoren Wang, Ying Yin 0001
J. Comput. Sci. Technol.3
2006 Mining Maximal Local Conserved Gene Clusters from Microarray Data
Yuhai Zhao, Guoren Wang, Ying Yin 0001
ADMA3
2006 Mining Positive and Negative Co-regulation Patterns from Microarray Data
abstract
Currently, pattern-based and tendency-based models are very popular for clustering co-regulated genes. In this paper, we propose another novel model, namely g-Cluster. The proposed model has the following advantages: (1) find positive and negative co-regulated genes in a shot, (2) get away from the restriction of magnitude transformation relationship among genes, and (3) guarantee quality of clusters and significance of regulations using a novel similarity measurement gCode and two user-specified thresholds, called wave constraint threshold and regulation threshold respectively. We also design a novel tree-based clustering algorithm, FBTD, combined with efficient pruning rules to identify all maximal g-Clusters. The extensive experiments on real and synthetic datasets show that (1) our algorithm can effectively and efficiently find an amount of co-regulated gene clusters missed by previous models, which are potentially of high biological significance, and (2) our algorithm is superior to the existing approaches
Yuhai Zhao, Guoren Wang, Ying Yin 0001, Ge Yu 0001
BIBE3
2006 WWW Information Integration Oriented Classification Ontology Integrating Approach
Anxiang Ma, Kening Gao, Bin Zhang 0001, Ying Yin 0001
KSEM5