VLDB 2026 Research / reviewers in the wild / expert
Yutao Ma
dblp:07/4142
· DBLP profile ↗
54ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0003-4239-2009ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 26 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring and characterizing cross-service defects in microservice projects
Chaochao Wu, Ran Mo, Haopeng Song, Zengyang Li, Yutao Ma |
Inf. Softw. Technol. | 6 |
| 2026 | Unveiling code clones in the Eclipse IIoT software ecosystem
Zengyang Li, Binbin Huang 0005, Ran Mo, Peng Liang 0001, Hui Liu 0004, Yutao Ma |
J. Syst. Softw. | 7 |
| 2026 | Standardized evaluation of automatic methods for perivascular spaces segmentation in MRI - MICCAI 2024 challenge resultsabstractPerivascular spaces (PVS), when abnormally enlarged and visible in magnetic resonance imaging (MRI) structural sequences, are important imaging markers of cerebral small vessel disease and potential indicators of neurodegenerative conditions. Despite their clinical significance, automatic enlarged PVS (EPVS) segmentation remains challenging due to their small size, variable morphology, similarity with other pathological features, and limited annotated datasets. This paper presents the EPVS Challenge organized at MICCAI 2024, which aims to advance the development of automated algorithms for EPVS segmentation across multi-site data. We provided a diverse dataset comprising 100 training, 50 validation, and 50 testing scans collected from multiple international sites (UK, Singapore, and China) with varying MRI protocols and demographics. All annotations followed the STRIVE protocol to ensure standardized ground truth and covered the full brain parenchyma. Seven teams completed the full challenge, implementing various deep learning approaches primarily based on U-Net architectures with innovations in multi-modal processing, ensemble strategies, and transformer-based components. Performance was evaluated using dice similarity coefficient, absolute volume difference, recall, and precision metrics. The winning method employed MedNeXt architecture with a dual 2D/3D strategy for handling varying slice thicknesses. The top solutions showed relatively good performance on test data from seen datasets, but significant degradation of performance was observed on the previously unseen Shanghai cohort, highlighting cross-site generalization challenges due to domain shift. This challenge establishes an important benchmark for EPVS segmentation methods and underscores the need for the continued development of robust algorithms that can generalize in diverse clinical settings. Yilei Wu, Zijian Dong 0001, An Sen Tan, Gifford Tan, Sizhao Tang, Huijuan Chen, Zijiao Chen, Eric Kwun Kei Ng, José Bernal, Hang Min, Ines Vati, Liz Cooper, Yuchen Pei, Yutao Ma, Victor Nozais, Ami Tsuchida, Pierre-Yves Hervé, Philippe Boutinaud, Marc Joliot, Junghwa Kang, Wooseung Kim, Dayeon Bak, Rachika E. Hamadache, Valeriia Abramova, Xavier Lladó, Yuntao Zhu, Zhenyu Gong, John McFadden, Pek Lan Khong, Roberto Duarte Coello, Hongwei Li 0004, Woon Puay Koh, Christopher Chen, Joanna M. Wardlaw, Maria del C. Valdés Hernández, Juan Helen Zhou |
Medical Image Anal. | 18 |
| 2026 | RGPRec: A RAG-Enhanced GNN for Personalized Task Recommendations in Open-Source CommunitiesabstractABSTRACT Context Open‐source communities have become a crucial driver of technological innovation and developer growth. While numerous deep learning‐based recommender systems exist, they often fail to provide accurate recommendations for identifying suitable developers who meet the task requirements of project development due to project popularity imbalance and developer ability bias. Recent advances in retrieval‐augmented generation (RAG) have shown promise in enhancing data augmentation and personalization for recommender systems in scenarios with data sparsity. Objective This study aims to develop RGPRec, a RAG‐enhanced graph neural network (GNN) for personalized, debiased task recommendations in open‐source communities. The model targets task recommendation quality, fairness, and coverage capabilities for long‐tail developers. Methods RGPRec leverages RAG‐enhanced large language models (LLMs) to retrieve project‐related documents and enriched developer attributes to improve node representations through ego graph structural learning. It also recognizes the importance of long‐tail developers—less involved in projects but still contributing significantly—often overlooked by traditional recommender systems. Moreover, RGPRec reasonably evaluates each developer's ability through a multilabel assessment based on an ego GNN to calculate a debiased rating, generating more personalized and debiased task recommendations. Results Evaluation using real‐world data from two famous open‐source communities of SourceForge and GitHub, indicates that RGPRec significantly outperforms state‐of‐the‐art (SOTA) approaches in rating and ranking performances. In addition, ablation studies demonstrate the necessity of each component in the RGPRec model. Conclusion RGPRec effectively enhances task recommendation quality in open‐source communities. By integrating RAG and LLMs with GNNs, RGPRec achieves superior data augmentation and personalization compared to traditional approaches. Therefore, RGPRec can be a valuable tool for promoting inclusiveness and efficiency in collaborations within open‐source communities. Shiyu He, Yuqi Zhao 0001, Qibo Li, Yutao Ma |
Softw. Pract. Exp. | 4 |
| 2025 | Joint Structure-Texture Representation Learning for Generalizable Cervical OCT Diagnosis Across Medical CentersabstractOptical coherence tomography (OCT) has emerged as a clinically viable tool for detecting early cervical diseases due to its non-invasive, high-resolution imaging capabilities. While computer-aided OCT diagnosis systems have shown promising potential, they confront three primary challenges: limited annotated data, class imbalance, and cross-center bias. These factors collectively compromise model generalizability across medical centers. Therefore, we propose a joint structure-texture representation learning framework to enhance the generalization of cervical OCT diagnosis in multi-center, small-sample scenarios by leveraging histomorphological invariance. The framework synergizes hierarchical structure features from tissue segmentation with histomorphology-enhanced texture representations learned through lesion-focused contrastive reconstruction. We design a semi-supervised segmentation strategy that achieves promising segmentation performance for layered tissue architecture with minimal annotation burden. An attention-based feature fusion module dynamically combines complementary structural and textural features, generating enriched representations optimized for classification robustness. External cross-center validations demonstrate the superior generalization performance and interpretability of our approach over competitive baseline methods. Moreover, low model complexity and high inference efficiency make our approach well-suited for adoption in low-resource clinical settings. Yutao Ma, Yuchen Pei |
BIBM | 1 |
| 2025 | Evidential Deep Learning with Reweighted Margin Adjustment for Uncertainty-Driven Cervical OCT Image DiagnosisabstractCervical optical coherence tomography (OCT) is crucial in diagnosing cervical diseases due to its high-resolution imaging and non-invasive characteristics. However, recent improvements in classifying cervical OCT images have mainly focused on accuracy, often neglecting the importance of uncertainty and robustness in disease predictions. Uncertainty estimation faces challenges in collecting evidence for the minority classes, leading to fake confidence in evidential deep learning. There are similar issues with cervical OCT image classification. This study proposes a novel evidential learning process for cervical disease screening using OCT. It provides a measure of decision margin for each class with two novel mechanisms to address the hinge posted by class imbalance. Specifically, we introduce a reweighted margin adjustment strategy to learn less biased evidence among classes and make reliable uncertainty estimations. Extensive experiments on internal and external cervical OCT datasets demonstrate the proposed method outperforms competitive baseline methods, providing more reliable and robust predictions, particularly in out-of-distributions scenarios. By improving uncertainty estimation for class imbalance, our method offers valuable insights for practical clinical applications, highlighting the importance of measuring uncertainty in medical AI-aided diagnosis systems. Hanfeng Zhu, Yutao Ma, Yuchen Pei |
ICASSP | 3 |
| 2025 | TDIC: Time-Aware Disentanglement of Interest and Conformity in Mobile App RecommendationsabstractPopularity bias in mobile APP recommendations skews results by over-prioritizing trending APPs, obscuring niche yet relevant alternatives, and conflating user interest with social conformity. To address this problem, we propose TDIC (short for Time-aware Disentanglement of Interest and Conformity), a novel causal framework tailored for personalized, debiased mobile APP suggestions. TDIC employs causal graph analysis to isolate user interest from conformity within interaction patterns, incorporates item quality to refine popularity judgments, and integrates temporal awareness to track the fluid nature of trends. TDIC disentangles authentic user intents from dynamic social influences, fostering unbiased and precise recommendations. Evaluated on two real-world datasets, MobileRec and Myket, TDIC demonstrates superior performance over state-of-the-art baselines, achieving gains of up to 17.24% in Recall@20 and 7.14% in NDCG@20 on MobileRec, and 2.55% in NDCG@50 on Myket. These outcomes highlight the crucial role of temporal dynamics and quality adjustments in mitigating popularity bias for mobile APP recommendations. Our code is available at https://github.com/ssea-lab/TDIC. Qibo Li, Yuqi Zhao 0001, Yutao Ma |
ICWS | 3 |
| 2025 | Activation Map-based Knowledge Distillation for Real-time Cervical OCT Image ClassificationabstractCervical cancer is a significant global health concern for women. Optical coherence tomography (OCT) offers a non-invasive, high-resolution imaging method for cervical examinations. The clinical need for real-time AI-aided diagnosis in low-resource settings necessitates model compression for deep learning models. Therefore, we develop a novel activation map-based knowledge distillation (AMKD) framework to address this issue. The AMKD framework ensures that the compressed (student) model achieves classification performance approximating that of the teacher model while maintaining the same activation area, enhancing interpretability for gynecologists. Due to the lack of high-quality annotated data, we also utilize unlabeled images for self-supervised distillation pre-training to improve model performance. On an internal dataset, the compressed models based on ResNet, ConvNeXt, and Swin-Transformer outperformed existing knowledge distillation frameworks while demonstrating better interpretability. On two external validation sets, the best compressed (ResNet) model surpassed the average performance of four medical experts in sensitivity and negative predictive value for cervical OCT volume classification, showcasing enhanced lesion detection capabilities. For a cervical OCT image of 760 × 1200 pixels, the compressed ResNet model with 1.37 M parameters uses 1.13 GB GPU memory during inference and predicts the input image in 0.06 seconds, potentially improving real-time detection of cervical lesions in resource-limited clinical settings. Qingbin Wang, Yuchen Pei, Wai Chon Wong, Xuefeng Mu, Yan Zhang 0123, Yutao Ma |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | Leveraging Modular Architecture for Bug Characterization and Analysis in Automated Driving SoftwareabstractWith the rapid advancement of automated driving technology, numerous manufacturers deploy vehicles with auto-driving features. This highlights the importance of ensuring the quality of automated driving software. To achieve this, characterizing bugs in automated driving software is important, as it can facilitate bug detection and bug fixes, thereby ensuring software quality. Automated driving software typically has a modular architecture, where software is divided into multiple modules, each designed for its own functionality for automated driving. This may lead to varying bug characteristics. Additionally, our recent study has shown a correlation between bugs caused by code clones and the functionalities of modules in automated driving software. Hence, we consider the modular structure when analyzing bug characteristics. In this article, we analyze 3,078 bugs from two representative open-source Level-4 automated driving systems, Apollo and Autoware. By analyzing the bug report description, title, and developers’ discussions, we have identified 20 bug symptoms and 17 bug-fixing strategies and analyzed their relationships with the respective modules. Our analysis achieves 12 main findings offering a comprehensive view of bug characteristics in automated driving software. We believe our findings can help developers better understand and manage bugs in automated driving software, thereby improving software quality and reliability. Yingjie Jiang, Ran Mo, Wenjing Zhan, Zengyang Li, Yutao Ma |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | Assessing and Analyzing the Correctness of GitHub Copilot's Code SuggestionsabstractAI programming has become a popular topic in recent years. Code suggestion, with code suggestion being a key capability of AI programming. Copilot, an “AI programmer” that provides code suggestions from natural language descriptions, has been launched by GitHub and OpenAI. By far, Copilot has been widely used by millions of developers. However, little work has systematically evaluated the correctness of Copilot’s suggestions. We conducted an empirical study on all 2,033 LeetCode problems to assess Copilot’s code generation across four mainstream languages: C, Java, JavaScript, and Python. We have found that: (1) 70.0% of problems received at least one correct suggestion, with language-specific rates of 29.7% (C), 57.7% (Java), 54.1% (JavaScript), and 41.0% (Python); (2) correctness decreases as problem difficulty increases, with acceptance rates of 89.3% (easy), 72.1% (medium), and 43.4% (hard); (3) acceptance rates vary across problem domains from 49.5% to 90.1%, while Graph problems challenge C and Python most, and Prefix Sum and Heap challenge Java and JavaScript most; (4) for the incorrect suggestions, we further summarize 17 types of error reasons accounting for their incorrectness and analyzed possible causes for why these errors occur. We believe our study can provide valuable insights into Copilot’s capabilities and limitations. Ran Mo, Wenjing Zhan, Yingjie Jiang, Yepeng Wang, Yuqi Zhao 0001, Zengyang Li, Yutao Ma |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2025 | Toward the Fractal Dimension of ClassesabstractThe fractal property has been regarded as a fundamental property of complex networks, characterizing the self-similarity of a network. Such a property is usually numerically characterized by the fractal dimension metric, and it not only helps the understanding of the relationship between the structure and function of complex networks but also finds a wide range of applications in complex systems. The existing literature shows that class-level software networks (i.e., class dependency networks) are complex networks with the fractal property. However, the fractal property at the feature (i.e., methods and fields) level has never been investigated, although it is useful for measuring class complexity and predicting bugs in classes. Furthermore, existing studies on the fractal property of software systems were all performed on un-weighted software networks and have not been used in any practical quality assurance tasks such as bug prediction. Generally, considering the weights on edges can give us more accurate representations of the software structure and thus help us obtain more accurate results. The illustration of an approach’s practical use can promote its adoption in practice. In this article, we examine the fractal property of classes by proposing a new metric. Specifically, we build a Feature-Level Software Network (FLSN) for each class to represent the methods/fields and their couplings (including coupling frequencies) within the class and propose a new metric, Fractal Dimension for Classes (FDC) , to numerically describe the fractal property of classes using FLSNs, which captures class complexity. We evaluate FDC theoretically against Weyuker’s nine properties, and the results show that FDC adheres to eight of the nine properties. Empirical experiments performed on a set of 12 large open source Java systems show that (i) for most classes (larger than \(96\%\) ), there exists the fractal property in their FLSNs, (ii) FDC is capable of capturing additional aspects of class complexity that have not been addressed by existing complexity metrics, (iii) FDC significantly correlates with both the existing class-level complexity metrics and the number of bugs in classes, and (iv) FDC , when used together with existing class-level complexity metrics, can significantly improve bug prediction in classes in three scenarios (i.e., bug-count , bug-classification , and effort-aware ) of the cross-project context, but in the within-project context, it cannot. Weifeng Pan 0001, Ming Hua 0003, Dae-Kyoo Kim, Zijiang Yang 0006, Yutao Ma |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | Classifying Cervical OCT Images using Masked Autoencoders with VMambaabstractOptical coherence tomography (OCT) emerged as a promising, non-invasive, real-time, high-resolution imaging technology for cervical cancer detection. The scarcity of high-quality annotated cervical OCT images severely hinders the predictive performance of deep learning models. In contrast, self-supervised learning (SSL) methods offer a viable solution to the above challenge. This study aims to develop a cost-effective SSL framework that efficiently classifies high-resolution cervical OCT images to meet the clinical "see-and-diagnose" requirement for cervical lesions. Therefore, we propose COVE, a novel SSL framework with masked autoencoders based on VMamba with linear complexity. To our knowledge, COVE is the first SSL framework using masked image modeling for VMamba to leverage large unlabeled cervical OCT datasets. It has two specific designs: (1) 2D-Selective-Scan for visible patches (VP-SS2D): the VMamba encoder’s core module processes only visible patches for efficient pre-training and addressing the inconsistency between pre-training and fine-tuning; (2) visible patch feature-preserving (VPFP): visible patch features in each decoder block are replaced with encoder-extracted features, decoupling the encoder’s feature extraction from the decoder’s pixel reconstruction tasks. Experimental results showed that COVE outperformed existing SSL frameworks in five-fold cross-validation on a 1,452-subject cervical OCT dataset from a multi-center study and two external validation sets from top Chinese hospitals. Additionally, COVE demonstrated the highest pre-training efficiency, with significantly faster speed than existing SSL frameworks. Qingbin Wang, Yuchen Pei, Jian Wang 0018, Yutao Ma |
BIBM | 4 |
| 2024 | Human-Machine Integration to Enhance Clinical Typing and Fine-Grained Interpretation of Cervical OCT ImagesabstractCervical cancer is the fourth most common cancer among women worldwide, posing a severe threat to female reproductive health. The cervical optical coherence tomography (OCT) technology, known for its high-resolution imaging and non-invasive characteristics, holds significant potential in gynecology applications. Annotating cervical OCT images is based on pathological typing of cervical tissue from matching OCT images with the corresponding pathological slice images. It depends mostly on domain-specific skills and experience, lacking comprehensive, straightforward interpretation of subtypes for cervical OCT images. The accuracy of computer-aided diagnosis for cervical OCT images is always limited, failing to meet gynecologists’ clinical requirements. Therefore, our work aims to establish a fine-grained subtyping method for cervical OCT images to complement high-level pathological categories. The proposed method is implemented in a self-supervised pre-training style to explore OCT image characteristics fully. Unlike data-driven clustering methods, our method constructs a form of anchor points to leverage medical prior knowledge to avoid the mismatch between the discovered rule and existing pathological knowledge. We assessed our method with competitive baseline approaches on internal and external datasets. Besides, we collaborated with medical experts to demonstrate that the image feature patterns discovered are normative and clinically valuable for gynecologists. Mi Yin, Yuchen Pei, Yutao Ma |
BIBM | 3 |
| 2024 | MobileEdgeSim: A Tool for Simulating Microservice-Oriented Mobile Edge ComputingabstractMobile edge computing (MEC) is an emerging computing paradigm receiving growing attention. MEC significantly reduces latency by processing user requests on edge servers rather than cloud centers, making it ideal for real-time applications. However, due to resource limitations and user mobility, microservice requests may fail, especially when users move at high speeds. This paper introduces a new tool, MobileEdgeSim, to simulate microservice-oriented MEC environments. MobileEdgeSim integrates mobility prediction and service composition to enhance the pre-deployment of microservices. To evaluate MobileEdgeSim, we conducted a series of experiments comparing it to several state-of-the-art baseline approaches. We also conducted a user study to evaluate the tool’s effectiveness in real-world scenarios. Our results indicate that MobileEdgeSim significantly improves the success rate of both user requests and responses while reducing resource costs. MobileEdgeSim is available at https://github.com/ssea-lab/MobileEdgesim. Yuqi Zhao 0001, Shiyu He, Qibo Li, Yuchen Pei, Yutao Ma |
Internetware | 5 |
| 2024 | Prior Activation Map Guided Cervical OCT Image Classification
Qingbin Wang, Wai Chon Wong, Mi Yin, Yutao Ma |
MICCAI (3) | 4 |
| 2024 | SciCode: A Research Coding Benchmark Curated by ScientistsabstractSince language models (LMs) now outperform average humans on many challenging tasks, it is becoming increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this by examining LM capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we create a scientist-curated coding benchmark, SciCode. The problems naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems, and it offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. OpenAI o1-preview, the best-performing model among those tested, can solve only 7.7\% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards realizing helpful scientific assistants and sheds light on the building and evaluation of scientific AI in the future. Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Shengyan Liu, Yutao Ma, Kha Trinh, Zihan Wang 0010, Bohao Wu, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanxin Lu, Genglin Liu, Yufeng Du, Tianhua Tao, Ofir Press, Jamie Callan, Eliu A. Huerta, Hao Peng 0009 |
NeurIPS | 13 |
| 2024 | CLHHN: Category-aware Lossless Heterogeneous Hypergraph Neural Network for Session-based RecommendationabstractIn recent years, session-based recommendation (SBR), which seeks to predict the target user’s next click based on anonymous interaction sequences, has drawn increasing interest for its practicality. The key to completing the SBR task is modeling user intent accurately. Due to the popularity of graph neural networks (GNNs), most state-of-the-art (SOTA) SBR approaches attempt to model user intent from the transitions among items in a session with GNNs. Despite their accomplishments, there are still two limitations. First, most existing SBR approaches utilize limited information from short user–item interaction sequences and suffer from the data sparsity problem of session data. Second, most GNN-based SBR approaches describe pairwise relations between items while neglecting complex and high-order data relations. Although some recent studies based on hypergraph neural networks have been proposed to model complex and high-order relations, they usually output unsatisfactory results due to insufficient relation modeling and information loss. To this end, we propose a category-aware lossless heterogeneous hypergraph neural network (CLHHN) in this article to recommend possible items to the target users by leveraging the category of items. More specifically, we convert each category-aware session sequence with repeated user clicks into a lossless heterogeneous hypergraph consisting of item and category nodes as well as three types of hyperedges, each of which can capture specific relations to reflect various user intents. Then, we design an attention-based lossless hypergraph convolutional network to generate sessionwise and multi-granularity intent-aware item representations. Experiments on three real-world datasets indicate that CLHHN can outperform the SOTA models in making a better tradeoff between prediction performance and training efficiency. An ablation study also demonstrates the necessity of CLHHN’s key components. Yutao Ma, Zesheng Wang 0005, Liwei Huang, Jian Wang 0018 |
ACM Trans. Web | 1 |
| 2023 | Cross-Attention Based Multi-Resolution Feature Fusion Model for Self-Supervised Cervical OCT Image ClassificationabstractCervical cancer seriously endangers the health of the female reproductive system and even risks women's life in severe cases. Optical coherence tomography (OCT) is a non-invasive, real-time, high-resolution imaging technology for cervical tissues. However, since the interpretation of cervical OCT images is a knowledge-intensive, time-consuming task, it is tough to acquire a large number of high-quality labeled images quickly, which is a big challenge for supervised learning. In this study, we introduce the vision Transformer (ViT) architecture, which has recently achieved impressive results in natural image analysis, into the classification task of cervical OCT images. Our work aims to develop a computer-aided diagnosis (CADx) approach based on a self-supervised ViT-based model to classify cervical OCT images effectively. We leverage masked autoencoders (MAE) to perform self-supervised pre-training on cervical OCT images, so the proposed classification model has a better transfer learning ability. In the fine-tuning process, the ViT-based classification model extracts multi-scale features from OCT images of different resolutions and fuses them with the cross-attention module. The ten-fold cross-validation results on an OCT image dataset from a multi-center clinical study of 733 patients in China indicate that our model achieved an AUC value of 0.9963 ± 0.0069 with a 95.89 ± 3.30% sensitivity and 98.23 ± 1.36 % specificity, outperforming some state-of-the-art classification models based on Transformers and convolutional neural networks (CNNs) in the binary classification task of detecting high-risk cervical diseases, including high-grade squamous intraepithelial lesion (HSIL) and cervical cancer. Furthermore, our model with the cross-shaped voting strategy achieved a sensitivity of 92.06% and specificity of 95.56% on an external validation dataset containing 288 three-dimensional (3D) OCT volumes from 118 Chinese patients in a different new hospital. This result met or exceeded the average of four medical experts who have used OCT for over one year. In addition to promising classification performance, our model has a remarkable ability to detect and visualize local lesions using the attention map of the standard ViT model, providing good interpretability for gynecologists to locate and diagnose possible cervical diseases. Qingbin Wang, Kaiyi Chen, Wanrong Dou, Yutao Ma |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Position-Enhanced and Time-aware Graph Convolutional Network for Sequential RecommendationsabstractThe sequential recommendation (also known as the next-item recommendation), which aims to predict the following item to recommend in a session according to users’ historical behavior, plays a critical role in improving session-based recommender systems. Most of the existing deep learning-based approaches utilize the recurrent neural network architecture or self-attention to model the sequential patterns and temporal influence among a user's historical behavior and learn the user's preference at a specific time. However, these methods have two main drawbacks. First, they focus on modeling users’ dynamic states from a user-centric perspective and always neglect the dynamics of items over time. Second, most of them deal with only the first-order user-item interactions and do not consider the high-order connectivity between users and items, which has recently been proved helpful for the sequential recommendation. To address the above problems, in this article, we attempt to model user-item interactions by a bipartite graph structure and propose a new recommendation approach based on a Position-enhanced and Time-aware Graph Convolutional Network (PTGCN) for the sequential recommendation. PTGCN models the sequential patterns and temporal dynamics between user-item interactions by defining a position-enhanced and time-aware graph convolution operation and learning the dynamic representations of users and items simultaneously on the bipartite graph with a self-attention aggregator. Also, it realizes the high-order connectivity between users and items by stacking multi-layer graph convolutions. To demonstrate the effectiveness of PTGCN, we carried out a comprehensive evaluation of PTGCN on three real-world datasets of different sizes compared with a few competitive baselines. Experimental results indicate that PTGCN outperforms several state-of-the-art sequential recommendation models in terms of two commonly-used evaluation metrics for ranking. In particular, it can make a better trade-off between recommendation performance and model training efficiency, which holds great potential for online session-based recommendation scenarios in the future. Liwei Huang, Yutao Ma, Bohong Danny Du, Shuliang Wang 0001, Deyi Li |
ACM Trans. Inf. Syst. | 2 |
| 2022 | Multi-information fusion based few-shot Web service classification
Bing Li 0010, Jian Wang 0018, Duantengchuan Li, Yutao Ma |
Future Gener. Comput. Syst. | 5 |
| 2022 | A spatial-temporal graph neural network framework for automated software bug triaging
Hongrun Wu, Yutao Ma, Zhenglong Xiang, Chen Yang 0007, Keqing He 0002 |
Knowl. Based Syst. | 2 |
| 2021 | DAN-SNR: A Deep Attentive Network for Social-aware Next Point-of-interest RecommendationabstractNext (or successive) point-of-interest (POI) recommendation, which aims to predict where users are likely to go next, has recently emerged as a new research focus of POI recommendation. Most of the previous studies on next POI recommendation attempted to incorporate the spatiotemporal information and sequential patterns of user check-ins into recommendation models to predict the target user's next move. However, few of the next POI recommendation approaches utilized the social influence of each user's friends. In this study, we discuss a new topic of next POI recommendation and present a deep attentive network for social-aware next POI recommendation called DAN-SNR. In particular, the DAN-SNR makes use of the self-attention mechanism instead of the architecture of recurrent neural networks to model sequential influence and social influence in a unified manner. Moreover, we design and implement two parallel channels to capture short-term user preference and long-term user preference as well as social influence, respectively. By leveraging multi-head self-attention, the DAN-SNR can model long-range dependencies between any two historical check-ins efficiently and weigh their contributions to the next destination adaptively. We also carried out a comprehensive evaluation using large-scale real-world datasets collected from two popular location-based social networks, namely, Gowalla and Brightkite. Experimental results indicate that the DAN-SNR outperforms seven competitive baseline approaches regarding recommendation performance and is highly efficient among six neural-network-based methods, four of which utilize the attention mechanism. Liwei Huang, Yutao Ma, Keqing He 0002 |
ACM Trans. Internet Techn. | 2 |
| 2021 | An Attention-Based Spatiotemporal LSTM Network for Next POI RecommendationabstractNext point-of-interest (POI) recommendation, also known as a natural extension of general POI recommendation, is recently proposed to predict user's next destination and has attracted considerable research interest. It focuses on learning users’ sequential patterns of check-in behavior and on training personalized recommendation models using different types of contextual information. Unfortunately, most of the previous studies failed to incorporate the spatiotemporal contextual information, which plays a critical role in analyzing user check-in behavior, into recommending the next POI. In recent years, embedding learning and recurrent neural network (RNN) based approaches show promising performance for modeling sequential patterns of check-in behavior in next POI recommendation. However, not all of the historical check-in records contribute equally to the next-step check-in behavior. To provide better next POI recommendation performance, we first proposed a spatiotemporal long and short-term memory (ST-LSTM) network. By feeding the spatiotemporal contextual information into the LSTM network in each step, ST-LSTM can model the spatial and temporal information better. Also, we developed an attention-based spatiotemporal LSTM (ATST-LSTM) network for next POI recommendation. By using the attention mechanism, ATST-LSTM can focus on the relevant historical check-in records in a check-in sequence selectively using the spatiotemporal contextual information. Besides, we conducted a comprehensive performance evaluation using large-scale real-world datasets collected from two popular location-based social networks, namely Gowalla and Brightkite. Experimental results indicated that the proposed ATST-LSTM network outperformed two state-of-the-art next POI recommendation approaches regarding three commonly-used evaluation metrics. Liwei Huang, Yutao Ma |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | Multi-modal Bayesian embedding for point-of-interest recommendation on location-based cyber-physical-social networks
Liwei Huang, Yutao Ma, Arun Kumar Sangaiah |
Future Gener. Comput. Syst. | 2 |
| 2020 | Computer-Aided Diagnosis in Histopathological Images of the Endometrium Using a Convolutional Neural Network and Attention MechanismsabstractUterine cancer (also known as endometrial cancer) can seriously affect the female reproductive system, and histopathological image analysis is the gold standard for diagnosing endometrial cancer. Due to the limited ability to model the complicated relationships between histopathological images and their interpretations, existing computer-aided diagnosis (CAD) approaches using traditional machine learning algorithms often failed to achieve satisfying results. In this study, we develop a CAD approach based on a convolutional neural network (CNN) and attention mechanisms, called HIENet. In the ten-fold cross-validation on ∼3,300 hematoxylin and eosin (H&E) image patches from ∼500 endometrial specimens, HIENet achieved a 76.91 ± 1.17% (mean ± s. d.) accuracy for four classes of endometrial tissue, i.e., normal endometrium, endometrial polyp, endometrial hyperplasia, and endometrial adenocarcinoma. Also, HIENet obtained an area-under-the-curve (AUC) of 0.9579 ± 0.0103 with an 81.04 ± 3.87% sensitivity and 94.78 ± 0.87% specificity in a binary classification task that detected endometrioid adenocarcinoma. Besides, in the external validation on 200 H&E image patches from 50 randomly-selected female patients, HIENet achieved an 84.50% accuracy in the four-class classification task, as well as an AUC of 0.9829 with a 77.97% (95% confidence interval, CI, 65.27%∼87.71%) sensitivity and 100% (95% CI, 97.42%∼100.00%) specificity. The proposed CAD method outperformed three human experts and five CNN-based classifiers regarding overall classification performance. It was also able to provide pathologists better interpretability of diagnoses by highlighting the histopathological correlations of local pixel-level image features to morphological characteristics of endometrial tissue. Xianxu Zeng, Gang Peng 0002, Yutao Ma |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Mining Domain Knowledge on Service Goals from Textual Service DescriptionsabstractWith the rapid development of service-oriented computing, a large number of software applications have been developed based on the services computing framework. It is well known that software engineering is a knowledge-intensive activity, and thus the effective management of service-related knowledge facilitates service-oriented software development. Although many methodologies have been proposed for service-oriented knowledge management, little attention has been paid to mining knowledge (especially domain-specific functionalities) from service resources. To address this issue, we propose an approach to mine domain knowledge on service goals (i.e., service functionalities) from textual descriptions of services. The approach consists of two components: service goal extraction from textual service descriptions based on linguistic analysis and domain service goal construction that merges semantically similar service goals within a domain. The effectiveness of the proposed approach is validated by a series of experiments conducted on a real-world dataset crawled from the ProgrammableWeb. Neng Zhang 0001, Jian Wang 0018, Yutao Ma |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | Predicting the Fixer of Software Bugs via a Collaborative Multiplex Network: Two Case Studies
Jinxiao Huang, Yutao Ma |
CollaborateCom | 2 |
| 2019 | An integrated service recommendation approach for service-based system development
Jian Wang 0018, Ruibin Xiong, Neng Zhang 0001, Yutao Ma, Keqing He 0002 |
Expert Syst. Appl. | 5 |
| 2018 | A Top-k Learning to Rank Approach to Cross-Project Software Defect PredictionabstractCross-project defect prediction (CPDP) has recently attracted increasing attention in the field of Software Engineering. Most of the previous studies, which treated it as a binary classification problem or a regression problem, are not practical for software testing activities. To provide developers with a more valuable ranking of the most severe entities (e.g., classes and modules), in this paper, we propose a top-k learning to rank (LTR) approach in the scenario of CPDP. In particular, we first convert the number of defects into graded relevance to a specific query according to the three-sigma rule; then, we put forward a new data resampling method called SMOTE-PENN to tackle the imbalanced data problem. An empirical study on the PROMISE dataset shows that SMOTE-PENN outperforms the other six competitive resampling algorithms and RankNet performs the best for the proposed approach framework. Thus, our work could lay a foundation for efficient search engines for top-ranked defective entities in real software testing activities without local historical data for a target project. Jinxiao Huang, Yutao Ma |
APSEC | 3 |
| 2018 | Deep hybrid collaborative filtering for Web service recommendation
Ruibin Xiong, Jian Wang 0018, Neng Zhang 0001, Yutao Ma |
Expert Syst. Appl. | 4 |
| 2018 | Analyzing the structure of Java software systems by weighted K-core decomposition
Weifeng Pan 0001, Bing Li 0010, Jing Liu 0033, Yutao Ma, Bo Hu 0013 |
Future Gener. Comput. Syst. | 4 |
| 2018 | Web service discovery based on goal-oriented query expansionabstractWith the broad adoption of service-oriented architecture, many software systems have been developed by composing loosely-coupled Web services. Service discovery, a critical step of building service-based systems (SBSs), aims to find a set of candidate services for each functional task to be performed by an SBS. The keyword-based search technology adopted by existing service registries is insufficient to retrieve semantically similar services for queries. Although many semantics-aware service discovery approaches have been proposed, they are hard to apply in practice due to the difficulties in ontology construction and semantic annotation. This paper aims to help service requesters (e.g., SBS designers) obtain relevant services accurately with a keyword query by exploiting domain knowledge about service functionalities (i.e., service goals) mined from textual descriptions of services. We firstly extract service goals from services’ textual descriptions using an NLP-based method and cluster service goals by measuring their semantic similarities. A query expansion approach is then proposed to help service requesters refine initial queries by recommending similar service goals. Finally, we develop a hybrid service discovery approach by integrating goal-based matching with two practical approaches: keyword-based and topic model-based. Experiments conducted on a real-world dataset show the effectiveness of our approach. Neng Zhang 0001, Jian Wang 0018, Yutao Ma, Keqing He 0002, Xiaoqing Frank Liu |
J. Syst. Softw. | 3 |
| 2018 | Guest Editorial: Cloud Services Meet Big DataabstractThe papers in this special issue are designed to solicit innovative and promising methods and techniques related closely to cloud services in the era of Big Data. The concept of Cloud Service represents a prime facility and feature of services in a cloud computing environment that can be made available to users on demand. Due to the flexibility of cloud computing in scaling IT resources up and down, cloud services gradually become valuable to attract the gaze of researchers and engineers from both academia and industry when they are faced with dynamically changing business requirements. Different stakeholders, such as consumers, providers, and operators, are generating a vast amount of data on such services per minute on the Internet, which increasingly comes to show the “4V” characteristics of big data. Therefore, new methodologies and techniques are urgently required for designing, validating, developing, testing, and deploying cloud services on demand in this specific scenario based on big data, as well as for efficiently being adaptive to business dynamics and users’ explicit and implicit requirements. Keqing He 0002, Liang-Jie Zhang, Schahram Dustdar, Yutao Ma |
IEEE Trans. Serv. Comput. | 4 |
| 2018 | Guest Editorial: Cloud Services Meet Big Data - Part IIabstractThis is the second part of a special issue on "Cloud Services Meet Big Data" organized to solicit innovative and promising methods and techniques related closely to cloud services in the era of Big Data. The seven papers included in this special investigate the most challenging issues in the areas of IaaS (Infrastructure as a Service) design and operations, requirements engineering and knowledge engineering for cloud services development, and service search and discovery. Keqing He 0002, Liang-Jie Zhang, Schahram Dustdar, Yutao Ma |
IEEE Trans. Serv. Comput. | 4 |
| 2016 | A Ranking-Oriented Approach to Cross-Project Software Defect Prediction: An Empirical StudyabstractIn recent years, cross-project defect prediction (CPDP) has become very popular in the field of software defect prediction.It was treated as a binary classification or regression problem in most of previous studies.However, the existing methods to solve this problem may be not suitable for those projects with limited manpower and time.In this paper, we revisit the issue and treat it as a ranking problem.Inspired by the idea of the Point-wise approach to Learning to Rank, we propose a ranking-oriented CPDP approach called ROCPDP.The empirical results obtained based on AEEEM show that the defect predictor built with our method under a specific CPDP context, in general, outperforms those predictors trained by using the benchmark methods in both CPDP and WPDP (within-project defect prediction) scenarios in terms of two common evaluation metrics for rank correlation.So, our work could be an initial attempt to construct new rankingoriented CPDP models for newly created or inactive projects. Guoan You, Yutao Ma |
SEKE | 2 |
| 2016 | An Empirical Study of Ranking-Oriented Cross-Project Software Defect PredictionabstractCross-project defect prediction (CPDP) has recently become very popular in the field of software defect prediction. It was generally treated as a binary classification problem or a regression problem in most of previous studies. However, these existing CPDP methods may be not suitable for those software projects that have limited manpower and budget. To address the issue of priority estimation for buggy software entities, in this paper CPDP is formulated as a ranking problem. Inspired by the idea of the pointwise approach to learning to rank, we propose a ranking-oriented CPDP approach called ROCPDP. A case study conducted on the datasets collected from AEEEM and PROMISE shows that ROCPDP outperforms the eight baseline methods in two CPDP scenarios, namely one-to-one and many-to-one. Besides, in the many-to-one scenario ROCPDP is, by and large, comparable to the best baseline method performed in a specific within-project defect prediction scenario. Guoan You, Yutao Ma |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2015 | Common Topic Group Mining for Web Service Discovery
Jian Wang 0018, Panpan Gao, Yutao Ma, Keqing He 0002 |
APSCC | 3 |
| 2015 | A Learning Approach to the Prediction of Reliability Ranking for Web ServicesabstractService computing is a popular development paradigm in information technology. The functional properties of Web services assure correct functionality of cloud applications, while the nonfunctional properties such as reliability might significantly influence the user-perceived availability evaluation. Reliability rankings provide valuable information for making optimal cloud service selection from a set of functionally-equivalent candidate services. There existed several approaches that can conduct reliability ranking prediction for Web services. Those approaches acquire different rankings with different preference functions. It is arduous to determine whether there exists the best one in them, and what is the best one if not. This paper proposes a learning approach to reliability ranking prediction for Web services which utilizes past service invocation logs to train preference function. To validate the proposed approach, large-scale experiments are conducted based on a real-world Web service dataset, WSDream. The results show that our proposed approach achieves higher prediction accuracy than the existing approaches. Bing Li 0010, Xiaohui Cui, Yutao Ma, Rong Yang 0007 |
ICWS | 4 |
| 2015 | An empirical study on predicting defect numbersabstractDefect prediction is an important activity to make software testing processes more targeted and efficient.Many methods have been proposed to predict the defect-proneness of software components using supervised classification techniques in within-and cross-project scenarios.However, very few prior studies address the above issue from the perspective of predictive analytics.How to make an appropriate decision among different prediction approaches in a given scenario remains unclear.In this paper, we empirically investigate the feasibility of defect numbers prediction with typical regression models in different scenarios.The experiments on six open-source software projects in PROMISE repository show that the prediction model built with Decision Tree Regression seems to be the best estimator in both of the scenarios, and that for all the prediction models, the results yielded in the cross-project scenario can be comparable to (or sometimes better than) those in the within-project scenario when choosing suitable training data.Therefore, the findings provide a useful insight into defect numbers prediction for those new and inactive projects. Yutao Ma |
SEKE | 2 |
| 2015 | Web Service Clustering Using Relational Database ApproachabstractIn the era of service-oriented software engineering (SOSE), service clustering is used to organize Web services, and it can help to enhance the efficiency and accuracy of service discovery. In order to improve the efficiency and accuracy of service clustering, this paper uses the self-join operation in relational database (RDB) to realize Web service clustering. Based on storing service information, it does the self-join operation towards the Input, Output, Precondition, Effect (IOPE) tables of Web services, which can enhance the efficiency of computing services similarity. The semantic reasoning relationship between concepts and the concept status path are used to do the calculation, which can improve the calculation accuracy. Finally, we use experiments to validate the effectiveness of the proposed methods. Jianxiao Liu, Keqing He 0002, Yutao Ma, Jian Wang 0018 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2015 | An empirical study on software defect prediction with a simplified metric set
Bing Li 0010, Jun Chen 0001, Yutao Ma |
Inf. Softw. Technol. | 5 |
| 2014 | An Empirical Study on Inter-Commit Times in SVN
Qiuju Hou, Yutao Ma, Jianxun Chen, Youwei Xu |
SEKE | 2 |
| 2013 | Empirical Evidence on Developer's Commit Activity for Open-Source Software Projects
Sihai Lin, Yutao Ma, Jianxun Chen |
SEKE | 2 |
| 2011 | Taxonomy for Evolution of Service-Based SystemabstractWith the rapid development of service computing related technology, the number of application of service based system (SBS) in enterprise is increasing rapidly. Research on evolution of SBS is becoming more and more important. In this paper, we propose a taxonomy framework for evolution of SBS, which is illustrated from five perspectives: (a) motivations of SBS evolutionary changes (why), (b) stakeholders of SBS evolutionary changes (who), (c) locations of SBS evolutionary changes happening (where), (d) times of SBS evolutionary changes happening (when), (e) support mechanisms in the process of SBS evolutionary changes (how). Furthermore, propagation of evolutionary changes of SBS is analyzed in the paper. Zaiwen Feng, Keqing He 0002, Rong Peng, Yutao Ma |
SERVICES | 4 |
| 2010 | A Hybrid Set of Complexity Metrics for Large-Scale Object-Oriented Software Systems
Yutao Ma, Keqing He 0002, Bing Li 0010, Jing Liu 0033, Xiao-Yan Zhou |
J. Comput. Sci. Technol. | 1 |
| 2010 | Measuring Structural Quality of Object-Oriented Softwares via Bug Propagation Analysis on Weighted Software Networks
Weifeng Pan 0001, Bing Li 0010, Yutao Ma, Yeyi Qin, Xiao-Yan Zhou |
J. Comput. Sci. Technol. | 3 |
| 2009 | Towards Merging Goal Models of Networked Software
Zaiwen Feng, Keqing He 0002, Rong Peng, Jian Wang 0018, Yutao Ma |
SEKE | 5 |
| 2009 | Class structure refactoring of object-oriented softwares using community detection in dependency networks
Weifeng Pan 0001, Bing Li 0010, Yutao Ma, Jing Liu 0033, Yeyi Qin |
Frontiers Comput. Sci. China | 3 |
| 2006 | Co-expression Gene Discovery from Microarray for Integrative Systems Biology
Yutao Ma, Yonghong Peng |
ADMA | 1 |
| 2006 | Scale Free in Software MetricsabstractSoftware has become a complex piece of work by the collective efforts of many. And it is often hard to predict what the final outcome will be. This transition poses new challenge to the software engineering (SE) community. By employing methods from the study of complex network, we investigate the object oriented (OO) software metrics from a different perspective. We incorporate the weighted methods per class (WMC) metric into our definition of the weighted OO software coupling network as the node weight. Empirical results from four open source OO software demonstrate power law distribution of weight and a clear correlation between the weight and the out degree. According to its definition, it suggests uneven distribution of function among classes and a close correlation between the functionality of a class and the number of classes it depending on. Further experiment shows similar distribution also exists between average LCOM and WMC as well as out degree. These discoveries will help uncover the underlying mechanisms of software evolution and will be useful for SE to cope with the emerged complexity in software as well as efficient test cases design Jing Liu 0033, Keqing He 0002, Yutao Ma, Rong Peng |
COMPSAC (1) | 3 |
| 2006 | Towards an Identification Framework for Software Drifts: A Case StudyabstractSoftware drift is a common phenomenon in software development processes, which may lead to process deviation and then affect software quality. Effectively controlling negative software drifts is the key to achieving acceptable, predictable, and dependable software evolution in the model-driven development. In this paper we put forward a general taxonomy to identify different drifts in development processes according to their effects on software development. Based on the taxonomy, categories of process drift and quality drift are detailedly introduced. Moreover, we propose an integration framework for different drifts to realize the integration between development process and product quality. Eventually, a case study from practical project development is shown to prove the validity of our identification framework. Yutao Ma, Keqing He 0002, Jianghua Wu, Jianxun Chen |
ICSEA | 1 |
| 2005 | A Qualitative Method for Measuring the Structural Complexity of Software Systems Based on Complex NetworksabstractHow can we effectively measure the complexity of a modern complex software system has been a challenge for software engineers. Complex networks as a branch of complexity science are recently studied across many fields of science, and many large-scale software systems are proved to represent an important class of artificial complex networks. So, we introduce the relevant theories and methods of complex networks to analyze the topological/structural complexity of software systems, which is the key to measuring software complexity. Primarily, basic concepts, operational definitions, and measurement units of all parameters involved are presented respectively. Then, we propose a qualitative measure based on the structure entropy that measures the amount of uncertainty of the structural information, and on the linking weight that measures the influences of interactions or relationships between components of software systems on their overall topologies/structures. Eventually, some examples are used to demonstrate the feasibility and effectiveness of our method. Yutao Ma, Keqing He 0002, Dehui Du |
APSEC | 1 |
| 2001 | New strategy of modeling inversion layer characteristics in MOS structure for ULSI applications
Yutao Ma, Litian Liu |
Sci. China Ser. F Inf. Sci. | 1 |
| 2001 | Analytical charge-control and I-V model for submicrometer anddeep-submicrometer MOSFETs fully comprising quantum mechanical effectsabstractA new analytical current-voltage (I-V) model for submicrometer and deep-submicrometer metal-oxide-semiconductor field-effect transistors (MOSFETs) is developed based on a newly developed charge-control model for the metal-oxide-semiconductor structure. Threshold-voltage shift due to quantum mechanical effects, finite inversion layer thickness effects (inversion layer capacitance), as well as increased depletion layer charge density after the strong inversion point are incorporated in the model. Inversion layer charge density with respect to the gate voltage from depletion through weak inversion to strong inversion regions with smooth transition between different regions is given by one expression. Two-dimensional short channel effects such as channel length modulation, drain-induced barrier lowering, mobility degradation, and carrier velocity saturation, as well as polysilicon depletion effects are included in the I-V model. Model results are compared with both numerical results of carrier sheet density and surface potential in the channel, and experimental results of I-V data for submicrometer and deep-submicrometer MOSFETs down to 0.09-/spl mu/m effective gate length and the accuracy of the model are demonstrated. Yutao Ma, Litian Liu, Lilin Tian, Zhiping Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |