EDBT 2026 Demo / reviewers in the wild / expert
Yuchi Ma
dblp:21/11039
· DBLP profile ↗
29ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) has proven effective in integrating external knowledge into large language models (LLMs) for solving question-answer (QA) tasks. The state-of-the-art RAG approaches often use the graph data as the external data since they capture the rich semantic information and link relationships between entities. However, existing graph-based RAG approaches cannot accurately identify the relevant information from the graph and also consume large numbers of tokens in the online retrieval process. To address these issues, we introduce a novel graph-based RAG approach, called Attributed Community-based Hierarchical RAG (ArchRAG), by augmenting the question using attributed communities, and also introducing a novel LLM-based hierarchical clustering method. To retrieve the most relevant information from the graph for the question, we build a novel hierarchical index structure for the attributed communities and develop an effective online retrieval method. Experimental results demonstrate that ArchRAG outperforms existing methods in both accuracy and token cost. Yixiang Fang, Yingli Zhou, Xilin Liu 0001, Yuchi Ma |
AAAI | 5 |
| 2026 | Peer-aided repairer: empowering large language models to repair advanced student assignments
Qianhui Zhao, Li Zhang 0029, Fang Liu 0032, Yang Liu 0003, Jing Jiang 0005, Ge Li 0001, Zian Sun, Zhong-Qi Li, Yuchi Ma |
Empir. Softw. Eng. | 12 |
| 2026 | Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
Fang Liu 0032, Yang Liu 0003, Lin Shi 0006, Zhen Yang 0022, Li Zhang 0029, Xiaoli Lian, Zhong-Qi Li, Yuchi Ma |
IEEE Trans. Software Eng. | 8 |
| 2026 | EffiReasonTrans: RL-Optimized Reasoning for Code TranslationabstractCode translation is a crucial task in software development and maintenance. While recent advancements in Large Language Models (LLMs) have improved automated code translation accuracy, these gains often come at the cost of increased inference latency–hindering real-world development workflows that involve human-in-the-loop inspection. To address this tradeoff, we propose EffiReasonTrans, a training framework designed to improve translation accuracy while balancing inference latency. We first construct a high-quality reasoning-augmented dataset by prompting a stronger language model DeepSeek-R1 to generate intermediate reasoning and target translations. Each (source code, reasoning, target code) triplet undergoes automated syntax and functionality checks to ensure reliability. Based on this dataset, we employ a two-stage training strategy: supervised fine-tuning on reasoning-augmented samples, followed by reinforcement learning to further enhance accuracy, which also helps balance inference latency. We evaluate EffiReasonTrans on six translation pairs. Experimental results show that EffiReason-Trans consistently improves translation accuracy (up to +49.2% CA and +27.8% CodeBLEU compared to the base model), while reducing the number of generated tokens (up to -19.3%) and lowering inference latency in most cases (up to -29.0%). Ablation studies further confirm the complementary benefits of the two-stage training framework. Additionally, EffiReasonTrans shows improvements of translation accuracy when integrated into agent-based frameworks. Our code and data are available athttps://github.com/DeepSoftwareAnalytics/EffiReasonTrans. Yanlin Wang 0001, Rongyi Ou, Yanli Wang 0001, Mingwei Liu 0002, Jiachi Chen, Ensheng Shi, Xilin Liu 0001, Yuchi Ma, Zibin Zheng |
IEEE Trans. Software Eng. | 8 |
| 2026 | RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code TranslationabstractRepository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark’s challenges and agent’s effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs’ deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data athttps://github.com/DeepSoftwareAnalytics/RepoTransBench. Yanli Wang 0001, Yanlin Wang 0001, Suiquan Wang, Daya Guo, Jiachi Chen, John C. Grundy, Xilin Liu 0001, Yuchi Ma, Mingzhi Mao, Hongyu Zhang 0002, Zibin Zheng |
IEEE Trans. Software Eng. | 8 |
| 2025 | RLCoder: Reinforcement Learning for Repository-Level Code CompletionabstractRepository-level code completion aims to generate code for unfinished code snippets within the context of a specified repository. Existing approaches mainly rely on retrievalaugmented generation strategies due to limitations in input sequence length. However, traditional lexical-based retrieval methods like BM25 struggle to capture code semantics, while model-based retrieval methods face challenges due to the lack of labeled data for training. Therefore, we propose RLCoder, a novel reinforcement learning framework, which can enable the retriever to learn to retrieve useful content for code completion without the need for labeled data. Specifically, we iteratively evaluate the usefulness of retrieved content based on the perplexity of the target code when provided with the retrieved content as additional context, and provide feedback to update the retriever parameters. This iterative process enables the retriever to learn from its successes and failures, gradually improving its ability to retrieve relevant and high-quality content. Considering that not all situations require information beyond code files and not all retrieved context is helpful for generation, we also introduce a stop signal mechanism, allowing the retriever to decide when to retrieve and which candidates to retain autonomously. Extensive experimental results demonstrate that RLCoder consistently outperforms state-of-the-art methods on CrossCodeEval and RepoEval, achieving 12.2% EM improvement over previous methods. Moreover, experiments show that our framework can generalize across different programming languages and further improve previous methods like RepoCoder. We provide the code and data at https://github.com/DeepSoftwareAnalytics/RLCoder. Yanlin Wang 0001, Yanli Wang 0001, Daya Guo, Jiachi Chen, Ruikai Zhang, Yuchi Ma, Zibin Zheng |
ICSE | 6 |
| 2025 | HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code GenerationabstractTo evaluate the repository-level code generation capabilities of Large Language Models (LLMs) in complex real-world software development scenarios, many evaluation methods have been developed. These methods typically leverage contextual code from the latest version of a project to assist LLMs in accurately generating the desired function. However, such evaluation methods fail to consider the dynamic evolution of software projects over time, which we refer to as evolution-ignored settings. This in turn results in inaccurate evaluation of LLMs' performance. In this paper, we conduct an empirical study to deeply understand LLMs' code generation performance within settings that reflect the evolution nature of software development. To achieve this, we first construct an evolution-aware repository-level code generation dataset, namely HumanEvo, equipped with an automated execution-based evaluation tool. Second, we manually categorize HumanEvo according to dependency levels to more comprehensively analyze the model's performance in generating functions with different dependency levels. Third, we conduct extensive experiments on HumanEvo with seven representative and diverse LLMs to verify the effectiveness of the proposed benchmark. We obtain several important findings through our experimental study. For example, we find that previous evolution-ignored evaluation methods result in inflated performance of LLMs, with performance overestimations ranging from 10.0% to 61.1% under different context acquisition methods, compared to the evolution-aware evaluation approach. Based on the findings, we give actionable suggestions for more realistic evaluation of LLMs on code generation. We also build a shared evolution-aware code generation toolbox to facilitate future research. The replication package including source code and datasets is anonymously available at https://github.Com/DeepSoftwareAnalytics/HumanEvo. Dewu Zheng, Yanlin Wang 0001, Ensheng Shi, Ruikai Zhang, Yuchi Ma, Hongyu Zhang 0002, Zibin Zheng |
ICSE | 5 |
| 2025 | AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionabstractRepository-level code completion remains a challenging task for existing code large language models (code LLMs) due to their limited understanding of repository-specific context and domain knowledge. While retrieval-augmented generation (RAG) approaches have shown promise by retrieving relevant code snippets as cross-file context, they suffer from two fundamental problems: misalignment between the query and the target code in the retrieval process, and the inability of existing retrieval methods to effectively utilize the inference information. To address these challenges, we propose AlignCoder, a repository-level code completion framework that introduces a query enhancement mechanism and a reinforcement learning based retriever training method. Our approach generates multiple candidate completions to construct an enhanced query that bridges the semantic gap between the initial query and the target code. Additionally, we employ reinforcement learning to train an AlignRetriever that learns to leverage inference information in the enhanced query for more accurate retrieval. We evaluate AlignCoder on two widely-used benchmarks (CrossCodeEval and RepoEval) across five backbone code LLMs, demonstrating an 18.1% improvement in EM score compared to baselines on the CrossCodeEval benchmark. The results show that our framework achieves superior performance and exhibits high generalizability across various code LLMs and programming languages. Tianyue Jiang, Yanlin Wang 0001, Yanli Wang 0001, Daya Guo, Ensheng Shi, Yuchi Ma, Jiachi Chen, Zibin Zheng |
ASE | 6 |
| 2025 | DrainCode: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context PoisoningabstractLarge language models (LLMs) have demonstrated impressive capabilities in code generation, by leveraging retrieval-augmented generation (RAG) methods. However, the computational costs associated with LLM inference, particularly in terms of latency and energy consumption, have received limited attention in the security context. This paper introduces DrainCode, the first adversarial attack targeting the computational efficiency of RAG-based code generation systems. By strategically poisoning retrieval contexts through mutation-based approach, DrainCode forces LLMs to produce significantly longer outputs, thereby increasing GPU latency and energy consumption. We evaluate the effectiveness of DrainCode across multiple models. Our experiments show that DrainCode achieves up to a 85% increase in latency, a 49% increase in energy consumption, and more than a 3× increase in output length compared to the baseline. Furthermore, we demonstrate the generalizability of the attack across different prompting strategies and its effectiveness compared to different defenses. The results highlight DrainCode as a potential method for increasing the computational overhead of LLMs, making it useful for evaluating LLM security in resource-constrained environments. We provide code and data at https://github.com/DeepSoftwareAnalytics/DrainCode. Yanli Wang 0001, Jiadong Wu, Tianyue Jiang, Mingwei Liu 0002, Jiachi Chen, Chong Wang 0013, Ensheng Shi, Xilin Liu 0001, Yuchi Ma, Zibin Zheng |
ASE | 9 |
| 2025 | Agents in software engineering: survey, landscape, and vision
Yanlin Wang 0001, Wanjun Zhong, Yanxian Huang, Ensheng Shi, Min Yang 0002, Jiachi Chen, Hui Li 0057, Yuchi Ma, Qianxiang Wang, Zibin Zheng |
Autom. Softw. Eng. | 8 |
| 2025 | In-depth Analysis of Graph-based RAG in a Unified Framework
Yingli Zhou, Yaodong Su, Youran Sun, Taotao Wang, Runyuan He, Sicong Liang, Xilin Liu 0001, Yuchi Ma, Yixiang Fang |
Proc. VLDB Endow. | 10 |
| 2024 | CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained ModelsabstractCode generation models based on the pre-training and fine-tuning paradigm have been increasingly attempted by both academia and industry, resulting in well-known industrial models such as Codex, CodeGen, and PanGu-Coder. To evaluate the effectiveness of these models, multiple existing benchmarks (e.g., HumanEval and AiXBench) are proposed, including only cases of generating a standalone function, i.e., a function that may invoke or access only built-in functions and standard libraries. However, non-standalone functions, which typically are not included in the existing benchmarks, constitute more than 70% of the functions in popular open-source projects, and evaluating models' effectiveness on standalone functions cannot reflect these models' effectiveness on pragmatic code generation scenarios (i.e., code generation for real settings of open source or proprietary code). Hao Yu 0016, Dezhi Ran, Jiaxin Zhang 0029, Qi Zhang 0020, Yuchi Ma, Guangtai Liang, Ying Li 0012, Qianxiang Wang, Tao Xie 0001 |
ICSE | 6 |
| 2024 | Deep Learning or Classical Machine Learning? An Empirical Study on Log-Based Anomaly DetectionabstractWhile deep learning (DL) has emerged as a powerful technique, its benefits must be carefully considered in relation to computational costs. Specifically, although DL methods have achieved strong performance in log anomaly detection, they often require extended time for log preprocessing, model training, and model inference, hindering their adoption in online distributed cloud systems that require rapid deployment of log anomaly detection service. Boxi Yu, Qiuai Fu, Zhiqing Zhong, Haotian Xie, Yaoliang Wu, Yuchi Ma, Pinjia He |
ICSE | 7 |
| 2024 | When to Stop? Towards Efficient Code Generation in LLMs with Excess Token PreventionabstractCode generation aims to automatically generate code snippets that meet given natural language requirements and plays an important role in software development. Although Code LLMs have shown excellent performance in this domain, their long generation time poses a signification limitation in practice use. In this paper, we first conduct an in-depth preliminary study with different Code LLMs on code generation task and identify a significant efficiency issue, i.e., continual generation of excess tokens. It harms the developer productivity and leads to huge computational wastes. To address it, we introduce CodeFast, an inference acceleration approach for Code LLMs on code generation. The key idea of CodeFast is to terminate the inference process in time when unnecessary excess tokens are detected. First, we propose an automatic data construction framework to obtain training data. Then, we train a unified lightweight model GenGuard applicable to multiple programming languages to predict whether to terminate inference at the current step. Finally, we enhance Code LLM with GenGuard to accelerate its inference in code generation task. We conduct extensive experiments with CodeFast on five representative Code LLMs across four widely used code generation datasets. Experimental results show that (1) CodeFast can significantly improve the inference speed of various Code LLMs in code generation, ranging form 34% to 452%, without compromising the quality of generated code. (2) CodeFast is stable across different parameter settings and can generalize to untrained datasets. Our code and data are available at https://github.com/DeepSoftwareAnalytics/CodeFast. Lianghong Guo, Yanlin Wang 0001, Ensheng Shi, Wanjun Zhong, Hongyu Zhang 0002, Jiachi Chen, Ruikai Zhang, Yuchi Ma, Zibin Zheng |
ISSTA | 8 |
| 2024 | FastFixer: An Efficient and Effective Approach for Repairing Programming AssignmentsabstractProviding personalized and timely feedback for student's programming assignments is useful for programming education. Automated program repair (APR) techniques have been used to fix the bugs in programming assignments, where the Large Language Models (LLMs) based approaches have shown promising results. Given the growing complexity of identifying and fixing bugs in advanced programming assignments, current fine-tuning strategies for APR are inadequate in guiding the LLM to identify bugs and make accurate edits during the generative repair process. Furthermore, the autoregressive decoding approach employed by the LLM could potentially impede the efficiency of the repair, thereby hindering the ability to provide timely feedback. To tackle these challenges, we propose FastFixer, an efficient and effective approach for programming assignment repair. To assist the LLM in accurately identifying and repairing bugs, we first propose a novel repair-oriented fine-tuning strategy, aiming to enhance the LLM's attention towards learning how to generate the necessary patch and its associated context. Furthermore, to speed up the patch generation, we propose an inference acceleration approach that is specifically tailored for the program repair task. The evaluation results demonstrate that FastFixer obtains an overall improvement of 20.46% in assignment fixing when compared to the state-of-the-art baseline. Considering the repair efficiency, FastFixer achieves a remarkable inference speedup of 16.67× compared to the autoregressive decoding algorithm. Fang Liu 0032, Qianhui Zhao, Jing Jiang 0005, Li Zhang 0029, Zian Sun, Ge Li 0001, Zhong-Qi Li, Yuchi Ma |
ASE | 9 |
| 2024 | MRCA: Metric-level Root Cause Analysis for Microservices via Multi-Modal DataabstractDue to the complexity and dynamic nature of large-scale microservice systems, manual troubleshooting is time-consuming and impractical. Therefore, automated Root Cause Analysis (RCA) is essential. However, existing RCA approaches face significant challenges. (1) Multi-modal data (e.g. traces, logs, and metrics) record the status of microservice systems, but most existing RCA approaches rely on single-source data, failing to understand the system fully. (2) Existing RCA approaches ignore the services' anomaly state and their anomaly intensity. (3) The service-level RCAs lack detailed information for quick issue resolution. To tackle these challenges, we propose MRCA, a metric-level RCA approach using multi-modal data. Our key insight is that using multi-modal data allows for a comprehensive understanding of the system, enabling the localization of root causes across more anomaly scenarios. MRCA first utilizes traces and logs to obtain the ranking list of abnormal services based on reconstruction probability. It further builds causal graphs from services with high anomaly probability to discover the order in which abnormal metrics of different services occur. By incorporating a reward mechanism, MRCA terminates the excessive expansion of the causal graph and significantly reduces the time taken for causal analysis. Finally, MRCA can prune the ranking list based on the causal graph and identify metric-level root causes. Experiments on two widely-used microservice benchmarks demonstrate that MRCA outperforms state-of-the-art approaches in terms of both accuracy and efficiency. Zhouruixing Zhu, Qiuai Fu, Yuchi Ma, Pinjia He |
ASE | 4 |
| 2024 | On Efficient Large Sparse Matrix Chain MultiplicationabstractSparse matrices are often used to model the interactions among different objects and they are prevalent in many areas including e-commerce, social network, and biology. As one of the fundamental matrix operations, the sparse matrix chain multiplication (SMCM) aims to efficiently multiply a chain of sparse matrices, which has found various real-world applications in areas like network analysis, data mining, and machine learning. The efficiency of SMCM largely hinges on the order of multiplying the matrices, which further relies on the accurate estimation of the sparsity values of intermediate matrices. Existing matrix sparsity estimators often struggle with large sparse matrices, because they suffer from the accuracy issue in both theory and practice. To enable efficient SMCM, in this paper we introduce a novel row-wise sparsity estimator (RS-estimator), a straightforward yet effective estimator that leverages matrix structural properties to achieve efficient, accurate, and theoretically guaranteed sparsity estimation. Based on the RS-estimator, we propose a novel ordering algorithm for determining a good order of efficient SMCM. We further develop an efficient parallel SMCM algorithm by effectively utilizing multiple CPU threads. We have conducted experiments by multiplying various chains of large sparse matrices extracted from five real-world large graph datasets, and the results demonstrate the effectiveness and efficiency of our proposed methods. In particular, our SMCM algorithm is up to three orders of magnitude faster than the state-of-the-art algorithms. Chunxu Lin, Wensheng Luo 0002, Yixiang Fang, Chenhao Ma 0001, Xilin Liu 0001, Yuchi Ma |
Proc. ACM Manag. Data | 6 |
| 2024 | Revisiting Log Parsing: The Present, the Future, and the UncertaintiesabstractIn the recent decade, the amount of software runtime logs has increased rapidly and spawned a line of automated log analysis research using machine learning or data mining algorithms. In the typical workflow of log analysis, log parsing, which aims to transform unstructured or semistructured logs into structured logs, is crucial to various downstream algorithms. While the state-of-the-art (SOTA) parsers achieve extremely high accuracy, recent research shows that these parsers are far from being useful under stricter evaluation metrics. Thus, researchers and practitioners are unclear about the current state of log parsing research and what might be important to explore in the future. To this end, we conduct an empirical study to revisit log parsing by running extensive experiments of SOTA parsers on 16 widely used log datasets under five evaluation metrics with different preprocessing settings. Our results show that the performance of log parsers varies significantly under different evaluation metrics. In addition, preprocessing plays an important role in the evaluation. In particular, preprocessing with common regular expressions can cause a 0.38 performance difference in group accuracy, highlighting the importance of reporting preprocessing details in parsing research. We also generalize the word-level regular expressions in preprocessing and try to use them to parse the whole logs, which leads to surprisingly decent accuracy. These results imply that formulating log parsing as a word-level classification task is a feasible future direction. Moreover, we find out that the most widely used dataset (i.e., LogHub) contains labeling errors. To address this issue, we make an extensive manual effort to fix the errors in the log dataset, providing a revised ground truth for future log parsing research. On the revised log dataset, our simple parser (word-level regular expression-based) achieves 0.97 precision-template accuracy on the Spark dataset and an average recall-template accuracy of 0.93 on 16 datasets, which outperforms all existing parsers. Zhijing Li 0007, Qiuai Fu, Zhijun Huang, Jianbo Yu 0003, Yiqian Li, Yuanhao Lai, Yuchi Ma |
IEEE Trans. Reliab. | 7 |
| 2023 | Flame: A Centralized Cache Controller for Serverless ComputingabstractCaching function is a promising way to mitigate coldstart overhead in serverless computing. However, as caching also increases the resource cost significantly, how to make caching decisions is still challenging. We find that the prior "local cache control" designs are insufficient to achieve high cache efficiency due to the workload skewness across servers. Laiping Zhao, Yuechan Hao, Yuchi Ma, Keqiu Li |
ASPLOS (4) | 6 |
| 2023 | Robust Log-Based Anomaly Detection with Hierarchical Contrastive LearningabstractLogs are widely employed in modern systems to record critical information and serve as an important source for anomaly detection, which has attracted increasing research interests. However, logs usually suffer from perturbations and it makes the existing log-based anomaly detection methods unstable. In this paper, we aim to solve this problem from the perspective of contrastive learning, by which the intrinsic and robust representations of logs are learned for anomaly detection. We propose two data augmentation methods to generate different views at different granularity for log data and design a deep hierarchical contrastive model for anomaly detection. In the contrastive semantic embedding module, we fine-tune a language model with a message-level contrastive loss. And in the contrastive anomaly detection module, we apply a sequence-level contrastive constraint to assist the detection model to learn robust embeddings for log sequences. Experiments on three datasets verify the effectiveness of our proposed method. Ruichun Yang, Qiuai Fu, Yuchi Ma |
ICASSP | 6 |
| 2023 | Hue: A User-Adaptive Parser for Hybrid LogsabstractLog parsing, which extracts log templates from semi-structured logs and produces structured logs, is the first and the most critical step in automated log analysis. While existing log parsers have achieved decent results, they suffer from two major limitations by design. First, they do not natively support hybrid logs that consist of both single-line logs and multi-line logs (Java Exception and Hadoop Counters). Second, they fall short in integrating domain knowledge in parsing, making it hard to identify ambiguous tokens in logs. This paper defines a new research problem, hybrid log parsing, as a superset of traditional log parsing tasks, and proposes Hue, the first attempt for hybrid log parsing via a user-adaptive manner. Specifically, Hue converts each log message to a sequence of special wildcards using a key casting table and determines the log types via line aggregating and pattern extracting. In addition, Hue can effectively utilize user feedback via a novel merge-reject strategy, making it possible to quickly adapt to complex and changing log templates. We evaluated Hue on three hybrid log datasets and sixteen widely-used single-line log datasets (Loghub). The results show that Hue achieves an average grouping accuracy of 0.845 on hybrid logs, which largely outperforms the best results (0.563 on average) obtained by existing parsers. Hue also exhibits SOTA performance on single-line log datasets. Junjielong Xu, Qiuai Fu, Zhouruixing Zhu, Yutong Cheng, Zhijing Li 0007, Yuchi Ma, Pinjia He |
ESEC/SIGSOFT FSE | 6 |
| 2023 | Multisource Maximum Predictor Discrepancy for Unsupervised Domain Adaptation on Corn Yield PredictionabstractRecently, with the advent of satellite missions and artificial intelligence techniques, supervised machine learning (ML) methods have been more and more used for analyzing remote sensing (RS) observation data for crop yield prediction. However, due to the domain shift between heterogeneous regions, supervised ML models tend to have poor spatial transferability. As a result, models trained with labeled data from one spatial region (i.e., source domain) often lose their validity when directly applied to another region (i.e., target domain). To address this issue, we proposed a multisource maximum predictor discrepancy (MMPD) neural network that is an unsupervised domain adaptation (UDA) approach for corn yield prediction at the county level. The novelties of this study include that: 1) we proposed to maximize the discrepancy between two source-specific yield predictors and align source and target domains by considering crop yield response in the target domain and 2) we adopted the strategy of multisource UDA to avoid negative interference between labeled samples from different sources. Case studies in the U.S. corn belt and Argentina demonstrated that the proposed MMPD model had effectively reduced domain shifts and outperformed several other state-of-the-art deep learning (DL) and UDA methods. Yuchi Ma, Zhengwei Yang 0002, Zhou Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Multitask Learning of Alfalfa Nutritive Value From UAV-Based Hyperspectral ImagesabstractAlfalfa is a valuable and widely adapted forage crop, and its nutritive value directly affects animal performance and ultimately affects the profitability of livestock production. Traditional nutritive value measurement method is labor-intensive and time-consuming and thus hinders the determination of alfalfa nutritive values over large fields. The adoption of unmanned aerial vehicles (UAVs) facilitates the generation of images with high spatial and temporal resolutions for field-level agricultural research. Additionally, compared with other imaging modalities, hyperspectral data usually consist of hundreds of narrow spectral bands and allow the accurate detection, identification, and quantification of crop quality. Although various machine-learning methods have been developed for alfalfa quality prediction, they were all single-task models that learned independently for each quality trait and failed to utilize the underlying relatedness between each task. Inspired by the idea of multitask learning (MTL), this study aims to develop an approach that simultaneously predicts multiple quality traits. The algorithm first extracts shared information through a long short-term memory (LSTM)-based common hidden layer. To enhance the model flexibility, it is then divided into multiple branches, each containing the same or different number of task-specific fully connected hidden layers. Through comparison with multiple mainstream single-task machine-learning models, the effectiveness of the model is illustrated based on the measured alfalfa quality data and multitemporal UAV-based hyperspectral imagery. Luwei Feng, Zhou Zhang 0001, Yuchi Ma, Yazhou Sun, Qingyun Du, Parker Williams, Jessica L. Drewry, Brian D. Luck |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Bayesian Domain Adversarial Neural Network for Corn Yield PredictionabstractCorn is the most widely grown crop in the U.S. and makes up a significant part of the American diet. Under the pressure of feeding a growing population, accurate and timely estimation of corn yield before the harvest is of great importance to supply chain management and regional food security. Recently, machine learning in conjunction with satellite remote sensing has been used for developing corn yield prediction models. Despite the success, a major bottleneck of training a reliable supervised machine learning model is the need for representative ground truth labels (e.g., yield records) which may be limited or even not available due to financial and manpower reasons. Also, due to domain shift, a machine learning model trained with labeled data from a label-rich region (i.e., source domain) could experience a significant performance decrease when directly applied to the region of interest (i.e., target domain). To address this issue, we proposed a Bayesian Domain Adversarial Neural Network (BDANN) for unsupervised domain adaptation on county-level corn yield prediction. By applying adversarial learning and Bayesian inference, BDANN was trained to reduce domain shift and accurately predict corn yield by extracting domain-invariant and task-informative features from both source and target domains. Moreover, the results also demonstrated that the BDANN model generalized well on small training sets. Experiments in two ecoregions in the U.S. corn belt have shown the effectiveness of the proposed BDANN and its superiority over other state-of-the-art methods. Yuchi Ma, Zhou Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | A meta-feature based unified framework for both cold-start and warm-start explainable recommendations
Ning Yang 0001, Yuchi Ma, Philip S. Yu |
World Wide Web | 2 |
| 2017 | Spatial and semantical label inference for social media - A cross-network data fusion approach
Yuchi Ma, Ning Yang 0001, Lei Zhang 0005, Philip S. Yu |
Knowl. Inf. Syst. | 1 |
| 2017 | Predicting neighbor label distributions in dynamic heterogeneous information networks
Yuchi Ma, Ning Yang 0001, Lei Zhang 0005, Philip S. Yu |
World Wide Web | 1 |
| 2016 | Explicable Location Prediction Based on Preference Tensor Model
Duoduo Zhang, Ning Yang 0001, Yuchi Ma |
WAIM (1) | 3 |
| 2015 | Predicting Neighbor Distribution in Heterogeneous Information NetworksabstractRecently, considerable attention has been devoted to the prediction problems arising from heterogeneous information networks. In this paper, we present a new prediction task, Neighbor Distribution Prediction (NDP), which aims at predicting the distribution of the labels on neighbors of a given node and is valuable for many different applications in heterogeneous information networks. The challenges of NDP mainly come from three aspects: the infinity of the state space of a neighbor distribution, the sparsity of available data, and how to fairly evaluate the predictions. To address these challenges, we first propose an Evolution Factor Model (EFM) for NDP, which utilizes two new structures proposed in this paper, i.e. Neighbor Distribution Vector (NDV) to represent the state of a given node's neighbors, and Neighbor Label Evolution Matrix (NLEM) to capture the dynamics of a neighbor distribution, respectively. We further propose a learning algorithm for Evolution Factor Model. To overcome the problem of data sparsity, the learning algorithm first clusters all the nodes and learns an NLEM for each cluster instead of for each node. For fairly evaluating the predicting results, we propose a new metric: Virtual Accuracy (VA), which takes into consideration both the absolute accuracy and the predictability of a node. Extensive experiments conducted on three real datasets from different domains validate the effectiveness of our proposed model EFM and metric VA. Yuchi Ma, Ning Yang 0001, Chuan Li 0002, Lei Zhang 0005, Philip S. Yu |
SDM | 1 |