Siyuan Jiang

dblp:84/9918 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
7 papers
Program synthesis and code generation · 40% Requirements engineering and software design · 26% Software maintenance and evolution · 19%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code completion
0.912025
Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion · ASE 2025
Program synthesis and code generation › code completion
repository-level code completion
0.912025
Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion · ASE 2025
Empirical software engineering
mining software repositories
0.622018
Towards Prioritizing Documentation Effort · IEEE Trans. Software Eng. 2018
Detecting user story information in developer-client conversations to generate extractive summaries · ICSE 2017
Program synthesis and code generation
code summarization
0.412019
A neural model for generating natural language summaries of program subroutines · ICSE 2019
Software maintenance and evolution
program comprehension
0.412019
A neural model for generating natural language summaries of program subroutines · ICSE 2019
Requirements engineering and software design › formal specification
requirements formalization
0.412019
Prema: A Tool for Precise Requirements Editing, Modeling and Analysis · ASE 2019
Requirements engineering and software design
requirements specification
0.412019
Prema: A Tool for Precise Requirements Editing, Modeling and Analysis · ASE 2019
Requirements engineering and software design › requirements engineering
requirements verification and validation
0.412019
Prema: A Tool for Precise Requirements Editing, Modeling and Analysis · ASE 2019
Software maintenance and evolution
software documentation
0.312018
Towards Prioritizing Documentation Effort · IEEE Trans. Software Eng. 2018
Software maintenance and evolution › software documentation
commit message generation
0.312017
Automatically generating commit messages from diffs using neural machine translation · ASE 2017
Requirements engineering and software design
requirements elicitation
0.312017
Detecting user story information in developer-client conversations to generate extractive summaries · ICSE 2017
Natural language and speech › Language models and text generation
alignment
0.312025
Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion · ASE 2025
Natural language and speech › Language models and text generation › alignment
long-context alignment
0.312025
Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion · ASE 2025
Program analysis › static analysis
program slicing
0.212013
Quantitative program slicing: separating statements by relevance · ICSE 2013
Program synthesis and code generation
code generation with language models
0.112017
Automatically generating commit messages from diffs using neural machine translation · ASE 2017
Software maintenance and evolution
change impact analysis
0.012013
Quantitative program slicing: separating statements by relevance · ICSE 2013

Methods — techniques the papers use, named apart from their topics

large language model · 1.7fine-tuning · 1.7data-driven alignment · 1.7neural machine translation · 0.7parsing · 0.4model checking · 0.4AST-based code structure · 0.4user study · 0.3textual analysis · 0.3static source code analysis · 0.3
YearPublicationVenuePosition
2025 Aligning LLMs to Fully Utilize the Cross-file Context in Repository-level Code Completion
abstract
Large Language Models (LLMs) have shown promising results in repository-level code completion, which completes code based on the in-file and cross-file context of a repository. The cross-file context typically contains different types of information (e.g., relevant APIs and similar code) and is lengthy. In this paper, we found that LLMs struggle to fully utilize the information in the cross-file context. We hypothesize that one of the root causes of the limitation is the misalignment between pre-training (i.e., relying on nearby context) and repo-level code completion (i.e., frequently attending to long-range cross-file context).To address the above misalignment, we propose Code Long-context Alignment - CoLA, a purely data-driven approach to explicitly teach LLMs to focus on the cross-file context. Specifically, CoLA constructs a large-scale repo-level code completion dataset - CoLA-132K, where each sample contains the long cross-file context (up to 128K tokens) and requires generating context-aware code (i.e., cross-file API invocations and code spans similar to cross-file context). Through a two-stage training pipeline upon CoLA-132K, LLMs learn the capability of finding relevant information in the cross-file context, thus aligning LLMs with repo-level code completion. We apply CoLA to multiple popular LLMs (e.g., aiXcoder-7B) and extensive experiments on CoLA-132K and a public benchmark - CrossCodeEval. Our experiments yield the following results. ❶ Effectiveness. CoLA substantially improves the performance of multiple LLMs in repo-level code completion. For example, it improves aiXcoder-7B by up to 19.7% in exact match. ❷ Generalizability. The capability learned by CoLA can generalize to new languages (i.e., languages not in training data). ❸ Enhanced Context Utilization Capability. We design two probing experiments, which show CoLA improves the capability of LLMs in utilizing the information (i.e., relevant APIs and similar code) in cross-file context. Our datasets and model weights are released in [1].
Jia Li 0011, Huanyu Liu 0001, Xianjie Shi, He Zong, Yihong Dong, Kechi Zhang, Siyuan Jiang, Zhi Jin 0001, Ge Li 0001
ASE8
2025 Efficient DOA Estimation for Coprime Array via Bi-Nuclear Schatten-p Norm Minimization
abstract
In this letter, we propose the Bi-nuclear Schatten-p norm minimization (BSNM) for coprime array to achieve the efficient direction of arrival (DOA) estimation. Specifically, BSNM factorizes a large Hermitian Toeplitz covariance matrix as the product of two small matrices, which not only improves the computational efficiency but also captures the low-rank property of the interpolated covariance matrix. This BSNM problem of the Hermitian Toeplitz matrix is then solved by a parallel alternating optimization algorithm. Finally, the recovered covariance matrix is applied to estimate the DOA via the MUSIC algorithm. Numerical simulation illustrates the superiority of the proposed BSNM method over state-of-the-art methods in terms of computational efficiency.
Siyuan Jiang, Shuai Liu 0005, Ming Jin 0004, Fenggang Yan, Zhiping Lin 0001
IEEE Signal Process. Lett.1
2024 A Nyström-based low-rank unitary MVDR beamforming scheme
Siyuan Jiang, Ming Jin 0004, Shuai Liu 0005, Zhiping Lin 0001
Signal Process.1
2023 Research on Torque Distribution Function Method for DSEM Based on H-Bridge Converter
abstract
The traditional control strategies for doubly salient electro-magnetic motor (DSEM) have problems such as large torque ripple, low current utilization, and difficulty in commutation during operation. To address these problems, a torque distribution function (TDF) control method based on H-bridge converter is proposed by this paper, which reduces the commutation area by three-phase independent chopping and increases the motor output by using the TDF combined with torque closed loop to suppress torque ripple. Firstly, the advantages of the proposed method in improving torque performance are analyzed according to the basic theory of DSEM, the principle of torque distribution is proposed, and the TDF is derived by combining the actual nonlinear back electromotive force (EMF) curve. Then, the existence of a better torque distribution ratio is demonstrated and the calculation formula is derived with the aim of improving the torque-ampere ratio. Finally, the TDF method was verified to be effective in suppressing torque ripple and increasing motor output by an 18/12-pole DSEM experimental platform.
Bo Zhou 0014, Siyuan Jiang, Hongjun Shi
IECON4
2023 High-dimensional MVDR beamforming based on a second unitary transformation
Siyuan Jiang
Signal Process.1
2019 A neural model for generating natural language summaries of program subroutines
abstract
Source code summarization -- creating natural language descriptions of source code behavior -- is a rapidly-growing research topic with applications to automatic documentation generation, program comprehension, and software maintenance. Traditional techniques relied on heuristics and templates built manually by human experts. Recently, data-driven approaches based on neural machine translation have largely overtaken template-based systems. But nearly all of these techniques rely almost entirely on programs having good internal documentation; without clear identifier names, the models fail to create good summaries. In this paper, we present a neural model that combines words from code with code structure from an AST. Unlike previous approaches, our model processes each data source as a separate input, which allows the model to learn code structure independent of the text in code. This process helps our approach provide coherent summaries in many cases even when zero internal documentation is provided. We evaluate our technique with a dataset we created from 2.1m Java methods. We find improvement over two baseline techniques from SE literature and one from NLP literature.
Alexander LeClair, Siyuan Jiang, Collin McMillan
ICSE2
2019 Research on Control Strategy of Excitation-loss for DSEM based on Full-bridge Converter
abstract
The application of doubly salient electromagnetic machine(DSEM) to more electric aircraft starter/generator system has broad prospects. And the more electric aircraft are extremely demanding on operational reliability, which is mainly determined by the operating stability of the starter/generator system. Therefore, DSEM's fault-tolerant control technology plays a significant role in improving the operational reliability for more electric aircraft. At present, the researches on fault-tolerant control of DSEM mostly focus on the fault-tolerant design of body structure. There are few studies on fault-tolerant operation control under excitation fault. In this paper, a three-phase six-state control method based on the full-bridge converter is proposed to enable the DSEM in generating operation to continue generating electricity without field current. And the experiments are carried out to verify the correctness of the proposed control strategy.
Bo Zhou 0014, Kaimiao Wang, Siyuan Jiang
IECON5
2019 Prema: A Tool for Precise Requirements Editing, Modeling and Analysis
abstract
We present Prema, a tool for Precise Requirement Editing, Modeling and Analysis. It can be used in various fields for describing precise requirements using formal notations and performing rigorous analysis. By parsing the requirements written in formal modeling language, Prema is able to get a model which aptly depicts the requirements. It also provides different rigorous verification and validation techniques to check whether the requirements meet users' expectation and find potential errors. We show that our tool can provide a unified environment for writing and verifying requirements without using tools that are not well inter-related. For experimental demonstration, we use the requirements of the automatic train protection (ATP) system of CASCO signal co. LTD., the largest railway signal control system manufacturer of China. The code of the tool cannot be released here because the project is commercially confidential. However, a demonstration video of the tool is available at https://youtu.be/BX0yv8pRMWs.
Yihao Huang 0001, Jincao Feng, Hanyue Zheng, Jiayi Zhu 0002, Siyuan Jiang, Weikai Miao, Geguang Pu
ASE6
2019 Automated Classification of Amyotrophic Lateral Sclerosis Using Multi-level Whole-brain Volumes from Structural Magnetic Resonance Imaging
abstract
We proposed and validated a fully-automated classification procedure for amyotrophic lateral sclerosis (ALS) using structural magnetic resonance imaging; specifically, T1-weighted images from 28 ALS subjects and 28 healthy control (HC) subjects were used. The raw features were obtained from a validated multi-granularity whole-brain analysis pipeline, providing multi-level whole-brain segmentation volumes. We employed the support vector machine as our classification algorithm with several feature selection techniques analyzed. According to our leave-one-out cross validation experiment results, the whole-brain structural volumes from Level 4, followed by a feature selection utilizing the standardized Wilcoxon two-sample rank sum statistic, yielded the best classification performance; overall accuracy: 83.93%, sensitivity: 85.71%, specificity: 82.14%, and the area under the receiver operating characteristic curve: 0.8380. The feature selection procedure revealed that the volumes of the thalamus, especially that on the left hemisphere, are the most important (of highest ranking) in the ALS-vs-HC discrimination.
Yuanyuan Wei 0001, Siyuan Jiang, Yuanyuan Qin, Xiaoying Tang 0001
SMC2
2019 NiPred: Need Predictor for Hurricane Disaster Relief
abstract
It is of paramount importance to know the situations of people who undergone disaster events and be aware of their updates, yet it is not an easy job to accomplish in the chaos of a disaster. To facilitate advanced disaster relief organization and efficient supplement distribution, we develop NiPred, a social media based need prediction prototype that predicts needs for victims across the affected area. NiPred first extracts problems and concerns posted by victims of disaster-hurricane, in our case study; then displays the statistics to offer an overview for awareness and further analysis; and last, predicts the needs such as "diaper", "boat", "canoe" and "shelter" etc. for disaster relief planning.
Long Hoang Nguyen 0002, Siyuan Jiang, Hashim Abu-gellban, Hanxiang Du, Fang Jin
SSTD2
2018 Towards Prioritizing Documentation Effort
abstract
Programmers need documentation to comprehend software, but they often lack the time to write it. Thus, programmers must prioritize their documentation effort to ensure that sections of code important to program comprehension are thoroughly explained. In this paper, we explore the possibility of automatically prioritizing documentation effort. We performed two user studies to evaluate the effectiveness of static source code attributes and textual analysis of source code towards prioritizing documentation effort. The first study used open-source API Libraries while the second study was conducted using closed-source industrial software from ABB. Our findings suggest that static source code attributes are poor predictors of documentation effort priority, whereas textual analysis of source code consistently performed well as a predictor of documentation effort priority.
Paul W. McBurney, Siyuan Jiang, Marouane Kessentini, Nicholas A. Kraft, Ameer Armaly, Mohamed Wiem Mkaouer, Collin McMillan
IEEE Trans. Software Eng.2
2017 Detecting user story information in developer-client conversations to generate extractive summaries
abstract
User stories are descriptions of functionality that a software user needs. They play an important role in determining which software requirements and bug fixes should be handled and in what order. Developers elicit user stories through meetings with customers. But user story elicitation is complex, and involves many passes to accommodate shifting and unclear customer needs. The result is that developers must take detailed notes during meetings or risk missing important information. Ideally, developers would be freed of the need to take notes themselves, and instead speak naturally with their customers. This paper is a step towards that ideal. We present a technique for automatically extracting information relevant to user stories from recorded conversations between customers and developers. We perform a qualitative study to demonstrate that user story information exists in these conversations in a sufficient quantity to extract automatically. From this, we found that roughly 10.2% of these conversations contained user story information. Then, we test our technique in a quantitative study to determine the degree to which our technique can extract user story information. In our experiment, our process obtained about 70.8% precision and 18.3% recall on the information.
Paige Rodeghero, Siyuan Jiang, Ameer Armaly, Collin McMillan
ICSE2
2017 TraceLab Components for Generating Extractive Summaries of User Stories
abstract
This artifact is a reproducibility package for experiments in user stories summarization. We implemented and packaged the artifact as a set of reusable TraceLab components. The existing implementation of the artifact was relatively difficult to use because it required the user to coordinate several different programming languages and dependencies. This artifact, available via our online appendix, provides the components, a detailed tutorial with screenshots that show exactly where to click and what to enter, and an example virtual machine image.
Rrezarta Krasniqi, Siyuan Jiang, Collin McMillan
ICSME2
2017 Docio: documenting API input/output examples
abstract
When learning to use an Application Programming Interface (API), programmers need to understand the inputs and outputs (I/O) of the API functions. Current documentation tools automatically document the static information of I/O, such as parameter types and names. What is missing from these tools is dynamic information, such as I/O examples-actual valid values of inputs that produce certain outputs. In this paper, we demonstrate Docio, a prototype toolset we built to generate I/O examples. Docio logs I/O values when API functions are executed, for example in running test suites. Then, Docio puts I/O values into API documents as I/O examples. Docio has three programs: 1) funcWatch, which collects I/O values when API developers run test suites, 2) ioSelect, which selects one I/O example from a set of I/O values, and 3) ioPresent, which embeds the I/O examples into documents. In a preliminary evaluation, we used Docio to generate four hundred I/O examples for three C libraries: ffmpeg, libssh, and protobuf-c. Docio is open-source and available at: http://www3.nd.edu/~sjiang1/docio/.
Siyuan Jiang, Ameer Armaly, Collin McMillan, Qiyu Zhi, Ronald A. Metoyer
ICPC1
2017 Towards automatic generation of short summaries of commits
abstract
Committing to a version control system means submitting a software change to the system. Each commit can have a message to describe the submission. Several approaches have been proposed to automatically generate the content of such messages. However, the quality of the automatically generated messages falls far short of what humans write. In studying the differences between auto-generated and human-written messages, we found that 82% of the human-written messages have only one sentence, while the automatically generated messages often have multiple lines. Furthermore, we found that the commit messages often begin with a verb followed by an direct object. This finding inspired us to use a "verb+object" format in this paper to generate short commit summaries. We split the approach into two parts: verb generation and object generation. As our first try, we trained a classifier to classify a diff to a verb. We are seeking feedback from the community before we continue to work on generating direct objects for the commits.
Siyuan Jiang, Collin McMillan
ICPC1
2017 Automatically generating commit messages from diffs using neural machine translation
abstract
Commit messages are a valuable resource in comprehension of software evolution, since they provide a record of changes such as feature additions and bug repairs. Unfortunately, programmers often neglect to write good commit messages. Different techniques have been proposed to help programmers by automatically writing these messages. These techniques are effective at describing what changed, but are often verbose and lack context for understanding the rationale behind a change. In contrast, humans write messages that are short and summarize the high level rationale. In this paper, we adapt Neural Machine Translation (NMT) to automatically "translate" diffs into commit messages. We trained an NMT algorithm using a corpus of diffs and human-written commit messages from the top 1k Github projects. We designed a filter to help ensure that we only trained the algorithm on higher-quality commit messages. Our evaluation uncovered a pattern in which the messages we generate tend to be either very high or very low quality. Therefore, we created a quality-assurance filter to detect cases in which we are unable to produce good messages, and return a warning instead.
Siyuan Jiang, Ameer Armaly, Collin McMillan
ASE1
2017 Do Programmers do Change Impact Analysis in Debugging?
Siyuan Jiang, Collin McMillan, Raúl A. Santelices
Empir. Softw. Eng.1
2016 Prioritizing Change-Impact Analysis via Semantic Program-Dependence Quantification
abstract
Software is constantly changing. To ensure the quality of this process, when preparing to change a program, developers must first identify the main consequences and risks of modifying the program locations they intend to change. This activity is called change-impact analysis. However, existing impact analysis suffers from two major problems: coarse granularity and large size of the resulting impact sets. Finer-grained analyses such as slicing give more detailed impact sets which, however, are also even larger in size. While various impact-set reduction approaches have been proposed at different levels of granularity, the challenge persists as very-large impact sets are still produced, impeding the adoption of impact analysis due to the great costs of inspecting those impact sets. To address these challenges, we present a novel dynamic-analysis technique called SensA which combines sensitivity analysis and execution differencing. SensA not only provides fine-grained (statement-level) impact sets but also prioritizes potential impacts via semantic-dependence quantification for program slices. We evaluated the benefits of impact prioritization using SensA with respect to static and dynamic forward slicing via an extensive empirical study of open-source Java applications and three case studies. Our results show that SensA can offer much better cost-effectiveness than slicing in assisting developers with impact inspection and fault cause-effect understanding.
Haipeng Cai, Raúl A. Santelices, Siyuan Jiang
IEEE Trans. Reliab.3
2014 SENSA: Sensitivity Analysis for Quantitative Change-Impact Prediction
abstract
Sensitivity analysis determines how a system responds to stimuli variations, which can benefit important software-engineering tasks such as change-impact analysis. We present SENSA, a novel dynamic-analysis technique and tool that combines sensitivity analysis and execution differencing to estimate the dependencies among statements that occur in practice. In addition to identifying dependencies, SENSA quantifies them to estimate how much or how likely a statement depends on another. Quantifying dependencies helps developers prioritize and focus their inspection of code relationships. To assess the benefits of quantifying dependencies with SENSA, we applied it to various statements across Java subjects to find and prioritize the potential impacts of changing those statements. We found that SENSA predicts the actual impacts of changes to those statements more accurately than static and dynamic forward slicing. Our SENSA prototype tool is freely available for download.
Haipeng Cai, Siyuan Jiang, Raúl A. Santelices, Ying-Jie Zhang, Yiji Zhang
SCAM2
2014 On the Accuracy of Forward Dynamic Slicing and Its Effects on Software Maintenance
abstract
Dynamic slicing is a practical and popular analysis technique used in various software-engineering tasks. Dynamic slicing is known to be incomplete because it analyzes only a subset of all possible executions of a program. However, it is less known that its results may inaccurately represent the dependencies that occur in those executions. Some researchers have identified this problem and developed extensions such as relevant slicing, which incorporates static information. Yet, dynamic slicing continues to be widely used, even though the extent of its inaccuracy is not well understood, which can affect the benefits of this analysis. In this paper, we present an approach to assess the accuracy of forward dynamic slices, which are used in software maintenance and evolution tasks. Because finding all actual dependencies is an undecidable problem, our approach instead computes bounds of the precision and recall of forward dynamic slices. Our approach uses sensitivity analysis and execution differencing to find a subset of all program statements that truly depend at runtime on another statement. Using this approach, we studied the accuracy of many forward dynamic slices from a variety of Java applications. Our results show that forward dynamic slicing can have low recall -- for dependencies in the analyzed executions -- and some potential imprecision. We also conducted a case study that shows how this inaccuracy affects a software maintenance task. To the best of our knowledge, ours is the first work that quantifies the intrinsic limitations of dynamic slicing.
Siyuan Jiang, Raúl A. Santelices, Mark Grechanik, Haipeng Cai
SCAM1
2013 Quantitative program slicing: separating statements by relevance
abstract
Program slicing is a popular but imprecise technique for identifying which parts of a program affect or are affected by a particular value. A major reason for this imprecision is that slicing reports all program statements possibly affected by a value, regardless of how relevant to that value they really are. In this paper, we introduce quantitative slicing (q-slicing), a novel approach that quantifies the relevance of each statement in a slice. Q-slicing helps users and tools focus their attention first on the parts of slices that matter the most. We present two methods for quantifying slices and we show the promise of q-slicing for a particular application: predicting the impacts of changes.
Raúl A. Santelices, Yiji Zhang, Siyuan Jiang, Haipeng Cai, Ying-Jie Zhang
ICSE3