VLDB 2026 Research / reviewers in the wild / expert
Ji Wu 0003
dblp:91/4957-3
· DBLP profile ↗
22ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KDG-Rec: Enhanced Dual-GNN Programming Exercise Recommendation via LLM-Powered Knowledge Annotation and Preference-DecouplingabstractThe rapid expansion of online programming exercise platforms has brought abundant learning resources for programming education, but also presents challenges for personalized exercise recommendation due to missing or imprecise knowledge annotations and the sparsity of learner-exercise interaction records, which together lead to reduced accuracy and a lack of explainability in the recommendation results. Existing natural language processing and large language models (LLM)-based annotation methods struggle to capture implicit knowledge and often generate redundant or inconsistent results. Moreover, mainstream recommendation systems are ineffective at handling the severe sparsity of interaction sequences in programming exercise datasets, and typically model student preferences in a single dimension—overlooking key educational factors such as knowledge gaps, difficulty tolerance, and preferred learning rhythm. To address these challenges, we propose KDG-Rec, a novel framework that first introduces AgentCo-KAS, a multiagent LLM-based collaborative annotation method with role-specific fine-tuning and cross-agent verification for high-precision, fine-grained knowledge annotation. Building on these enriched annotations, we develop a disentangled graph neural network model that constructs dual exercise-interaction graphs to effectively capture learning patterns at multiple granularities in interaction sequences and explicitly decouples student preferences into four interpretable dimensions for adaptive fusion. Extensive experiments on real-world datasets demonstrate that KDG-Rec outperforms nine state-of-the-art methods across multiple metrics, significantly advancing personalized programming exercise recommendation. Bangqi Li, Qing Sun 0004, Ji Wu 0003, Wenge Rong |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | TROI: Cross-Subject Pretraining with Sparse Voxel Selection for Enhanced fMRI Visual DecodingabstractfMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually labeled ROIs (Regions of Interest) to select brain voxels. However, these ROIs can contain redundant information and noise, reducing decoding performance. Additionally, the lack of automated ROI labeling methods hinders the practical application of fMRI visual decoding technology, especially for new subjects. This work presents TROI (Trainable Region of Interest), a novel two-stage, data-driven ROI labeling method for cross-subject fMRI decoding tasks, particularly when subject samples are limited. TROI leverages labeled ROIs in the dataset to pretrain an image decoding backbone on a cross-subject dataset, enabling efficient optimization of the input layer for new subjects without retraining the entire model from scratch. In the first stage, we introduce a voxel selection method that combines sparse mask training and low-pass filtering to quickly generate the voxel mask and determine input layer dimensions. In the second stage, we apply a learning rate rewinding strategy to fine-tune the input layer for downstream tasks. Experimental results on the same small sample dataset as the baseline method for brain visual retrieval and reconstruction tasks show that our voxel selection method surpasses the state-of-the-art method MindEye2 with an annotated ROI mask. Tengyu Pan, Zhenyu Li 0008, Ji Wu 0003, Xiuxing Li, Jianyong Wang 0001 |
ICASSP | 4 |
| 2025 | DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenariosabstract3D object detection from multi-view images in traffic scenarios has garnered significant attention in recent years. Many existing approaches rely on object queries that are generated from 3D reference points to localize objects. However, a limitation of these methods is that some reference points are often far from the target object, which can lead to false positive detections. In this paper, we propose a depth-guided query generator for 3D object detection (DQ3D) that leverages depth information and 2D detections to ensure that reference points are sampled from the surface or interior of the object. Furthermore, to address partially occluded objects in current frame, we introduce a hybrid attention mechanism that fuses historical detection results with depth-guided queries, thereby forming hybrid queries. Evaluation on the nuScenes dataset demonstrates that our method outperforms the baseline by 6.3% in terms of mean Average Precision (mAP) and 4.3% in the NuScenes Detection Score (NDS). Ji Wu 0003 |
IJCNN | 3 |
| 2025 | BFGen: Basic Flow Generation for Refining Requirements via LLM and Relational Graph Attention NetworksabstractIdentifying interaction scenarios between a system and its actors from the high-level requirements and forming use case basic flows is crucial in requirement refinement. Traditional manual methods often yield incomplete or inaccurate flows due to engineers' limited domain expertise, while rulebased methods-relying on predefined parsing rules-suffer from linguistic ambiguities and domain-dependent limitations. Although large language model (LLM) approaches leverage rich domain knowledge and robust natural language processing, they are constrained by input length, generation instability, and the risk of out-of-system outputs, frequently resulting in context-unaware or irrelevant flows. To overcome these challenges, this paper proposes BFGen to generate contextcompliant basic flows strictly adhering to domain constraints and requirement boundaries. BFGen employs LLMs to accurately extract domain-specific terms and interactions, and integrates a Relational Graph Attention Network with attention preservation factors to model logical dependencies and domain constraints effectively. Empirical evaluations on 13 public and 7 industrial datasets show that BFGen outperforms leading baselines by$\approx 14 \%$in Precision,$\approx 7-25 \%$in Recall,$\approx 11- 30 \%$in F1 Score, and$\approx 10-19 \%$in AUC. Furthermore, our evaluations confirm the effectiveness of both the LLM module and the attention preservation factors, and assess the impact of requirement completeness on the performance of BFGen. Bangqi Li, Ji Wu 0003, Zhijun Shao |
QRS | 3 |
| 2025 | Advancing Software Project Effort Estimation: Leveraging a NIVIM for Enhanced PreprocessingabstractABSTRACT Software development effort estimation (SDEE) is essential for effective project planning and relies heavily on data quality affected by incomplete datasets. Missing data (MD) are a prevalent problem in machine learning, yet many models treat it arbitrarily despite its significance. Inadequate handling of MD may introduce bias into the induced knowledge. It can be challenging to choose optimal imputation approaches for software development projects. This article presents a novel incomplete value imputation model (NIVIM) that uses a variational autoencoder (VAE) for imputation and synthetic data. By combining contextual and resemblance components, our approach creates an SDEE dataset and improves the data quality using contextual imputation. The key feature of the proposed model is its applicability to a wide variety of datasets as a preprocessing unit. Comparative evaluations demonstrate that NIVIM outperforms existing models such as VAE, generative adversarial imputation network (GAIN), ‐nearest neighbor (K‐NN), and multivariate imputation by chained equations (MICE). Our proposed model NIVIM produces statistically substantial improvements on six benchmark datasets, that is, ISBSG, Albrecht, COCOMO81, Desharnais, NASA, and UCP, with an average improvement in RMSE of 11.05% to 17.72% and MAE of 9.62% to 21.96%. Syed Sarmad Ali, Jian Ren 0004, Ji Wu 0003, Chao Liu 0002 |
J. Softw. Evol. Process. | 3 |
| 2024 | Test Architecture Generation by Leveraging BERT and Control and Data Flows
Ji Wu 0003, Qing Sun 0004, Tao Yue 0002 |
ICECCS | 2 |
| 2022 | Editorial to theme section on open environmental software systems modeling
Tao Yue 0002, Paolo Arcaini, Ji Wu 0003, Xiaowei Huang 0001 |
Softw. Syst. Model. | 3 |
| 2022 | Does PageRank apply to service ranking in microservice regression testing?
Lizhe Chen, Ji Wu 0003 |
Softw. Qual. J. | 2 |
| 2022 | SRTEF: Test Function Recommendation With Scenarios and Latent Semantic for Implementing Stepwise Test CaseabstractImplementing test cases as programs to automate test execution is a popular testing practice. Current industrial practices usually use test functions to implement the test steps of a test case and then to compose the executable test case by choosing the test functions to call manually. It is time-consuming and could lead to invalid test results by selecting inappropriate test functions. In this article, we propose an automatic test function recommendation approach named Scenario-based Recommendation of TEst Function (SRTEF). Given a test step of a test case, SRTEF uses the weighted description similarity and the scenario similarity to recommend test functions. The description similarity utilizes the deep structured semantic model (DSSM) to measure the relatedness between a test step and a test function by their literal descriptions. The test scenario and the test function usage scenario are considered to calculate the scenario similarity. SRTEF has been successfully applied in Huawei. The systematic experiments have been conducted to evaluate SRTEF by using the dataset from Huawei and comparing with BiInformation source-based KnowledgE Recommendation (BIKER), reported as the best approach so far. The results show that SRTEF outperforms BIKER with significant positive ratios consistently in all the three selection strategies, i.e., Top-3, Top-5, and Top-10. The DSSM shows its advantage over word embedding by the double performance of capturing the semantic relatedness in SRTEF. Ji Wu 0003, Qing Sun 0004, Ruiyuan Wan |
IEEE Trans. Reliab. | 2 |
| 2021 | SRTEF: Automatic Test Function Recommendation with Scenarios for Implementing Stepwise Test CaseabstractImplementing test cases to automate test execution is a popular testing practice currently. A stepwise test case consists of several sequential test steps. Given a test function library, the typical way to implement a test case is calling the existing test functions in the library to reduce test cost. How to find the appropriate test function(s) to implement a test step in a given test case thus becomes an important problem. However, in current testing practices, test engineers usually select the appropriate test function manually by experience. It is time-consuming and could lead to invalid test results by selecting inappropriate or wrong test functions to call. In this paper, we propose an automatic test function recommendation approach with scenario named SRTEF (Scenario-based Recommendation of TEst Function). Given a test step, SRTEF uses two levels of similarities to recommend test functions, description similarity and scenario similarity. The description similarity measures the semantic relatedness between the test step and test function by their literal descriptions. To calculate the scenario similarity, SRTEF at first retrieves a set of historical test cases that contains test step(s) semantically similar to the given test step; then the scenario similarity between test step and test function is calculated according to the calling relation between retrieved test case and test function, and the co-occurrence relation among test functions. SRTEF has been successfully applied in Huawei. We evaluate SRTEF by using the dataset from Huawei and comparing with BIKER, reported as the best recommendation approach so far. The results show that SRTEF outperforms the BIKER approach by at least 49% in Mean Average Precision, 33% in Mean Reciprocal Rank, and 25% in Mean Recall. Ji Wu 0003, Qing Sun 0004, Ruiyuan Wan |
QRS | 2 |
| 2017 | A Restricted Natural Language Based Use Case Modeling Methodology for Real-Time SystemsabstractTime-related properties are a critical type of extrafunctional requirements for designing real-time systems. Modeling and validating time-related properties at the requirements specification and analysis phases is important for the successful development of real-time systems in terms of cost, quality and productivity. In the literature and practice, timing analyses (e.g., Worst Case Execution Time) are often performed to ensure that the design of a real-time system fully conforms to its time-related constraints. However, such analyses are mostly performed at the design and implementation stages, but not at the requirements level. This paper presents a restricted, natural language based, use case modeling methodology (named as RUCM4RT) to specify functional requirements of real-time systems as use case models, along with associated time-related constraints. RUCM4RT was proposed based on the UML profile for Modeling and Analysis of Real-Time and Embedded Systems (MARTE). In addition, in this paper, we also propose a metamodel-based formalization mechanism named as UCMeta4RT to automatically formalize use case models. We have conducted two real-world case studies to evaluate our solution and 40 use cases were modeled, among which 27 realtime use cases, 118 time-related constraints and 47 other extrafunctional (also commonly called non-functional) constraints were specified. Results show that RUCM4RT was able to handle all the real-time related elements (e.g., time-related constraints) of the use case models. Huihui Zhang 0003, Tao Yue 0002, Shaukat Ali 0001, Ji Wu 0003, Chao Liu 0002 |
MiSE@ICSE | 4 |
| 2017 | Assessing the quality of industrial avionics software: an extensive empirical evaluation
Ji Wu 0003, Shaukat Ali 0001, Tao Yue 0002, Chao Liu 0002 |
Empir. Softw. Eng. | 1 |
| 2015 | A modeling methodology to facilitate safety-oriented architecture design of industrial avionics softwareabstractSummary Ensuring that avionics software meets safety requirements at each development stage is very important to warrant the safe operation of an avionics system. Many safety requirements are imposed by various standards and industrial regulations that must be met by avionics software. One of such standards is DO‐178B/C, which provides guidelines (e.g., development process and objectives to satisfy in development activities) for meeting safety requirements. This paper presents a modeling methodology including a UML profile for specifying safety requirements on a component‐based architecture model and a set of design guidelines on avionics software. These safety requirements were identified from both standards (mainly DO‐178B/C) and current engineering practices in the domain of avionics systems. The methodology automatically enforces these safety requirements. We have applied the methodology on an industrial autopilot system, and several previously uncaught faults were revealed. Copyright © 2014 John Wiley & Sons, Ltd. Ji Wu 0003, Tao Yue 0002, Shaukat Ali 0001, Huihui Zhang 0003 |
Softw. Pract. Exp. | 1 |
| 2013 | A Model-Driven Approach for Evaluating System of SystemsabstractTo reduce the cost and risk of the development of system of systems (SoS) by pre-evaluation before the SoS is built, a model-driven approach is proposed to evaluate the SoS based on its architecture, especially focused on measures of performance and effectiveness. In order to implement the pre-evaluation, the system architecture needs to be transformed to the simulation model of the system under the directing of the evaluation requirement model, then the evaluation can be done based on the evaluation model and the data from simulation. The architecture, evaluation requirement model, simulation model and evaluation model involved in the above process constitute the model system (a set of models related with each other) of our approach. Through the modeling activities based on the model system, we can implement the systematically pre-evaluation of SoS. The model system and its meta-model, which are the core of our approach, are detailed after the introduction of the framework of our approach. Then the transformations from DoDAF architecture to evaluation requirement model and simulation model are studied for accelerating the evaluation process. The case study shows that our approach can not only ensure the standardization and systematization of the evaluating process, but can improve the evaluating efficiency and creditability based on the proposed transformation method. Xiaokai Xia, Ji Wu 0003, Chao Liu 0002 |
ICECCS | 2 |
| 2013 | Experience report: Assessing the reliability of an industrial avionics software: Results, insights and recommendationsabstractReal Time Operating System for Avionics (RTOS4A) is responsible for providing an operating environment for avionics application software. Avionics software being safety-critical in nature poses several safety and reliability requirements on RTOS4A in addition to the requirements imposed by standards, for instance, DO-178B. Due to this reason, reliability assessment of RTOS4A is very critical to demonstrate confidence about its reliability to its relevant stakeholders. One common way of assessing reliability is by systematic analyses of testing data such as number of tests, number of failures, and coverage using appropriate statistical tests. In this paper, we report our experience of assessing the reliability of an industrial RTOS4A based on testing data collected for 17 months on eight continuous releases. We studied correlation among various measures including: Testing Effort Measures (e.g., complexity of test cases), Testing Effectiveness Measures (e.g., number of failures), and Complexity Measures (e.g., number of functions in a release) and provide in this paper a set of recommendations to assess the reliability of RTOS4A, which serve as guidelines to practitioners in the domain of RTOS4A. Ji Wu 0003, Shaukat Ali 0001, Tao Yue 0002 |
ISSRE | 1 |
| 2011 | Guest editors' introduction to the special section from the international symposium on web systems evolution
Ying Zou 0001, Ji Wu 0003, Kenny Wong |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2008 | Jata: A Language for Distributed Component TestingabstractDistributed component requires test automation more than other components. Test language plays an important role in test automation. This paper proposes a new language, Jata, for testing distributed component in a systematic way by integrating the advantages of Junit and TTCN-3. To test a distributed component, multiple test clients are needed to emulate users to request services from the component under test. Those test clients should be deployed in different machines and can collaborate to finish the testing. A distributed component, as a kind of software component, often has rich data types in its interfaces and the data is usually marshaled to transmit via a network. By inheriting the U2TP (UML2 Testing Profile) concepts, Jata is designed with rich constructs to specify test behavior and test data. By test component, Jata enables the building of distributed test clients; by test case and multi-threading test evaluation, Jata enables the development of any complicated test scenario in flexible way; by test data and its coding and decoding utilities, Jata enables the development of test data in the same way as a programmer does. The script meta-model and system architecture are presented to give a holistic view of Jata. To show the effectiveness of Jata, a banking service testing case study is illustrated. Ji Wu 0003 |
APSEC | 1 |
| 2008 | Mining Open Source Component Behavior for Reuse Evaluation
Ji Wu 0003, Xiao-xia Jia, Chao Liu 0002 |
ICSR | 1 |
| 2006 | A Framework of Model-Driven Web Application TestingabstractWeb applications have become complex and crucial in many fields. In order to assure their quality, a high demand for systematic methodologies of Web application testing is emerging. In this paper, a methodology of model-driven testing (MDT) for Web application is presented. Model is the core of this method. Web application model is built to describe the system under testing. Test case models are generated based-on the Web application model. Test deployment model and test control model are designed to describe the environment and process of test execution. The test engine executes test cases based-on the test deployment and control model automatically. After that, testing results are reflected to test case models. A framework is designed for supporting this methodology. In order to get better extensibility and flexibility, it is loose-coupled by a modeler and a tester. The modeler enables developers to design meta-models, and is responsible for creating, visualizing and saving models. The tester takes in charge of recovering the tested Web application model, generating test cases, and executing test Qin-qin Ma, Ji Wu 0003, Maozhong Jin, Chao Liu 0002 |
COMPSAC (2) | 3 |
| 2006 | Java Object Behavior Modeling and VisualizationabstractJava developers need to know what a specific object did during a program run. Object behavior visualization can fulfill this requirement. This paper presents a novel object behavior model, a Lifetime Behavior Model (LBM) and visualization methods to provide deductive and inductive visualizations of Java object behavior. For the deductive visualization, this paper visualizes the object behavior by three different LBMTrees from thread, object interaction and method invocation view respectively. For the inductive visualization, this paper presents an Activity Spectrum Model (ASM) and a set of performance measurements based on the LBM. The visualization prototype is developed to access object behavior events by JVMPI, construct the models and visualize the models. Experiment shows that the results proposed here can provide comprehensive and clear understanding of Java object behaviors. Ji Wu 0003, Xiao-xia Jia, Yong Po Liu, Guo-huan Li |
ICSEA | 1 |
| 2004 | Application of Maximum Entropy Principle to Software Failure PredictionabstractPredicting failures from software input is still a tough issue. Two models, namely the surface model and structure model, are presented in this paper to predict failure by applying the maximum entropy principle. The surface model forecasts a failure from the statistical co-occurrence between input and failure, while the structure model does from the statistical cause-effect between fault and failure. To evaluate the models, precision is applied and 17 testing experiments are conducted on 5 programs. Based on the experiments, the surface model and structure model get an average precision of 0.876 and 0.858, respectively Ji Wu 0003, Xiao-xia Jia, Chao Liu 0002, Maozhong Jin |
COMPSAC | 1 |
| 2004 | A Statistical Model to Locate Faults at Input Level
Ji Wu 0003, Xiao-xia Jia, Chao Liu 0002, Maozhong Jin |
ASE | 1 |