VLDB 2026 Research / reviewers in the wild / expert
Huan Jin
dblp:25/1667
· DBLP profile ↗
23ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HGTRTracer: Advancing Requirements-Code Traceability with LLM-Augmented Attribution Reasoning and Heterogeneous Graph TransformerabstractTraceability links between requirements and source code are crucial for maintaining software quality, yet traditional traceability methods exhibit critical limitations. Current methods predominantly rely on semantic similarity among artifacts, failing to account for link transitivity and the underlying rationale governing traceability link formation. To address these challenges, this study introduces HGTRTracer, a novel framework that integrates Large Language Models (LLMs) with Graph Neural Networks (GNNs) to advance traceability analysis. The proposed method innovatively incorporates attribution reasoning via LLMs to derive explicit “Reason nodes” that justify traceability links, with a Heterogeneous Graph Transformer (HGT) captures structural dependencies among nodes and edges to improve link prediction accuracy. Experiments conducted on 10 open-source projects demonstrate that HGTRTracer outperforms the state-of-the-art baseline (HGNNLink) by 12.90% in F1-score on average. Beyond performance gains, the framework enhances interpretability through its attribution-aware design, providing actionable insights for traceability decisions. Huan Jin, Yingkai Yuan, Bangchao Wang, Hongyan Wan, Zhiyuan Zou |
APSEC | 1 |
| 2025 | DesDD: A Design-Enabled Framework with Dual-Layer Debugging for LLM-based Iterative API OrchestratingabstractIn contemporary Software Engineering (SE), coordinated API calls are necessary to perform complex data retrieval operations as well as tasks.Large Language Models (LLMs) offer highly potential capabilities for natural language parsing and automation of tasks, which sparked research into integrating APIs orchestration with LLMs.Nonetheless, while existing LLM-based frameworks have developed considerably, yet they experience challenges in tackling complex tasks which tend to involve iterative, step-by-step problemsolving.Current frameworks lack structured guidance, relying on LLMs' own capabilities, resulting in blind iterations, inefficient error correction, and inefficient token utilization.This work introduces DesDD (Design-enabled framework with Dual-layer Debugging), a structured framework for LLM-driven iterative API orchestration.By applying software engineering design principles, we organize API orchestration workflows into distinct design and coding phases.Our dual-layer debugging mechanism detects and corrects errors in * Z. Zhou Zou, Zhenchang Xing, Xueting Yi, Huan Jin, Zhaojin Lu |
Internetware | 8 |
| 2025 | How to enhance requirements-to-code traceability? From the perspective of project artifactsabstractCurrent research on requirements-to-code traceability predominantly focuses on improving the performance of algorithms or models, while neglecting the quality of project artifacts themselves.Nevertheless, high-quality data can significantly improve the performance of a model, and disregarding data quality can be challenging to rectify even with sophisticated models.Often, high-quality data contributes substantially more to model performance enhancement than improvements to the model's structure itself.In real-world software development, requirements and code artifacts are the primary data sources for establishing traceability links.However, there is a lack of effective guidance on what factors of requirements and code artifacts are conducive to traceability.To address this issue, this study proposes five metrics that are verified to have a close impact on the requirementsto-code traceability: cyclomatic complexity, annotation ratio, co-occurrence ratio, noun-verb ratio, and noise ratio.Based on a systematic mapping study, we selected 11 open-source projects from 20 relevant publications to conduct experiments, and employed Spearman correlation analysis to validate the association between the proposed metrics and traceability.The experimental results show that the proposed five metrics can significantly affect the traceability of requirements-to-code, of which the cyclomatic complexity and the annotation ratio are more influential. Huan Jin, Yingkai Yuan, Hongyan Wan, Zhiyuan Zou, Bangchao Wang |
SEKE | 1 |
| 2025 | HGNNLink: recovering requirements-code traceability links with text and dependency-aware heterogeneous graph neural networks
Bangchao Wang, Zhiyuan Zou, Xuanxuan Liang, Huan Jin, Peng Liang 0001 |
Autom. Softw. Eng. | 4 |
| 2025 | CLIP prior-guided 3D open-vocabulary occupancy prediction
Zongkai Zhang, Jingrui Ye, Huan Jin, Lihui Jiang, Wenming Yang |
Pattern Recognit. | 4 |
| 2025 | Throughput Improvement for RIS-Empowered Wireless Powered Anti-Jamming Communication Networks (WPAJCN)abstractIn this paper, we propose a reconfigurable intelligent surface (RIS)-aided wireless powered anti-jamming communication network (WPAJCN), where the RIS is utilized to participate in downlink wireless power transfer (WPT), as well as uplink anti-jamming wireless information transfer (AJ-WIT). To evaluate the network anti-jamming performance, we maximize a sum anti-jamming throughput, with the constraints of downlink WPT and uplink AJ-WIT time scheduling, and unit-modulus RIS phase shifts. The formulated problem is not convex in terms of these two types of coupled variables, which cannot be directly solved. To address this problem, the Lagrange dual method and Karush-Kuhn-Tucker conditions are presented to transform its sum-of-logarithmic objective function into the logarithmically fractional counterpart, which reformulate the original problem into that with respect to RIS phase shift vectors and WPT time scheduling. Next, we propose to apply the Dinkelback algorithm to solve a non-linear fractional programming with respect to the downlink WPT and uplink AJ-WIT RIS phase shifts in an alternating fashion, each of which is derived into a semi-closed solution by utilizing theRiemannian Manifold Optimization(RMO). In addition, the optimal WPT time scheduling is obtained by numerical search. Finally, the numerical results are demonstrated to confirm the improved performance of the proposed approach compared to the benchmark counterparts, which highlights the that RIS can effectively enhance the uplink anti-jamming WIT capability as well as the downlink WPT efficiency. Zheng Chu 0001, David Chieng, Chiew Foong Kwong, Huan Jin, Zhengyu Zhu 0001, Chongwen Huang, Chau Yuen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | HANTracer: Leveraging Heterogeneous Graph Attention Network for Large-Scale Requirements-Code Traceability Link RecoveryabstractIn the task of requirements-to-code traceability link recovery, the continues growth of software scale has led to diminishing differences between indices and more complex nonlinear relationships within the data. This results in the performance decline of the most widely used information retrieval methods and machine learning methods when handling this task. Therefore, we propose a requirement traceability method based on heterogeneous graph attention networks, named HANTracer. The model integrates the high-dimensional vectors generated by a pre-trained model as node features to deepen the dif-ferentiation between nodes and enhances node representations with contextual information learned from the graph structure. Additionally, it utilizes the properties of heterogeneous edges to construct various edge features, such as code calling, code inheritance, and text similarity relationships, aiding the model in understanding and utilizing the relationships between different types of data elements. By incorporating average pooling layers and multiple fully connected layers, the HANTracer model is further improved to enhance its ability to extract nonlinear features. Experimental results indicate that HANTracer achieves an average F1 performance higher than the state-of-the-art methods TAROT by 100.30% and DF4RT by 46.27% on seven real-world open source software (OSS) datasets, demonstrating significant performance advantages in large-scale and complex data environments. Zhiyuan Zou, Bangchao Wang, Hongyan Wan, Huan Jin, Yukun Cao |
APSEC | 4 |
| 2024 | A pattern-based algorithm with fuzzy logic bin selector for online bin packing problemabstractThe online bin packing problem is a well-known optimization challenge that finds application in a wide range of real-world scenarios. In the paper, we propose a novel algorithm called FuzzyPatternPack(FPP), which leverages fuzzy inference and pattern-based predictions of the distribution of item sizes in online bin packing. In comparison to traditional heuristics like BestFit(BF) and FirstFit(FF), as well as the more recent PatternPack(PaP) and ProfilePacking(PrP) algorithm based on online predictions, FPP demonstrates competitive and superior performance in solving various benchmark problems. Particularly, it excels in addressing problems with evolving distributions, making it a promising solution for real-world applications where the item sizes may change over time. This research unveils the promising potential of employing fuzzy logic to effectively address uncertainty in scheduling and planning problems. Bingchen Lin, Jiawei Li 0001, Tianxiang Cui, Huan Jin, Ruibin Bai, Rong Qu, Jonathan M. Garibaldi |
Expert Syst. Appl. | 4 |
| 2024 | Let's Discover More API Relations: A Large Language Model-Based AI Chain for Unsupervised API Relation InferenceabstractAPIs have intricate relations that can be described in text and represented as knowledge graphs to aid software engineering tasks. Existing relation extraction methods have limitations, such as limited API text corpus, and are affected by the characteristics of the input text. To address these limitations, we propose utilizing large language models (LLMs) (e.g., GPT-3.5) as a neural knowledge base for API relation inference. This approach leverages the entire Web used to pre-train LLMs as a knowledge base and is insensitive to the context and complexity of input texts. To ensure accurate inference, we design an AI chain consisting of three AI modules: API Fully Qualified Name (FQN) Parser, API Knowledge Extractor, and API Relation Decider. The accuracy of the API FQN Parser and API Relation Decider is 0.81 and 0.83, respectively. Using the generative capacity of the LLM and our approach’s inference capability, we achieve an average F1 value of 0.76 under the three datasets, significantly higher than the state-of-the-art method’s average F1 value of 0.40. Compared to the original CoT and modularized CoT methods, our AI chain design has improved the performance of API relation inference by 71% and 49%, respectively. Meanwhile, the prompt ensembling strategy enhances the performance of our approach by 32%. The API relations inferred by our method can be further organized into structured forms to provide support for other software engineering tasks. Yanbang Sun, Zhenchang Xing, Yuanlong Cao, Jieshan Chen, Xiwei Xu 0001, Huan Jin |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2024 | $\mathbf{A^{3}}$A3-CodGen: A Repository-Level Code Generation Framework for Code Reuse With Local-Aware, Global-Aware, and Third-Party-Library-AwareabstractLLM-based code generation tools are essential to help developers in the software development process. Existing tools often disconnect with the working context, i.e., the code repository, causing the generated code to be not similar to human developers. In this paper, we propose a novel code generation framework, dubbed$A^{3}$-CodGen, to harness information within the code repository to generate code with fewer potential logical errors, code redundancy, and library-induced compatibility issues. We identify three types of representative information for the code repository: local-aware information from the current code file, global-aware information from other code files, and third-party-library information. Results demonstrate that by adopting the$A^{3}$-CodGen framework, we successfully extract, fuse, and feed code repository information into the LLM, generating more accurate, efficient, and highly reusable code. The effectiveness of our framework is further underscored by generating code with a higher reuse rate, compared to human developers. This research contributes significantly to the field of code generation, providing developers with a more powerful tool to address the evolving demands in software development in practice. Dianshu Liao, Shidong Pan, Xiaoyu Sun 0002, Xiaoxue Ren, Zhenchang Xing, Huan Jin, Qinying Li |
IEEE Trans. Software Eng. | 7 |
| 2022 | UCC: Uncertainty guided Cross-head Cotraining for Semi-Supervised Semantic SegmentationabstractDeep neural networks (DNNs) have witnessed great successes in semantic segmentation, which requires a large number of labeled data for training. We present a novel learning framework called Uncertainty guided Cross-head Cotraining (UCC) for semi-supervised semantic segmentation. Our framework introduces weak and strong augmentations within a shared encoder to achieve cotraining, which naturally combines the benefits of consistency and self-training. Every segmentation head interacts with its peers and, the weak augmentation result is used for supervising the strong. The consistency training samples' diversity can be boosted by Dynamic Cross-Set Copy-Paste (DCSCP), which also alleviates the distribution mismatch and class imbalance problems. Moreover, our proposed Uncertainty Guided Re-weight Module (UGRM) enhances the self-training pseudo labels by suppressing the effect of the low-quality pseudo labels from its peer via modeling uncertainty. Extensive experiments on Cityscapes and PASCAL VOC 2012 demonstrate the effectiveness of our UCC. Our approach significantly outperforms other state-of-the-art semi-supervised semantic segmentation methods. It achieves 77.17%, 76.49% mIoU on Cityscapes and PASCAL VOC 2012 datasets respectively under 1/16 protocols, which are + 10.1%, + 7.91% better than the supervised baseline. Jiashuo Fan, Huan Jin, Lihui Jiang |
CVPR | 3 |
| 2022 | Occlusion-Invariant Representation Alignment for Entity Re-IdentificationabstractEntity re-identification is the foundation of tracking- and matching-based computer vision tasks, which are widely employed in a variety of applications. However, when trained exclusively on clear images, the models capacity to generalize is significantly affected by the presence of occlusion at referencing time, whereas data argumentation-based approaches are costly to construct without guaranteeing a test-time improvement. To tackle this problem, we propose a domain adaptation framework based on learning representations that generates occlusion-invariant feature representations by aligning the clean image embedding distribution with the occluded one, using a disparity discrepancy metric derived from the siamese network architecture. Without the need for additional processing modules during the inference stage or an expensive occlusion-augmentation-enlarged dataset during the training stage, we could obtain occlusion invariant embeddings that are free of the impact of occluders. Extensive experimental results for two tasks across three datasets indicate the proposed method’s robustness and effectiveness to a variety of occlusions at all levels. Zhanghao Jiang, Heshan Du, Huan Jin, Zheng Lu 0002, Qian Zhang 0018 |
ICIP | 4 |
| 2022 | An Empirical Study on Source Code Feature Extraction in Preprocessing of IR-Based Requirements TraceabilityabstractIn information retrieval-based (IR-based) requirements traceability research, a great deal of researches have focused on establishing trace links between requirements and source code. However, as the description styles of source code and requirements are very different, how to better preprocess the code is crucial for the quality of trace link generation. This paper aims to draw empirical conclusions about code feature extraction, annotation importance assessment, and annotation redundancy removal through comprehensive experiments, which impact the quality of trace links generated by IR-based methods between requirements and source code. The results show that when the average annotaion density is higher than 0.2, feature extraction is recommended. Removing redundancy from code with high annotation redundancy can enhance the quality of trace links. The above experiences can help developers to improve the quality of trace link generation and provide them with advice on writing code. Bangchao Wang, Ruiqi Luo, Huan Jin |
QRS | 4 |
| 2022 | LSNet: Real-time attention semantic segmentation network with linear complexity
Pengpeng Sheng, Yanli Shi, Huan Jin |
Neurocomputing | 4 |
| 2021 | PBMC Cell Classification from Single Cell mRNA Expression by Artificial Neural Networks, Profiles, Gene Markers, and Protein MarkersabstractWe performed classification of healthy Peripheral Blood Mononuclear Cells cell types using four methods Artificial Neural Network (ANN), Profiles, Protein Markers (PMs), and RNA markers (RNAMs). Profiles represent patterns of gene expressions characteristic of the subtypes of cells. PMs are protein found exclusively in certain types or subtypes of cells, or represent particular cell states, RNAMs are genes which demonstrate significant differential expressions between cell types. A total of 109 datasets from four different sources containing $\sim$ 120,000 single cells gene expression were used to train and test prediction models. We combined the methods which perform prediction using the whole set of gene features (ANN and Profiles), and those that used specific gene features (PMs and RNAMs) to predict the cell type. The overall classification accuracy was 94.8% for ANN, 94.5% for Profiles, 90.7% for PMs, 67.9% for RNAMs. The combination of four methods showed accuracy of 90.9% with high confidence of positive predictions. The combination of four methods allowed identification of mislabeled cell types in test data sets. Minjie Lyu, Yihan Zhang 0003, Luning Yang, Huan Jin, Anthony Bellotti, Nenad S. Mitic, Vladimir Brusic |
BIBM | 6 |
| 2021 | Focus on Local: Detecting Lane Marker From Bottom Up via Key PointabstractMainstream lane marker detection methods are implemented by predicting the overall structure and deriving parametric curves through post-processing. Complex lane line shapes require high-dimensional output of CNNs to model global structures, which further increases the demand for model capacity and training data. In contrast, the locality of a lane marker has finite geometric variations and spatial coverage. We propose a novel lane marker detection solution, FOLOLane, that focuses on modeling local patterns and achieving prediction of global structures in a bottom-up manner. Specifically, the CNN models low-complexity local patterns with two separate heads, the first one predicts the existence of key points, and the second refines the location of key points in the local range and correlates key points of the same lane line. The locality of the task is consistent with the limited FOV of the feature in CNN, which in turn leads to more stable training and better generalization. In addition, an efficiency-oriented decoding algorithm was proposed as well as a greedy one, which achieving 36% runtime gains at the cost of negligible performance degradation. Both of the two decoders integrated local information into the global geometry of lane markers. In the absence of a complex network architecture design, the proposed method greatly outperforms all existing methods on public datasets while achieving the best state-of-the-art results and real-time processing simultaneously. Huan Jin, Zhen Yang 0008, Wei Zhang 0196 |
CVPR | 2 |
| 2019 | Moiety modeling framework for deriving moiety abundances from mass spectrometry measured isotopologuesabstractAbstract Background Stable isotope tracing can follow individual atoms through metabolic transformations through the detection of the incorporation of stable isotope within metabolites. This resulting data can be interpreted in terms related to metabolic flux. However, detection of a stable isotope in metabolites by mass spectrometry produces a profile of isotopologue peaks that requires deconvolution to ascertain the localization of isotope incorporation. Results To aid the interpretation of the mass spectroscopy isotopologue profile, we have developed a moiety modeling framework for deconvoluting metabolite isotopologue profiles involving single and multiple isotope tracers. This moiety modeling framework provides facilities for moiety model representation, moiety model optimization, and moiety model selection. The moiety_modeling package was developed from the idea of metabolite decomposition into moiety units based on metabolic transformations, i.e. a moiety model. The SAGA-optimize package, solving a boundary-value inverse problem through a combined simulated annealing and genetic algorithm, was developed for model optimization. Additional optimization methods from the Python scipy library are utilized as well. Several forms of the Akaike information criterion and Bayesian information criterion are provided for selecting between moiety models. Moiety models and associated isotopologue data are defined in a JSONized format. By testing the moiety modeling framework on the timecourses of 13 C isotopologue data for uridine diphosphate N-acetyl-D-glucosamine (UDP-GlcNAc) in human prostate cancer LnCaP-LN3 cells, we were able to confirm its robust performance in isotopologue deconvolution and moiety model selection. Conclusions SAGA-optimize is a useful Python package for solving boundary-value inverse problems, and the moiety_modeling package is an easy-to-use tool for mass spectroscopy isotopologue profile deconvolution involving single and multiple isotope tracers. Both packages are freely available on GitHub and via the Python Package Index. Huan Jin, Hunter N. B. Moseley |
BMC Bioinform. | 1 |
| 2019 | Team orienteering with uncertain rewards and service times with an application to phlebotomist intrahospital routingabstractAbstract This study focuses on the intrahospital routing of phlebotomists at the University of Iowa Hospitals and Clinics (UIHC). Phlebotomists are responsible for drawing specimens from patients based on doctors' orders. The results of the analysis of these specimens play an important role in determining patient treatment. However, the demand for phlebotomists is likely to outpace supply over the next few years. Therefore, it is important to improve the efficiency of phlebotomists. In this paper, we formulate the phlebotomist intrahospital routing problem as a team orienteering problem with stochastic rewards and service times. The rewards and service times are particularly interesting as they are the result of a queueing process. We present an a priori solution approach and derive a method for efficiently sampling the value of a solution, a value that cannot be determined analytically. Finally, we demonstrate that our proposed approach outperforms the current practice at UIHC. Huan Jin, Barrett W. Thomas |
Networks | 1 |
| 2009 | People location and orientation tracking in multiple viewsabstractThis paper presents a multi-view approach to the tracking of people location and orientation. To achieve efficient and accurate likelihood evaluation, a novel likelihood computation method is proposed. Mixtures of Gaussian (MoG) are used to represent the color models of subjects. The scaled unscented transformation is used to project the MoG color models onto the image plane to predict the color distribution for a motion sample. The efficacy of the proposed approach is demonstrated by experiment results obtained using real videos. Huan Jin, Gang Qian |
ICASSP | 1 |
| 2009 | C-MAC: a MAC protocol supporting cooperation in wireless LANsabstractCooperative diversity is a transmission technique, where multiple terminals forms a virtual antenna array that realizes spatial diversity gain in a distributed fashion. The concept of cooperation has already been introduced to MAC layer to design MAC protocol. However, it's much different with that at physical layer. In this paper, we present a new MAC protocol based on IEEE 802.11, called C-MAC, that can support the basic building block of cooperative system. That is, in C-MAC, source would invite a relay node into data transmission if there exits an available one. During data transmission, source sends the signal to destination at first. The relay node will retransmit the overheard information to the destination at the second time slot. The destination combines two signals from source and helper, thus creating spatial diversity and robustness against channel fading. The C-MAC is backward compatible with legacy IEEE 802.11 system. The performance of C-MAC mainly depends on the physical layer's performance as it just provides the support for cooperation at MAC layer. If the physical layer works well, C-MAC would outperform IEEE 802.11 considering packet error rate. We also do the simulation using ns-2 with assumptive physical parameters. The result shows C-MAC would outperform 802.11 if packet error rate is a little high, and C-MAC would lead to some unfairness to nodes without relay. Huan Jin, Xinbing Wang, Hui Yu 0002, Youyun Xu, Yunfeng Guan 0001, Xinbo Gao 0001 |
WCNC | 1 |
| 2008 | Real-Time Multi-view Object Tracking in Mediated Environments
Huan Jin, Gang Qian, David Birchfield |
MMM | 1 |
| 2007 | Robust Multi-Camera 3D People Tracking with Partial Occlusion HandlingabstractThis paper presents an approach to robust 3D people tracking using multiple synchronized and calibrated cameras. The goal is to improve people tracking accuracy when the subjects being tracked partially occlude each other in some of the camera views. To achieve this goal, Monte Carlo fine-tuning is deployed to rectify 3D people locations obtained from partially occluded image observations. In our approach, Gaussian mixture models and axis-parallel ellipsoids are used to represent the appearance and the 3D body structures of the subjects, respectively. Related parameters are learned off-line. Experimental results obtained using real videos illustrate that the proposed approach is capable of accurate and robust 3D people tracking under partial or complete occlusions. Huan Jin, Gang Qian |
ICASSP (1) | 1 |
| 2003 | A Generic Visualisation and Editing Tool for Hierarchical and Object-Oriented SystemsabstractWe present a data retrieval and information visualisation tool, named the object-oriented graph editor (OOGE). The OOGE is designed to represent information of special kinds of system, hierarchical systems or object-oriented systems. These systems can be represented by object-oriented graph, in which nodes have attributes and their connections are constraint-based. The editor presented provides a friendly graphical user interface for users to define any class of nodes of the OO systems and to draw the OO graph based upon the user-defined node classes. The validity of the proposed system can be checked against the set of constraints that are based on the attributes of the user-defined classes. Huan Jin, Carsten Maple |
IV | 1 |