EDBT 2026 Demo / reviewers in the wild / expert
Li Kuang
dblp:86/3715
· DBLP profile ↗
102ranked-venue papers
19as first author
59since 2021 · last 2026
0000-0003-4975-034XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 27 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 15 since 2021Computer networks · 15 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Security and privacy · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating-Filtering-Ranking: A Three-Stage MultiModal Data Augmentation Framework Under Partial Modality MissingabstractMultimodal data significantly improves the performance of pretrained models, but its practical application is often limited by missing or incomplete data across modalities. There are two key challenges that existing methods of synthesizing missing data face: (1) semantic inaccuracies due to model hallucinations and (2) discrepancies in distribution preferences between generated and original data. To address these challenges, we propose a novel three-stage multimodal data augmentation framework (GFR), which Generate, Filter, and Rank missing modality data. Our framework leverages multimodal large models for diverse data generation, designs a scene graph matching-based filtering algorithm to ensure semantic consistency, and constructs a preference-aware ranking model to align the generated data with both the original distribution and task relevance. Our framework not only enhances semantic diversity and consistency in data generation but also effectively captures the implicit characteristics of the original dataset and the target model. We demonstrate the effectiveness of GFR across multiple datasets by testing different missing types and missing ratios. Zhirui Kuai, Mingjing Huang, Ning Gui, Li Kuang |
AAAI | 7 |
| 2026 | LC3: Long Cross-Language Code Clone Detection Enhanced by Opcode Sequences and Affinity AggregationabstractCross-language code clone detection, which identifies functionally similar code across programming languages, is critical for ensuring synchronized evolution and reducing maintenance costs in multi-platform software development. While zero-shot approaches have emerged as a practical solution to data scarcity, state-of-the-art methods still face two major limitations: an insufficiency in learning language-agnostic representations and information loss during the processing of long code. To address these challenges, we propose LC3, a novel framework for robust zero-shot cross-language code clone detection. To overcome the language-agnostic representation insufficiency, LC3 fuses source code with its underlying opcode sequences, leveraging a bimodal architecture and adversarial training to learn a language-agnostic representation. To resolve long-code information loss, LC3 introduces a semantic affinity aggregation strategy. This strategy synthesizes a robust clone score from a complete pairwise similarity matrix computed between segmented code blocks, overcoming the limitations of both simple truncation and aggregation. Extensive experiments show that LC3 significantly outperforms state-of-the-art zero-shot baselines, especially in challenging long-code scenarios. Xilin Lan, Chengwu Xue, Li Kuang |
AAAI | 5 |
| 2026 | VerilogLAVD: LLM-Aided Pattern Generation for Verilog CWE DetectionabstractLLMs often fail in hardware vulnerability detection due to the intrinsic semantic concurrency of HDLs (Hardware Description Language), where vulnerabilities arise from the interaction of multiple concurrent execution statements rather than a single sequential execution path.Existing LLM-based methods struggle to capture the concurrency features.To address the problem, we propose VerilogLAVD, a LLM-Aided Vulnerability Detection framework by generating executable Traversal Detection Patterns (TDPs), i.e. the rules describing how to find the evidence of vulnerabilities in Verilog HDL.We first introduce a Unified Verilog Property Graph (VeriPG) that explicitly models parallel semantics by combining AST, CFG, and DDG.Furthermore, a semantic validation mechanism is designed to constrain and filter the LLM-generated TDPs.By executing these validated TDPs on VeriPG, our method produces stable and deterministic detection results.Experiments demonstrate that VerilogLAVD improves the F1 score by 133% compared to LLM-based methods.Furthermore, the framework successfully identifies real-world hardware vulnerabilities in open-source hardware design repositories.The code and datasets of this study are available at https://github.com/ Chip-Security-Lab/VerilogLAVD Xiang Long, Yingjie Xia, Li Kuang, Yao Wan 0001 |
ACL (1) | 3 |
| 2026 | CCL-Diff: Representation-Consistent Diffusion with Intrinsic Contrastive Learning for Recommender Systems
Wanyu Ling, Shuwen Daizhou, Li Kuang, Kehua Guo |
WWW | 4 |
| 2026 | LLMUpdater: Automatic comment synchronization via edit model guided LLMs
Haiyang Yang, Qi Xie 0010, Shixiang Cai, Li Kuang, Yingjie Xia |
Empir. Softw. Eng. | 5 |
| 2026 | PV3M-YOLO: A triple attention-enhanced model for detecting pedestrians and vehicles in UAV-enabled smart transport networks
Noor Ul Ain Tahir, Li Kuang, Melikamu Liyih Sinishaw, Muhammad Asim 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Transfer learning from 2D natural images to 4D fMRI brain images via geometric mapping
Kai Gao 0011, Liang Li 0006, Yu-Wei Wang, Xue-Ying Li, Hui-Xian Li, Yi-Fan Liao, Li-Ping Cao, Guan-Mao Chen, Jian-Shan Chen, Tao-Lin Chen, Yan-Rong Chen, Yu-Qi Cheng, Zhao-Song Chu, Shi-Xian Cui, Xi-Long Cui, Zhao-Yu Deng, Qing-Lin Gao, Qi-Yong Gong, Wen-Bin Guo, Can-Can He, Zheng-Jia-Yi Hu, Xin-Lei Ji, Feng-Nan Jia, Li Kuang, Bao-Juan Li, Tao Lian, Xiao-Yun Liu, Yan-Song Liu, Zhe-Ning Liu, Yi-Cheng Long, Jian-Ping Lu, Jiang Qiu, Xiao-Xiao Shan, Tian-Mei Si, Peng-Feng Sun, Chuan-Yue Wang, Han-Lin Wang, Ying Wang 0007, Chen-Nan Wu, Xiao-Ping Wu, Xin-Ran Wu, Yan-Kun Wu, Chun-Ming Xie, Guang-Rong Xie, Xiu-Feng Xu, Zhen-Peng Xue, Jian Yang 0003, Yong-Qiang Yu, Min-Lan Yuan, Yong-Gui Yuan, Ai-Xia Zhang, Ke-Rang Zhang, Wei Zhang 0090, Zi-Jing Zhang, Jing-Ping Zhao, Jia-Jia Zhu, Xi-Nian Zuo, Hua-Ning Wang, Chaogan Yan, Yufeng Zang, Dewen Hu |
Medical Image Anal. | 30 |
| 2026 | SODM-YOLOv9: an efficient architectural refinement for small object detection in aerial imagery
Noor Ul Ain Tahir, Li Kuang, Muhammad Asim 0002, Zhe Long |
Multim. Tools Appl. | 2 |
| 2025 | CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long ContextsabstractZhaowen Wang, Xiang Wei, Kangshao Du, Yiting Zhang, Libo Qin, Yingjie Xia, Li Kuang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kangshao Du, Libo Qin 0001, Yingjie Xia, Li Kuang |
ACL (1) | 7 |
| 2025 | Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingabstractZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan, Li Kuang, Lu Wang, De Wen Soh. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Renke Shan, Li Kuang, De Wen Soh |
ACL (1) | 5 |
| 2025 | HDRec: Hierarchical Distillation for Enhanced LLM-based Recommendation SystemsabstractLarge Language Models (LLMs) have shown significant potential in recommendation systems by enhancing the semantic reasoning capabilities derived from user-item interactions. However, existing methods often rely on original reviews as ground truth explanations, with limited attention to uncovering the underlying rationales behind each interaction, which hampers the reasoning performance of LLMs. In this paper, we propose a novel Hierarchical Distillation for Recommendation (HDRec) model that effectively specifies user and item profiles by hierarchically distilling interaction rationales from reviews using LLMs. Additionally, we introduce a review summary task that condenses distilled information, such as user preferences, personality traits, item attributes, and target audience, improving both model training and interpretability. Extensive experiments demonstrate that HDRec achieves state-of-the-art performance on three real-world datasets in both sequential and Top-N recommendation tasks. The source code for HDRec is publicly available at https://github.com/linglingl635/HDRec. Wanyu Ling, Shuwen Daizhou, Li Kuang |
ICASSP | 4 |
| 2025 | Keep Your Friends Close, and Your Enemies Farther: Distance-Aware Voxel-Wise Contrastive Learning for Semi-Supervised Multi-Organ Segmentation
Jianwei Niu 0002, Xuefeng Liu 0001, Xiaozheng Xie, Li Kuang, Bin Dai 0009 |
ICCV | 5 |
| 2025 | IS-ANED: Dual-Module Graph Learning with Hybrid Attention for Edge Anomaly Detection
Genwei Zhang, Li Kuang, Yiman Xie |
ICIC (8) | 4 |
| 2025 | DBE: Dual Branch re-Extraction for Unseen Diffusion-Generated Image DetectionabstractThe rapid development of generation models has brought potential risks, necessitating generated image detection. In the meantime, evaluating the generalization to unseen diffusion models has become the detectors’ major challenge. Existing methods attempt to train detectors using the reconstruction features obtained from diffusion reconstruction, but they use these features in a limited way, resulting in a lack of generalization ability. Therefore, we propose that there are some extra forgery clues implicit in the input image and reconstruction features, and design a Dual Branch re-Extraction (DBE) module to extract them. To obtain a more generalized feature representation, spatial and frequency features are extracted through the calculation of neighboring pixel relationships and the DWT-based high-frequency feature extractor, respectively. Extensive experiments demonstrate the superior performance of our method, with an average improvement of 5%/7% in AUROC/AP compared to recent state-of-the-art works. Shixiang Cai, Liangzhen Liu, Zhirui Kuai, Li Kuang |
ICME | 4 |
| 2025 | Causal Reasoning on Temporal Knowledge Graph with Fuzzy Logic and Graph Attention NetworkabstractCausal reasoning algorithms can evaluate the strength of causal relationships between events and play an essential role in revealing and understanding the underlying mechanisms of them. However, traditional algorithms find it difficult to handle the complex interaction effects between numerous variables, can only consider event information at a single moment, and are challenging to solve the uncertainty problem in temporal evolution. To address these challenges, this paper proposes a temporal knowledge graph causal reasoning model (TKGR) that combines fuzzy logic and graph attention networks. Graph attention networks are used to capture the complex interaction effects between numerous variables at a single moment. At the same time, a multi-head self-attention mechanism is used to aggregate temporal information from multiple time points. In response to the uncertainty problem in time series evolution, we also designed a fuzzy logic module to comprehensively consider the importance of features from different time points. Finally, this paper carried out multiple rounds of comparative experiments with traditional causal reasoning algorithms and ablation experiments on the simulation dataset of satellite reconnaissance missions. The results show that the TKGR model can effectively reason the strength of causal relationships, thus providing necessary reference for decisionmaking. Shixiang Cai, Zhirui Kuai, Li Kuang, Zhifang Liao |
ICWS | 3 |
| 2025 | EvoAPR: Enhancing Large Language Models for Automatic Program Repair with Genetic Algorithm and Dynamic LoRAabstractAutomated Program Repair (APR) aims to automate the patch generation for buggy code and is vital in software devel-opment and maintenance. While large language models (LLMs) excel in various tasks, our empirical study shows they still face challenges in APR. LLMs take infilling templates with different qualities as input, and low-quality templates may misguide LLMs in constantly generating incorrect patches. Additionally, LLMs lack project-specific knowledge and struggle to leverage bug context fully. Therefore, we propose EvoAPR, which integrates the genetic algorithm and dynamic LoRA technique to enhance LLMs for better APR. First, to generate high-quality infilling templates, buggy codes are encoded to genetic representations, and genetic operators and multi-angle evaluation are designed to produce better templates. Then, to effectively utilize bug context, a bug-context-aware LoRA fine-tuning method is proposed to fuse various bug contexts by dynamically activating certain blocks of the LoRA module according to the bug index. Finally, these high-quality infilling templates are fed the fine-tuned LLMs for patch generation. Experimental results demonstrate that EvoAPR significantly improves LLMs' performance, surpassing state-of-the-art methods. Qingyang Yan, Weihuan Min, Li Kuang, Yingjie Xia |
ICWS | 5 |
| 2025 | SecV: LLM-based Secure Verilog Generation with Clue-Guided Exploration on Hardware-CWE Knowledge GraphabstractVerilog is specified as the primary Register Transfer Level (RTL) hardware description language, which designs the logical functions between registers for digital circuit systems. Recently, there emerges much cutting-edge research in leveraging Large Language Models (LLMs) to generate Verilog, aiming at effectively reducing errors and costs in the logic design of chips. However, these works mainly focus on logical correctness or PPA (Power, Performance, Area) measurement of the generated results, while neglecting the security problems in Verilog. In this study, we propose SecV, a novel and unified framework to generate secure Verilog by clue-guided exploration on Common Weakness Enumeration (CWE) knowledge graph (KG) for chips. First, the builder of the KG utilizes the instance-adapted chain of thought (COT) to extract entities and their relationships from raw Hardware-CWE corpora. Then, a fine-tuned BERT model is employed to verify the Hardware-CWE KG and collaborate with builder iteratively to achieve the precise KG. Based on Hardware-CWE KG, a clue-guided graph exploration paradigm is designed to facilitate collaborative inference of knowledge to generate secure Verilog by LLMs. Experiments demonstrate that SecV achieves 82.6% secure Verilog code without specified CWE in the generated functionally correct Verilog, with superior performance of a 21.7% performance improvement compared to SOTA. Fanghao Fan, Yingjie Xia, Li Kuang |
IJCAI | 3 |
| 2025 | APIMig: A Project-Level Cross-Multi-Version API Migration Framework Based on Evolution Knowledge GraphabstractAPI migration is essential for software maintenance due to the rapid evolution of third-party libraries where API elements may change continuously through updates. There are two main challenges for API migration at the project level, especially across multiple versions: 1) lack of specific library evolution knowledge across multi-version; 2) difficulty in identifying the chain of changes at the project level. This paper proposes a project-level cross-multi-version API migration framework APIMig. We first construct an API evolution knowledge graph (KG) to capture changes between adjacent library versions and then derive coherent cross-version API evolution knowledge by KG reasoning. Second, we design a chain exploration algorithm to track the chain of changes and aggregate the affected code segments. Finally, a large language model is employed in completing API migration by providing the API evolution knowledge and the chain of changes. We construct an evolution KG for the Lucene library from version 4.0.0 to 10.1.0 and evaluate our approach through project migration pairs that depend on different major versions. Our framework shows improvements over the baseline in migrating projects across 7 major versions, achieving average increases of 16.52% in CodeBLEU scores and 28.49% in VCEU scores in GPT-4o. Li Kuang, Qi Xie 0010, Haiyang Yang, HaoYue Kang, Yingjie Xia |
IJCAI | 1 |
| 2025 | DLCoG: A Novel Framework for Dual-Level Code Comment Generation Based on Semantic Segmentation and In-Context LearningabstractIn large software projects with collaborative development, comprehensive code comments are crucial for code readability and maintainability. Code comments mainly include method comments and inline comments, where the former describes the functionality globally, and the latter describes the implementation details locally. Existing methods typically generate these two kinds of comments with specific locations independently, which results in weak correlations between comments and code context, as well as high model inference costs due to long token inputs. To address these issues, we define the combination of inline comments and method comments as Dual-Level Code Comments. We formulate the novel task of automatically generate dual-level code comments based on given code and propose an approach named DLCoG (DualLevel Code Comment Generation) to automate this task. First, a Semantic Segmentation and Identification multi-task model based on CodeBERT, termed Se2Iden (Semantic Segmentation and Identification model), is proposed to identify code segments requiring inline comments. Next, we retrieve similar samples to adopting the in-context learning paradigm, which can enhance the generation quality of large language models (LLMs) in specific domains. Finally, the LLM is guided to generate duallevel code comments using Chain-of-Thought (CoT) prompts that first produce inline comments, followed by method comments. We manually constructed a high-quality clean Java dataset consisting of*> based on open-source Java projects by (i) determining comments type and (ii) manually associating inline comments with their corresponding code. Then, we trained a multi-task learning model based on CodeBERT to automatically take the two steps needed, termed ICSA (Inline Comment Classification and Scope Association), thus to expand to a dataset containing 80k dual-level code comments. Experimental results on clean and extended datasets show that DLCoG outperforms all baselines by substantial margins. The contextual information provided by DLCoG can effectively improve the inline comments generated by LLM. Coordinated generation of dual-level comment also brings effective improvements to method comments, which is particularly significant when there are few contextual examples. Our work fills the long-standing gap in the dual-level code comment generation field, and can provide insights for future research in this direction. We provide open-source datasets and source code for future research. Haiyang Yang, Qingyang Yan, Weihuan Min, Zhao Wei, Li Kuang, Yingjie Xia |
ICPC | 7 |
| 2025 | SeMi: When Imbalanced Semi-Supervised Learning Meets Mining Hard ExamplesabstractSemi-Supervised Learning (SSL) can leverage abundant unlabeled data to boost model performance. However, the class-imbalanced data distribution in real-world scenarios poses great challenges to SSL, resulting in performance degradation. Existing class-imbalanced semi-supervised learning (CISSL) methods mainly focus on rebalancing datasets but ignore the potential of using hard examples to enhance performance, making it difficult to fully harness the power of unlabeled data even with sophisticated algorithms. To address this issue, we propose a method that enhances the performance of Imbalanced Semi-Supervised Learning by Mining Hard Examples (SeMi). This method distinguishes the entropy differences among logits of hard and easy examples, thereby identifying hard examples and increasing the utility of unlabeled data, better addressing the imbalance problem in CISSL. In addition, we maintain a class-balanced memory bank with confidence decay for storing high-confidence embeddings to enhance the pseudo-labels' reliability. Although our method is simple, it is effective and seamlessly integrates with existing approaches. We perform comprehensive experiments on standard CISSL benchmarks and experimentally demonstrate that our proposed SeMi outperforms existing state-of-the-art methods on multiple benchmarks, especially in reversed scenarios, where our best result shows approximately a 54.8% improvement over the baseline methods. Our code is available at https://github.com/pywin/SeMi. Yin Wang 0004, Hao Lu 0009, Zhen Qin 0004, Hailiang Zhao, Guanjie Cheng, Xin Du 0002, Ge Su, Li Kuang, MengChu Zhou, Shuiguang Deng |
ACM Multimedia | 9 |
| 2025 | FCLLM-DT: Enpowering Federated Continual Learning With Large Language Models for Digital-Twin-Based Industrial IoTabstractThe Industrial Internet of Things (IIoT) represents a sophisticated technology designed to enhance production management and predict output in industrial settings, including machinery fault diagnostics. The precision of fault diagnosis is contingent upon the training efficacy of diagnostic models and their interoperability with models from other industrial facilities. Nonetheless, several critical challenges persist in maintaining these diagnostic models: 1) machinery sensors may generate abnormal data, resulting in suboptimal quality in model training; 2) sensor malfunctions may lead to interruptions in continuous data flow, thus impeding model training; and 3) collaborative interactions with other factories aiming at improving model performance may pose risks of privacy breaches. In this study, we introduce the FCLLM-DT scheme, which integrates the digital twin (DT) methodology to create a physical model of bearing for fixing abnormal sensor data. Additionally, retrieval-augmented generation (RAG)-assisted large language models (LLMs) are utilized to generate virtual datasets in instances of sensor failure. Moreover, for IIoT applications across distributed industrial environments, federated continual learning (FCL) is employed to enhance global model training by aggregating localized models from diverse facilities, thereby improving the accuracy of bearing fault diagnosis while safeguarding data privacy. The experiments on the accuracy of DT for abnormal data fix, RAG-assisted LLM for virtual data generation, and FCL for bearing fault diagnosis are conducted in comparison with three alternative methods across two datasets. The results indicate that our proposed scheme surpasses existing methods in both the enhancement of sensing data quality and the accuracy of bearing fault diagnosis. Yingjie Xia, Yunxiao Zhao, Li Kuang, Xuejiao Liu 0002, Ji Hu 0002, Zhiquan Liu 0001 |
IEEE Internet Things J. | 4 |
| 2025 | ChatDL: An LLM-Based Defect Localization Approach for Software in IIoT Flexible ManufacturingabstractWith the rapid advancement of flexible manufacturing in the Industrial Internet of Things (IIoT), there has been a significant increase in the number of IIoT devices and application software aimed at meeting various needs. The software defects may lead to delays or crashes in flexible manufacturing system, thereby affecting the production schedule. Automated software defect localization based on code changes can significantly reduce development and maintenance time costs, thereby maintaining the competitive edge of flexible manufacturing in the IIoT. Current efforts in software defect localization are primarily based on deep learning models or information retrieval models. This article investigates the performance of large language models (LLMs) in software defect localization and optimizes localization accuracy by combining it with an information retrieval model. Our empirical study reveals that GPT, given a software defect description, is unable to determine whether specific code changes are relevant. The model is unable to provide accurate answers, which aligns with the generative nature of LLMs where responses are generated according to probability distributions. However, the combined framework of LLMs and information retrieval models proposed in this article outperforms the current state-of-the-art models on public datasets. We conclude that LLMs can enhance localization performance when used as side information in conjunction with existing information retrieval models. The effectiveness of the framework has been validated through experiments conducted on publicly available datasets and in practical applications within IIoT projects. This offers valuable insights into the application and development of LLMs for defect localization in the software development and maintenance processes in the IIoT flexible manufacturing. Haiyang Yang, Yulu Zhou, Li Kuang |
IEEE Internet Things J. | 4 |
| 2025 | Edge intelligence in wireless networks
Xiaoxian Yang, Li Kuang |
Wirel. Networks | 2 |
| 2024 | RegGPT: A Tool for Cross-Domain Service Regulation Language ConversionabstractDigital services have become essential to the modern service industry, offering great convenience to consumers. Because of the lack of regulation, it has posed an unprecedented challenge to the regulation of digital services. With the development of Regulatory Technology, many regulatory platforms have emerged. However, most of these platforms focus on a single service domain and are difficult to migrate to other domains to meet cross-domain regulatory demands. By extracting common elements from multi-domain regulatory rules, we propose a regulatory language called Cross-Domain Service Regulation Language (CDSRL), which aims to improve the comprehension of rules for machines, so the automated regulation can be achieved. Meanwhile, we construct the fine-tuned datasets for the regulatory domain and train the RegGPT based on the large language model, which can identify and classify natural language rules and automatically convert them into CDSRL. Experiments show that the language is highly comprehensible, scalable, and suitable for expressing rules in different domains and categories. The RegGPT shows a strong ability in the conversion process and improves regulatory efficiency. It provides a new scheme for automatically converting regulatory rule into rule language. Qi Xie 0010, Weihuan Min, Li Kuang |
ICWS | 5 |
| 2024 | Towards Cross-Domain Multimodal Automated Service Regulation SystemsabstractThe proliferation of modern digital services has been rapid, prompting the impracticality and inefficiency of developing individual regulatory systems for each domain. The existing regulatory frameworks are notably deficient in providing a succinct abstraction of the commonalities over cross-domain service regulations, thus encountering three primary challenges: the absence of a unified language to represent cross-domain regulation rules, the deficiency of regulation models capable of generalizing over cross-domain regulation rules, and the lack of a reliable mechanism for detecting regulation violations over multi-modal runtime data. This paper introduces an innovative framework, the Automated Service Regulation System (ASRS), which comprises four pivotal components: Representation for Regulation Rules, Data Governance for ASRS, AI for ASRS, and ASRS & Evaluation. A case study in the domain of Urban Management is conducted in accordance with the framework proposed in this work, demonstrating that the urban management regulation process can be effectively streamlined and yield superior results in alignment with the four components of our proposed framework. We contend that our framework amalgamates the hitherto fragmented research on automating service regulation, thereby elucidating the path toward a unified ASRS system and furnishing a resilient and adaptable solution to the dynamic challenges of service regulation. Jianwei Yin, Li Kuang |
ICWS | 3 |
| 2024 | Improving AST-Level Code Completion with Graph Retrieval and Multi-Field AttentionabstractCode completion, which provides code suggestions by generating code snippets or structures, has become an essential feature of integrated development environments (IDEs). Recently, some studies have begun to use graph neural networks to complete AST-level code, and shown that it is promising to introduce GNNs into ASTlevel completion. However, these methods do not fully exploit the potential of reference codes with similar structures nor solve out-of-vocabulary (OOV). We propose Retrieval-Assisted Graph Code Completion (ReGCC) to enhance AST-level code completion further. ReGCC integrates a retrieval model that searches for similar code graphs to generate graph nodes and a completion model that leverages information from multiple domains. The key component of both the retrieval and completion models is the Multi-field Graph Attention Block, which consists of three layers of stacked attention: (1) Neighborhood Attention: preserves the heterogeneity and local dependency of the graph, enabling nodes to exchange information within their neighborhood. (2) Global & Memory Attention: addresses the long-distance dependency problem by providing nodes with a global view and the ability to extract information from the memory domain. (3) Reference Attention: lets nodes obtain valuable information from structurally similar reference code graphs. Furthermore, we tackle the OOV issue by employing feature matching and copying values from existing nodes. Specifically, we predict edges between nodes beyond the vocabulary, enabling effective information transfer. Experimental results demonstrate the superiority of our approach over state-of-the-art AST-level completion methods and generative language models. Yu Xia 0010, Weihuan Min, Li Kuang |
ICPC | 4 |
| 2024 | ASKDetector: An AST-Semantic and Key Features Fusion based Code Comment Mismatch DetectorabstractCode comments are essential for programming comprehension. Nevertheless, developers often neglect to update comments after modifying the source code. Wrong code comments may lead to bugs in the maintenance process, thus affecting the reliability of the software. So, timely comment mismatch detection is crucial for software development and maintenance. However, existing works have the following two limitations: 1) the lack of use of code structural and sequential information, and 2) the ignorance of existing associations between code and comments. In this paper, we propose a new model called ASKDetector (AST-Semantic and Key features fusion based mismatch Detector). For the first limitation, we encode code with an attention-based preorder traversal abstract syntax tree sequence to obtain both order and structural information. And CodeBERT is utilized to capture contextual semantic features further. For the second one, we encode extracted association information between the code snippets and comments to reduce the semantic gap. The correlations between the encoders are learned through a fusion layer and a multi-layer perceptron. The experimental results prove that our detector outperforms the state-of-the-art model in evaluation metrics, where our F1 and accuracy exceed an average of 3.4%. Haiyang Yang, Hao Chen 0116, Zhirui Kuai, Shuyuan Tu, Li Kuang |
ICPC | 5 |
| 2024 | A Just-in-time Software Defect Localization Method based on Code Graph RepresentationabstractTraditional software defect localization aims to locate defective files, methods, or code lines based on symptoms such as defect reports. In comparison, Just-In-Time (JIT) software defect localization focuses on identifying defective code lines when a defective code change is initially submitted. It can identify issues at the code line level before the defect becomes apparent, preventing it from adversely affecting the software. Although researchers have proposed various methods for JIT defect localization, existing methods still have the following shortcomings: (1) Most methods rely heavily on tokens from single code lines to calculate naturalness for defect localization, which makes it challenging to effectively distinguish between code lines that have the same content but different labels (defective code lines or non-defective code lines) - termed Duplicate Lines with Different Labels (DLDL). (2) Existing methods represent code in the form of sequences, neglecting the structural information of the code. Therefore, we propose a JIT defect localization method based on code graph representation. First, we construct code linelevel code graphs for code changes to distinguish DLDL explicitly. Next, to extract sequential and structural information from the code, we propose a code graph representation model with contrastive learning to generate graph feature vectors and node scores with rich semantics. Finally, we calculate the naturalness of code lines based on the graph feature vectors and node scores. Using this naturalness, we identify defective code lines. Experimental results show that our JIT defect localization method outperforms the state-of-the-art methods. Huan Zhang 0017, Weihuan Min, Zhao Wei, Li Kuang, Honghao Gao, Huaikou Miao |
ICPC | 4 |
| 2024 | Multi-Source Augmentation and Composite Prompts for Visual Recognition with Missing ModalityabstractIn multimodal learning for visual recognition, missing modality is a common issue that can significantly impact the performance and robustness of vision-language models. Most existing approaches have only considered the situation where a single modality-either image or text-is missing and then use a data augmentation method to recover the missing modality data. However, in reality, it is common for either text or image to be missing, and in such cases, a data augmentation method that is effective for one modality might not be suitable for the other, thereby necessitating distinct methods for text and image data augmentation. There are also approaches aimed at enhancing the robustness of vision-language models to handle missing data inputs. However since most of these approaches often involve significant modifications to complex model structures and require extensive retraining, these solutions would be impractical with limited computational resources. To address the abovementioned limitations, we develop a Multi-source Augmentation and Composite Prompts method (MACP) to alleviate the performance degradation due to missing modalities from both data and model levels. On the data level, we designed a multi-source data augmentation framework that integrates different data augmentation methods and a data selector to restore the missing data for each image-text sample as well as possible. On the model level, we designed a method for generating prompt vectors that simultaneously indicate the missing modalities in the model input and the source of augmentation data. The prompts will enhance the ability of the vision-language model to handle different input types in low-resource situations by applying prompt tuning. Experimental results demonstrate the effectiveness of our approach in mitigating the impact of modality missing on three vision-language datasets. Code is available. Zhirui Kuai, Yulu Zhou, Qi Xie 0010, Li Kuang |
ICMR | 4 |
| 2024 | rPPG-HiBa: Hierarchical Balanced Framework for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) is a promising technique for non-contact physiological signal measurement. It has great potential applications in human health monitoring and emotion analysis. However, existing methods for the rPPG task ignore the long-tail phenomenon of physiological signal data, especially on multi-domain joint training. In addition, we find that the long-tail problem of the physiological label (phys-label) exists in different datasets, and the long-tail problem of some domain exists under the same phys-label. To tackle these problems, we propose a hierarchical balanced framework, to mitigate the bias caused by domain and phys-label imbalance. Specifically, we propose anti-spurious domain center learning tailored to learning domain-balanced embeddings space. Then, we adopt compact-aware continuity regularization to estimate phys-label-wise imbalances and construct continuity between embeddings. Extensive experiments demonstrate that our method outperforms the state-of-the-art in cross-dataset and intra-dataset settings. Our code is available at https://github.com/pywin/HiBa. Yin Wang 0004, Hao Lu 0009, Ying-Cong Chen, Li Kuang, MengChu Zhou, Shuiguang Deng |
ACM Multimedia | 4 |
| 2024 | Unsupervised Graph Representation Learning Beyond Aggregated ViewabstractUnsupervised graph representation learning aims to condense graph information into dense vector embeddings to support various downstream tasks. To achieve this goal, existing UGRL approaches mainly adopt the message-passing mechanism to simultaneously incorporate graph topology and node attribute with an aggregated view. However, recent research points out that this direct aggregation may lead to issues such as over-smoothing and/or topology distortion, as topology and node attribute of totally different semantics. To address this issue, this paper proposes a novel Graph Dual-view AutoEncoder framework (GDAE) which introduces the node-wise view for an individual node beyond the traditional aggregated view for aggregation of connected nodes. Specifically, the node-wise view captures the unique characteristics of individual node through a decoupling design, i.e., topology encoding by multi-steps random walk while preserving node-wise individual attribute. Meanwhile, the aggregated view aims to better capture the collective commonality among long-range nodes through an enhanced strategy, i.e., topology masking then attribute aggregation. Extensive experiments on 5 synthetic and 11 real-world benchmark datasets demonstrate that GDAE achieves the best results with up to 49.5% and 21.4% relative improvement in node degree prediction and cut-vertex detection tasks and remains top in node classification and link prediction tasks. Li Kuang, Ning Gui |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Integrating Transient and Long-term Physical States for Depression Intelligent Diagnosis
Li Kuang, Huaiqian Ye, Yuanbo He |
BMVC | 3 |
| 2023 | Origin-Destination Convolution Recurrent Network: A Novel OD Matrix Prediction Framework
Jiayu Chang, Wanzhi Xiao, Li Kuang |
CollaborateCom (2) | 4 |
| 2023 | JARAD: An Approach for Java API Mention Recognition and Disambiguation in Stack Overflow
Qingmi Liang, Qi Xie 0010, Li Kuang, Yu Sheng |
CollaborateCom (1) | 4 |
| 2023 | CTTSR: A Hybrid CNN-Transformer Network for Scene Text Image Super-ResolutionabstractThe accuracy of scene text recognition has been significantly improved, which can be attributed to the development of deep learning. However, the blurring and low-resolution text images usually lead to unsatisfactory results in text recognition. Several researchers design super-resolution models that adopt convolutional neural networks (CNNs) to relieve the image blurring, while these models are limited to the receptive field of the convolution kernel and fail to extract the long-distance semantic relations of text images enough. In this paper, we propose a CNN-Transformer Text Super Resolution Network (CTTSR) to capture the semantic features of text images by the multi-head attention mechanism of the transformer. Furthermore, we propose the text position loss to optimize the network and make the text regions of images more effectively detectable. Experimental results demonstrate that our model can improve the quality of images and outperform the existing methods in text recognition tasks. Kaiwei Dai, Nan Kang, Li Kuang |
ICASSP | 3 |
| 2023 | Enhancing intelligent IoT services development by integrated multi-token code completion
Yu Xia 0010, Weihuan Min, Li Kuang, Honghao Gao |
Comput. Commun. | 4 |
| 2023 | Just-in-time defect prediction enhanced by the joint method of line label fusion and file filteringabstractAbstract Just‐In‐Time (JIT) defect prediction aims to predict the defect proneness of software changes when they are initially submitted. It has become a hot topic in software defect prediction due to its timely manner and traceability. Researchers have proposed many JIT defect prediction approaches. However, these approaches cannot effectively utilise line labels representing added or removed lines and ignore the noise caused by defect‐irrelevant files. Therefore, a JIT defect prediction model enhanced by the joint method of line label Fusion and file Filtering (JIT‐FF) is proposed. Firstly, to distinguish added and removed lines while preserving the original software changes information, the authors represent the code changes as original, added, and removed codes according to line labels. Secondly, to obtain semantics‐enhanced code representation, a cross‐attention‐based line label fusion method to perform complementary feature enhancement is proposed. Thirdly, to generate code changes containing fewer defect‐irrelevant files, the authors formalise the file filtering as a sequential decision problem and propose a reinforcement learning‐based file filtering method. Finally, based on generated code changes, CodeBERT‐based commit representation and multi‐layer perceptron‐based defect prediction are performed to identify the defective software changes. The experiments demonstrate that JIT‐FF can predict defective software changes more effectively. Huan Zhang 0017, Li Kuang, Aolang Wu, Qiuming Zhao, Xiaoxian Yang |
IET Softw. | 2 |
| 2023 | Deep learning to mobile hypermedia and multimedia
Xiaoxian Yang, Li Kuang |
Wirel. Networks | 2 |
| 2022 | Retrieve-Guided Commit Message Generation with Semantic Similarity And DisparityabstractHigh quality commit messages are important for program understanding and maintenance, which describes the content of code changes. Neural-based methods are most popular ways to generate commit messages, but they neglect retrieval results including retrieval diffs and retrieval messages. The existing combination models of neural-based methods and retrievalbased methods have two major limitations: a) Only use retrieval diffs but ignore retrieval messages. b) Seldom consider the similarity and disparity between the retrieval results and given diff. To address the above two issues, we propose a retrieveguided method named ReGenSD to generate commit messages, which consists of three steps. Firstly, we apply a similaritybased IR technique to get retrieval diff and retrieval messages. Secondly, we introduce a selective mechanism to decide whether to use retrieve-guided model based on lexical similarity between retrieval diff and given diff. Lastly, for retrieve-guided model, we design a novel seq2seq network with Bi-LSTM that takes given diff, retrieval diff and retrieval message as input. We introduce a relation gate in encoder to leverage retrieval message adaptively based on semantic similarity, and a difference vector in decoder to refine the utilization of retrieval message based on semantic disparity. Experimental results on an open source dataset demonstrate that retrieval messages guidance can facilitate commit message generation task. Besides, ablation experiments prove the effectiveness of our proposed mechanisms on adjusting the use of retrieval results. Haiyang Yang, Li Kuang |
APSEC | 4 |
| 2022 | Combining Global and Local Representations of Source Code for Method NamingabstractCode is a kind of complex data. Recent models learn code representation using global or local aggregation. Global encoding allows all tokens of code to be connected directly and neglects the graph structure. Local encoding focuses on the neighbor nodes when capturing the graph structure but fails to capture long dependencies. In this work, we gather both encoding strategies and investigate different models that combine both global and local representations of code in order to learn code representation better. Specifically, we modify the layer structure based on the sequence-to-sequence model to incorporate a structured model in the encoder and decoder parts, respectively. To further consider different integration ways, we propose four models for method naming. In an extensive evaluation, we demonstrate that our models have a significant improvement on a well-studied dataset of method naming, achieving ROUGE-1 score of 54.1, ROUGE-2 score of 26.7, and ROUGE-L score of 54.3, outperforming state-of-the-art models by 2.7, 1.7, and 4.3 points, respectively. Our data and code are available at https://github.com/zc-work/CGLNaming. Li Kuang |
ICECCS | 2 |
| 2022 | MisuseHint: A Service for API Misuse Detection Based on Building Knowledge Graph from Documentation and CodebaseabstractDevelopers often call APIs to improve development efficiency, but they misuse APIs due to lack of understanding of source code logic and other unavoidable reasons, resulting in serious consequences such as program crashes. Many studies that extract API usage constraints from API documentation or codebases expect to get out of this dilemma through API misuse detection. However, low recall remains a hurdle for researchers to overcome. In this work, we make full use of API documentation and codebases to construct constraint knowledge graph, and propose a new API misuse detector, MisuseHint. We precisely define API constraints into seven categories, utilize API caveat knowledge in documentation and API usage patterns in codebases, and fuse knowledge from both to build knowledge graph with rich constraints. To detect API misuses, we obtain API usage constraints in the knowledge graph and analyze static code to propose different strategies to determine whether API misuses exist. Through defect pattern analysis, object variable tracking, and Z3 SAT solver, our detector can identify various complex situations of code at a fine-grained level, especially solving various complex problems of Call Order and State Checking constraints. Experimental results on MUBench show that our recall reaches 39.78%, demonstrating the validity and theoretical feasibility of fusing documentation and codebases using knowledge graphs. MisuseHint achieves a recall of 76.34% when it is always given sufficient API constraints. This detector can practically help developers program effectively. Qingmi Liang, Zhirui Kuai, Yangqi Zhang, Li Kuang |
ICWS | 5 |
| 2022 | CSRS: code search with relevance matching and semantic matchingabstractDevelopers often search and reuse existing code snippets in the process of software development. Code search aims to retrieve relevant code snippets from a codebase according to natural language queries entered by the developer. Up to now, researchers have already proposed information retrieval (IR) based methods and deep learning (DL) based methods. The IR-based methods focus on keyword matching, that is to rank codes by relevance between queries and code snippets, while DL-based methods focus on capturing the semantic correlations. However, the existing methods do not consider capturing two matching signals simultaneously. Therefore, in this paper, we propose CSRS, a code search model with relevance matching and semantic matching. CSRS comprises (1) an embedding module containing convolution kernels of different sizes which can extract n-gram embeddings of queries and codes, (2) a relevance matching module that measures lexical matching signals, and (3) a co-attention based semantic matching module to capture the semantic correlation. We train and evaluate CSRS on a dataset with 18.22M and 10k code snippets. The experimental results demonstrate that CSRS achieves an MRR of 0.614, which outperforms two state-of-the-art models DeepCS and CARLCS-CNN by 33.77% and 18.53% respectively. In addition, we also conducted several experiments to prove the effectiveness of each component of CSRS. Li Kuang |
ICPC | 2 |
| 2022 | Multiple Biological Granularities Network for Person Re-IdentificationabstractThe task of person re-identification is to retrieve images of a specific pedestrian among cross-camera person gallery captured in the wild. Previous approaches commonly concentrate on the whole person images and local pre-defined body parts, which are ineffective with diversity of person poses and occlusion. In order to alleviate the problem, researchers began to implement attention mechanisms to their model using local convolutions with limited fields. However, previous attention mechanisms focus on the local feature representations ignoring the exploration of global spatial relation knowledge. The global spatial relation knowledge contains clustering-like topological information which is helpful for overcoming the situation of diversity of person poses and occlusion. In this paper, we propose the Multiple Biological Granularities Network (MBGN) based on Global Spatial Relation Pixel Attention (GSRPA) taking the human body structure and global spatial relation pixels information into account. First, we design an adaptive adjustment algorithm (AABS) based on human body structure, which is complementary to our MBGN. Second, we propose a feature fusion strategy taking multiple biological granularities into account. Our strategy forces the model to learn diversity of person poses by balancing the local semantic human body parts and global spatial relations. Third, we propose the attention mechanism GSRPA. GSRPA enhances the weight of spatial relational pixels, which digs out the person topological information for overcoming occlusion problem. Extensive evaluations on the popular datasets Market-1501 and CUHK03 demonstrate the superiority of MBGN over the state-of-the-art methods. Shuyuan Tu, Tianzhen Guan, Li Kuang |
ICMR | 3 |
| 2022 | CGMBL: Combining GAN and Method Name for Bug LocalizationabstractDevelopers often need to locate buggy code files in the software quality maintenance process. Bug localization aims to automatically identify potentially buggy source code files from the project codes for developers based on the bug reports. Up to now, researchers have proposed many methods to advance this task. However, the early studies only focus on the accuracy of capturing text features or the efficiency of calculating relevance scores, which do not consider the semantic gap between bug reports in natural language and codes in programming language. In this paper, we propose a novel adversarial learning model to bridge the semantic gap. Due to the different characteristics of natural language and programming language, we propose two different representation models for bug reports and code files respectively, and regards the two representation models as the generators. Then we construct adversarial learning by adding a discriminator to distinguish the source of representations so that the model can learn the public features of different texts. In addition, method name is the summary of the code function, and the relevant method name often appears in the bug report. We consider the method name information according to whether the method name appears in the report. Our model can dynamically integrate the information to improve the model effect. We evaluate our model on three open-source java project datasets and compare it with four state-of-the-art methods. The experimental results show that our model outperforms the baseline models and has a significant improvement in evaluation metrics. Besides, we conduct ablation experiments to explain each module’s contribution to the model. Haiyang Yang, Zilun Yan, Li Kuang |
QRS | 4 |
| 2022 | Code comment generation based on graph neural network enhanced transformer model for code understanding in open-source software ecosystems
Li Kuang, Xiaoxian Yang |
Autom. Softw. Eng. | 1 |
| 2022 | Suggesting method names based on graph neural network with salient information modellingabstractAbstract Descriptive method names have a great impact on improving program readability and facilitating software maintenance. Recently, due to high similarity between the task of method naming and text summarization, large amount of research based on natural language processing has been conducted to generate method names. However, method names are much shorter compared to long source code sequences. The salient information of the whole code snippet account for an relatively small part. Additionally, unlike natural language, source code has complicated structure information. Thus, modelling the salient information from highly structured input presents a great challenge. To tackle this problem, we propose a graph neural network (GNN)‐based model with a novel salient information selection layer. Specifically, to comprehensively encode the tokens of the source code, we employ a GNN‐based encoder, which can be directly applied to the code graph to ensure that the syntactic information of code structure and semantic information of code sequence can be modelled sufficiently. To effectively discriminate the salient information, we introduce an information selection layer which contains two parts: a global filter gate used to filter irrelevant information, and a semantic‐aware convolutional layer used to focus on the semantic information contained in code sequence. To improve the precision of the copy mechanism when decoding, we introduce a salient feature enhanced attention mechanism to facilitate the accuracy of copying tokens from input. Experimental results on an open source dataset indicate that our proposed model, equipped with the salient information selection layer, can effectively improve method naming performance compared to other state‐of‐the‐art models. Li Kuang, Fan Ge |
Expert Syst. J. Knowl. Eng. | 1 |
| 2022 | Editorial: Collaborative Computing in AI Empowered Mobile Networks
Xiaoxian Yang, Li Kuang |
Mob. Networks Appl. | 2 |
| 2022 | An Information Fusion Approach to Intelligent Traffic Signal Control Using the Joint Methods of Multiagent Reinforcement Learning and Artificial Intelligence of ThingsabstractWith the development of communication technology and artificial intelligence of things (AIoT), transportation systems have become much smarter than ever before. However, the volume of vehicles and traffic flows have rapidly increased. Optimizing and improving urban traffic signal control is a potential way to relieve traffic congestion. In general, traffic signal control is a sequential decision process that conforms to the characteristics of reinforcement learning, in which an agent constantly interacts with its environment, thus providing strategy for optimizing behavior in accordance with feedback in response. In this paper, we propose multiagent reinforcement learning for traffic signals (MARL4TS) to support the control and deployment of traffic signals. First, information on traffic flows and multiple intersections is formalized as input environments for performing reinforcement learning. Second, we design a new reward function to continuously select the most appropriate strategy as control during multiagent learning to track actions for traffic signals. Finally, we use a supporting tool, Simulation of Urban MObility (SUMO), to simulate the proposed traffic signal control process and compare it with other methods. The experimental results show that our proposed MARL4TS method is superior to the baselines. In particular, our method can reduce vehicle delay. Xiaoxian Yang, Yueshen Xu, Li Kuang, Honghao Gao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Crowdturfing Detection in Online Review System: A Graph-Based Modeling
Qilong Feng, Li Kuang |
CollaborateCom (2) | 3 |
| 2021 | Backdoor Attack of Graph Neural Networks Based on Subgraph Trigger
Yu Sheng, Guanyu Cai, Li Kuang |
CollaborateCom (2) | 4 |
| 2021 | Traffic Flow Prediction Through the Fusion of Spatial-Temporal Data and Points of Interest
Wanzhi Xiao, Li Kuang, Ying An |
DEXA (1) | 2 |
| 2021 | CCMC: Code Completion with a Memory Mechanism and a Copy MechanismabstractCode completion tools are increasingly important when developing modern software. Recently, statistical language modeling techniques have achieved great success in the code completion task. However, two major issues with these techniques severely affect the performance of neural language models (NLMs) of code completion. a) Long-range dependences are common in program source code. b) New and rare vocabulary in code is much higher than natural language. To address the challenges above, in this paper, we propose code completion with a memory mechanism and a copy mechanism (CCMC). To capture the long-range dependencies in the program source code, we employ Transformer-XL as our base model. To utilize the locally repeated terms in program source code, we apply the pointer network into our base model and design CopyMask to improve the training efficiency, which is inspired by masked multihead attention in the transformer decoder. To combine the long-range dependency modeling ability from Transformer-XL and the ability to copy the input token to output from the pointer network, we design a memory mechanism and a copy mechanism. Through our memory mechanism, our model can uniformly manage the context used by Transformer-XL and pointer network. Through our copy mechanism, our model can either generate a within-vocabulary token or copy an out-of-vocabulary (OOV) token from inputs. Experiments on a real-world dataset demonstrate the effectiveness of our CCMC on the code completion task. Li Kuang |
EASE | 2 |
| 2021 | KG2Code: Correct Code Examples Mining Service Based on Knowledge Graph for Fixing API Misuses
Yangqi Zhang, Zhirui Kuai, Wenjin Yao, Li Kuang |
ICSOC | 5 |
| 2021 | Keywords Guided Method Name GenerationabstractHigh quality method names are descriptive and readable, which are helpful for code development and maintenance. The majority of recent research suggest method names based on the text summarization approach. They take the token sequence and abstract syntax tree of the source code as input, and generate method names through a powerful neural network based model. However, the tokens composing the method name are closely related to the entity name within its method implementation. Actually, high proportions of the tokens in method name can be found in its corresponding method implementation, which makes it possible for incorporating these common shared token information to improve the performance of method naming task. Inspired by this key observation, we propose a two-stage keywords guided method name generation approach to suggest method names. Specifically, we decompose the method naming task into two subtasks, including keywords extraction task and method name generation task. For the keywords extraction task, we apply a graph neural network based model to extract the keywords from source code. For the method name generation task, we utilize the extracted keywords to guide the method name generation model. We apply a dual selective gate in encoder to control the information flow, and a dual attention mechanism in decoder to combine the semantics of input code sequence and keywords. Experiment results on an open source dataset demonstrate that keywords guidance can facilitate method naming task, which enables our model to outperform the competitive state-of-the-art models by margins of 1.5%-3.5% in ROUGE metrics. Especially when programs share one common token with method names, our approach improves the absolute ROUGE-1 score by 7.8%. Fan Ge, Li Kuang |
ICPC | 2 |
| 2021 | Multiple Wavelet Convolutional Neural Network for Short-Term Load ForecastingabstractAlthough the accuracy of load forecasting has been studied by many works, the actual deployability of a model is rarely considered. In this work, we consider the actual deployability of a model from four aspects: 1) the prediction performance of the model; 2) the robustness of the model; 3) the dependence of the model on external data; and 4) the storage size of the model. From these four aspects, we propose a multiple wavelet convolutional neural network (MWCNN) for load forecasting. On two public data sets, we verified the performance and robustness of the MWCNN. The MWCNN only uses load data, and the storage size of the model is only 497 kB, which shows that MWCNN has good deployability. In addition, our MWCNN prediction results are interpretable. The experimental results show that the MWCNN can effectively capture the periodic characteristics of load data. Zhifang Liao, Haihui Pan, Xiaoping Fan, Yan Zhang 0047, Li Kuang |
IEEE Internet Things J. | 5 |
| 2021 | Special Issue on Deep Learning in Mobile and Wireless Networks: Algorithms, Models and Techniques
Yueshen Xu, Yuyu Yin, Li Kuang |
Mob. Networks Appl. | 3 |
| 2021 | SDABS: A Flexible and Efficient Multi-Authority Hybrid Attribute-Based Signature Scheme in Edge EnvironmentabstractThe explosive growth of the Internet of Things and modern networking technologies lay the foundation for the development of intelligent transportation systems and smart cities. To analyzing massive data under the required time for transportation issues, the edge computing paradigm is applied, which pre-processing large amounts of data at the network edge to save bandwidth and improve response time. However, data reliability and security are still facing many challenges in the edge environment. In this article, we propose a multi-authority hybrid attribute-based signature scheme (SDABS). It is composed of four phases: system initialization, signature generation, signature verification, and attribute revocation phases. To better describe frequently changing features in the transportation systems like location, the dynamic attribute is introduced in building the signature. The multi-layer policy tree is applied to support flexible and various access policies, which also naturally form user groups and help data searching. Besides, the multi-authority structure is more suitable for the distributed edge environment. We evaluate SDABS from both theoretical analysis and practical analysis. Compared with two classical signature schemes (MABS and ODMA-ABS), experimental results demonstrate that the proposed SDABS can achieve better performance at an acceptable cost in the terms of attributes and attribute authorities. Youhuizi Li, Xu Chen 0048, Yuyu Yin, Jian Wan 0001, Li Kuang, Zeyong Dong |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Intelligent Traffic Signal Control Based on Reinforcement Learning with State Reduction for Smart CitiesabstractEfficient signal control at isolated intersections is vital for relieving congestion, accidents, and environmental pollution caused by increasing numbers of vehicles. However, most of the existing studies not only ignore the constraint of the limited computing resources available at isolated intersections but also the matching degree between the signal timing and the traffic demand, leading to high complexity and reduced learning efficiency. In this article, we propose a traffic signal control method based on reinforcement learning with state reduction. First, a reinforcement learning model is established based on historical traffic flow data, and we propose a dual-objective reward function that can reduce vehicle delay and improve the matching degree between signal time allocation and traffic demand, allowing the agent to learn the optimal signal timing strategy quickly. Second, the state and action spaces of the model are preliminarily reduced by selecting a proper control phase combination; then, the state space is further reduced by eliminating rare or nonexistent states based on the historical traffic flow. Finally, a simplified Q-table is generated and used to optimize the complexity of the control algorithm. The results of simulation experiments show that our proposed control algorithm effectively improves the capacity of isolated intersections while reducing the time and space costs of the signal control algorithm. Li Kuang, Jianbo Zheng, Kemu Li, Honghao Gao |
ACM Trans. Internet Techn. | 1 |
| 2021 | Social media data mining and knowledge discovery under wireless network
Xiaoxian Yang, Li Kuang |
Wirel. Networks | 2 |
| 2020 | API Misuse Detection Based on Stacked LSTM
Shuyin OuYang, Fan Ge, Li Kuang, Yuyu Yin |
CollaborateCom (1) | 3 |
| 2020 | How to Construct Software Knowledge Graph: A Case StudyabstractKnowledge graph has been more and more widely used, but software knowledge graph is still in the initial stage. To explore the way of constructing a software knowledge graph, this paper proposes a case study by using source code, related documents and relations among the information to build software knowledge graph. The source code comes from the open source project Tomcat, the related documents come from various software communities. Different approaches are designed for each type of resources to achieve entity and attribute extraction, and to categorize the relation into internal and external relations to extract. Internal relation is contained in each resource, external relation exists between different resources. And then subtree comparison and mailbox matching are used to align entities, completing the construction of software knowledge graph. Software knowledge graph is applied to knowledge retrieval and association visualization to improve the efficiency of knowledge acquisition. Wuqian Lv, Zhifang Liao, Yan Zhang 0047, Li Kuang, Shuixiu Bi |
SERVICES | 4 |
| 2020 | COSINE: a software development model integrating collective intelligence, service and ecosystemabstractWith the development of the internet technology, a large amount of softwares have emerged to meet users' increasing needs. At the mean time, software systems have been faced with a problem that they must adapt to the dynamic network environment. It is obvious that a variety of software development models have been proposed in the past few decades. However, the majority of these methods are gradually unadaptable to new circumstances. In this paper, we proposed a new software development model integrating collective intelligence, service and ecosystem. On the one hand, we have introduced the model in detail. On the other hand, We took a practical example to demonstrate the effectiveness of the proposed model. Tianjing Hong, Jian Cao 0001, Haijun Zhang 0002, Changhai Nie, Bo Cheng 0001, Yangfan He, Li Kuang, Dun-Wei Gong, Wuhui Chen, Yuliang Shi, Deyi Huang |
SERVICES | 8 |
| 2020 | A spam worker detection approach based on heterogeneous network embedding in crowdsourcing platforms
Li Kuang, Ruyi Shi, Zhifang Liao, Xiaoxian Yang |
Comput. Networks | 1 |
| 2020 | Offloading decision methods for multiple users with structured tasks in edge computing for smart cities
Li Kuang, Shuyin OuYang, Honghao Gao, Shuiguang Deng |
Future Gener. Comput. Syst. | 1 |
| 2020 | Mining consuming Behaviors with Temporal Evolution for Personalized Recommendation in Mobile Marketing Apps
Honghao Gao, Li Kuang, Yuyu Yin, Kai Dou |
Mob. Networks Appl. | 2 |
| 2020 | Traffic Volume Prediction Based on Multi-Sources GPS Trajectory Data by Temporal Convolutional Network
Li Kuang, Chunbo Hua, Jiagui Wu, Yuyu Yin, Honghao Gao |
Mob. Networks Appl. | 1 |
| 2019 | Predicting Traffic Flow Based on Encoder-Decoder Framework
Xiaosen Zheng, Liwen Liu, Li Kuang |
CollaborateCom | 4 |
| 2019 | Feature Engineering for Deep Reinforcement Learning Based RoutingabstractRecent advances in Deep Reinforcement Learning (DRL) techniques are providing a dramatic improvement in decision-making and automated control problems. As a result, we are witnessing a growing number of research works that are proposing ways of applying DRL techniques to network-related problems such as routing. However, such proposals failed to achieve good results, often under-performing traditional routing techniques. We argue that successfully applying DRL-based techniques to networking requires finding good representations of the network parameters: feature engineering. DRL agents need to represent both the state (e.g., link utilization) and the action space (e.g., changes to the routing policy). In this paper, we show that existing approaches use straightforward representations that lead to poor performance. We propose a novel representation of the state and action that outperforms existing ones and that is flexible enough to be applied to many networking use-cases. We test our representation in two different scenarios: (i) routing in optical transport networks and (ii) QoS-aware routing in IP networks. Our results show that the DRL agent achieves significantly better performance compared to existing state/action representations. José Suárez-Varela, Albert Mestres, Junlin Yu, Li Kuang, Haoyu Feng, Pere Barlet-Ros, Albert Cabellos-Aparicio |
ICC | 4 |
| 2019 | An Outlier Degree Shilling Attack Detection Algorithm Based on Dynamic Feature SelectionabstractRecommender system is widely used in various fields for dealing with information overload effectively, and collaborative filtering plays a vital role in the system. However, recommender system suffers from its vulnerabilities by malicious attacks significantly, especially, shilling attacks because of the open nature of recommender system and the dependence on data. Therefore, detecting shilling attack has become an important issue to ensure the security of recommender system. Most of the existing methods of detecting shilling attack are based on user ratings, and one limitation is that they are likely to be interfered by obfuscation techniques. Moreover, traditional detection algorithms cannot handle different types of shilling attacks flexibly. In order to solve the problems, we proposed an outlier degree shilling attack detection algorithm by using dynamic feature selection. Considering the differences when users choose items, we combined rating-based indicators with user popularity, and utilized the information entropy to select detection indicators dynamically. Therefore, a variety of shilling attack models can be dealt with flexibility in this way. The experiments show that the proposed algorithm can achieve better detection performance and interference immunity. Gaofeng Cao, Jianbo Zheng, Li Kuang |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2019 | A privacy-preserving multimedia recommendation in the context of social network based on weighted noise injection
Kai Dou, Li Kuang |
Multim. Tools Appl. | 3 |
| 2018 | Predicting Duration of Traffic Accidents Based on Ensemble Learning
Lina Shan, Ruyi Shi, Li Kuang |
CollaborateCom | 5 |
| 2018 | Finding Shilling Attack in Recommender System based on Dynamic Feature SelectionabstractRecommender system is widely used as an important tool in various fields for effectively dealing with information overload, and collaborative filtering algorithm plays a vital role in the system.However, such system is highly vulnerable to malicious attacks, especially shilling attack because of data openness and independence.Therefore, detecting shilling attack has become an important issue to ensure the security of recommender system.Most of existing methods for detecting shilling attack are based on rating classification features and their limitation is that they are easily to be interfered by obfuscation techniques.Moreover, traditional detection algorithms can not handle multiple types of shilling attack flexibly.In order to solve these problems, in this paper, we propose an outlier degree shilling attack detection algorithm based on dynamic feature selection.By considering the differences of user choosing items and taking user popularity as a detection metric, as well as using information entropy to select detection metrics dynamically, a variety of shilling attack models can be dealt with flexibly.Experiments show that the algorithm has stronger detection performance and interference immunity in shilling attack detection. Gaofeng Cao, Yuyou Fan, Li Kuang |
SEKE | 4 |
| 2018 | Structure convolutional extreme learning machine and case-based shape template for HCC nucleus segmentation
Siqi Li 0002, Huiyan Jiang, Yu-Dong Yao, Wenbo Pang, Qingjiao Sun, Li Kuang |
Neurocomputing | 6 |
| 2018 | A Privacy Protection Model of Data Publication Based on Game TheoryabstractWith the rapid development of sensor acquisition technology, more and more data are collected, analyzed, and encapsulated into application services. However, most of applications are developed by untrusted third parties. Therefore, it has become an urgent problem to protect users’ privacy in data publication. Since the attacker may identify the user based on the combination of user’s quasi-identifiers and the fewer quasi-identifier fields result in a lower probability of privacy leaks, therefore, in this paper, we aim to investigate an optimal number of quasi-identifier fields under the constraint of trade-offs between service quality and privacy protection. We first propose modelling the service development process as a cooperative game between the data owner and consumers and employing the Stackelberg game model to determine the number of quasi-identifiers that are published to the data development organization. We then propose a way to identify when the new data should be learned, as well, a way to update the parameters involved in the model, so that the new strategy on quasi-identifier fields can be delivered. The experiment first analyses the validity of our proposed model and then compares it with the traditional privacy protection approach, and the experiment shows that the data loss of our model is less than that of the traditional k-anonymity especially when strong privacy protection is applied. Li Kuang, Yujia Zhu, Xuejin Yan, Shuiguang Deng |
Secur. Commun. Networks | 1 |
| 2018 | Predicting Short-Term Electricity Demand by Combining the Advantages of ARMA and XGBoost in Fog Computing EnvironmentabstractWith the rapid development of IoT, the disadvantages of Cloud framework have been exposed, such as high latency, network congestion, and low reliability. Therefore, the Fog Computing framework has emerged, with an extended Fog Layer between the Cloud and terminals. In order to address the real‐time prediction on electricity demand, we propose an approach based on XGBoost and ARMA in Fog Computing environment. By taking the advantages of Fog Computing framework, we first propose a prototype‐based clustering algorithm to divide enterprise users into several categories based on their total electricity consumption; we then propose a model selection approach by analyzing users’ historical records of electricity consumption and identifying the most important features. Generally speaking, if the historical records pass the test of stationarity and white noise, ARMA is used to model the user’s electricity consumption in time sequence; otherwise, if the historical records do not pass the test, and some discrete features are the most important, such as weather and whether it is weekend, XGBoost will be used. The experiment results show that our proposed approach by combining the advantage of ARMA and XGBoost is more accurate than the classical models. Chuanbin Li, Xiaosen Zheng, Li Kuang |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | An Improved Privacy-Preserving Framework for Location-Based Services Based on Double Cloaking Regions with Supplementary Information ConstraintsabstractWith the rapid development of location-based services in the field of mobile network applications, users enjoy the convenience of location-based services on one side, while being exposed to the risk of disclosure of privacy on the other side. Attacker will make a fierce attack based on the probability of inquiry, map data, point of interest (POI), and other supplementary information. The existing location privacy protection techniques seldom consider the supplementary information held by attackers and usually only generate single cloaking region according to the protected location point, and the query efficiency is relatively low. In this paper, we improve the existing LBSs system framework, in which we generate double cloaking regions by constraining the supplementary information, and then k-anonymous task is achieved by the cooperation of the double cloaking regions; specifically speaking, k dummy points of fixed dummy positions in the double cloaking regions are generated and the LBSs query is then performed. Finally, the effectiveness of the proposed method is verified by the experiments on real datasets. Li Kuang, Yin Wang 0004, Pengju Ma, Chuanbin Li, Mengyao Zhu 0003 |
Secur. Commun. Networks | 1 |
| 2016 | A Time-Aware Weighted-SVM Model for Web Service QoS Prediction
Kai Dou, Li Kuang |
CollaborateCom | 3 |
| 2016 | Identifying Core Users Based on Trust Relationships and Interest Similarity in Recommender SystemabstractWith the rapid development of Internet, the explosive growth of information challenges people's capability on finding out items fitting to their own interests. The emergence of recommender system helps users to make decisions to a certain degree. So far, most of the studies pay much attention to designing or improving recommendation algorithms. However, few works consider the extraction of core users with whom recommender systems can generate satisfactory recommendation. In this paper, we propose new approaches to identifying core users based on trust relationships and interest similarity. The trust degree and interest similarity between all user pairs are calculated and sorted first, and two strategies based on frequency and weight of location are used to select core users. Experiments show the effectiveness of the extraction of core users and prove that 20% of core users enable recommender systems to achieve more than 90% of the accuracy of the top-N recommendation. Gaofeng Cao, Li Kuang |
ICWS | 2 |
| 2016 | Life stage based recommendation in e-commerceabstractThe rapid development of e-commerce has greatly changed the lifestyle of people. Nowadays, people are used to buying various kinds of things online, and recommender systems become more and more necessary since users are overwhelmed by a large amount of information. However, that a user's consuming behavior would change with his life stage has not been taken into consideration in most existing recommender systems. In this paper, we find the obvious correlation between temporal evolution and consuming behavior from a large amount of data. Motivated by this, we propose to obtain the relationship between items and user's life stage first. And then based on the relationship model, we can predict user's current stage according to his/her recent consuming behavior. Finally, we can recommend appropriate items to the user according to the prediction result of his/her life stage. The experimental results show that the introduction of temporal evolution of consuming behavior plays a significant part in improving the effectiveness of recommendation. Kai Dou, Li Kuang |
IJCNN | 3 |
| 2016 | TMR: Towards an efficient semantic-based heterogeneous transportation media big data retrieval
Kehua Guo, Ruifang Zhang, Li Kuang |
Neurocomputing | 3 |
| 2016 | Markov Chain-Like Model for Prediction Service Based on Improved Hierarchical Particle Swarm Optimization Cluster AlgorithmabstractSince web pages visited by users contain a variety of data resources and the clustering algorithms frequently used for web data do not take the heterogeneous nature into account when processing the heterogeneous data, this paper proposes a new algorithm, namely IHPSOC algorithm, to cluster web log data on the basis of web log mining. Based on particle swarm optimization (PSO), IHPSOC algorithm clusters the web log data through particle swarm iteration. Based on clustering results, this paper establishes Markov chain-like models which create a corresponding Markov chain for users in each different category so as to predict the web resources in users’ need. The results of the experiments show that the proposed model gives better predication. Zhifang Liao, Tianhui Song, Li Kuang, Yan Zhang 0047, Zhining Liao |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2016 | Combined retrieval: A convenient and precise approach for Internet image retrieval
Kehua Guo, Ruifang Zhang, Zhurong Zhou, Yayuan Tang, Li Kuang |
Inf. Sci. | 5 |
| 2016 | Multimedia services quality prediction based on the association mining between context and QoS properties
Li Kuang, Zhifang Liao, Wentao Feng, Haoneng He |
Signal Process. | 1 |
| 2016 | Interactive differential evolution for user-oriented image retrieval system
Fei Yu 0004, Yuanxiang Li 0001, Bo Wei 0004, Li Kuang |
Soft Comput. | 4 |
| 2015 | The Searching Ranking Model Based on the Sharing and Recommending Mechanism of Social Network
Hongxiao Fei, Tianchi Mo, Zequan Wu, Yihuan Liu, Li Kuang |
APSCC | 6 |
| 2015 | A discrete particle swarm optimization box-covering algorithm for fractal dimension on complex networksabstractResearchers have widely investigated the fractal property of complex networks, in which the fractal dimension is normally evaluated by box-covering method. The crux of box-covering method is to find the solution with minimum number of boxes to tile the whole network. Here, we introduce a particle swarm optimization box-covering (PSOBC) algorithm based on discrete framework. Compared with our former algorithm, the new algorithm can map the search space from continuous to discrete one, and reduce the time complexity significantly. Moreover, because many real-world networks are weighted networks, we also extend our approach to weighted networks, which makes the algorithm more useful on practice. Experiment results on multiple benchmark networks compared with state-of-the-art algorithms show that this PSOBC algorithm is effective and promising on various network structures. Li Kuang, Feng Wang 0048, Yuanxiang Li 0001, Haiqiang Mao, Fei Yu 0004 |
CEC | 1 |
| 2015 | Image retrieval based on interactive differential evolutionabstractDifferent from text-based image retrieval, the content-based image retrieval (CBIR) uses low-level visual features to retrieve images. The question of how to reduce semantic gap between the low level visual features and the high level image semantics is still a difficult problem. This paper uses a comparison-based mechanism based on interactive differential evolution (IDE) to help users retrieve their preferred images in a user-oriented way. The effect of the proposed framework is evaluated, and the performance of the technique is better than that of the relevance feedback (RF) based on feature re-weighting method. Fei Yu 0004, Yuanxiang Li 0001, Bo Wei 0004, Li Kuang |
CEC | 4 |
| 2015 | A fractal and scale-free model of complex networks with hub attraction behaviors
Li Kuang, Bojin Zheng, Deyi Li, Yuanxiang Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2014 | Identifying Users' Interest Similarity Based on Clustering Hot Vertices in Social NetworksabstractIdentifying users' similarity is a very important researching point because its result can be applied to many application systems. In social networks, the user circles are built not only based on their relationships in real-life, but also on common interests. Some existing approaches cannot fully capture users' similarity from the perspective of their common interests, while some other approaches are too time-consuming or space-consuming. In this paper, we propose a method of identifying users' interest similarity based on clustering Hot Vertices (HotV). A hot vertex in a social network is an account which has a large number of fans. The approach extracts users' common interests by mining and clustering the hot vertices that the two users are following simultaneously. Both the experiment and theoretical analysis have proved that the proposed approach makes a significant improvement on the precision of similarity measuring with a relatively low time and space complexity. Tianchi Mo, Hongxiao Fei, Li Kuang, Qifei Qin |
APSCC | 3 |
| 2014 | A differential evolution box-covering algorithm for fractal dimension on complex networksabstractThe fractality property are discovered on complex networks through renormalizaiton procedure, which is implemented by box-covering method. The unsolved problem of box-covering method is finding the minimum number of boxes to cover the whole network. Here, we introduce a differential evolution box-covering algorithm based on greedy graph coloring approach. We apply our algorithm on some benchmark networks with different structures, such as a E.coli metabolic network, which has low clustering coefficient and high modularity; a Clustered scale-free network, which has high clustering coefficient and low modularity; and some community networks (the Politics books network, the Dolphins network, and the American football games network), which have high clustering coefficient. Experimental results show that our algorithm can get better results than state of art algorithms in most cases, especially has significant improvement in clustered community networks. Li Kuang, Feng Wang 0048, Yuanxiang Li 0001, Fei Yu 0004 |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Artificial immune system application for solving dynamic optimization problemsabstractFor the purpose of adaptation to a changing environment, immune mutation and memory mechanism in the immune system are introduced in thermodynamic genetic algorithm, which helps to prevent the diversity loss and rapidly track the optimum in dynamic environments. Experimental results on 0/1 dynamic knapsack problems demonstrate the merits of the proposed immune thermodynamic genetic algorithm (ITDGA). Compared with the existing classical primal-dual genetic algorithm (PDGA), this algorithm can maintain better diversity and be more suitable to solve 0-1 dynamic problems. Yuanxiang Li 0001, Li Kuang, Fei Yu 0004 |
IJCNN | 3 |
| 2012 | Personalized Services Recommendation Based on Context-Aware QoS PredictionabstractWith the increase of published Web services, it has become a great challenge to recommend service consumers the best services with regard to the quality of services (QoS). Collaborative filtering is often employed to predict the QoS of a specific service to a certain consumer. However, in existing collaborative filtering based service recommendation approaches, the context under which consumers submit a recommendation request is seldom taken into account when filtering similar recommenders and their corresponding experience. In this paper, we propose a new method dubbed CASR (Context-Aware Services Recommendation) by referring to previous service invocation experiences under similar context with the current consumer, which is of great importance in the personalized service recommendation system. First, the proposed algorithm clusters the service invocation records according to the similarity on context properties and selects the cluster that is most similar to the context of current consumer. Then it predicts the QoS of an unused service for current consumer based on the filtered recommendation records by Bayesian inference. Experimental results demonstrate that the proposed approach can significantly improve the accuracy of QoS prediction and service recommendation. Li Kuang, Yingjie Xia, Yuxin Mao |
ICWS | 1 |
| 2011 | A Parallel Fusion Method for Heterogeneous Multi-sensor Transportation Data
Yingjie Xia, Chengkun Wu, Qing-Jie Kong, Zhenyu Shan, Li Kuang |
MDAI | 5 |
| 2011 | Accelerating geospatial analysis on GPUs using CUDAabstractInverse distance weighting (IDW) interpolation and viewshed are two popular algorithms for geospatial analysis. IDW interpolation assigns geographical values to unknown spatial points using values from a usually scattered set of known points, and viewshed identifies the cells in a spatial raster that can be seen by observers. Although the implementations of both algorithms are available for different scales of input data, the computation for a large-scale domain requires a mass amount of cycles, which limits their usage. Due to the growing popularity of the graphics processing unit (GPU) for general purpose applications, we aim to accelerate geospatial analysis via a GPU based parallel computing approach. In this paper, we propose a generic methodological framework for geospatial analysis based on GPU and its programming model Compute Unified Device Architecture (CUDA), and explore how to map the inherent parallelism degrees of IDW interpolation and viewshed to the framework, which gives rise to a high computational throughput. The CUDA-based implementations of IDW interpolation and viewshed indicate that the architecture of GPU is suitable for parallelizing the algorithms of geospatial analysis. Experimental results show that the CUDA-based implementations running on GPU can lead to dataset dependent speedups in the range of 13–33-fold for IDW interpolation and 28–925-fold for viewshed analysis. Their computation time can be reduced by an order of magnitude compared to classical sequential versions, without losing the accuracy of interpolation and visibility judgment. Yingjie Xia, Li Kuang, Xiu-mei Li |
J. Zhejiang Univ. Sci. C | 2 |
| 2010 | Analyzing Behavioral Substitution of Web Services Based on Pi-calculusabstractThe behavioral analysis for Web services provides a priori detection of errors to ensure successful interactions in services invocation and composition, and the behavioral substitution of Web services is one of the most important issues in such analysis. In this paper, we propose to formalize the behavior of a Web service by π-calculus. Based on the formalization, we introduce two notions of behavioral substitution of Web services namely strong and weak simulation. Furthermore, we propose a derivative approach to analyzing the behavioral substitution of services according to the given notions, which is implemented based on an existing tool of π-calculus. The proposed approach takes advantage of formalization and theory of π-calculus, so that the formalized services can be naturally analyzed and the behavioral substitution of them can be easily determined. Li Kuang, Yingjie Xia, Shuiguang Deng, Jian Wu 0001 |
ICWS | 1 |
| 2009 | Towards Adaptation of Service Interface SemanticsabstractInteroperability promised by Web service makes it a most promising technology for the development of next generation distributed heterogeneous software systems. Services should be compliant at signature, behavioral and semantic level to make the interoperation successful and correct. Service adaptation provides an effective approach to bridge the incompatibility of services to make them interoperate as well as possible. In this paper, we aim to contribute to the definition of a methodology to develop adaptors that are capable of making two incompatible services interoperate not only successfully but also correctly at semantic level. To achieve this goal, we proposed service specifications for both atomic and composite services with semantic dependency between outputs and inputs specified; then we proposed adaptor specification consisting of three parts, which are message mapping, action mapping and treatment for non-mapping messages. Based on service and adaptor specifications, an incremental derivation approach of a concrete adaptor is given. Li Kuang, Shuiguang Deng, Jian Wu 0001, Ying Li 0001 |
ICWS | 1 |
| 2008 | Service Behavioral Adaptation Based on Dependency GraphabstractService adaptation is one of the most important issues in SOC (Service Oriented Computing). This paper focuses on the issue of service behavioral adaptation and proposes an adapting method based on dependency graph. It can be divided into three sequential sub-problems: (1) service description-the foundation of service adaptation. We propose a formal approach to describing service behavior protocols; (2) mismatch definition-the identification of service mismatches. We define several kinds of behavior mismatches based on dependency graph; (3) service adaptation-the adaptor construction process. We detect all possible behavior mismatches and then generate different adaptors correspondingly. Shuiguang Deng, Jian Wu 0001, Ying Li 0001, Li Kuang, Jianwei Yin |
APSCC | 5 |
| 2007 | Inverted Indexing for Composition-Oriented Service DiscoveryabstractService discovery becomes a key to hastening the evolution of web services as the number of services is expected to increase dramatically. In this paper, we propose to index all the ontology-annotated outputs in registered services. For each ontology-annotated output, there is a service list which records all the services in the registry that deliver the output. Based on the indexing, we propose a composition-oriented service discovery algorithm, which greatly accelerates the filtering of irrelevant atomic services by making use of the inverted indexing, and increases the likelihood of finding a possible candidate by exploring service composition. Experimental results show that the proposed algorithm provides a better performance on response time than the sequential matchmaking, and a better recall rate than the algorithms without the exploration of composition. Li Kuang, Ying Li 0001, Jian Wu 0001, Shuiguang Deng, Zhaohui Wu 0001 |
ICWS | 1 |
| 2006 | Service Classification Using Adaptive Back-Propagation Neural Network and Semantic SimilarityabstractWith the growing population of Web services, the discovery of services is a key to the development of Web services. While extensive researches focus mainly on service matchmaking algorithms, service classification that is also a meaningful approach to accelerating service discovery only receives little attention. In this paper, we propose to use adaptive back-propagation neural network model (BPM) to perform service recognition. During the training process, the feature vectors of training services and their categories are learned by the BPM. The element in the feature vector is the semantic similarity between the feature word in the system dictionary and the occurrence in the feature set for a service. During the recognition process, the characteristics of the test service are analyzed by the BPM and the output shows the category of the test service. Furthermore, the BPM is adapted with correctly recognized test services, which results in better modeling over time. Based on extensive experiments, we show that using the adaptive BPM is a promising way to realize automatic service classification, and the adaptive BPM using semantic similarity as the element in feature vector provides better performance than the general BPM using word frequency Li Kuang, Jian Wu 0001, Shuiguang Deng, Ying Li 0001, Zhaohui Wu 0001 |
CSCWD | 1 |
| 2006 | Expressing Service and Query Behavior Using pi-Calculus for MatchmakingabstractService discovery becomes a key to accelerating the evolution of Web services as the number of services is expected to increase dramatically. Foregoing work on service discovery is primarily based on the interfaces of services through the use of ontology. Ongoing work targets at service behavior, with not only individual message exchanges being captured, but also constraints between these message exchanges. In this paper, we propose a formal approach to expressing the service and query behavior using pi-calculus for service matchmaking. The resulting pi-calculus expressions of services and queries are precise in defining single operations involving message exchanges as well as execution sequence between operations. Based on the formalizations, service matchmaking between a service query and a service description is reasoned through the capability of pi-calculus. Expressing service behavior using pi-calculus is expected to be a promising way to realize intelligent service discovery Li Kuang, Ying Li 0001, Shuiguang Deng, Jian Wu 0001, Zhaohui Wu 0001 |
Web Intelligence | 1 |
| 2006 | Intelligent Transportation Information Sharing and Service Integration in Semantic Grid EnvironmentabstractITSGrid is an undergoing joint engineering project designed and developed by advanced computing and system (CCNT) lab in Zhejiang University and Hangzhou Enjoyor Electronics Co. Ltd (Enjoyor). The new features of ITSGrid are originated from two important research projects - DartGrid and DartFlow, and one key engineering project - JTang application server, in CCNT lab. Its goal is to build an integrated intelligent transportation information and service platform (ITISP), to integrate traffic data resources collected by Enjoyor and cooperate existing ITS subsystems and services deployed by Enjoyor, finally serve for transportation construction in China. During building this project, we utilize systematically the grid technology, the semantic Web technology, the Web service technology, the messaging oriented middleware technology Jian Wu 0001, Ying Li 0001, Li Kuang |
Web Intelligence | 4 |
| 2004 | Management of Serviceflow in a Flexible Way
Shuiguang Deng, Zhaohui Wu 0001, Li Kuang, Yueping Jin, Shifeng Yan, Ying Li 0001 |
WISE | 3 |