VLDB 2026 Research / reviewers in the wild / expert
Zhifang Liao
dblp:42/2640
· DBLP profile ↗
45ranked-venue papers
18as first author
27since 2021 · last 2026
0000-0002-5525-904XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Software engineering, systems software and programming languages · 10 · 5 first-author · 6 since 2021Computer networks · 9 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PURE: Probabilistic Utility for Robust Evidence in Multimodal RecommendationabstractModern recommender systems increasingly rely on multimodal content to mitigate data sparsity. However, existing methods typically treat multimodal signals as deterministic features, indiscriminately fusing them without assessing their reliability. This overlooks the inherent ambiguity of content, causing low-quality modalities to degrade performance. To address this, we propose PURE (Probabilistic Utility for Robust Evidence), a framework that reconceptualizes multimodal signals as uncertain evidence. Instead of operating on raw, noisy features, PURE first constructs a hybrid hypergraph to derive robust content evidence, harmonizing local proximity and global community structures to rectify structural inconsistencies. Grounded on this refined topology, we introduce an Uncertainty-Aware Fusion mechanism driven by evidential learning. Unlike prior methods focusing on prediction uncertainty, PURE explicitly models utility uncertainty via Dirichlet parameterization. This allows the model to dynamically regulate information injection: high-confidence semantics are utilized, while low-confidence ones are bypassed by reverting to collaborative ID signals, effectively securing a robust performance baseline. Extensive experiments on three large-scale benchmarks demonstrate that PURE consistently outperforms state-of-the-art baselines while functioning as a transparent data quality diagnostician. Jinguo Hu, Yuelan Qi, Zhifang Liao |
ICMR | 4 |
| 2025 | FEDDGCN: A Frequency-Enhanced Decoupling Dynamic Graph Convolutional Network for Traffic Flow PredictionabstractAs a core task in Intelligent Transportation Systems (ITS), traffic flow prediction is essential for resource allocation and real-time route planning. Effectively capturing complex temporal correlations and dynamic spatial dependencies in traffic flow data is critical yet challenging for accurate prediction. However, existing approaches are still limited by the insufficient capability for spatial-temporal pattern decoupling and the underutilization of frequency domain information. To address these issues, we propose a novel Frequency-Enhanced Dynamic Decoupling Graph Convolutional Network (FEDDGCN), which introduces a gated decoupling mechanism integrating temporal and spatial embeddings to decouple traffic flow into prominent periodic and perturbative component. It also achieves effective pattern separation by incorporating frequency domain analysis with Fourier filters. Furthermore, a dual-branch spatial-temporal learning module, employing a divide-and-conquer strategy, is designed to achieve separate modeling for the two distinct components. Specially, the dynamic graph convolution modules are utilized to learn spatial dependencies and temporal and frequency attention mechanisms further capture complex temporal correlations for prominent periodic and perturbative components.Extensive experiments on multiple real-world datasets demonstrate that FEDDGCN achieves superior predictive performance compared with state-of-the-art methods. Ruobai Xiang, Zhifang Liao, Peng Lan, Qihao Liang |
CIKM | 3 |
| 2025 | FGML-DG: Feynman-Inspired Cognitive Science Paradigm for Cross-Domain Medical Image SegmentationabstractIn medical image segmentation across multiple modalities (e.g., MRI, CT, etc.) and heterogeneous data sources (e.g., different hospitals and devices), Domain Generalization (DG) remains a critical challenge in AI-driven healthcare. This challenge primarily arises from domain shifts, imaging variations, and patient diversity, which often lead to degraded model performance in unseen domains. To address these limitations, we identify key issues in existing methods, including insufficient simplification of complex style features, inadequate reuse of domain knowledge, and a lack of feedback-driven optimization. To tackle these problems, inspired by Feynman’s learning techniques in educational psychology, this paper introduces a cognitive science-inspired meta-learning paradigm for medical image domain generalization segmentation. We propose, for the first time, a cognitive-inspired Feynman-Guided Meta-Learning framework for medical image domain generalization segmentation (FGML-DG), which mimics human cognitive learning processes to enhance model learning and knowledge transfer. Specifically, we first leverage the ‘concept understanding’ principle from Feynman’s learning method to simplify complex features across domains into style information statistics, achieving precise style feature alignment. Second, we design a meta-style memory and recall method (MetaStyle) to emulate the human memory system’s utilization of past knowledge. Finally, we incorporate a Feedback-Driven Re-Training strategy (FDRT), which mimics Feynman’s emphasis on targeted relearning, enabling the model to dynamically adjust learning focus based on prediction errors. Experimental results demonstrate that our method outperforms other existing domain generalization approaches on two challenging medical image domain generalization tasks. Yucheng Song, Haokang Ding, Zhining Liao, Zhifang Liao |
ECAI | 5 |
| 2025 | Exploring Boundary-Aware Spatial-Frequency Fusion for Camouflaged Object DetectionabstractCamouflaged Object Detection is challenging due to the high degree of similarity between camouflaged objects and their surrounding backgrounds. Current COD methods mainly rely on edge extraction in the spatial domain and local pixel-level information, neglecting the importance of global structural features. Additionally, they fail to effectively leverage the importance of phase spectrum information within frequency domain features. To this end, we propose a COD framework BASFNet based on boundary-aware frequency domain and spatial domain fusion. This method uses dual guided integration of frequency domain and spatial domain features. A phase-spectrum-based frequency-enhanced edge exploration module (FEEM) and a spatial core segmentation module (SCSM) are introduced to jointly capture the boundary and object features of camouflaged objects. These features are then effectively integrated through a spatial-frequency fusion interaction module (SFFIM). Furthermore, the boundary detection is further optimized through an boundary-aware training strategy. BASFNet outperforms existing state-of-the-art methods on three benchmark datasets, validating the effectiveness of the fusion of frequency and spatial domain information in COD tasks. Haokang Ding, Zhifang Liao, Yucheng Song |
ECAI | 4 |
| 2025 | Causal Reasoning on Temporal Knowledge Graph with Fuzzy Logic and Graph Attention NetworkabstractCausal reasoning algorithms can evaluate the strength of causal relationships between events and play an essential role in revealing and understanding the underlying mechanisms of them. However, traditional algorithms find it difficult to handle the complex interaction effects between numerous variables, can only consider event information at a single moment, and are challenging to solve the uncertainty problem in temporal evolution. To address these challenges, this paper proposes a temporal knowledge graph causal reasoning model (TKGR) that combines fuzzy logic and graph attention networks. Graph attention networks are used to capture the complex interaction effects between numerous variables at a single moment. At the same time, a multi-head self-attention mechanism is used to aggregate temporal information from multiple time points. In response to the uncertainty problem in time series evolution, we also designed a fuzzy logic module to comprehensively consider the importance of features from different time points. Finally, this paper carried out multiple rounds of comparative experiments with traditional causal reasoning algorithms and ablation experiments on the simulation dataset of satellite reconnaissance missions. The results show that the TKGR model can effectively reason the strength of causal relationships, thus providing necessary reference for decisionmaking. Shixiang Cai, Zhirui Kuai, Li Kuang, Zhifang Liao |
ICWS | 4 |
| 2025 | SECMG: Style Enhancement Commit Message GenerationabstractThe commit message (CM) is a summary used in software development to understand code changes (Diff) quickly. A good CM helps improve the efficiency of software development and maintenance. Most existing methods generate CM based on Diff files, but these methods have some limitations. Traditional methods neglect to retrieve the highly similar semantic information that may be embedded in the Commit Message; they didn’t fully consider the similarities and differences between the retrieved CM and the query Diff. To better retrieve relevant commit messages and assist in CM generation, this paper proposes a Style Enhancement Commit Message Generation (SECMG) model. The model divides the CM generation task into two modules, including the Commit Message retrieval module and the Style-Enhancement CM generation module. To maximize the capture of highly relevant retrieval results, we employ a combination of semantic and keyword-based retrieval methods. Additionally, we optimize the retrieval outcomes by training a re-ranking model. In the generation module, we first augment the Diff with structural information at the word level via difflib. Then we use the CodeT5 model as a base to process the Diff and retrieve Commit Messages separately using Bi-Encoder, which are fused through a collaborative attention mechanism. Finally, we introduce a Gram-based style loss function to align the outputs and generate the commit message closest to the target message. We conducted extensive experiments on a real-world dataset, and the experimental results demonstrate the superiority of the present method over other models, improving by 17.50% and 8.73% on BLEU and B-Norm metrics, respectively. Zhifang Liao, Peng Lan |
IJCNN | 1 |
| 2025 | CMSPA-Former: Causal Multi-Scale Pyramid Attention Former for Estimating Counterfactual OutcomesabstractCausal effect estimation is a crucial research area in causal analysis. Conventional methods include Randomized Controlled Trials (RCT), Propensity Score Matching (PSM), and counterfactual inference based on Potential Outcome Framework, with counterfactual inference being the most important approach. Existing counterfactual estimation methods suffer from several limitations, such as over-reliance on handcrafted metric functions, difficulty in effectively modeling complex long-range dependencies, high computational costs, and insufficient consideration of the multi-scale dynamic characteristics of time series. To address these issues, this paper proposes a Causal Multi-Scale Pyramid Attention Former (CMSPA-Former) model based on a multi-scale pyramid attention mechanism. The model employs a multi-scale time-series construction module to perform hierarchical feature extraction by downsampling the input sequence layer by layer, capturing dependencies at different scales. Meanwhile, to reduce computational complexity, the model introduces a Pyramid Attention Mechanism (PAM) to efficiently process sequence data and model representations. Additionally, this paper replaces the traditional Attentional Dropout method with Top-k contribution scores to retain critical information and remove redundant features, further improving the accuracy of counterfactual outcome predictions. Experimental results demonstrate that the proposed model significantly outperforms baseline models on both synthetic and real-world datasets, validating its effectiveness. Haibin Peng, Zhifang Liao |
IJCNN | 3 |
| 2025 | Cmdesum: A Cross-Modal Deliberation Network for Code SummarizationabstractThe task of Code Summarization aims to generate concise natural language descriptions for given source code snippets, thereby assisting developers in reducing the cognitive load of comprehending the code. Existing learning-based models, employing single-pass encoder-decoder frameworks, are unable to harness global information to optimize local content, while those utilizing multi-pass encoder-decoder architectures fail to consider the structural information of the code. To tackle this issue, we propose CMDeSum, a novel deliberation framework that injects cross-modal information in a staged manner, aiming to better balance the code sequence information and Abstract Syntax Tree (AST) structural information. Specifically, we first retrieve comments from code segments similar to the given one as drafts and extract method names and ASTs from the code. Then, in the First-Pass stage, we utilize code sequences and method names to generate initial comments and refine them based on the drafts. In the Second-Pass stage, building upon the results from the First-Pass stage, we utilize additional AST information to modify the comments, producing the final comments. To evaluate our approach, we conducted experiments on existing Java and Python datasets. The experimental results indicate that compared with the state-of-the-art models for code summarization generation, our model has improved by at least 6.3%, 3.0%, and 5.8% in BLEU, ROUGE-L, and METEOR. Zhifang Liao, Peng Lan |
ICPC | 1 |
| 2025 | GitHub project recommendation based on knowledge graph and developer similarityabstractAbstract Finding and recommending projects that match developer’s interests is always an urgent problem in open-source community. There are some problems in the existing project recommendation methods, such as insufficient use of information, ignoring the relationship between projects, one-sided consideration, and so on. To solve the above problems, we propose a project recommendation model based on project knowledge graph and developer similarity, called knowledge graphs and developer interest similarity (KGDS). KGDS mines developer interest from project similarity and developer similarity. For project similarity, we first construct the project knowledge graph. Then, content features and potential features are extracted from the project Readme document and knowledge graph, respectively, and the two features are merged to enrich the developer embedding and project embedding, which solves the problem of insufficient utilization of information. For developer similarity, we first construct a developer-project matrix, then obtain the historical developers related to candidate project, and then calculate the similarity between the historical developers and the target developer, which solves the problem of one-sided consideration. Finally, we combine the two part information to recommend projects that meet the interests of developers. We have conducted experiments on the GitHub dataset, and the results show that KGDS outperforms the baseline model. Hannan Wu, Zhifang Liao |
Comput. J. | 4 |
| 2025 | Multiple teachers are beneficial: A lightweight and noise-resistant student model for point-of-care imaging classification
Yucheng Song, Anqi Song, Jincan Wang, Zhifang Liao |
Expert Syst. Appl. | 4 |
| 2024 | ModWaveMLP: MLP-Based Mode Decomposition and Wavelet Denoising Model to Defeat Complex Structures in Traffic ForecastingabstractTraffic prediction is the core issue of Intelligent Transportation Systems. Recently, researchers have tended to use complex structures, such as transformer-based structures, for tasks such as traffic prediction. Notably, traffic data is simpler to process compared to text and images, which raises questions about the necessity of these structures. Additionally, when handling traffic data, researchers tend to manually design the model structure based on the data features, which makes the structure of traffic prediction redundant and the model generalizability limited. To address the above, we introduce the ‘ModWaveMLP’—A multilayer perceptron (MLP) based model designed according to mode decomposition and wavelet noise reduction information learning concepts. The model is based on simple MLP structure, which achieves the separation and prediction of different traffic modes and does not depend on additional features introduced such as the topology of the traffic network. By performing experiments on real-world datasets METR-LA and PEMS-BAY, our model achieves SOTA, outperforms GNN and transformer-based models, and outperforms those that introduce additional feature data with better generalizability, and we further demonstrate the effectiveness of the various parts of the model through ablation experiments. This offers new insights to subsequent researchers involved in traffic model design. The code is available at: https://github.com/Kqingzheng/ModWaveMLP. Zhifang Liao |
AAAI | 4 |
| 2024 | DupLLM:Duplicate Pull Requests Detection Based on Large Language ModelabstractAs the scale of projects expands, the concurrent development model adopted by the open source community leads to an increasingly prominent problem of repetitive pull requests (PRs). The large number of rejections caused by duplicate pull requests increases the review workload of project maintainers and reduces the efficiency of pull request review. Therefore, it is very necessary to conduct automated duplicate PR detection. In this study, we propose DupLLM, a framework designed to detect duplicate PRs. The framework generates refined sum-maries by feeding the content of individual PRs into a large language model (LLM). Subsequently, the resulting summary is vectorized, converting the textual content into a numerical representation. The similarity between PRs is evaluated by calculating the similarity score between PR summary vectors. Ultimately, the model showed better performance than the best existing model, achieving an effect of 0.929 on P@l. This confirms that LLM can also achieve equivalent results in the field of duplicate PR detection as deep learning is used to train on this task, providing a new direction for the application of LLM in the field of software engineering. Zhifang Liao, Peng Lan |
APSEC | 1 |
| 2024 | Exploring the Potential of Large Language Models in Automatic Pull Request Title Generation: An Empirical StudyabstractPull Requests (PRs) are a collaborative mechanism in GitHub, allowing developers to merge their code changes into another branch of the software repository. The PR title serves as a summary of the PR and needs to accurately and concisely describe the specific changes made, which is useful for reviewers and other developers to review and understand. There are many existing methods for automatically generating PR titles, most of which are based on pre-trained models. Although these methods are effective, pre-trained models often require extensive fine-tuning for specific tasks. Compared to pre-trained models, large language models (LLMs) possess superior semantic understanding capabilities. As a foundational model, they can solve most tasks directly without relying on fine-tuning, providing an alternative solution for PR title generation. However, the capabilities of LLMs in the automatic PR title generation have not been fully explored. To fill this gap, we conducted an empirical study to understand the capabilities of LLMs in PR title generation. Initially, the direct application of LLMs to generate PR titles did not yield satisfactory results. We found that using similar PRs from the dataset as auxiliary information can effectively enhance the title generation capability of LLMs. When the number of most similar PRs used as input increased from 0 to 5, the ROUGE-L F1 score of the titles generated by LLMs increased by an average of 23.48 %, with improvements in other metrics as well. In further experiments, we discovered that setting a lower temperature for the LLMs can bring better performance. We then selected the best parameter configuration and compared it with the existing state-of-the-art methods. Our experimental results show that LLMs outperform the state of the art methods in Precision, Recall, and METEOR metrics on the PRTiger dataset. Additionally, human evaluation results indicate that PR titles generated by LLMs receive higher scores in Correctness, Naturalness, and Comprehensibility. Yitao Zuo, Peng Lan, Zhifang Liao |
APSEC | 3 |
| 2024 | Alternate Interaction Multi-Task Hierarchical Model for Medical ImagesabstractDeep learning models in medical image segmentation and classification tasks can automatically extract features and perform high-performance inference, thereby assisting doctors in achieving efficient and accurate automated decision support. However, these models are typically trained for a single task, which leads to limitations such as ignoring task relevance and poor scalability. In response to the above problems, we hope to apply one model to different tasks to enhance the model’s scalability. Hence, this paper proposes an Alternating Multi-Task Hierarchical Network (AMTH-Net) for medical image segmentation and classification. The model is divided into three hierarchical modules: Pathological Region Clarity (PRC) serves as an auxiliary module to improve segmentation and classification capabilities, the Multi-Resolution Attention (MRA) segmentation module focuses on image information at different resolution levels through deep supervision to enhance segmentation accuracy, and the Cascading Multi-Scale Information (CMSI) classification module employs a cascading multi-scale mechanism to gradually integrate discrete information from different network layers, thereby enhancing classification performance. Additionally, we propose a novel Alternating Interaction Loss (AI-Loss) based on a Multi-Level Gradient Information Feedback (MGIF) algorithm to further improve the model’s segmentation and diagnostic performance. We validated our approach on two datasets: the public COVID CXR dataset and our newly proposed F_BUSI breast ultrasound dataset. Experimental results demonstrate that AMTH-Net achieves excellent performance in both segmentation and classification, and it outperforms existing methods in terms of segmentation and classification capabilities. Yucheng Song, Yifan Ge, Zhifang Liao, Lifeng Li |
BIBM | 4 |
| 2024 | RMIB: Representation Matching Information Bottleneck for Matching Text RepresentationsabstractRecent studies have shown that the domain matching of text representations will help improve the generalization ability of asymmetrical domains text matching tasks. This requires that the distribution of text representations should be as similar as possible, similar to matching with heterogeneous data domains, in order to make the data after feature extraction indistinguishable. However, how to match the distribution of text representations remains an open question, and the role of text representations distribution match is still unclear. In this work, we explicitly narrow the distribution of text representations by matching them with the same prior distribution. We theoretically prove that narrowing the distribution of text representations in asymmetrical domains text matching is equivalent to optimizing the information bottleneck (IB). Since the interaction between text representations plays an important role in asymmetrical domains text matching, IB does not restrict the interaction between text representations. Therefore, we propose the adequacy of interaction and the incompleteness of a single text representation on the basis of IB and obtain the representation matching information bottleneck (RMIB). We theoretically prove that the constraints on text representations in RMIB is equivalent to maximizing the mutual information between text representations on the premise that the task information is given. On four text matching models and five text matching datasets, we verify that RMIB can improve the performance of asymmetrical domains text matching. Our experimental code is available at https://github.com/chenxingphh/rmib. Haihui Pan, Zhifang Liao, Wenrui Xie |
ICML | 2 |
| 2024 | Growing with the Help of Multiple Teachers: Lightweight and Noise-Resistant Student Model for Medical Image Classification
Yucheng Song, Jincan Wang, Yifan Ge, Zhifang Liao, Peng Lan, Lifeng Li |
PRCV (14) | 4 |
| 2024 | Self-Admitted Technical Debts Identification: How Far Are We?abstractSelf-admitted technical debt (SATD) is a kind of technical debt that is already acknowledged by the developers and needs additional work or resources to address in the future. In recent years, though many methods have been proposed to detect SATDs, these methods have mainly focused on Java-type code comments published by Maldonado et al. It is unclear whether these methods trained on Maldonado's code comments dataset can find SATD in other programming languages or other software artifacts, such as issue trackers, pull requests, and commit messages effectively. In order to answer the above confusion and investigate how far our community has progressed in the field of SATD identification, we first collect a comprehensive dataset that contains SATDs in code comments from four different programming languages (java, python, docker file, XML) and SATDs in other different artifacts (issue tracker, pull requests, commit messages) from previous papers working in the field of SATD. Then, we re-train the existing models with Maldonado's code comments dataset and test all the models on other programming languages and other artifacts. The results show that existing SATD identification methods can find SATDs in other non-Java languages, but perform poorly in identifying SATDs from three other different artifacts. In addition, in order to simultaneously identify four different artifacts of SATDs, we develop a Multi-Task Learning model utilizing BERT for SATD identification (MT-BERT-SATD). Considering four different artifacts and the SATD identification tasks, MT-BERT-SATD achieves an average F1-score of 0.712 (0.625-0.859), which is superior to existing models from 4.6% to 30.4%. Results show that MT-BERT-SATD can effectively identify SATD instances across explored programming languages and software artifacts, indicating its capability to identify SATD instances in new and unexplored programming languages and software artifacts. Zhifang Liao, David Lo 0001 |
SANER | 4 |
| 2024 | Classification of open source software bug report based on transfer learningabstractAbstract Currently, the feature richness of text encoding vectors in the bug report classification model based on deep learning is limited by the size of the domain dataset and the quality of the text. However, it is difficult to further enrich the features of text encoding vectors. At the same time, most existing bug report classification methods ignore the submitter's personal information. To solve these problems, we construct nine personal information characteristics of bug report submitters in GitHub by survey. Then, we propose a GitHub bug report classification method named personal information fine‐tuning network (PIFTNet) based on transfer learning and the submitter's personal information. PIFTNet transfers the general text feature vectors in bidirectional encoder representation from transformers (BERT) to the domain of bug report classification by fine‐tuning the pre‐training parameters in BERT. It also combines the text characteristics and the characteristics of the submitter's personal information to construct the classification model. In addition, we propose a two‐stage training method to alleviate the catastrophic changes in the pre‐training parameters and loss of the initially learned knowledge caused by direct training of PIFTNet. We verify the proposed PIFTNet on the dataset extracted from GitHub and empirical results prove the effectiveness of PIFTNet. Zhifang Liao, Zeng Qi, Shengzong Liu, Yan Zhang 0047 |
Expert Syst. J. Knowl. Eng. | 1 |
| 2024 | Medical image classification: Knowledge transfer via residual U-Net and vision transformer-based teacher-student model with knowledge distillation
Yucheng Song, Jincan Wang, Yifan Ge, Lifeng Li, Quanxing Dong, Zhifang Liao |
J. Vis. Commun. Image Represent. | 7 |
| 2023 | Knowledge Distillation of Attention and Residual U-Net: Transfer from Deep to Shallow Models for Medical Image Classification
Zhifang Liao, Quanxing Dong, Yifan Ge, Huaiyi Chen, Yucheng Song |
PRCV (13) | 1 |
| 2023 | Blockchain-based mobile crowdsourcing model with task security and task assignment
Zhifang Liao, Jincheng Ai, Shaoqiang Liu, Yan Zhang 0047, Shengzong Liu |
Expert Syst. Appl. | 1 |
| 2021 | Detecting Duplicate Questions in Stack Overflow via Semantic and Relevance ApproachesabstractStack Overflow is a popular online Q&A website related to programming. Although Stack Overflow has detailed questioning guidance, duplicate questions still appear frequently, and a large number of duplicate questions make the quality of the community degraded. To solve this problem, Stack Overflow allows users with high reputations to manually mark duplicate questions. However, this method is inefficient and causes many duplicate questions to remain undiscovered. Therefore, this paper proposes a duplicate questions detection model based on semantic and relevance. The model employs Siamese BiLSTM to encode question pairs and captures the semantic interaction information of title and body through soft align attention and inference composition. The soft term match captures the relevance information in the title. We evaluate the effectiveness of the model in six question groups on Stack Overflow. Compared with the latest deep learning model, the F1-Score and ACC of our model increased by 9.401% and 8.901%, respectively. Experimental results show that our model outperforms the baselines and achieves competitive performance. Zhifang Liao, Yan Zhang 0047 |
APSEC | 1 |
| 2021 | CMF Net: Detecting Objects in Infrared Traffic Image with Combination of Multiscale FeaturesabstractInfrared image target detection has always been a hot topic of research, but there is still little research on infrared image target detection in the field of transportation. In this paper, we use the idea of transfer learning to transfer the target detection framework in the visible domain of deep learning to the infrared domain, and propose the target detection model CMF Net based on multi-scale feature fusion. CMF Net uses two multi-scale feature extraction mechanisms and features fusion, so that the final output feature map of the backbone network contains not only low-level visual features which are beneficial to target localization, but also high-level semantic features which are beneficial to target recognition, and can adapt to multi-scale features of the target. The experiment verified the advantages of CMF Net, and its mAP on the test data of the infrared image data set FLIR reached about 71%. This result is an increase of about 13% compared to Faster R-CNN, an increase of about 6% compared to YOLO3, and an increase of about 17% compared to SSD. Zhifang Liao, Yiqi Zhao, Xuechun Huang, Jinsong Wu 0001 |
GLOBECOM | 1 |
| 2021 | Short-term load forecasting with dense average network
Zhifang Liao, Haihui Pan, Xuechun Huang, Ronghui Mo, Xiaoping Fan, Huanwen Chen |
Expert Syst. Appl. | 1 |
| 2021 | GRBMC: An effective crowdsourcing recommendation for workers groupsabstractCrowdsourcing as an important computing resource for the Crowd-based Cooperative, crowdsourcing workers, due to the complexity of their personnel composition and the ambiguity of group characteristics, make it difficult to accurately recommend suitable worker groups for the task. Recommending crowdsourcing tasks to suitable workers can greatly improve the efficiency and quality of the execution of crowdsourcing tasks. The current crowdsourcing recommendation algorithm is based on the single characteristics of the workers, ignoring the multi-dimensional characteristics of the workers and the community characteristics of the crowd. This paper proposes a worker group recommendation method based on multi-community collaboration (GRMBC) by utilizing the worker characteristics information extracted from crowdsourcing platform. Based on the characteristics of workers'reputation, preference and activity, the method divides the worker group into several characteristics communities with similar behavior to discover the potential multi-community structure in the crowdsourcing worker group, and then selects Top-N worker group for recommendation through the interaction among multi-communities. At the same time, we also proposed two mitigation strategies for the cold start problem of data in crowdsourcing. This paper uses the public data collected by AMT to do the experiments, and compares the aggregation results of the recommendations generated by different algorithms. The results show the recommendations generated by the GRBMC algorithm proposed in this paper performs the best comprehensively. When the GRBMC algorithm that considers both worker attributes and communi-ty characteristics has an accuracy improvement of 0.03 ~ 0.04 compared with not using the recommendation algorithm on each index, the accuracy of 0.01 ~ 0.02 is improved com-pared with the method considering single characteristic. Moreover, compared with the individual recommendation method of the workers, the recommendation method of the workers' group is more in line with the platform requirements, and can well reflect the wisdom of crowd wisdom. Zhifang Liao, Xiaoping Fan, Yan Zhang 0047 |
Expert Syst. Appl. | 1 |
| 2021 | Multiple Wavelet Convolutional Neural Network for Short-Term Load ForecastingabstractAlthough the accuracy of load forecasting has been studied by many works, the actual deployability of a model is rarely considered. In this work, we consider the actual deployability of a model from four aspects: 1) the prediction performance of the model; 2) the robustness of the model; 3) the dependence of the model on external data; and 4) the storage size of the model. From these four aspects, we propose a multiple wavelet convolutional neural network (MWCNN) for load forecasting. On two public data sets, we verified the performance and robustness of the MWCNN. The MWCNN only uses load data, and the storage size of the model is only 497 kB, which shows that MWCNN has good deployability. In addition, our MWCNN prediction results are interpretable. The experimental results show that the MWCNN can effectively capture the periodic characteristics of load data. Zhifang Liao, Haihui Pan, Xiaoping Fan, Yan Zhang 0047, Li Kuang |
IEEE Internet Things J. | 1 |
| 2021 | NSDH: A Nonlinear Supervised Discrete Hashing framework for large-scale cross-modal retrieval
Zhan Yang 0001, Liu Yang 0015, Osolo Ian Raymond, Lei Zhu 0005, Wenti Huang, Zhifang Liao |
Knowl. Based Syst. | 6 |
| 2020 | ServeNet: A Deep Neural Network for Web Services ClassificationabstractAutomated service classification plays a crucial role in service discovery, selection, and composition. Machine learning has been widely used for service classification in recent years. However, the performance of conventional machine learning methods highly depends on the quality of manual feature engineering. In this paper, we present a novel deep neural network to automatically abstract low-level representation of both service name and service description to high-level merged features without feature engineering and the length limitation, and then predict service classification on 50 service categories. To demonstrate the effectiveness of our approach, we conduct a comprehensive experimental study by comparing 10 machine learning methods on 10,000 real-world web services. The result shows that the proposed deep neural network can achieve higher accuracy in classification and more robust than other machine learning methods. Yilong Yang 0001, Nafees Qamar, Peng Liu 0070, Katarina Grolinger, Weiru Wang 0001, Zhi Li 0017, Zhifang Liao |
ICWS | 7 |
| 2020 | How to Construct Software Knowledge Graph: A Case StudyabstractKnowledge graph has been more and more widely used, but software knowledge graph is still in the initial stage. To explore the way of constructing a software knowledge graph, this paper proposes a case study by using source code, related documents and relations among the information to build software knowledge graph. The source code comes from the open source project Tomcat, the related documents come from various software communities. Different approaches are designed for each type of resources to achieve entity and attribute extraction, and to categorize the relation into internal and external relations to extract. Internal relation is contained in each resource, external relation exists between different resources. And then subtree comparison and mailbox matching are used to align entities, completing the construction of software knowledge graph. Software knowledge graph is applied to knowledge retrieval and association visualization to improve the efficiency of knowledge acquisition. Wuqian Lv, Zhifang Liao, Yan Zhang 0047, Li Kuang, Shuixiu Bi |
SERVICES | 2 |
| 2020 | WLeidenRDF: RDF Data Query Method based on Semantic-Enhanced Graph-Clustering AlgorithmabstractGraph-clustering algorithms are designed to split large-scale Resource Description Framework (RDF) graphs into subgraphs to improve RDF query performance. However, triple semantics and graph structures embedded in RDF graphs are often ignored during the RDF graph partition. Accordingly, this study utilizes the Leiden algorithm for uncovering community structure to cluster RDF graphs. We propose an optimized WLeidenRDF algorithm to uncover RDF communities and cluster vertex with strong semantics. Different weights are set to predicates in RDF triples to identify semantic-relevance degree and ensure semantic connections after RDF clustering and segmentation. The experiments on WatDiv data sets demonstrate that our WLeidenRDF algorithm obtains better partitions in accordance with modularity than those of the Leiden algorithm. We implement our algorithm on Presto distributed SQL query engine. Experimental results indicate that our algorithm can substantially reduce query time by clustering RDF graphs than with other RDF query methods. Liu Yang 0015, Yiqing Feng, Zhifang Liao, Zhigang Hu 0001 |
TASE | 4 |
| 2020 | A spam worker detection approach based on heterogeneous network embedding in crowdsourcing platforms
Li Kuang, Ruyi Shi, Zhifang Liao, Xiaoxian Yang |
Comput. Networks | 4 |
| 2020 | Scalable deep asymmetric hashing via unequal-dimensional embeddings for image similarity search
Zhan Yang 0001, Osolo Ian Raymond, Wenti Huang, Zhifang Liao, Lei Zhu 0005 |
Neurocomputing | 4 |
| 2020 | Core-reviewer recommendation based on Pull Request topic model and collaborator social networkabstractPull Request (PR) is a major contributor to external developers of open-source projects in GitHub. PR reviewing is an important part of open-source software developments to ensure the quality of project. Recommending suitable candidates of reviewer to the new PRs will make the PR reviewing more efficient. However, there is not a mechanism of automatic reviewer recommendation for PR in GitHub. In this paper, we propose an automatic core-reviewer recommendation approach, which combines PR topic model with collaborators in the social network. First PR topics will be extracted from PRs by the latent Dirichlet allocation, and then the collaborator–PR network will be constructed with the connection between collaborators and PRs, and the influence of each collaborator will be calculated via the improved PageRank algorithm which combines with HITS. Finally, the relationship between topics and collaborators will also be built by the history of PR reviewing. When a new PR presents, a collaborator will be chosen as a core reviewer according to the influence of collaborators and the relationship between the new PR and collaborators. The experiment results show in the matching score calculation processing, the influence of collaborators shows higher than that with the expert, and the recommendation precision is better than 70%. Zhifang Liao, Zexuan Wu, Yanbing Li, Yan Zhang 0047, Xiaoping Fan, Jinsong Wu 0001 |
Soft Comput. | 1 |
| 2019 | TIRR: A Code Reviewer Recommendation Algorithm with Topic Model and Reviewer InfluenceabstractCode review is an important way to improve software quality and ensure project security. Pull Request (PR), as an important method of collaborative code modification in GitHub open source software community platform, is very important to find a suitable code reviewer to improve code modification efficiency for Pull Request submitted by code modifiers. In order to solve this problem, we have proposed a review recommendation algorithm based on Pull Request topic model and reviewer's influence. This algorithm has not only extracted the topic information of PR through Latent Dirichlet Allocation (LDA) method, but also analyzed the professional knowledge influence of reviewers through influence network. What’s more, it has combined the topic information of reviewers to find the appropriate PR reviewers. The experimental results based on GitHub show that the algorithm is more efficient, which can effectively reduce the time of code review and improve the recommendation accuracy. Zhifang Liao, Zexuan Wu, Jinsong Wu 0001, Yan Zhang 0047 |
GLOBECOM | 1 |
| 2019 | A Recommendation of Crowdsourcing Workers Based on Multi-community Collaboration
Zhifang Liao, Peng Lan, Yan Zhang 0047 |
ICSOC | 1 |
| 2019 | Local Core Members Aided Community Structure Detection
Xiaoping Fan, Jinsong Wu 0001, Shengzong Liu, Zhining Liao, Zhifang Liao |
Mob. Networks Appl. | 7 |
| 2019 | A Prediction Model of the Project Life-Span in Open Source Software Ecosystem
Zhifang Liao, Benhong Zhao, Shengzong Liu, Haozhi Jin, Dayu He, Liu Yang 0015, Yan Zhang 0047, Jinsong Wu 0001 |
Mob. Networks Appl. | 1 |
| 2018 | NBSL: A Supervised Classification Model of Pull Request in GithubabstractA lot of Pull Requests (PRs) appear in Github everyday, and thus it is a very important work to review these PRs quickly in Github. Labeling PRs according to the PRs classification can improve the success rate and the review efficiency. However, recent research works have shown that most of the PRs are not labeled, and if the PR is labeled, it is done manually. To solve this problem, we propose a supervised classification model combined with supervised topics model and Naive Bayes classifier, which can make the PR be classified automatically. The method creates a one-one relationship between labels and PRs, and the approach classifies most PRs automatically with the only label which record the closest topic of PRs. The experimental results show that the proposed model can reach a precision of 60% in majority situation. The proposed model may support a better result via adjusting the parameters in case. Yan Zhang 0047, Jinsong Wu 0001, Zhifang Liao, Yanbing Li |
ICC | 5 |
| 2017 | DevRank: Mining Influential Developers in GithubabstractAs the social coding is becoming increasingly popular, understanding the influence of developers can benefit various applications, such as advertisement for new projects and innovations. However, most existing works have focused only on ranking influential nodes in non-weighted and homogeneous networks, which are not able to transfer proper importance scores to the real important node. To rank developers in Github, we define developer's influence on the capacity of attracting attention which can be measured by the number of followers obtained in the future. We further defined a new method, DevRank, which ranks the developers by influence propagation through heterogeneous network constructed according to user behaviors, including "commit" and "follow". Our experiment compares the performance between DevRank and some other link analysis algorithms, the results have shown that DevRank can improve the ranking accuracy. Zhifang Liao, Haozhi Jin, Benhong Zhao, Jinsong Wu 0001, Shengzong Liu |
GLOBECOM | 1 |
| 2017 | Topic-Based Integrator Matching for Pull RequestabstractPull Request (PR) is the main method for code contributions from the external contributors in GitHub. PR review is an essential part of open source software developments to maintain the quality of software. Matching a new PR for an appropriate integrator will make the PR reviewing more effective. However, PR and integrator matching are now organized manually in GitHub. To make this process more efficient, we propose a Topic-based Integrator Matching Algorithm (TIMA) to predict highly relevant collaborators(the core developers) as the integrator to incoming PRs . TIMA takes full advantage of the textual semantics of PRs. To define the relationships between topics and collaborators, TIMA builds a relation matrix about topic and collaborators. According to the relevance between topics and collaborators, TIMA matches the suitable collaborators as the PR integrator. Zhifang Liao, Yanbing Li, Dayu He, Jinsong Wu 0001, Yan Zhang 0047, Xiaoping Fan |
GLOBECOM | 1 |
| 2016 | Markov Chain-Like Model for Prediction Service Based on Improved Hierarchical Particle Swarm Optimization Cluster AlgorithmabstractSince web pages visited by users contain a variety of data resources and the clustering algorithms frequently used for web data do not take the heterogeneous nature into account when processing the heterogeneous data, this paper proposes a new algorithm, namely IHPSOC algorithm, to cluster web log data on the basis of web log mining. Based on particle swarm optimization (PSO), IHPSOC algorithm clusters the web log data through particle swarm iteration. Based on clustering results, this paper establishes Markov chain-like models which create a corresponding Markov chain for users in each different category so as to predict the web resources in users’ need. The results of the experiments show that the proposed model gives better predication. Zhifang Liao, Tianhui Song, Li Kuang, Yan Zhang 0047, Zhining Liao |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2016 | Multimedia services quality prediction based on the association mining between context and QoS properties
Li Kuang, Zhifang Liao, Wentao Feng, Haoneng He |
Signal Process. | 2 |
| 2011 | Tensor Fold-in Algorithms for Social Tagging PredictionabstractSocial tagging predictions involve the co occurrence of users, items and tags. The tremendous growth of users require the recommender system to produce tag recommendations for millions of users and items at any minute. The triplets of users, items and tags are most naturally described by a 3D tensor, and tensor decomposition-based algorithms can produce high quality recommendations. However, each day, thousands of new users are added to the system and the decompositions must be updated daily in a online fashion. In this paper, we provide analysis of the new user problem, and present fold-in algorithms for Tucker, Para Fac, and Low-order tensor decompositions. We show that these algorithm can very efficiently compute the needed decompositions. We evaluate the fold-in algorithms experimentally on several datasets and the results demonstrate the effectiveness of these algorithms. Chris Ding, Zhifang Liao |
ICDM | 3 |
| 2011 | A Weighted Cluster Ensemble Algorithm Based on GraphabstractCluster ensemble is an effective method to improve the effect in data clustering, but the results of the existing cluster ensemble algorithms are usually not so good when they process the mixed attributes datas, the main reason is that the results of the algorithms are still dispersed. To solve this problem, this paper presents a new weighted cluster ensemble algorithm based on graph theory. It first clusters the datasets and gets cluster members, and then sets weights to each data object with a proposed ensemble function, and determines the relationship between the data-pair by setting weights to the edges between them, so it can get a weighted nearest neighbor graph. At last it does a last-clustering based on graph theory. Experiments show that the accuracy and stability of this cluster ensemble algorithm is better than other clustering ensemble algorithms. Xiaoping Fan, Yueshan Xie, Zhifang Liao, Li-min Liu |
TrustCom | 3 |
| 2003 | Determining Remote System Contention States in Query Processing over the Internet abstractIn the environment of data integration over the Internet, three major factors affect the cost of a query: network congestion situation, server contention states (workload), and data/query complexity. We concentrate on system contention states. For a remote data source, we first determine the total number of contention states of the system through applying clustering techniques to the costs of sample queries. We then develop a set of cost formulae for each of the contention states using a multiple regression process. Finally, we estimate the system's current contention state when a query is issued and using either a time slides method or a statistical method depending on the information we have about the system. Our method can accurately predict the system contention state so that the effect of the contention states on the cost of queries can be estimated precisely. Weiru Liu, Zhining Liao, Jun Hong 0001, Zhifang Liao |
Web Intelligence | 4 |