VLDB 2026 Research / reviewers in the wild / expert
Chen Zhao 0024
dblp:81/3-24
· DBLP profile ↗
14ranked-venue papers
0as first author
11since 2021 · last 2026
0009-0005-3386-0335ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Graph-Based Code Generation: Competition-Level Coding Agents with Graph-Based Reasoning and Backtracking
Huidi Zhu, JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Chen Zhao 0024 |
KSEM (3) | 5 |
| 2026 | RepoCBench: A benchmark for C-oriented repository-level code generation with large language models and agents
Kaichun Yao, Libo Zhang 0001, Chen Zhao 0024 |
Expert Syst. Appl. | 4 |
| 2026 | Detecting Malicious Packages in PyPI and NPM by Clustering Installation ScriptsabstractSoftware repositories such as PyPI and npm are vital for software development but expose users to serious security risks from malicious packages. The malicious packages often execute their payloads immediately upon installation, leading to rapid system compromise. Existing detection methods are heavily dependent on difficult-to-obtain explicit knowledge, rendering them susceptible to overlooking emergent malicious packages.In this paper, we present a lightweight and effective method, namely EMPHunter, to detect malicious packages without requiring any explicit prior knowledge. EMPHunter is founded upon two fundamental and insightful observations. First, malicious packages are considerably rarer than benign ones, and second, the functionality of installation scripts for malicious packages diverges significantly from those of benign packages, with the latter frequently forming clusters. Consequently, EMPHunter utilizes the clustering technique to group the unique installation scripts of new-uploaded packages and identifies outliers as candidate malicious packages. It then ranks the outliers according to their deviate degrees and the distance between each of them and known malicious instances, effectively highlighting potential malicious packages.With EMPHunter, we successfully identified 122 previously unknown malicious packages from a pool of 267,009 newly-uploaded PyPI and npm packages, achieving an mAP (Mean Average Precision) of 0.813 and an exceptional recall of 0.992 when auditing the top-10 rankings. All detected packages have been officially confirmed as genuine malicious package by PyPI and npm. We assert that EMPHunter offers a valuable and advantageous supplement to existing detection tools, augmenting the arsenal of software supply chain security analysis. Wentao Liang, Xiang Ling 0001, Chen Zhao 0024, JingZheng Wu, Tianyue Luo |
IEEE Trans. Software Eng. | 3 |
| 2025 | Towards Efficient Compiler Auto-tuning: Leveraging Synergistic Search SpacesabstractDetermining the optimal sequence of compiler optimization passes is challenging due to the extensive and intricate search space. Traditional auto-tuning techniques, such as iterative compilation and machine learning methods, are often limited by high computational costs and difficulties in generalizing to new programs. These approaches can be inefficient and may not fully address the varying optimization needs across different programs. This paper introduces a novel approach that leverages the synergistic relationships between optimization passes to effectively reduce the search space. By focusing on chained synergy pass pairs that jointly optimize a specific target, our method uses K-means clustering to capture common optimization patterns across programs and forms these pairs into coresets. Leveraging a supervised learning model trained on these coresets, we effectively predict the most beneficial coreset for new programs, streamlining the search for optimal sequences. By integrating various search strategies, our method quickly converges to near-optimal solutions. Our approach achieves state-of-the-art performance on ten benchmark datasets, including MiBench, CBench, NPB, and CHStone, demonstrating an average reduction of 7.5% in Intermediate Representation (IR) instruction count compared to Oz. Furthermore, this set of chained synergy pass pairs is also well-suited for iterative search studies by other researchers, as it enables achieving an average codesize reduction of 13.9% compared to Oz with a simple search strategy that takes only about 5 seconds, outperforming existing search-based techniques in the initial pass search space across five datasets. Haolin Pan, Yuanyu Wei, Mingjie Xing, Chen Zhao 0024 |
CGO | 5 |
| 2025 | HyperSF: A Hypergraph Representation Learning Method Based on Structural FusionabstractHypergraph Neural Networks (HNNs) have recently gained attention as a powerful approach for capturing high-order correlations through hypergraph-structured encoding and learning techniques. However, despite their potential, existing HNN methods often encounter over-smoothing issues, which limit their ability to effectively integrate global information while maintaining high-order structural details. This limitation compromises the overall effectiveness of these models. To tackle this challenge, we introduce a novel HNN framework called Hypergraph Structural Fusion (HyperSF). HyperSF combines the structural characteristics of both hypergraphs and graphs to effectively integrate global and local information while preserving the complex high-order structures inherent in hypergraphs. This structural fusion mechanism significantly improves model performance by ensuring that both types of information are utilized in a balanced manner. Comprehensive evaluations show that our method outperforms state-of-the-art approaches, demonstrating its effectiveness in hypergraph representation learning. Xiangfei Fang, Chengying Huan, Boying Wang, Shaonan Ma, Heng Zhang 0005, Chen Zhao 0024 |
ICASSP | 6 |
| 2025 | HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold NetworksabstractHypergraph representation learning has garnered increasing attention across various domains due to its capability to model high-order relationships. Traditional methods often rely on hypergraph neural networks (HNNs) employing messagepassing mechanisms to aggregate vertex and hyperedge features. However, these methods are constrained by their dependence on hypergraph topology, leading to the challenge of imbalanced information aggregation, where high-degree vertices tend to aggregate redundant features, while low-degree vertices often struggle to capture sufficient structural features. To overcome the above challenges, we introduce HyperKAN, a novel framework for hypergraph representation learning that transcends the limitations of message-passing techniques. Hyper- KAN begins by encoding features for each vertex and then leverages Kolmogorov-Arnold Networks (KANs) to capture complex nonlinear relationships. By adjusting structural features based on similarity, our approach generates refined vertex representations that effectively addresses the challenge of imbalanced information aggregation. Experiments conducted on the real-world datasets demonstrate that HyperKAN significantly outperforms stateof-the-art HNN methods, achieving nearly a 9% performance improvement on the Senate dataset. Xiangfei Fang, Boying Wang, Chengying Huan, Shaonan Ma, Heng Zhang 0005, Chen Zhao 0024 |
ICASSP | 6 |
| 2025 | OTM: Efficient k-Order-Based Core Maintenance in Large-Scale Dynamic HypergraphsabstractThe k -core model has garnered widespread adoption for preserving essential cohesive subgraphs owing to its linear-time computability, making it particularly suitable for hypergraph analysis. However, considering the continuously evolving characteristics of real-world hypergraphs, recent research efforts have focused on developing efficient algorithms that can maintain the core value of each vertex amid structural alterations. Despite these efforts, frequent insertions and deletions in dynamic hypergraphs continue to pose significant inefficiencies, primarily due to the increased traversal overhead incurred by hyperedge insertion algorithms. This exacerbates performance disparities between handling hyperedge insertions and deletions, underscoring the persistent challenge of effective k -core analysis in hypergraphs. To effectively address these challenges, we have gained key insights that enable us to define a specific order, termed the hypergraph k -order, which significantly reduces redundant vertex traversal and narrows down the search space during hyperedge insertions. Based on the proposed hypergraph k -order, we define two indices, the order index and the pivotal index, aimed at minimizing traversal costs and expediting the hyperedge insertion algorithm. Moreover, it is essential to recognize that the recomputation of the support degree ( sd ) for all vertices following each hyperedge deletion can significantly diminish the performance efficiency of deletion algorithms. To address this, we introduce an optimized approach that leverages the incremental maintenance of the support degree ( sd ) value to expedite the hyperedge deletion process. By leveraging these optimizations, we introduce a novel Order-based Traversal core Maintenance methodology, designated as OTM , which markedly enhances the efficiency of core maintenance in dynamic hypergraphs. Our comprehensive evaluation, which covers 12 real-world hypergraph datasets and a synthetic dataset, reveals that OTM achieves staggering speedup, outperforming the state-of-the-art approach with a 41,420 \(\times\) speedup in the insertion algorithm and 8,284 \(\times\) speedup in the deletion algorithm, underscoring its remarkable efficiency and effectiveness. Xiangfei Fang, Chengying Huan, Heng Zhang 0005, Yongchao Liu 0004, Shaonan Ma, Chen Zhao 0024 |
ACM Trans. Knowl. Discov. Data | 7 |
| 2024 | BPDO: Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text DetectionabstractArbitrary shape scene text detection is of great importance in scene understanding tasks. Due to the complexity and diversity of text in natural scenes, existing scene text algorithms have limited accuracy for detecting arbitrary shape text. In this paper, we propose a novel arbitrary shape scene text detector through boundary points dynamic optimization(BPDO). The proposed model is designed with a text aware module (TAM) and a boundary point dynamic optimization module (DOM). Specifically, the model designs a text aware module based on segmentation to obtain boundary points describing the central region of the text by extracting a priori information about the text region. Then, based on the idea of deformable attention, it proposes a dynamic optimization model for boundary points, which gradually optimizes the exact position of the boundary points based on the information of the adjacent region of each boundary point. Experiments on CTW-1500, Total-Text, and MSRATD500 datasets show that the model proposed in this paper achieves a performance that is better than or comparable to the state-of-the-art algorithm, proving the effectiveness of the model. Jinzhi Zheng, Libo Zhang 0001, Chen Zhao 0024 |
ICASSP | 4 |
| 2024 | Text Region Multiple Information Perception Network for Scene Text DetectionabstractSegmentation-based scene text detection algorithms can handle arbitrary shape scene texts and have strong robustness and adaptability, so it has attracted wide attention. Existing segmentation-based scene text detection algorithms usually only segment the pixels in the center region of the text, while ignoring other information of the text region, such as edge information, distance information, etc., thus limiting the detection accuracy of the algorithm for scene text. This paper proposes a plug-and-play module called the Region Multiple Information Perception Module (RMIPM) to enhance the detection performance of segmentation-based algorithms. Specifically, we design an improved module that can perceive various types of information about scene text regions, such as text foreground classification maps, distance maps, direction maps, etc. Experiments on MSRA-TD500 and TotalText datasets show that our method achieves comparable performance with current state-of-the-art algorithms. Jinzhi Zheng, Libo Zhang 0001, Chen Zhao 0024 |
ICASSP | 4 |
| 2023 | CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
Jinzhi Zheng, Ruyi Ji, Libo Zhang 0001, Chen Zhao 0024 |
ICONIP (6) | 5 |
| 2021 | Multi-peak Graph-based Multi-instance Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD), aiming to detect objects with only image-level annotations, has become one of the research hotspots over the past few years. Recently, much effort has been devoted to WSOD for the simple yet effective architecture and remarkable improvements have been achieved. Existing approaches using multiple-instance learning usually pay more attention to the proposals individually, ignoring relation information between proposals. Besides, to obtain pseudo-ground-truth boxes for WSOD, MIL-based methods tend to select the region with the highest confidence score and regard those with small overlap as background category, which leads to mislabeled instances. As a result, these methods suffer from mislabeling instances and lacking relations between proposals, degrading the performance of WSOD. To tackle these issues, this article introduces a multi-peak graph-based model for WSOD. Specifically, we use the instance graph to model the relations between proposals, which reinforces multiple-instance learning process. In addition, a multi-peak discovery strategy is designed to avert mislabeling instances. The proposed model is trained by stochastic gradients decent optimizer using back-propagation in an end-to-end manner. Extensive quantitative and qualitative evaluations on two publicly challenging benchmarks, PASCAL VOC 2007 and PASCAL VOC 2012, demonstrate the superiority and effectiveness of the proposed approach. Ruyi Ji, Ze-Yu Liu 0011, Libo Zhang 0001, Jianwei Liu 0006, Chen Zhao 0024 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2020 | Attention Convolutional Binary Neural Tree for Fine-Grained Visual CategorizationabstractFine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for weakly supervised FGVC. Specifically, we incorporate convolutional operations along edges of the tree structure, and use the routing functions in each node to determine the root-to-leaf computational paths within the tree. The final decision is computed as the summation of the predictions from leaf nodes. The deep convolutional operations learn to capture the representations of objects, and the tree structure characterizes the coarse-to-fine hierarchical feature learning process. In addition, we use the attention transformer module to enforce the network to capture discriminative features. The negative log-likelihood loss is used to train the entire network in an end-to-end fashion by SGD with back-propagation. Several experiments on the CUB-200-2011, Stanford Cars and Aircraft datasets demonstrate that the proposed method performs favorably against the state-of-the-arts. Ruyi Ji, Longyin Wen, Libo Zhang 0001, Dawei Du, Chen Zhao 0024, Xianglong Liu 0001, Feiyue Huang |
CVPR | 6 |
| 2020 | Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu |
ECCV (13) | 6 |
| 2017 | EpCom: A parallel community detection approach for epidemic diffusion over social networksabstractDetecting community structure in epidemics networks is crucial for the assessment of epidemic dynamics and effective control of disease spread by targeting at the individuals bridging communities. Common community detection models (e.g., cut-criteria and modularity-criteria based model) are efficient in optimal quality of network partitions. However, most of the approaches fail to consider the dynamic infected possibility in person-to-person interactions. In addition, they present high computational complexity, which was limited by the scale of networks and the performance of hardware platform. In this paper, we propose a Jaccard distance based community detection model by considering both the quality of network partitions and the dynamics of infected interacts (i.e., edges) between two individuals in epidemic diffusion. Then, we design a novel parallel approach based on the high parallism of GPU, called EpCom, for boosting the performance and scalability of parallel community detection over large-scale epidemic networks. From the evaluation results, the proposed GPU-based implementation EpCom exhibits great performance and achieves maximum 604 million TEPS (traversed edges per second), which corresponds to up to 54.2 times and 15.6 times than CPU-based NCut and Louvain approaches separately. Heng Zhang 0005, Libo Zhang 0001, Da Cheng, Chen Zhao 0024 |
BIBM | 5 |