Zicheng Zhao

dblp:193/7982 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capability of Large Language Models (LLMs). Current RLVR approaches typically conduct training across all generated tokens, but neglect to explore which tokens (e.g., prefix tokens) actually contribute to reasoning. This uniform training strategy spends substantial effort on optimizing low-return tokens, which in turn impedes the potential improvement from high-return tokens and reduces overall training effectiveness. To address this issue, we propose a novel RLVR approach called Progressive Prefix-token Policy Optimization (PPPO), which highlights the significance of the prefix segment of generated outputs. Specifically, inspired by the well-established human thinking theory of Path Dependence, where early-stage thoughts substantially constrain subsequent thinking trajectory, we identify an analogous phenomenon in LLM reasoning termed Beginning Lock-in Effect (BLE). PPPO leverages this finding by focusing its optimization objective on the prefix reasoning process of LLMs. This targeted optimization strategy can positively influence subsequent reasoning processes, and ultimately improve final results. To improve the learning effectiveness of LLMs on how to start reasoning with high quality, PPPO introduces two training strategies: (a) Progressive Prefix Retention, which shapes a progressive learning process by increasing the proportion of retained prefix tokens during training; (b) Continuation Accumulated Reward, which mitigates reward bias by sampling multiple continuations for one prefix token sequence, and accumulating their scores as the reward signal. Extensive experimental results on various reasoning tasks (e.g., math, physics, chemistry, and biology) demonstrate that our proposed PPPO outperforms representative RLVR methods, with the accuracy improvements of 18.02% on only 26.17% training tokens.
Yiliu Sun, Zicheng Zhao, Yang Wei 0003, Yanfang Zhang 0001, Chen Gong 0002
AAAI2
2026 CogStream: Context-guided Streaming Video Question Answering
abstract
Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual information. Existing paradigms feed all available historical contextual information into Vid-LLMs, resulting in a significant computational burden for visual data processing. Furthermore, the inclusion of irrelevant context distracts models from key details. This paper introduces a challenging task called Context-guided Streaming Video Reasoning (CogStream), which simulates real-world streaming video scenarios, requiring models to identify the most relevant historical contextual information to deduce answers for questions about the current stream. To support CogStream, we present a densely annotated dataset featuring extensive and hierarchical question-answer pairs, generated by a semi-automatic pipeline. Additionally, we present CogReasoner as a baseline model. It effectively tackles this task by leveraging visual stream compression and historical dialogue retrieval. Extensive experiments prove the effectiveness of this method.
Zicheng Zhao, Kangyu Wang, Rui Qian 0001, Weiyao Lin, Huabin Liu 0001
AAAI1
2026 Joint structural priors and diversity for multi-view subspace clustering
Zicheng Zhao
Knowl. Based Syst.3
2026 NeRS: Negative Relational Smoothing for Graph Contrastive Learning
Sheng Wan, Shougang Ren, Zicheng Zhao, Chen Gong 0002
Pattern Recognit.3
2026 TraceCluster: A Lightweight and Adaptive Clustering-Based Subgraph Attention Network for APT Detection in Provenance Graphs
abstract
Provenance graph-based anomaly detection, particularly for Advanced Persistent Threat (APT) detection, addresses the issues of large-scale graphs and data imbalance. However, existing methods struggle with information loss, high computational complexity, and low detection accuracy. To address the above challenges, this paper proposes TraceCluster, a lightweight and adaptive clustering-based Subgraph Attention Network (SAN) for APT detection in provenance graph. TraceCluster mitigates the neighborhood explosion problem by clustering nodes to partition large-scale graphs, thus reducing reliance on the global graph while preserving local neighborhood information. Furthermore, the method dynamically models complex inter-node dependencies within subgraphs. It employs an attention mechanism to adaptively highlight the most relevant connections. This enhances node representations and improves overall feature extraction. This design substantially reduces memory consumption and avoids the high computational complexity of global graph processing. In addition, an adaptive category-weighting loss function assigns variable weights to different classes, improving the detection of rare and anomalous behaviors. Experimental results show that on the OpTC dataset, the currently faster method is 37-fold and 3-fold slower than our approach in terms of inference time respectively. Furthermore, in the nine real-world scenarios of four evaluated datasets, TraceCluster outperforms state-of-the-art (SOTA) approaches in terms of overall performance, especially in node-level APT detection tasks.
Lijuan Xu 0001, Zicheng Zhao, Dawei Zhao 0001, Zhen Wang 0004, Chunpeng Ge 0001
IEEE Trans. Inf. Forensics Secur.2
2026 Graph Stochastic Neural Process for Inductive Few-shot Knowledge Graph Completion
abstract
Knowledge graphs (KGs) store enormous facts as relationships between entities. Due to the long-tailed distribution of relations and the incompleteness of KGs, there is growing interest in few-shot knowledge graph completion (FKGC). Existing FKGC methods often assume the existence of all entities in KGs, which may not be practical since new relations and entities can emerge over time. Therefore, we focus on a more challenging task called inductive few-shot knowledge graph completion (I-FKGC), where both relations and entities during the test phase are unknown before. Inspired by the idea of inductive reasoning, we cast I-FKGC as an inductive reasoning problem. Specifically, we propose a novel Graph Stochastic Neural Process ( GS-NP ) approach, which consists of two major modules. In the first module, to obtain a generalized hypothesis (e.g., shared subgraph), we present a neural process-based hypothesis extractor that models the joint distribution of hypothesis, from which we can sample a hypothesis for predictions. In the second module, based on the hypothesis, we propose a graph stochastic attention-based predictor to test if the triple in the query set aligns with the extracted hypothesis. Meanwhile, the predictor can generate an explanatory subgraph identified by the hypothesis. Finally, the training of these two modules is seamlessly combined into a unified objective function, of which the effectiveness is verified by theoretical analyses as well as empirical studies. Extensive experiments on three public datasets demonstrate that our method outperforms existing methods and derives new state-of-the-art performance.
Zicheng Zhao, Linhao Luo, Shirui Pan, Chengqi Zhang, Chen Gong 0002
ACM Trans. Intell. Syst. Technol.1
2025 Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models
abstract
Large language models (LLMs) have demonstrated impressive reasoning abilities, but they still struggle with faithful reasoning due to knowledge gaps and hallucinations. To address these issues, knowledge graphs (KGs) have been utilized to enhance LLM reasoning through their structured knowledge. However, existing KG-enhanced methods, either retrieval-based or agent-based, encounter difficulties in accurately retrieving knowledge and efficiently traversing KGs at scale. In this work, we introduce graph-constrained reasoning (GCR), a novel framework that bridges structured knowledge in KGs with unstructured reasoning in LLMs. To eliminate hallucinations, GCR ensures faithful KG-grounded reasoning by integrating KG structure into the LLM decoding process through KG-Trie, a trie-based index that encodes KG reasoning paths. KG-Trie constrains the decoding process, allowing LLMs to directly reason on graphs and generate faithful reasoning paths grounded in KGs. Additionally, GCR leverages a lightweight KG-specialized LLM for graph-constrained reasoning alongside a powerful general LLM for inductive reasoning over multiple reasoning paths, resulting in accurate reasoning with zero reasoning hallucination. Extensive experiments on several KGQA benchmarks demonstrate that GCR achieves state-of-the-art performance and exhibits strong zero-shot generalizability to unseen KGs without additional training.
Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Yuan-Fang Li, Chen Gong 0002, Shirui Pan
ICML2
2025 GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) has proven effective in integrating knowledge into large language models (LLMs). However, conventional RAGs struggle to capture complex relationships between pieces of knowledge, limiting their performance in intricate reasoning that requires integrating knowledge from multiple sources. Recently, graph-enhanced retrieval augmented generation (GraphRAG) builds a graph structure to explicitly model these relationships, enabling more effective and efficient retrievers. Nevertheless, its performance is still hindered by the noise and incompleteness within the graph structure. To address this, we introduce GFM-RAG, a novel graph foundation model (GFM) for retrieval augmented generation. GFM-RAG is powered by an innovative graph neural network that reasons over graph structure to capture complex query-knowledge relationships. The GFM with 8M parameters undergoes a two-stage training process on large-scale datasets, comprising 60 knowledge graphs with over 14M triples and 700k documents. This results in impressive performance and generalizability for GFM-RAG, making it the first graph foundation model applicable to unseen datasets for retrieval without any fine-tuning required. Extensive experiments on three multi-hop QA datasets and seven domain-specific RAG datasets demonstrate that GFM-RAG achieves state-of-the-art performance while maintaining efficiency and alignment with neural scaling laws, highlighting its potential for further improvement.
Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Dinh Q. Phung, Chen Gong 0002, Shirui Pan
NeurIPS2
2025 ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
abstract
Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances, each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks.
Rui Xu 0026, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang 0001, Jiangjie Chen, Yinghui Xu 0001
NeurIPS3
2025 Graph Retrieval-Augmented LLM for Conversational Recommendation Systems
Zhangchi Qiu, Linhao Luo, Zicheng Zhao, Shirui Pan, Alan Wee-Chung Liew
PAKDD (3)3
2025 AJSAGE: A intrusion detection scheme based on Jump-Knowledge Connection To GraphSAGE
Lijuan Xu 0001, Zicheng Zhao, Dawei Zhao 0001, Xin Li 0002, Xiyu Lu, Dingyu Yan
Comput. Secur.2
2025 DSNet: Dynamic Stitchable Neural Network for Hyperspectral Image Classification
abstract
Hyperspectral image classification (HSIC) aims to identify land cover categories by leveraging the spectral and spatial information contained in hyperspectral images (HSI). Currently, many deep learning approaches utilize dual-branch networks to process spectral and spatial data separately, followed by the application of specialized modules to facilitate feature interaction or fusion. However, the design of these modules demands considerable time and effort from researchers and may not adequately capture the inherent relationships between independent spatial and spectral features in a dynamic manner. To address these issues, we propose the dynamic stitchable neural network (DSNet) for HSIC. While the DSNet maintains a dual-branch structure, it operates without traditional feature fusion or interaction. Instead, it employs a stitching network approach to integrate the two branches. Specifically, a spatial-spectral stitching module is presented to incorporates multiple stitching layers at various positions between the two network branches, creating new stitched networks that retain the strengths of both original networks. Additionally, a reinforcement learning-based strategy is designed for dynamically selecting stitching positions tailored to specific datasets, enabling the model to adaptively optimize the integration of spatial and spectral features. Recognizing the effectiveness of vision transformer (ViT) in learning spatial information and the capability of 1D convolutional neural network (1DCNN) in capturing spectral details, the DSNet directly stitches these two networks together. This fusion maximizes the utilization of both foundational networks, yielding a new hybrid network that delivers exceptional performance while also alleviating the burden on researchers to develop new architectures from scratch. Extensive experiments and analyses conducted on three public HSI datasets demonstrate the superiority of the proposed method, validating the effectiveness of our innovative modules. The codes of this work will be available from the website: https://github.com/ZZC/IEEE-TGRS-DSNet.
Shou Feng, Zicheng Zhao, Bobo Xi, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Identify gestational diabetes mellitus by deep learning model from cell-free DNA at the early gestation stage
abstract
Gestational diabetes mellitus (GDM) is a common complication of pregnancy, which has significant adverse effects on both the mother and fetus. The incidence of GDM is increasing globally, and early diagnosis is critical for timely treatment and reducing the risk of poor pregnancy outcomes. GDM is usually diagnosed and detected after 24 weeks of gestation, while complications due to GDM can occur much earlier. Copy number variations (CNVs) can be a possible biomarker for GDM diagnosis and screening in the early gestation stage. In this study, we proposed a machine-learning method to screen GDM in the early stage of gestation using cell-free DNA (cfDNA) sequencing data from maternal plasma. Five thousand and eighty-five patients from north regions of Mainland China, including 1942 GDM, were recruited. A non-overlapping sliding window method was applied for CNV coverage screening on low-coverage (~0.2×) sequencing data. The CNV coverage was fed to a convolutional neural network with attention architecture for the binary classification. The model achieved a classification accuracy of 88.14%, precision of 84.07%, recall of 93.04%, F1-score of 88.33% and AUC of 96.49%. The model identified 2190 genes associated with GDM, including DEFA1, DEFA3 and DEFB1. The enriched gene ontology (GO) terms and KEGG pathways showed that many identified genes are associated with diabetes-related pathways. Our study demonstrates the feasibility of using cfDNA sequencing data and machine-learning methods for early diagnosis of GDM, which may aid in early intervention and prevention of adverse pregnancy outcomes.
Zicheng Zhao, Yousheng Yan, Wentao Yue, Ruixia Liu, Hailong Feng, Yujiao Chen, Bairong Shen, Lijian Zhao, Chenghong Yin
Briefings Bioinform.3
2023 Towards Few-Shot Inductive Link Prediction on Knowledge Graphs: A Relational Anonymous Walk-Guided Neural Process Approach
Zicheng Zhao, Linhao Luo, Shirui Pan, Nguyen Quoc Viet Hung, Chen Gong 0002
ECML/PKDD (3)1
2021 Deep Transport Network for Unsupervised Video Object Segmentation
abstract
The popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between the input modalities while suppressing the distracting signals via optimal structural matching. Given a video frame, we extract the dense local features from the RGB image and optical flow, and treat them as two complex structured representations. The Wasserstein distance is then employed to compute the global optimal flows to transport the features in one modality to the other, where the magnitude of each flow measures the extent of the alignment between two local features. To plug the structural matching into a two-stream network for end-to-end training, we factorize the input cost matrix into small spatial blocks and design a differentiable long-short Sinkhorn module consisting of a long-distant Sinkhorn layer and a short-distant Sinkhorn layer. We integrate the module into a dedicated two-stream network and dub our model TransportNet. Our experiments show that aligning motion-appearance yields the state-of-the-art results on the popular video object segmentation datasets.
Kaihua Zhang 0001, Zicheng Zhao, Dong Liu 0002, Qingshan Liu 0001, Bo Liu 0005
ICCV2
2020 LDscaff: LD-based scaffolding of de novo genome assemblies
abstract
BACKGROUND: Genome assembly is fundamental for de novo genome analysis. Hybrid assembly, utilizing various sequencing technologies increases both contiguity and accuracy. While such approaches require extra costly sequencing efforts, the information provided millions of existed whole-genome sequencing data have not been fully utilized to resolve the task of scaffolding. Genetic recombination patterns in population data indicate non-random association among alleles at different loci, can provide physical distance signals to guide scaffolding. RESULTS: In this paper, we propose LDscaff for draft genome assembly incorporating linkage disequilibrium information in population data. We evaluated the performance of our method with both simulated data and real data. We simulated scaffolds by splitting the pig reference genome and reassembled them. Gaps between scaffolds were introduced ranging from 0 to 100 KB. The genome misassembly rate is 2.43% when there is no gap. Then we implemented our method to refine the Giant Panda genome and the donkey genome, which are purely assembled by NGS data. After LDscaff treatment, the resulting Panda assembly has scaffold N50 of 3.6 MB, 2.5 times larger than the original N50 (1.3 MB). The re-assembled donkey assembly has an improved N50 length of 32.1 MB from 23.8 MB. CONCLUSIONS: Our method effectively improves the assemblies with existed re-sequencing data, and is an potential alternative to the existing assemblers required for the collection of new data.
Zicheng Zhao, Yingxiao Zhou, Shuai Wang 0036, Xiuqing Zhang, Changfa Wang, Shuaicheng Li 0001
BMC Bioinform.1
2016 Finding more effective microsatellite markers for forensics
abstract
Published by the Combined DNA Index System (CODIS) program of the Federal Bureau of Investigation (FBI) in 1997, the 13 core short tandem repeat (STR) loci are widely adopted as genetic markers in forensic applications, e.g., identity testing and paternity testing. However, these loci may be biased and suffer from reduced sensitivities towards specific population groups. In addition, the rapid growth of entries in forensic databases raises the chance of random hits, which can cause false judgments of innocents as criminals. A solution to these problems is to introduce more effective STR markers. The availability of whole genome sequencing enables us to identify more reliable STR markers for forensic applications computationally. In this study, we proposed an algorithm to identify STR markers with high discriminative abilities from the next-generation sequencing (NGS) data. Our algorithm could select a customized set of loci for a given population with pre-specified discriminative thresholds. We have applied the method to 320 Chinese individuals from the 1,000 Genomes Project and obtained various numbers of loci, which were able to statistically identify an individual worldwide and had higher combined powers of discrimination (CPD) and combined probabilities of exclusion (CPE) than the existing CODIS 13 loci. For identity testing, the mean frequency of DNA profile (FDP) with the selected 11 STRs was smaller than that with CODIS 13 STRs by student's t-test. With more loci, much smaller FDPs were obtained. The database matching probabilities (DMP) for selected loci were also lower than that for CODIS 13 STRs in a database with 10 billion entries. Moreover, the selected loci were able to provide considerably low chance of random profile matches so that statistically no false judgments could occur. The selected loci also reduced the risk of random allele matches when doing the familial search, with lower random allele matching probabilities. In addition, the selected STRs were statistically better than CODIS STRs for paternity testing in our simulated data, with lower probabilities of false inclusions and exclusions.
Bowen Tan, Zicheng Zhao, Shengbin Li, Shuaicheng Li 0001
BIBM3
2016 Eliminating heterozygosity from reads through coverage normalization
abstract
Heterozygosity has long plagued genome assembly. A wealth of sophisticated algorithms and procedures have been introduced for their treatment in genome assembly. In this paper, we propose a method called Hank (Heterozygosity assimilation through normalized k-mers) for this purpose. The method eliminates heterozygosity from reads by modifying the k-mers in heterozygous regions to increase their coverages to the levels of those in homozygous regions. When evaluated on simulated Illumina data at levels of heterozygosity from 0.1% to 2.0%, Hank was able to remove 80~96% of the heterozygosity. We also examined the effects of the corrections on de novo genome assembly using SOAPdenovo2, ALLPATH-LG and Platanus. All three methods improved in performance using the treated k-mers (we do not include these assembly results in this manuscript due to space constraint).
Zicheng Zhao, Yen Kaow Ng, Xiaodong Fang, Shuaicheng Li 0001
BIBM1