Zicheng Ma

dblp:309/8537 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
2 papers
Information extraction and text analysis · 60% Question answering and dialogue systems · 30% Language models and text generation · 10%
Software engineering, system software, and programming languages
1 paper
Program verification · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein function prediction
1.922026
Interleaved Tool-Call Reasoning for Protein Function Understanding · ACL (1) 2026
Prot2Chat: protein large language model with early fusion of text, sequence, and structure · Bioinform. 2025
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein function analysis
1.012026
Interleaved Tool-Call Reasoning for Protein Function Understanding · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
natural language interface
0.912025
AllHands :Ask Me Anything on Large-scale Verbatim Feedback via Large Language Models · ICDE 2025
Natural language and speech › Information extraction and text analysis
topic model
0.912025
AllHands :Ask Me Anything on Large-scale Verbatim Feedback via Large Language Models · ICDE 2025
Bioinformatics and computational biology › machine learning for biology
multimodal protein representation
0.912025
Prot2Chat: protein large language model with early fusion of text, sequence, and structure · Bioinform. 2025
Bioinformatics and computational biology › structural biology
protein structure and function
0.912025
Prot2Chat: protein large language model with early fusion of text, sequence, and structure · Bioinform. 2025
Program verification › temporal logic verification
liveness verification
0.812024
Anvil: Verifying Liveness of Cluster Management Controllers · OSDI 2024
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.812024
Anvil: Verifying Liveness of Cluster Management Controllers · OSDI 2024

Methods — techniques the papers use, named apart from their topics

tool-call reasoning · 2.0interleaved reasoning · 2.0large language model · 1.7formal verification · 1.5topic modeling · 0.9low-rank adaptation · 0.9code generation · 0.9ProteinMPNN · 0.9
YearPublicationVenuePosition
2026 Interleaved Tool-Call Reasoning for Protein Function Understanding
abstract
Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Ziqiang Cao, Guohong Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chuanliu Fan, Zicheng Ma, Huanran Meng, Jun Zhang 0069, Ziqiang Cao, Guohong Fu
ACL (1)2
2026 Bidirectional GPT
Chuanliu Fan, Zicheng Ma, Jun Zhang 0071, Yiqin Gao, Ziqiang Cao, Guohong Fu
Inf. Process. Manag.2
2025 LlaMol: A Unified Molecule Designer via Preference Ranking and Numerical Enhancement
abstract
Goal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints from scratch, is a crucial yet challenging task in drug discovery. Existing research often relies on separate predictors for distinct properties and struggles with integrating substructure constraints due to the complexities involved in modeling structural information via multitask learning. This separation necessitates a dedicated prediction model for each constraint, limiting the flexibility and posing challenges for realworld applications. To address these limitations, we propose a unified framework for molecular design that incorporates multiple property and substructure constraints, leveraging LLMs to handle diverse constraint settings within a single model. We first integrate feedback learning derived from preference ranking to eliminate the need for separate property predictors. Then, we enhance the model's ability to follow numerical instructions by introducing a unified numerical encoding into the prompt. We conduct extensive experiments across single-property, substructureproperty, and multi-property constrained tasks. Experimental results demonstrate that LlaMol consistently outperforms state-of-the-art baselines across various constraint settings. Notably, in the multi-objective binding affinity maximization task, LlaMol achieves a significantly lower$\mathrm{K}_{\mathrm{D}}$value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76 %. These results underscore the effectiveness and versatility of LLM-based frameworks for molecule generation under complex constraints.
Chuanliu Fan, Zicheng Ma, Jun Zhang 0069, Ziqiang Cao, Yiqin Gao, Guohong Fu
BIBM2
2025 AllHands :Ask Me Anything on Large-scale Verbatim Feedback via Large Language Models
abstract
Verbatim feedback constitutes a valuable repository of user experiences, opinions, and requirements, crucial for data engineering and software development. Extracting meaningful insights from large-scale feedback data presents a significant challenge. This paper introduces Allhands, an innovative ana-lytic framework that transforms traditional large-scale feedback analysis tasks through a natural language interface, leveraging large language models (LLMs). Allhands performs initial classification and topic modeling on feedback to convert it into a structurally augmented format, enhancing accuracy, robustness and generalization with the aid of LLMs. Subsequently, an LLM-based code-first agent interprets users' diverse natural language questions about the feedback, automatically translates them into executable call of analytic tools or code, and delivers comprehensive multi-modal responses, including text, code, tables, and images. This eliminates the need for developing individual feedback analytic tools for each request, reducing human effort and making the system more accessible and flexible to users. We evaluate Allhands across three diverse feedback datasets, demonstrating its superior efficacy in all stages of analysis, from classification and topic modeling to providing an “ask me anything” experience with comprehensive, accurate, and human-readable responses. To the best of our knowl-edge, Allhands is the first comprehensive feedback analysis framework supporting diverse and customized insight extraction requirements through a natural language interface.
Chaoyun Zhang, Zicheng Ma, Shilin He, Si Qin, Minghua Ma, Xiaoting Qin, Yu Kang 0006, Yuyi Liang, Xiaoyu Gou, Yajie Xue, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066
ICDE2
2025 Prot2Chat: protein large language model with early fusion of text, sequence, and structure
abstract
MOTIVATION: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of training parameters, limited flexibility of classification-based methods, and the lack of systematic evaluation metrics for protein question answering systems. To tackle these issues, we propose the Prot2Chat framework. RESULTS: We modified ProteinMPNN to encode protein sequence and structural information in a unified way. We used a large language model (LLM) to encode questions into vectors and developed a protein-text adapter to compress protein information into virtual tokens based on these vectors, achieving the early fusion of text and protein information. Finally, the same LLM reads the virtual tokens and the questions to generate answers. To optimize training efficiency, we froze the encoder and employed low-rank adaptation (LoRA) techniques for the LLM. Experiments on two datasets show that both automated metrics and expert evaluations demonstrate the superior performance of our model, and zero-shot prediction results highlight its generalization ability. We have developed an easy-to-use web interactive platform and a rapid installation option, allowing users to swiftly engage with Prot2Chat. AVAILABILITY AND IMPLEMENTATION: The models and codes are available at https://github.com/wangzc1233/Prot2Chat.
Zhicong Wang, Zicheng Ma, Ziqiang Cao, Changlong Zhou, Jun Zhang 0071, Yiqin Gao
Bioinform.2
2024 Anvil: Verifying Liveness of Cluster Management Controllers
Xudong Sun 0013, Jiawei Tyler Gu, Zicheng Ma, Tej Chajed, Jon Howell, Andrea Lattuada 0001, Oded Padon, Lalith Suresh 0001, Adriana Szekeres, Tianyin Xu
OSDI4
2024 MTSecurity: Privacy-Preserving Malicious Traffic Classification Using Graph Neural Network and Transformer
abstract
Encrypting network traffic is an effective means of safeguarding user privacy and sensitive information. However, it also introduces potential vulnerabilities that can be exploited by network attackers, posing significant security risks to the Internet. In response to the challenge of low accuracy in existing methods for classifying encrypted malicious traffic, we propose a novel approach named MTSecurity, which leverages Transformer and Graph Neural Network technologies. This method automatically extracts raw byte features and graph-based traffic interaction features from encrypted malicious flows, combining them to substantially enhance the classification accuracy of encrypted malicious traffic. Furthermore, we introduce a graph structure called the Malicious Traffic Interaction Graph (MTIG) for representing encrypted malicious traffic. MTIG is based on the client-server interaction process and incorporates multidimensional traffic features. Experimental results demonstrate that the proposed MTSecurity model consistently performs well across different datasets, surpassing state-of-the-art methods. It achieves an accuracy of 0.9946 and an F1 score of 0.9940 on the MCFP dataset, and an accuracy of 0.9948 with an F1 score of 0.9934 on the USTC-TFC dataset.
Xinyun Jiang, Yulin Lei, Weiheng Liang, Zicheng Ma
IEEE Trans. Netw. Serv. Manag.5
2021 A novel framework for image-based malware detection with a deep neural network
Yifei Jian, Hongbo Kuang, Chenglong Ren, Zicheng Ma, Haizhou Wang 0001
Comput. Secur.4