Zhen Bi

dblp:279/8441 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes
abstract
Detecting deepfakes has been an increasingly important topic, especially given the rapid development of AI generation techniques. In this paper, we ask: How can we build a universal detection framework that is effective for most facial deepfakes? One significant challenge is the wide variety of deepfake generators available, resulting in varying forgery artifacts (e.g., lighting inconsistency, color mismatch, etc). But should we ``teach" the detector to learn all these artifacts separately? It is impossible and impractical to elaborate on them all. So the core idea is to pinpoint the more common and general artifacts across different deepfakes. Accordingly, we categorize deepfake artifacts into two distinct yet complementary types: Face Inconsistency Artifacts (FIA) and Up-Sampling Artifacts (USA). FIA arise from the challenge of generating all intricate details, inevitably causing inconsistencies between the complex facial features and relatively uniform surrounding areas. USA, on the other hand, are the inevitable traces left by the generator's decoder during the up-sampling process. This categorization stems from the observation that all existing deepfakes typically exhibit one or both of these artifacts. To achieve this, we propose a new data-level pseudo-fake creation framework that constructs fake samples with only the FIA and USA, without introducing extra less-general artifacts. Specifically, we employ a super-resolution to simulate the USA, while utilise image-level self-blending on diverse facial regions to create the FIA. We surprisingly found that, with this intuitive design, a standard image classifier trained only with our pseudo-fake data can non-trivially generalize well to previously unseen deepfakes.
Yize Chen, Qinglang Guo, Zhen Bi, Yong Liao 0003
NeurIPS6
2024 When Do Program-of-Thought Works for Reasoning?
abstract
In the realm of embodied artificial intelligence, the reasoning capabilities of Large Language Models (LLMs) play a pivotal role. Although there are effective methods like program-of-thought prompting for LLMs which uses programming language to tackle complex reasoning tasks, the specific impact of code data on the improvement of reasoning capabilities remains under-explored. To address this gap, we propose complexity-impacted reasoning score CIRS, which combines structural and logical attributes, to measure the correlation between code and reasoning abilities. Specifically, we use the abstract syntax tree to encode the structural information and calculate logical complexity by considering the difficulty and the cyclomatic complexity. Through an empirical analysis, we find not all code data of complexity can be learned or understood by LLMs. Optimal level of complexity is critical to the improvement of reasoning abilities by program-aided prompting. Then we design an auto-synthesizing and stratifying algorithm, and apply it to instruction generation for mathematical reasoning and code data filtering for code generation tasks. Extensive results demonstrates the effectiveness of our proposed approach.
Zhen Bi, Ningyu Zhang 0001, Yinuo Jiang, Shumin Deng, Guozhou Zheng, Huajun Chen
AAAI1
2024 OceanGPT: A Large Language Model for Ocean Science Tasks
abstract
Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhen Bi, Ningyu Zhang 0001, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, Huajun Chen
ACL (1)1
2024 Relphormer: Relational Graph Transformer for Knowledge Graph Representations
Zhen Bi, Siyuan Cheng 0008, Jing Chen 0060, Xiaozhuan Liang, Feiyu Xiong, Ningyu Zhang 0001
Neurocomputing1
2024 CodeKGC: Code Language Model for Generative Knowledge Graph Construction
abstract
Current generative knowledge graph construction approaches usually fail to capture structural knowledge by simply flattening natural language into serialized texts or a specification language. However, large generative language model trained on structured data such as code has demonstrated impressive capability in understanding natural language for structural prediction and reasoning tasks. Intuitively, we address the task of generative knowledge graph construction with code language model: given a code-format natural language input, the target is to generate triples which can be represented as code completion tasks. Specifically, we develop schema-aware prompts that effectively utilize the semantic structure within the knowledge graph. As code inherently possesses structure, such as class and function definitions, it serves as a useful model for prior semantic structural knowledge. Furthermore, we employ a rationale-enhanced generation method to boost the performance. Rationales provide intermediate steps, thereby improving knowledge extraction abilities. Experimental results indicate that the proposed approach can obtain better performance on benchmark datasets compared with baselines. 1
Zhen Bi, Jing Chen 0060, Yinuo Jiang, Feiyu Xiong, Huajun Chen, Ningyu Zhang 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Multi-Modal Protein Knowledge Graph Construction and Applications (Student Abstract)
abstract
Existing data-centric methods for protein science generally cannot sufficiently capture and leverage biology knowledge, which may be crucial for many protein tasks. To facilitate research in this field, we create ProteinKG65, a knowledge graph for protein science. Using gene ontology and Uniprot knowledge base as a basis, we transform and integrate various kinds of knowledge with aligned descriptions and protein sequences, respectively, to GO terms and protein entities. ProteinKG65 is mainly dedicated to providing a specialized protein knowledge graph, bringing the knowledge of Gene Ontology to protein function and structure prediction. We also illustrate the potential applications of ProteinKG65 with a prototype. Our dataset can be downloaded at https://w3id.org/proteinkg65.
Siyuan Cheng 0008, Xiaozhuan Liang, Zhen Bi, Huajun Chen, Ningyu Zhang 0001
AAAI3
2023 Tele-Knowledge Pre-training for Fault Analysis
abstract
In this work, we share our experience on tele-knowledge pre-training for fault analysis, a crucial task in telecommunication applications that requires a wide range of knowledge normally found in both machine log data and product documents. To organize this knowledge from experts uniformly, we propose to create a Tele-KG (tele-knowledge graph). Using this valuable data, we further propose a tele-domain language pre-training model TeleBERT and its knowledge-enhanced version, a tele-knowledge re-training model KTeleBERT. which includes effective prompt hints, adaptive numerical data encoding, and two knowledge injection paradigms. Concretely, our proposal includes two stages: first, pre-training TeleBERT on 20 million tele-related corpora, and then re-training it on 1 million causal and machine-related corpora to obtain KTeleBERT. Our evaluation on multiple tasks related to fault analysis in tele-applications, including root-cause analysis, event association prediction, and fault chain tracing, shows that pretraining a language model with tele-domain data is beneficial for downstream tasks. Moreover, the KTeleBERT re-training further improves the performance of task models, highlighting the effectiveness of incorporating diverse tele-knowledge into the model.
Zhuo Chen 0007, Wen Zhang 0015, Mingyang Chen 0002, Yuxia Geng, Zhen Bi, Yichi Zhang 0009, Zhen Yao 0001, Wenting Song, Xinliang Wu, Zhaoyang Lian, Lei Cheng 0005, Huajun Chen
ICDE7
2022 Learning to Ask for Data-Efficient Event Argument Extraction (Student Abstract)
abstract
Event argument extraction (EAE) is an important task for information extraction to discover specific argument roles. In this study, we cast EAE as a question-based cloze task and empirically analyze fixed discrete token template performance. As generating human-annotated question templates is often time-consuming and labor-intensive, we further propose a novel approach called “Learning to Ask,” which can learn optimized question templates for EAE without human annotations. Experiments using the ACE-2005 dataset demonstrate that our method based on optimized questions achieves state-of-the-art performance in both the few-shot and supervised settings.
Hongbin Ye, Ningyu Zhang 0001, Zhen Bi, Shumin Deng, Chuanqi Tan, Hui Chen 0018, Fei Huang 0002, Huajun Chen
AAAI3
2022 CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
abstract
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ningyu Zhang 0001, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li 0040, Xin Shang, Kangping Yin, Chuanqi Tan, Fei Huang 0002, Luo Si, Yuan Ni, Guo Tong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan 0002, Jun Yan 0010, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen
ACL (1)3
2022 OntoProtein: Protein Pretraining With Gene Ontology Embedding
Ningyu Zhang 0001, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng 0008, Haosen Hong, Shumin Deng, Qiang Zhang 0026, Jiazhang Lian, Huajun Chen
ICLR2
2022 Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners
Ningyu Zhang 0001, Luoqiu Li, Xiang Chen 0016, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang 0002, Huajun Chen
ICLR5