Zefan Zhang

dblp:307/6777 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0001-5627-050XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Adam's Law: Textual Frequency Law on Large Language Models
abstract
While textual frequency has been validated as relevant to human cognition in reading speed, its relatedness to Large Language Models (LLMs) is seldom studied.We propose a novel research direction in terms of textual data frequency, which is an understudied topic, to the best of our knowledge.Our framework is composed of three units.First, this paper proposes Textual Frequency Law (TFL), which indicates that frequent textual data should be preferred for LLMs for both prompting and fine-tuning.Since many LLMs are closedsource in their training data, we propose using online resources to estimate the sentencelevel frequency.We then utilize an input paraphraser to paraphrase the input into a more frequent textual expression.Next, we propose Textual Frequency Distillation (TFD) by querying LLMs to conduct story completion by further extending the sentences in the datasets, and the resulting corpora are used to adjust the initial estimation.Finally, we propose Curriculum Textual Frequency Training (CTFT) that finetunes LLMs in an increasing order of sentencelevel frequency.Experiments are conducted on our curated dataset Textual Frequency Paired Dataset (TFPD) on math reasoning, machine translation, commonsense reasoning and agentic tool calling.Results show the effectiveness of our framework.1 * Equal Contribution. 1 https://github.com/HongyuanLuke/frequencylawReferences
Hongyuan Lu, Zefan Zhang, Bowen Cao, Wai Lam
ACL (1)3
2026 CogECI: Context Grounded Document-level Event Causality Identification via Large Language Models
Zefan Zhang, Xumeng Zhang, Shijie Jiang, Tian Bai 0002
Knowl. Based Syst.1
2026 TDP-DETR: Temporal dynamics perception framework for video moment retrieval and highlight detection
abstract
Video Moment Retrieval (VMR) and Highlight Detection (HD) aim to localize query-relevant temporal segments and evaluate clip-level saliency within untrimmed videos. Accurate temporal boundary perception is essential for VMR and HD. While current models have made significant progress, they still struggle to achieve precise action semantic alignment with temporally dynamic video content and are prone to boundary perception bias when action-related visual semantic cues experience fluctuations in specific frames. In this paper, we propose a Temporal Dynamics Perception DEtection TRansformer (TDP-DETR) that models action temporal dynamics from two complementary perspectives: temporal persistence and temporal progression. For temporal persistence, we introduce a dynamic masking strategy for action duration-aware temporal modeling, enabling the model to infer action persistence from query semantics and incorporate it as a temporal prior for boundary prediction. For temporal progression, we design an action state difference perception module that captures frame-to-frame action state variations, allowing the model to perceive action progression speed and thereby improve anticipation of action boundaries. Extensive experiments on three MR/HD benchmarks demonstrate that our method consistently outperforms existing state-of-the-art approaches. Our code will be public soon.
Huilin An, Zefan Zhang, Shijie Jiang, Kehua Zhu, Tian Bai 0002
Neural Networks2
2026 Dual-level dynamic heterogeneous graph network for video question answering
Zefan Zhang, Tian Bai 0002
Neural Networks1
2026 DAMON: Difference-Aware Medical Visual Question Answering via Multimodal Large Language Model
abstract
Difference-aware Medical Visual Question Answering (MVQA) aims to answer questions regarding disease-related content and the visual differences between the paired medical images, which is crucial for assessing disease progression and guiding further treatment planning. Although current medical Multimodal Large Language Models (MLLMs) have shown promising results in MVQA, they still exhibit poor generalization performance in difference-aware MVQA due to two key challenges. Firstly, existing difference-aware MVQA datasets are biased toward temporal variations of individual diseases, limiting their ability to model multi-disease coexistence and overlapping symptoms in real-world clinical scenarios. Secondly, disease-level semantic alignment becomes more challenging with multi-image inputs, as they introduce more redundant and interfering visual features. To address the first challenge, we introduce DAMON-QA, a large-scale difference-aware MVQA dataset designed to support visual difference analysis across multiple diseases. Leveraging this dataset, we train MLLMs and propose a Difference-Aware Medical visual questiON answering (DAMON) model. To tackle the second challenge, we further propose a Disease-driven Prompt Module (DPM) to identify the relevant diseases and guide the disease difference analysis process. Experiments on MIMIC-Diff-VQA show that our DAMON model achieves state-of-the-art (SOTA) performance.
Zefan Zhang, Ruihong Zhao, Tian Bai 0002
IEEE J. Biomed. Health Informatics1
2025 Prototype-Guided Multimodal Relation Extraction based on Entity Attributes
abstract
Multimodal Relation Extraction (MRE) aims to predict relations between head and tail entities based on the context of sentence-image pairs. Most existing MRE methods progressively incorporate textual and visual inputs to dominate the learning process, assuming both contribute significantly to the task. However, the diverse visual appearances and text with ambiguous semantics contain less-informative contexts for the corresponding relation. To tackle these challenges, we highlight the importance of semantically invariant entity attributes that encompass fine-grained categories. Towards this, we propose a novel Prototype-Guided Multimodal Relation Extraction (PG-MRE) framework based on Entity Attributes. Specifically, we first generate detailed entity explanations using Large Language Models (LLMs) to supplement the attribute semantics. Then, the Attribute Prototype Module (APM) refines attribute categories and condenses scattered entity attribute features into cluster-level prototypes. Furthermore, prototype-aligned attribute features guide diverse visual appearance features to produce compact and distinctive multimodal representations in the Relation Prototype Module (RPM). Extensive experiments demonstrate that our method gains superior relation classification capability (especially in scenarios involving various unseen entities), achieving new state-of-the-art performances on MNRE dataset.
Zefan Zhang, Tian Bai 0002
AAAI1
2025 SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models
abstract
There are more than 7,000 languages around the world, and current Large Language Models (LLMs) only support hundreds of languages.Dictionary-based prompting methods can enhance translation on them, but most methods use all the available dictionaries, which could be expensive.Instead, it will be flexible to have a trade-off between token consumption and translation performance.This paper proposes a novel task called Automatic Dictionary Selection (ADS).The goal of the task is to automatically select which dictionary to use to enhance translation.We propose a novel and effective method which we call Select Lowfrequency Words! (SLoW) which selects those dictionaries that have a lower frequency.Our methods have unique advantages.First, there is no need for access to the training data for frequency estimation (which is usually unavailable).Second, it inherits the advantage of dictionary-based methods, where no additional tuning is required on LLMs.Experimental results on 100 languages from FLORES indicate that SLoW surpasses strong baselines, and it can obviously save token usage, with many languages even surpassing the translation performance of the full dictionary baseline.
Hongyuan Lu, Zefan Zhang, Wai Lam
EMNLP3
2025 Towards Efficient General Feature Prediction in Masked Skeleton Modeling
Shengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng 0001, Xiaojun Chang, Meng Wang 0001
ICCV2
2025 Video-Level Multimodal Relation Extraction with Event-Entity Semantic Consistency
abstract
Previous research on Multimodal Relation Extraction (MRE) has primarily focused on identifying textual relations enhanced by static visual clues from images, benefiting fields such as multimedia analysis and knowledge graphs. With the rapid rise of video content on social media platforms, Multimodal Relation Extraction (MRE) systems face new challenges. To bridge this gap, we introduce Video-level Multimodal Relation Extraction (VMRE), a novel task aimed at extracting relational facts from videos. To advance this research, we present Vid-MRE, a new dataset containing 32 relation types and 12,402 multimodal relational facts, annotated across 3,970 pairs of textual news titles and corresponding videos. Since this task demands precise event and entity grounding to filter out excessive noise in the video, we propose an Event-Entity Semantic Consistency Network (E2SCN) to capture relational clues in the video effectively. Experimental results demonstrate that incorporating video content into the model significantly improves relation identification performance but also introduces more noise. Our E2SCN method effectively reduces the noise, enhancing fine-grained multimodal event and entity alignments while achieving state-of-the-art (SOTA) performance.
Zefan Zhang, Kailong Suo, Tian Bai 0002
ACM Multimedia1
2025 Prompting visual dialog with implicit logical knowledge
Zefan Zhang, Tian Bai 0002
Knowl. Inf. Syst.1
2024 Estimation of Single Tree Height Based on Improved K-Means Method for Unmanned Aerial Vehicle Point Cloud Data
abstract
The technology of identifying and measuring parameters of tree plays an important role in city 3D modeling and statistical analysis. This article focuses on tree height estimation based on UAV point cloud data, designs a method to achieve the separation of single trees from sparse tree clusters and tree-ground separation for accurate tree height estimation. The key of the algorithm is to obtain the single tree point and its corresponding ground point. To obtain accurate single tree points, this paper uses an improved k-means method. The highlight of the algorithm lies in the design of a validation algorithm based on tree structure to improve k-means method, making single tree separation results more accurate. After comparing the experimental results with actual measurement data, it is believed that this method is effective in extracting tree height.
Zefan Zhang
IGARSS1
2024 Caption-Aware Multimodal Relation Extraction with Mutual Information Maximization
abstract
Multimodal Relation Extraction (MRE) has achieved great improvements. However, modern MRE models are easily affected by irrelevant objects during multimodal alignment which are called error sensitivity issues. The main reason is that visual features are not fully aligned with textual features and the reasoning process may suppress redundant and noisy information at the risk of losing critical information. In light of this, we propose a Caption-Aware Multimodal Relation Extraction Network with Mutual Information Maximization (CAMIM). Specifically, we first generate detailed image captions through the Large Language Model (LLM). Then, the Caption-Aware Module (CAM) hierarchically aligns the fine-grained visual entities and textual entities for reasoning. In addition, for preserving crucial information within different modalities, we leverage a Mutual Information Maximization method to regulate the multimodal reasoning module. Experiments show that our model outperforms the state-of-the-art MRE models on the benchmark dataset MNRE. Further ablation studies prove the pluggable and effective performance of our Caption-Aware Module and Mutual Information Maximization method. Our code is available at https://github.com/zefanZhang-cn/CAMIM.
Zefan Zhang, Tian Bai 0002
ACM Multimedia1
2023 A Novel Drug-Drug Interaction Prediction Model Based on Line Subgraph Generation Strategy
abstract
Drug-Drug Interaction (DDI) prediction task is helpful for better-understanding drugs. In this paper, we propose a novel drug-drug interaction prediction model based on line subgraph generation strategy, named DDI-LSG model. Our DDI-LSG model consists of three main parts which include drug relation graph construction, line subgraph generation strategy, and graph-level classification. To consider more relationships among drugs, we propose a node feature-enhancing method to encode drug features in drug relation graph construction process. To consider drugs and DDI as equivalent factors of our DDI-LSG model, we introduce line graph transformation to integrate DDI with drug feature enhancing vector. Combining Jaccard similarity with cosine similarity, we propose a line subgraph generation strategy to evaluate node relation and extract key structures around target DDI in the line graph. Then, we reformulate the DDI prediction task into a graph-level classification task for the line subgraph of the target DDI. Therefore, in the final part of our DDI-LSG model, we use a graph-level classifier to classify the line subgraphs. Our DDI-LSG model outperforms better experiment results than baselines. Ablation results have validated the node feature enhancing method and line subgraph generation strategy.
Tian Bai 0002, Chu Li 0002, Xinyue Peng, Haotian Guan, Zefan Zhang, Guishen Wang
BIBM5
2023 Knowledge-Aware Causal Inference Network for Visual Dialog
abstract
The effective knowledge and interaction within multi-modalities are key to Visual Dialog. Classic graph-based framework with the direct connection between history dialog and answer fails to give the right answer for the spurious guidance and strong bias induced from history dialog. Recent causal inference framework without this direct connection improves the generalization while worse accuracy. In this work, we propose a novel Knowledge-Aware Causal Inference framework(KACI-Net) in which the commonsense knowledge is introduced into the causal inference framework to achieve both high accuracy and generalization. Specifically, the commonsense knowledge is first generated according to the entities extracted from the question and fused with language and visual features with the co-attention to get the final answer. Comparisons with knowledge-unaware framework and graph-based knowledge-aware framework on VisDial v1.0 dataset show the superiority of our proposed framework and verify the effectiveness the usage of the commonsense knowledge for a good reasoning in Visual Dialog. Both high NDCG and MRR metrics indicate a good trade-off between accuracy and generalization.
Zefan Zhang, Yi Ji 0001, Chunping Liu
ICMR1
2023 Infer unseen from seen: Relation regularized zero-shot visual dialog
Zefan Zhang, Yi Ji 0001, Chunping Liu
J. Vis. Commun. Image Represent.1
2023 Multi-view semantic understanding for visual dialog
Tianling Jiang, Zefan Zhang, Yi Ji 0001, Chunping Liu
Knowl. Based Syst.2
2022 Coupling Attention and Convolution for Heuristic Network in Visual Dialog
abstract
Visual Dialog is a typical AI-agent task on images, in which the agent interprets information from heterogeneous modalities and provides the correct answer. In this area, most approaches are based on the attention mechanism. When the agent enjoys the large-capacity advantage of attention, the lack of the right inductive bias compared with convolution hinders its success. Therefore, in order to utilize their advantages and compensate for their respective shortcomings, inspired by the paraventricular thalamus (PVT) in the brain, we couple convolution and attention, termed as Attention Convolution Enhanced (ACE) method to enhance the agent’s activation of key features and strengthening the semantic understanding of visual and textual data. Meanwhile, we propose Heuristic Adjustment (HA) module to globally strengthen the agent’s semantic understanding and reduce language bias that is easy to occur after using the enhanced features. Finally, we concatenate the ACE and the HA in our Coupling Attention and Convolution for Heuristic Network (CACH-Net) to train the agent for better semantic comprehension and generalization ability. Extensive experiments on the VisDial v1.0 benchmark show that our CACH-Net has a better performance.
Zefan Zhang, Tianling Jiang, Chunping Liu, Yi Ji 0001
ICIP1
2022 Spatio-Temporal Graph-based Semantic Compositional Network for Video Captioning
abstract
Video Captioning aims to generate natural language descriptions for given videos and is one of the challenging problems in computer vision's high-level understanding tasks. Existing methods are relatively lacking in the mining of object-level spatio-temporal relationships, which is important for generating captions with accurate object information. In this paper, we improve the existing SCN-LSTM method from the perspective of modeling spatio-temporal relationships and propose the Spatio-Temporal Graph-based Semantic Compositional Network for Video Captioning (STG-SCN). In terms of spatial-temporal relationships modeling, we propose the Spatial Relation Graph (SRG) and the Temporal Relation Graph (TRG) based on the Graph Attention Network, respectively. SRG is employed to establish the spatial relationships between spatially Neighboring objects within each keyframe conditioned on their correlation with the current keyframe. TRG is used to model the temporal relationship between all the objects at different time steps and incorporates the object-level information into frame-level features. Based on the proposed Semantics Guided Decoder, visual representations enhanced by object-level information are dynamically fused with high-level semantic concepts to generate captions that not only consider the global visual content but also have stronger language expressiveness. Extended experiments show that our proposed method achieves significant performance gains on Microsoft Video Description (MSVD) and Microsoft Research Video-to-Text (MSR-VTT) datasets, outperforming existing methods.
Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu
IJCNN2
2022 Scene graph generation with award-punishment strategy
Haiyan Gao, Dibo Shi, Tianling Jiang, Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu
Knowl. Based Syst.5