Bo Zhang 0096

dblp:36/2259-96 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-3555-8202ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Counterfactual Question Generation Uncovering Learner Contradictions
abstract
Conventional feedback, even when accompanied by brief explanations, rarely uncovers the hidden contradictions that trigger a learner's mistake. We bridge this gap with counterfactual question generation (CFQG): given a learner's answer, generate a follow-up question that deliberately contradicts it, compelling the learner to confront the underlying conflict. CFQG thus transforms assessment from passive scoring into an interactive and contradiction-centered dialogue that supports knowledge repair. To automate CFQG, we propose GapProbe, which probes the knowledge gap between a learner’s belief and curated facts through a knowledge graph (KG), then designs counterfactual questions (CFQs) that negate the belief. Identifying contradiction-aware triples, and more importantly, selecting those most likely to confuse the learner, are highly challenging in large-scale KGs. GapProbe tackles these challenges with an iterative ProConB cycle coupled with a schema-aware KGMap. By caching one- and multi-hop schema patterns of the KG, KGMap provides ``roadmap'' to guide LLMs jump to deep and contradiction-aware triples, beyond traditional step-wise graph traversal. We present the CFQG benchmark and corresponding metrics for evaluating how generated CFQs trigger, focus, and deepen learner reflection through explicit contradictions. Experiments on multiple datasets and LLMs show that GapProbe boosts LLM reasoning over KGs and generates follow-up questions that consistently promote deeper and more focused learner reflection.
Bo Zhang 0096, Yvhang Yang, Dezhuang Miao, Fengyi Song, Yanhui Gu, Xiaoming Zhang 0001, Junsheng Zhou
AAAI1
2026 DilatedTAD: Enhancing Adaptability to Actions of Varying Durations for Temporal Action Detection
abstract
Temporal Action Detection (TAD) aims to identify action boundaries and their corresponding categories in untrimmed videos, playing a crucial role in long-video understanding. Prior works often struggle to balance the trade-off between capturing long-range dependencies and ensuring computational efficiency. Recently, the state space model Mamba has exhibited impressive capabilities and efficiency in long-term sequence modeling. However, current methods based on Mamba generally lack a unified framework to simultaneously address the redundancy of long-duration actions and the boundary sensitivity of short-duration actions—limitations that largely stem from Mamba’s reliance on limited state representations and its unidirectional modeling. To tackle the aforementioned challenges, we propose DilatedTAD, a novel TAD framework with an expanded receptive field. DilatedTAD leverages the Inter-Parallel DIM component (InterDIM) to integrate multi-scale temporal information, enabling a better trade-off between short-duration and long-duration action detection. InterDIM is built upon our proposed Dilated Mamba (DIM), where multiple DIM branches with different dilation rates are designed to focus on actions of varying durations. Specifically, DIM introduces a novel use of dilation to skip redundant temporal information, thereby enhancing the model’s focus on crucial boundary features. Additionally, a bidirectional modeling design is adopted in DIM to compensate for the lack of future temporal context in the original Mamba architecture. Extensive experiments show that DilatedTAD outperforms state-of-the-art methods on multiple datasets, achieving mAPs of 74.9% (THUMOS14), 42.90% (ActivityNet 1.3), 45.0% (HACS), and 26.3% and 24.3% (EPIC-Kitchens 100). Our code will be publicly available.
Longyang Tang, Bo Zhang 0096, Rui Xu 0021, Junsheng Zhou, Yi Chen 0023
IEEE Trans. Circuits Syst. Video Technol.2
2025 What Is a Good Question? Assessing Question Quality via Meta-Fact Checking
abstract
Knowledge-based questions are typically employed to evaluate LLM's knowledge boundaries; meanwhile, numerous studies focus on question generation as a means to enhance the capabilities of both models and individuals. However, there is a lack of in-depth exploration about what constitutes a good question from the perspective of knowledge cognition. This paper proposes aligning the complete knowledge underlying questions with educational criteria effectively employed in physics courses, thereby developing novel knowledge-intensive metrics of question quality. To this end, we propose Meta-Fact Checking (MFC), which transforms questions into knowledge graph (KG) triples utilizing LLMs through few-shot prompting, thereby quantifying question quality based on the patterns observed within these triples. MFC introduces a novel interaction mechanism for KGs that communicates meta-facts, illustrating the types of knowledge that KGs can offer to the LLM for reasoning questions, rather than relying solely on the original triples. This strategy ensures that MFC remains unaffected by unexplored triples that LLM has not yet encountered within KGs compared to the retrieve-while-reasoning routine. Experiments across multiple datasets and LLMs demonstrate that MFC significantly improves the accuracy and efficiency of both question answering and assessing. This research marks a pioneering effort to automate the evaluation of question quality based on cognitive capabilities.
Bo Zhang 0096, Jianghua Zhu, Chaozhuo Li, Dezhuang Miao, Xiaoming Zhang 0001, Junsheng Zhou
AAAI1
2025 Reverse Chain-of-Thought and Causal Path Verification: A Modular Plugin for Aligning LLMs with Knowledge Graphs
abstract
Large language models (LLMs) exhibit strong language understanding capabilities, but encounter challenges when integrating structured knowledge from knowledge graphs (KGs) for complex reasoning tasks such as knowledge graph question answering (KGQA). Existing methods often rely on prompt engineering or fixed templates, which obscure the relational structure and limit generalization. To address these limitations, this paper introduces the Reverse Chain-of-Thought (R-CoT) and Causal Path Verification Plugin, a modular framework that reconstructs retrieved KG triples into reverse chains of sub-questions. Each reasoning step is aligned with a supporting triple, forming interpretable multi-hop paths. In particular, Semantic Causal Scoring (SCS) module is further incorporated to evaluate the causal alignment between each reverse sub-question and the original question through dynamic semantic vector matching. The SCS design avoids frequent interactions with LLMs and effectively filters irrelevant or unsupported reasoning steps. Based on the scoring results, a template-free, model-agnostic R-CoT input format is constructed as a semi-structured sequence. This design preserves the KG structure in natural language form and enables seamless integration with standard LLMs without fine-tuning. Experimental results demonstrate that the R-CoT Plugin consistently improves factual alignment, enhances reasoning stability, and outperforms conventional prompt-based methods in both accuracy and coherence.
Dezhuang Miao, Yibin Du, Xiang Li 0117, Jiahe Li 0007, Bo Zhang 0096, Bingyu Yan, Litian Zhang
CIKM6
2025 Moment matching of joint distributions for unsupervised domain adaptation
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Yancong Li, Feiran Huang
Inf. Process. Manag.1
2025 Vision-language representation learning with breadth and depth attention pre-training
abstract
The rapid advances in computer vision and natural language processing have led to increased attention toward the challenge of understanding vision and language together across multiple domains. Representation learning has become a major focus of research on cross-modal information understanding. However, current methods often fall short of providing comprehensive interaction and meaningful supervised guidance that would allow for effective learning of visual-linguistic joint representation. In this paper, we introduce the Breadth and Depth Attention Pre-training (BDAP) model for vision–language representation learning. Our model includes a breadth attention network designed to model feature associations between text sentences and image regions across different image levels. It uses fine-grained image features to promote more effective cross-modal feature interactions. Additionally, a depth attention network, which repeatedly calculates attention scores , is designed to deeply capture the complementarity between the image and text by gradually refining important image regions related to the text. Furthermore, we propose an attention pre-training network that leverages attention annotated distribution maps as prior knowledge to supervise the learning process of the breadth and depth attention networks, thereby enabling weight initialization of both types of attention networks. Extensive experiments on datasets of visual question answering and multi-modal sentiment analysis demonstrate the promising superiority of our BDAP model for vision–language representation learning.
Yun Liu 0017, Bo Zhang 0096, Chencheng Wang, Genglong Yan, Zhoujun Li 0001
Knowl. Based Syst.2
2024 LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly Detection
abstract
Log anomaly detection is a key component in the field of artificial intelligence for IT operations (AIOps). Considering log data of variant domains, retraining the whole network for unknown domains is inefficient in real industrial scenarios. However, previous deep models merely focused on extracting the semantics of log sequences in the same domain, leading to poor generalization on multi-domain logs. To alleviate this issue, we propose a unified Transformer-based framework for Log anomaly detection (LogFormer) to improve the generalization ability across different domains, where we establish a two-stage process including the pre-training and adapter-based tuning stage. Specifically, our model is first pre-trained on the source domain to obtain shared semantic knowledge of log data. Then, we transfer such knowledge to the target domain via shared parameters. Besides, the Log-Attention module is proposed to supplement the information ignored by the log-paring. The proposed method is evaluated on three public datasets and one real-world dataset. Experimental results on multiple benchmarks demonstrate the effectiveness of our LogFormer with fewer trainable parameters and lower training costs.
Hongcheng Guo, Jian Yang 0030, Jiaqi Bai 0001, Boyang Wang 0006, Zhoujun Li 0001, Tieqiao Zheng, Bo Zhang 0096, Junran Peng
AAAI8
2024 OWL: A Large Language Model for IT Operations
abstract
With the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition, machine translation, and dialogue systems. Recently, Large Language Models (LLMs) have achieved significant improvements across various domain-specific areas. However, there is a noticeable gap in the development of specialized Large Language Models (LLMs) tailored for IT operations. In this paper, we introduce the OWL, a large language model trained on our constructed Owl-Instruct with a wide range of IT-related information. Specifically, limited by the maximum input length, we propose the \textbf{H}omogeneous \textbf{M}arkov \textbf{C}ontext \textbf{E}xtension method (HMCE). The mixture-of-adapter strategy is leveraged to improve the parameter-efficient tuning across different domains or tasks. Further, we evaluate the performance of OWL on the Owl-Bench established by us and open IT-related benchmarks. OWL demonstrates superior performance results on IT tasks, which outperforms existing models by significant margins. Moreover, we hope that the findings of our work will provide more insights to revolutionize the techniques of IT operations with specialized LLMs.
Hongcheng Guo, Jian Yang 0030, Liqun Yang, Linzheng Chai, Jiaqi Bai 0001, Junran Peng, Xiaorong Hu, Dongfeng Zhang, Xu Shi 0005, Tieqiao Zheng, Liangfan Zheng, Bo Zhang 0096, Ke Xu 0001, Zhoujun Li 0001
ICLR14
2024 Cross-domain knowledge collaboration for blending-target domain adaptation
Bo Zhang 0096, Xiaoming Zhang 0001, Feiran Huang, Dezhuang Miao
Inf. Process. Manag.1
2024 SANe: Space adaptation network for temporal knowledge graph completion
Yancong Li, Xiaoming Zhang 0001, Bo Zhang 0096, Feiran Huang, Shuai Ma 0001
Inf. Sci.3
2023 LogLG: Weakly Supervised Log Anomaly Detection via Log-Event Graph Construction
Hongcheng Guo, Yuhui Guo, Jian Yang 0030, Zhoujun Li 0001, Tieqiao Zheng, Liangfan Zheng, Weichao Hou, Bo Zhang 0096
DASFAA (4)9
2023 Deep Kernel Network Embedding
abstract
This paper concerns the problem of network embedding (NE), whose aim is to learn a low-dimensional representation for each node in networks. We provide a new train to solve the sparsity problem where most of nodes including the new arrival nodes have little knowledge with respect to the network. A novel paradigm is proposed to integrate the multiple information from the subgraph covering the target node instead of only the target node. Particularly, to epxress the distinctive feature over the vertex domain, a probabiltiy distribution over subgraph space is constructed for each node. The distribution is more effective to express the distinctive characteristic and feature in a higher dimension compared to the latent representation vectors. One of the primary goals of this paradigm is to define the convolution operation over the distributions, which are efficient to evaluate and learn. Experiments on four real-world network datasets demonstrate that our approach significantly outperforms state-of-the-art methods, especially on the representation learning for the nodes newly joining in the network.
Bo Zhang 0096, Xiaoming Zhang 0001, Feiran Huang, Shuai Ma 0001
IEEE Trans. Knowl. Data Eng.1
2022 ALSA: Adversarial Learning of Supervised Attentions for Visual Question Answering
abstract
Visual question answering (VQA) has gained increasing attention in both natural language processing and computer vision. The attention mechanism plays a crucial role in relating the question to meaningful image regions for answer inference. However, most existing VQA methods: 1) learn the attention distribution either from free-form regions or detection boxes in the image, which is intractable in answering questions about the foreground object and background form, respectively and 2) neglect the prior knowledge of human attention and learn the attention distribution with an unguided strategy. To fully exploit the advantages of attention, the learned attention distribution should focus more on the question-related image regions, such as human attention for both the questions, about the foreground object and background form. To achieve this, this article proposes a novel VQA model, called adversarial learning of supervised attentions (ALSAs). Specifically, two supervised attention modules: 1) free form-based and 2) detection-based, are designed to exploit the prior knowledge for attention distribution learning. To effectively learn the correlations between the question and image from different views, that is, free-form regions and detection boxes, an adversarial learning mechanism is implemented as an interplay between two supervised attention modules. The adversarial learning reinforces the two attention modules mutually to make the learned multiview features more effective for answer inference. The experiments performed on three commonly used VQA datasets confirm the favorable performance of ALSA.
Yun Liu 0017, Xiaoming Zhang 0001, Zhiyun Zhao, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Cybern.4
2022 Cross-Attentional Spatio-Temporal Semantic Graph Networks for Video Question Answering
abstract
Due to the rich spatio-temporal visual content and complex multimodal relations, Video Question Answering (VideoQA) has become a challenging task and attracted increasing attention. Current methods usually leverage visual attention, linguistic attention, or self-attention to uncover latent correlations between video content and question semantics. Although these methods exploit interactive information between different modalities to improve comprehension ability, inter- and intra-modality correlations cannot be effectively integrated in a uniform model. To address this problem, we propose a novel VideoQA model called Cross-Attentional Spatio-Temporal Semantic Graph Networks (CASSG). Specifically, a multi-head multi-hop attention module with diversity and progressivity is first proposed to explore fine-grained interactions between different modalities in a crossing manner. Then, heterogeneous graphs are constructed from the cross-attended video frames, clips, and question words, in which the multi-stream spatio-temporal semantic graphs are designed to synchronously reasoning inter- and intra-modality correlations. Last, the global and local information fusion method is proposed to coalesce the local reasoning vector learned from multi-stream spatio-temporal semantic graphs and the global vector learned from another branch to infer the answer. Experimental results on three public VideoQA datasets confirm the effectiveness and superiority of our model compared with state-of-the-art methods.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Image Process.4
2021 Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation
abstract
Bo Zhang, Xiaoming Zhang, Yun Liu, Lei Cheng, Zhoujun Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Zhoujun Li 0001
ACL/IJCNLP (1)1
2021 Discriminative Feature Adaptation via Conditional Mean Discrepancy for Cross-Domain Text Classification
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017
DASFAA (2)1