Lu Zhang 0060

dblp:82/10609-60 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-9532-5219ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SemanticLog: Towards Effective and Efficient Large-Scale Semantic Log Parsing
abstract
Logs of large-scale cloud systems record diverse system events, ranging from routine statuses to critical errors. As the fundamental step of automated log analysis, log parsing is to transform unstructured logs into structured data for easier management and analysis. However, existing syntax-based and deep learning-based parsers struggle with complex real-world logs. Recent parsers based on large language models (LLMs) achieve higher accuracy, but they typically rely on online APIs (e.g., ChatGPT), raising privacy concerns and suffering from network latency. Moreover, with the rise of artificial intelligence for IT operations (AIOps), traditional parsers that focus on syntax-level templates fail to capture the semantics of dynamic log parameters, limiting their usefulness for downstream tasks. These challenges highlight the need for semantic log parsing that goes beyond template extraction to understand parameter semantics.This paper presents SemanticLog, an effective and efficient semantic log parser powered by open-source LLMs. SemanticLog adapts the structure of LLMs to the log parsing task, leveraging their rich knowledge while safeguarding log data privacy. It first extracts informative feature representations from log data, then refines them through fine-grained semantic perception to enable accurate template and parameter extraction together with semantic category prediction. To boost scalability, SemanticLog introduces the EffiParsing tree for faster inference on large-scale logs. Extensive experiments on the LogHub-2.0 dataset show that SemanticLog significantly outperforms the state-of-the-art log parsers in terms of accuracy. Moreover, it also surpasses existing LLM-based parsers in efficiency while showcasing advanced semantic parsing capability. Notably, SemanticLog employs much smaller open-source LLMs compared to existing LLM-based parsers (mainly based on ChatGPT), while maintaining better capability of log data privacy protection.
Chenbo Zhang, Wenying Xu, Jinbu Liu, Lu Zhang 0060, Guiyang Liu, Jihong Guan, Qi Zhou 0001, Shuigeng Zhou
IEEE Trans. Software Eng.4
2025 Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
abstract
Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on the models’ response generation, and overlooking the task question itself. This paper discusses the vulnerability of LVLMs in solving counterfactual presupposition questions (CPQs), where the models are prone to accept the presuppositions of counterfactual objects and produce severe hallucinatory responses. To this end, we introduce "Antidote", a unified, synthetic data-driven post-training framework for mitigating both types of hallucination above. It leverages synthetic data to incorporate factual priors into questions to achieve self-correction, and decouple the mitigation process into a preference optimization problem. Furthermore, we construct "CP-Bench", a novel benchmark to evaluate LVLMs’ ability to correctly handle CPQs and produce factual responses. Applied to the LLaVA series, Antidote can simultaneously enhance performance on CP-Bench by over 50%, POPE by 1.8-3.3%, and CHAIR & SHR by 30-50%, all without relying on external supervision from stronger LVLMs or human feedback and introducing noticeable catastrophic forgetting issues.
Yuanchen Wu, Lu Zhang 0060, Hang Yao 0002, Junlong Du, Shouhong Ding, Yunsheng Wu, Xiaoqiang Li 0002
CVPR2
2025 Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan
MICCAI (2)4
2025 Fine-grained Zero-Shot Object Detection
abstract
Zero-shot object detection (ZSD) aims to leverage semantic descriptions to localize and recognize objects of both seen and unseen classes. Existing ZSD works are mainly coarse-grained object detection, where the classes are visually quite different, thus are relatively easy to distinguish. However, in real life we often have to face fine-grained object detection scenarios, where the classes are too similar to be easily distinguished. For example, detecting different kinds of birds, fishes, and flowers. In this paper, we propose and solve a new problem called Fine-Grained Zero-Shot Object Detection (FG-ZSD for short), which aims to detect objects of different classes with minute differences in details under the ZSD paradigm. We develop an effective method called MSHC for the FG-ZSD task, which is based on an improved two-stage detector and employs a multi-level semantics-aware embedding alignment loss, ensuring tight coupling between the visual and semantic spaces. Considering that existing ZSD datasets are not suitable for the new FG-ZSD task, we build the first FG-ZSD benchmark dataset FGZSD-Birds, which contains 148,820 images falling into 36 orders, 140 families, 579 genera and 1432 species. Extensive experiments on FGZSD-Birds show that our method outperforms existing ZSD models.
Hongxu Ma 0001, Chenbo Zhang, Lu Zhang 0060, Jiaogen Zhou, Jihong Guan, Shuigeng Zhou
ACM Multimedia3
2024 Weakly Supervised Few-Shot Object Detection with DETR
abstract
In recent years, Few-shot Object Detection (FSOD) has become an increasingly important research topic in computer vision. However, existing FSOD methods require strong annotations including category labels and bounding boxes, and their performance is heavily dependent on the quality of box annotations. However, acquiring strong annotations is both expensive and time-consuming. This inspires the study on weakly supervised FSOD (WS-FSOD in short), which realizes FSOD with only image-level annotations, i.e., category labels. In this paper, we propose a new and effective weakly supervised FSOD method named WFS-DETR. By a well-designed pretraining process, WFS-DETR first acquires general object localization and integrity judgment capabilities on large-scale pretraining data. Then, it introduces object integrity into multiple-instance learning to solve the common local optimum problem by comprehensively exploiting both semantic and visual information. Finally, with simple fine-tuning, it transfers the knowledge learned from the base classes to the novel classes, which enables accurate detection of novel objects. Benefiting from this ``pretraining-refinement'' mechanism, WSF-DETR can achieve good generalization on different datasets. Extensive experiments also show that the proposed method clearly outperforms the existing counterparts in the WS-FSOD task.
Chenbo Zhang, Yinglu Zhang, Lu Zhang 0060, Jihong Guan, Shuigeng Zhou
AAAI3
2024 Tail Classes Matter: Long-Tailed Object Detection Revisited
abstract
Real-world data ubiquitously exhibit long-tailed distribution, which sparks the increasing interest in long-tailed object detection (LTOD). However, existing methods neglect that a lack of diverse data in tail classes will cause underrepresented tail class features, making their efforts for balancing foreground classes tend to over-fit tail classes and be less effective. In this paper, we propose a multi-class co-attention generation network to increase data diversity of tail classes by generating augmented samples. To alleviate imbalance, we develop a distribution-aware up-sampling strategy, performing differential up-sampling for different classes and design a bi-directional regulation loss to adjust both positive and negative gradients. Moreover, we construct a new dataset LVIS-X with more rare classes based on existing LTOD benchmark dataset LVIS. Experiments on LVIS and LVIS-X demonstrate the superiority of the proposed method.
Yinglu Zhang, Chenbo Zhang, Lu Zhang 0060, Tianying Liu, Jihong Guan, Xinkai Liang, Shuigeng Zhou
ICASSP3
2024 Weakly-Supervised Graph Classification with Even a Single Key Subgraph Per Class
abstract
Traditional graph classification requires large amounts of labeled data, which is expensive and time-consuming to acquire, especially in some special scenarios that domain knowledge is indispensable for labeling graphs. Observing that some key subgraphs can determine the properties of graphs (e.g. the toxicity of drug molecules depend on some toxic functional groups), in this paper we explore to classify graphs using unlabeled graphs plus a small number of key subgraphs for each class, which is called weakly-supervised graph classification. To this end, we develop the WeGraph method, where the graph classifier is trained with subgraph-based self-supervised learning and divergence- minimization based fine-tuning. Moreover, we design a key subgraph extraction algorithm to iteratively extract and update the key subgraphs, which makes the training process a closed loop. We conduct extensive experiments on different types of graph datasets to evaluate the effectiveness of WeGraph. Experimental results show that WeGraph can achieve high performance even when only one key subgraph is provided for each class.
Lu Zhang 0060, Chenbo Zhang, Jihong Guan, Shuigeng Zhou
ICDM1
2024 Neural Atoms: Propagating Long-range Interaction in Molecular Graphs through Efficient Communication Channel
abstract
Graph Neural Networks (GNNs) have been widely adopted for drug discovery with molecular graphs. Nevertheless, current GNNs mainly excel in leveraging short-range interactions (SRI) but struggle to capture long-range interactions (LRI), both of which are crucial for determining molecular properties. To tackle this issue, we propose a method to abstract the collective information of atomic groups into a few $\textit{Neural Atoms}$ by implicitly projecting the atoms of a molecular. Specifically, we explicitly exchange the information among neural atoms and project them back to the atoms’ representations as an enhancement. With this mechanism, neural atoms establish the communication channels among distant nodes, effectively reducing the interaction scope of arbitrary node pairs into a single hop. To provide an inspection of our method from a physical perspective, we reveal its connection to the traditional LRI calculation method, Ewald Summation. The Neural Atom can enhance GNNs to capture LRI by approximating the potential LRI of the molecular. We conduct extensive experiments on four long-range graph benchmarks, covering graph-level and link-level tasks on molecular graphs. We achieve up to a 27.32% and 38.27% improvement in the 2D and 3D scenarios, respectively. Empirically, our method can be equipped with an arbitrary GNN to help capture LRI. Code and datasets are publicly available in https://github.com/tmlr-group/NeuralAtom.
Xuan Li 0005, Zhanke Zhou, Jiangchao Yao, Yu Rong 0001, Lu Zhang 0060, Bo Han 0003
ICLR5
2024 AlignCLIP: Align Multi Domains of Texts Input for CLIP models with Object-IoU Loss
abstract
Since the release of the CLIP model by OpenAI, it has received widespread attention. However, categories in the real world often exhibit a long-tail distribution, and existing CLIP models struggle to effectively recognize rare, tail-end classes, such as an endangered African bird. An intuitive idea is to generate visual descriptions for these tail-end classes and use descriptions to create category prototypes for classification. However, experiments reveal that visual descriptions, image captions, and test prompt templates belong to three distinct domains, leading to distribution shifts. In this paper, we propose the use of caption object parsing to identify the objects set contained within captions. During training, the object sets is used to generate visual descriptions and test prompts, aligning these three domains and enabling the text encoder to generate category prototypes based on visual descriptions. Thanks to the acquired object sets, our approach can construct many-to-many relationships at a lower cost and derive soft labels, addressing the noise issues associated with traditional one-to-one matching. Extensive experimental results demonstrate that our method significantly surpasses the CLIP baseline and exceeds existing methods, achieving a new state-of-the-art (SOTA).
Lu Zhang 0060, Shouhong Ding
ACM Multimedia1
2023 DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations
abstract
AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-distribution (OOD) learning problem with noise. We present DrugOOD, a systematic OOD dataset curator and benchmark for AIDD. Particularly, we focus on the drug-target binding affinity prediction problem, which involves both macromolecule (protein target) and small-molecule (drug compound). DrugOOD offers an automated dataset curator with user-friendly customization scripts, rich domain annotations aligned with biochemistry knowledge, realistic noise level annotations, and rigorous benchmarking of SOTA OOD algorithms, as opposed to only providing fixed datasets. Since the molecular data is often modeled as irregular graphs using graph neural network (GNN) backbones, DrugOOD also serves as a valuable testbed for graph OOD learning problems. Extensive empirical studies have revealed a significant performance gap between in-distribution and out-of-distribution experiments, emphasizing the need for the development of more effective schemes that permit OOD generalization under noise for AIDD.
Yuanfeng Ji, Lu Zhang 0060, Jiaxiang Wu 0001, Bingzhe Wu, Lanqing Li, Long-Kai Huang, Tingyang Xu, Yu Rong 0001, Ding Xue, Houtim Lai, Wei Liu 0005, Junzhou Huang, Shuigeng Zhou, Ping Luo 0002, Peilin Zhao, Yatao Bian
AAAI2
2023 Meta-ZSDETR: Zero-shot DETR with Meta-learning
abstract
Zero-shot object detection aims to localize and recognize objects of unseen classes. Most of existing works face two problems: the low recall of RPN in unseen classes and the confusion of unseen classes with background. In this paper, we present the first method that combines DETR and meta-learning to perform zero-shot object detection, named Meta-ZSDETR, where model training is formalized as an individual episode based meta-learning task. Different from Faster R-CNN based methods that firstly generate class-agnostic proposals, and then classify them with visual-semantic alignment module, Meta-ZSDETR directly predict class-specific boxes with class-specific queries and further filter them with the predicted accuracy from classification head. The model is optimized with meta-contrastive learning, which contains a regression head to generate the coordinates of class-specific boxes, a classification head to predict the accuracy of generated boxes, and a contrastive head that utilizes the proposed contrastive-reconstruction loss to further separate different classes in visual space. We conduct extensive experiments on two benchmark datasets MS COCO and PASCAL VOC. Experimental results show that our method outperforms the existing ZSD methods by a large margin.
Lu Zhang 0060, Chenbo Zhang, Jihong Guan, Shuigeng Zhou
ICCV1
2023 Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive Loss
abstract
Zero-shot object detection (ZSD) aims to localize and recognize unseen objects in unconstrained images by leveraging semantic descriptions. Existing ZSD methods typically suffer from two drawbacks: 1) Due to the lack of data on unseen categories during the training phase, the model inevitably has a bias towards the seen categories, i.e., it prefers to subsume objects of unseen categories to seen categories; 2) It is usually very tricky for the feature extractor trained on data of seen categories to learn discriminative features that are good enough to help the model transfer the knowledge learned from data of seen categories to unseen categories. To tackle these problems, this paper proposes a novel zero-shot detection method based on a semantics-aware DETR and a class-wise adaptive contrastive loss. Concretely, to address the first problem, we develop a novel semantics-aware attention mechanism to mitigate the bias towards seen categories and integrate it into DETR, which results in a new end-to-end zero-shot object detection approach. Furthermore, to handle the second problem, a novel class-wise adaptive contrastive loss is proposed, which considers the relevance between each pair of categories according to their semantic description in order to learn separable features for better visual-semantic alignment. Extensive experiments and ablation studies on benchmark datasets demonstrate the effectiveness and superiority of the proposed method.
Lu Zhang 0060, Jihong Guan, Shuigeng Zhou
ACM Multimedia2
2023 Recent Few-shot Object Detection Algorithms: A Survey with Performance Comparison
abstract
The generic object detection (GOD) task has been successfully tackled by recent deep neural networks, trained by an avalanche of annotated training samples from some common classes. However, it is still non-trivial to generalize these object detectors to the novel long-tailed object classes, which have only few labeled training samples. To this end, the Few-Shot Object Detection (FSOD) has been topical recently, as it mimics the humans’ ability of learning to learn and intelligently transfers the learned generic object knowledge from the common heavy-tailed to the novel long-tailed object classes. Especially, the research in this emerging field has been flourishing in recent years with various benchmarks, backbones, and methodologies proposed. To review these FSOD works, there are several insightful FSOD survey articles [ 58 , 59 , 74 , 78 ] that systematically study and compare them as the groups of fine-tuning/transfer learning and meta-learning methods. In contrast, we review the existing FSOD algorithms from a new perspective under a new taxonomy based on their contributions, i.e., data-oriented, model-oriented, and algorithm-oriented. Thus, a comprehensive survey with performance comparison is conducted on recent achievements of FSOD. Furthermore, we also analyze the technical challenges, the merits and demerits of these methods, and envision the future directions of FSOD. Specifically, we give an overview of FSOD, including the problem definition, common datasets, and evaluation protocols. The taxonomy is then proposed that groups FSOD methods into three types. Following this taxonomy, we provide a systematic review of the advances in FSOD. Finally, further discussions on performance, challenges, and future directions are presented.
Tianying Liu, Lu Zhang 0060, Yang Wang 0100, Jihong Guan, Yanwei Fu 0001, Shuigeng Zhou
ACM Trans. Intell. Syst. Technol.2
2022 Hierarchical Few-Shot Object Detection: Problem, Benchmark and Method
abstract
Few-shot object detection (FSOD) is to detect objects with a few examples. However, existing FSOD methods do not consider hierarchical fine-grained category structures of objects that exist widely in real life. For example, animals are taxonomically classified into orders, families, genera and species etc. In this paper, we propose and solve a new problem called hierarchical few-shot object detection (Hi-FSOD), which aims to detect objects with hierarchical categories in the FSOD paradigm. To this end, on the one hand, we build the first large-scale and high-quality Hi-FSOD benchmark dataset HiFSOD-Bird, which contains 176,350 wild-bird images falling to 1,432 categories. All the categories are organized into a 4-level taxonomy, consisting of 32 orders, 132 families, 572 genera and 1,432 species. On the other hand, we propose the first Hi-FSOD method HiCLPL, where a hierarchical contrastive learning approach is developed to constrain the feature space so that the feature distribution of objects is consistent with the hierarchical taxonomy and the model's generalization power is strengthened. Meanwhile, a probabilistic loss is designed to enable the child nodes to correct the classification errors of their parent nodes in the taxonomy. Extensive experiments on the benchmark dataset HiFSOD-Bird show that our method HiCLPL outperforms the existing FSOD methods.
Lu Zhang 0060, Yang Wang 0100, Jiaogen Zhou, Chenbo Zhang, Yinglu Zhang, Jihong Guan, Yatao Bian, Shuigeng Zhou
ACM Multimedia1
2022 EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification
abstract
Minyi Zhao, Lu Zhang, Yi Xu, Jiandong Ding, Jihong Guan, Shuigeng Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Minyi Zhao, Lu Zhang 0060, Yi Xu 0003, Jiandong Ding, Jihong Guan, Shuigeng Zhou
NAACL-HLT2
2021 Accurate Few-Shot Object Detection With Support-Query Mutual Guidance and Hybrid Loss
abstract
Most object detection methods require huge amounts of annotated data and can detect only the categories that appear in the training set. However, in reality acquiring massive annotated training data is both expensive and time-consuming. In this paper, we propose a novel two-stage detector for accurate few-shot object detection. In the first stage, we employ a support-query mutual guidance mechanism to generate more support-relevant proposals. Concretely, on the one hand, a query-guided support weighting module is developed for aggregating different supports to generate the support feature. On the other hand, a supportguided query enhancement module is designed by dynamic kernels. In the second stage, we score and filter proposals via multi-level feature comparison between each proposal and the aggregated support feature based on a distance metric learnt by an effective hybrid loss, which makes the embedding space of distance metric more discriminative. Extensive experiments on benchmark datasets show that our method substantially outperforms the existing methods and lifts the SOTA of FSOD task to a higher level.
Lu Zhang 0060, Shuigeng Zhou, Jihong Guan, Ji Zhang 0001
CVPR1
2021 Weakly-supervised Text Classification Based on Keyword Graph
abstract
Weakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data.Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudolabels for unlabeled texts.However, existing methods treat keywords independently, thus ignore the correlation among them, which should be useful if properly exploited.In this paper, we propose a novel framework called ClassKG to explore keyword-keyword correlation on keyword graph by GNN.Our framework is an iterative process.In each iteration, we first construct a keyword graph, so the task of assigning pseudo labels is transformed to annotating keyword subgraphs.To improve the annotation quality, we introduce a self-supervised task to pretrain a subgraph annotator, and then finetune it.With the pseudo labels generated by the subgraph annotator, we then train a text classifier to classify the unlabeled texts.Finally, we re-extract keywords from the classified texts.Extensive experiments on both long-text and short-text datasets show that our method substantially outperforms the existing ones.
Lu Zhang 0060, Jiandong Ding, Yi Xu 0003, Yingyao Liu, Shuigeng Zhou
EMNLP (1)1
2021 DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples
abstract
The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, possibly due to the low quality of learned pseudo labels. In this paper, we propose a new SSL method called DP-SSL that adopts an innovative data programming (DP) scheme to generate probabilistic labels for unlabeled data. Different from existing DP methods that rely on human experts to provide initial labeling functions (LFs), we develop a multiple-choice learning~(MCL) based approach to automatically generate LFs from scratch in SSL style. With the noisy labels produced by the LFs, we design a label model to resolve the conflict and overlap among the noisy labels, and finally infer probabilistic labels for unlabeled samples. Extensive experiments on four standard SSL benchmarks show that DP-SSL can provide reliable labels for unlabeled data and achieve better classification performance on test sets than existing SSL methods, especially when only a small number of labeled samples are available. Concretely, for CIFAR-10 with only 40 labeled samples, DP-SSL achieves 93.82% annotation accuracy on unlabeled data and 93.46% classification accuracy on test data, which are higher than the SOTA results.
Yi Xu 0003, Jiandong Ding, Lu Zhang 0060, Shuigeng Zhou
NeurIPS3
2020 Learning Latent Semantic Attributes for Zero-Shot Object Detection
abstract
Zero-shot Object Detection (ZSD) aims to locate and classify instances of unseen categories. Existing methods focus on learning the mapping from visual space to semantic space, while the learning of discriminative representations for ZSD has not gained enough attention. In this paper, we demonstrate the necessity to learn discriminative semantic representations for ZSD, and propose a new end-to-end framework for this task. Our framework is able to learn discriminative semantic representations in an augmented space introduced for both user-defined and latent attributes, and refine the user-defined attributes with the help of unseen and external classes. The proposed method is extensively evaluated on two challenging ZSD datasets, and the experimental results show that our method significantly outperforms several existing methods.
Lu Zhang 0060, Yifan Tan, Shuigeng Zhou
ICTAI2