Zhaoqing Zhu

dblp:303/0785 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-2005-5645ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
abstract
Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particularly after training on document instruction datasets. An effective evaluation method for document instruction data is crucial in constructing instruction data with high efficacy, which, in turn, facilitates the training of LLMs and MLLMs for document VQA. However, most existing evaluation methods for instruction data are limited to the textual content of the instructions themselves, thereby hindering the effective assessment of document instruction datasets and constraining their construction. In this paper, we propose ProcTag, a data-oriented method that assesses the efficacy of document instruction data. ProcTag innovatively performs tagging on the execution process of instructions rather than the instruction text itself. By leveraging the diversity and complexity of these tags to assess the efficacy of the given dataset, ProcTag enables selective sampling or filtering of document instructions. Furthermore, DocLayPrompt, a novel semi-structured layout-aware document prompting strategy, is proposed for effectively representing documents. Experiments demonstrate that sampling existing open-sourced and generated document VQA/instruction datasets with ProcTag significantly outperforms current methods for evaluating instruction data. Impressively, with ProcTag-based sampling in the generated document datasets, only 30.5 percent of the document instructions are required to achieve 100 percent efficacy compared to the complete dataset.
Yufan Shen, Chuwei Luo, Zhaoqing Zhu, Qi Zheng 0002, Jiajun Bu, Cong Yao
AAAI3
2025 A Simple yet Effective Layout Token in Large Language Models for Document Understanding
abstract
Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks.
Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011
CVPR1
2025 Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
abstract
Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks.
Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu
EMNLP3
2024 LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
abstract
Recently, leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However, previous works that employ LLMs/MLLMs for document understanding have not fully explored and utilized the document layout information, which is vital for precise document understanding. In this paper, we propose LayoutLLM, an LLM/MLLM based method for document understanding. The core of LayoutLLM is a layout instruction tuning strategy, which is specially designed to enhance the comprehension and utilization of document layouts. The proposed layout instruction tuning strategy consists of two components: Layout-aware Pre-training and Layout-aware Supervised Fine-tuning. To capture the characteristics of document layout in Layout-aware Pre-training, three groups of pretraining tasks, corresponding to document-level, region-level and segment-level information, are introduced. Furthermore, a novel module called layout chain-of-thought (LayoutCoT) is devised to enable LayoutLLM to focus on regions relevant to the question and generate accurate answers. LayoutCoT is effective for boosting the performance of document understanding. Meanwhile, it brings a certain degree of interpretability, which could facilitate manual inspection and correction. Experiments on standard benchmarks show that the proposed LayoutLLM significantly outperforms existing methods that adopt open-source 7B LLMs/MLLMs for document understanding.
Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 0002, Cong Yao
CVPR3
2024 CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition
abstract
As one of the most informative behaviors of humans, facial expressions are often compound and variable, which is manifested by the fact that different people may express the same expression in very different ways. However, most facial expression recognition (FER) methods still use one-hot or soft labels as the supervision, which lack sufficient semantic descriptions of facial expressions and are less interpretable. Recently, contrastive vision-language pre-training models (e.g., CLIP) use text as the supervision and have injected new vitality into various computer vision tasks, benefiting from the rich semantics in text. Therefore, we propose CLIPER, a unified framework for both static and dynamic facial Expression Recognition based on CLIP. Besides, we introduce multiple expression text descriptors (METD) to learn fine-grained expression representations and a two-stage training paradigm to reserve the interpretability of CLIP. We conduct extensive experiments on several popular FER benchmarks to demonstrates the effectiveness of CLIPER. The source code will be available at https://github.com/muse1998/CLIPER.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
ICME3
2023 Intensity-Aware Loss for Dynamic Facial Expression Recognition in the Wild
abstract
Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natural expression recognition scene. However, DFER is often more challenging. One of the main reasons is that video sequences often contain frames with different expression intensities, especially for the facial expressions in the real-world scenarios, while the images in SFER frequently present uniform and high expression intensities. Nevertheless, if the expressions with different intensities are treated equally, the features learned by the networks will have large intra-class and small inter-class differences, which are harmful to DFER. To tackle this problem, we propose the global convolution-attention block (GCA) to rescale the channels of the feature maps. In addition, we introduce the intensity-aware loss (IAL) in the training process to help the network distinguish the samples with relatively low expression intensities. Experiments on two in-the-wild dynamic facial expression datasets (i.e., DFEW and FERV39k) indicate that our method outperforms the state-of-the-art DFER approaches. The source code will be available at https://github.com/muse1998/IAL-for-Facial-Expression-Recognition.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
AAAI3
2023 AFNet-M: Adaptive Fusion Network with Masks for 2D+3D Facial Expression Recognition
abstract
2D+3D facial expression recognition (FER) can effectively cope with illumination and pose changes by merging texture and robust depth information. Most deep learning-based approaches employ the simple fusion strategy that concatenates the multimodal features directly after fully-connected layers, without considering the different degrees of significance for each modality. Meanwhile, how to focus more on both 2D and 3D local features is still a great challenge. In this paper, we propose the adaptive fusion network with masks (AFNet-M) for 2D+3D FER. To enhance 2D and 3D local features, we take the masks annotating salient regions of the face as prior knowledge and design the mask attention module (MA) which can automatically learn two modulation vectors to scale the feature maps. We also introduce an adaptive fusion module (AF) at convolutional layers through the computed importance weights. Experimental results demonstrate that our AFNet-M achieves the state-of-the-art performance on BU-3DFE and Bosphorus datasets and requires fewer parameters in comparison with other models.
Mingzhe Sui, Hanting Li, Zhaoqing Zhu, Feng Zhao 0004
ICIP3
2022 RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive Learning
abstract
Conventional visual relationship detection models only use the numeric ids of relation labels for training, but ignore the semantic correlation between the labels, which leads to severe training biases and harms the generalization ability of representations.In this paper, we introduce compact language information of relation labels for regularizing the representation learning of visual relations.Specifically, we propose a simple yet effective visual Relationship prediction framework that transfers natural language knowledge learned from Contrastive Language-Image Pre-training (CLIP) models to enhance the relationship prediction, termed as RelCLIP.Benefiting from the powerful visual-semantic alignment ability of CLIP at image level, we introduce a novel Relational Contrastive Learning (RCL) approach that explores relation-level visual-semantic alignment via learning to match cross-modal relational embeddings.By collaboratively learning the semantic coherence and discrepancy from relation triplets, the model can generate more discriminative and robust representations.Experimental results on the Visual Genome dataset show that RelCLIP achieves significant improvements over strong baselines under full (providing accurate labels) and distant supervision (providing noise labels), demonstrating its powerful generalization ability in learning relationship representations.
Yi Zhu 0004, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang, Feng Zhao 0004, Jianzhuang Liu
EMNLP2
2022 CMANET: Curvature-Aware Soft Mask Guided Attention Fusion Network for 2D+3D Facial Expression Recognition
abstract
As 2D texture and 3D structural information can describe facial features complementarily, 2D+3D facial expression recognition (FER) has received widespread attention. Though recent methods for 2D+3D FER have reached excellent performance, they still face two challenges: the way for attending to critical face areas and the strategy for fusing multi-modal information. To address these issues, we propose a curvature-aware soft mask guided attention fusion network (CMANet), which mainly consists of two components: curvature-aware attention module and multi-modal attention fusion module. The former utilizes the curvature-aware soft mask guiding the homo-modal attention mechanism to focus on potentially important areas with soft weights, while the latter applies pixel-level fusion on multi-modal features to retain the significant information from different modalities and also allows multi-modal features to interact in a larger field of view. Extensive experimental results show that our CMANet achieves outstanding accuracies (90.24% on BU-3DFE and 89.36% on Bosphorus) and outperforms the state-of-the-art methods.
Zhaoqing Zhu, Mingzhe Sui, Hanting Li, Feng Zhao 0004
ICME1
2022 MMNet: Muscle Motion-Guided Network for Micro-Expression Recognition
abstract
Facial micro-expressions (MEs) are involuntary facial motions revealing people’s real feelings and play an important role in the early intervention of mental illness, the national security, and many human-computer interaction systems. However, existing micro-expression datasets are limited and usually pose some challenges for training good classifiers. To model the subtle facial muscle motions, we propose a robust micro-expression recognition (MER) framework, namely muscle motion-guided network (MMNet). Specifically, a continuous attention (CA) block is introduced to focus on modeling local subtle muscle motion patterns with little identity information, which is different from most previous methods that directly extract features from complete video frames with much identity information. Besides, we design a position calibration (PC) module based on the vision transformer. By adding the position embeddings of the face generated by the PC module at the end of the two branches, the PC module can help to add position information to facial muscle motion-pattern features for the MER. Extensive experiments on three public micro-expression datasets demonstrate that our approach outperforms state-of-the-art methods by a large margin. Code is available at https://github.com/muse1998/MMNet.
Hanting Li, Mingzhe Sui, Zhaoqing Zhu, Feng Zhao 0004
IJCAI3
2021 FFNet-M: Feature Fusion Network with Masks for Multimodal Facial Expression Recognition
abstract
Compared with 2D facial expression recognition (FER) and 3D FER, 2D+3D FER can handle the effects of illumination changes and pose variations. The combination of 2D texture and 3D attribute information can further improve the performance. However, most existing approaches still face two challenges: the selection of proper networks for extracting multimodal features, and the significance of local features in salient regions for expression classification. To address these challenges, we propose an efficient feature fusion network with masks (FFNet-M) for 2D+3D FER. Each 3D scan is rep-resented by three types of attribute maps (i.e., depth map, normal map, and texture image), which are then fed into FFNet-M with different networks to extract both 2D and 3D features. Moreover, we design two masks to make FFNet-M focus on 2D local features while paying attention to 3D local features in salient regions. Experimental results show that our FFNet-M outperforms state-of-the-art methods on BU-3DFE dataset and also achieves a high accuracy on Bosphorus dataset.
Mingzhe Sui, Zhaoqing Zhu, Feng Zhao 0004, Feng Wu 0001
ICME2