Guozhi Tang

dblp:248/6204 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
abstract
Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.
Weitao Jia, Jinghui Lu, Haiyang Yu 0004, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng 0009, Irene Li, Kun Yang 0010, Jingqun Tang, Teng Fu 0001, Changhong Jin, Xiaohui Lv, Can Huang 0002
AAAI5
2026 Proxy-Based Classification Boundary Alignment for Few-Shot Class-Incremental Learning
Hong-Wei Ge, Guozhi Tang, Jiulin Fan
IEEE Trans. Circuits Syst. Video Technol.3
2026 ModSolAgent: Automated Finite Element Code Generation for Abaqus via LLM-Based Agent
abstract
Finite element simulation and solution process represents a critical component in engineering analysis. While large language models (LLMs) have demonstrated remarkable capabilities in general-purpose code generation from textual descriptions, their application to generating structured and specialized finite element simulation scripts presents unique challenges. A key challenge is AI-generated hallucination, as these tasks require precise intent analysis, complex planning and reasoning, and strict adherence to logical consistency during execution. This often results in outputs that seem convincing but are ultimately erroneous. To address this challenge, we propose ModSolAgent, an LLM-based agent that generates Python scripts for Abaqus software to perform finite element modeling and solving tasks. ModSolAgent ensures the accuracy and logical coherence of code generation by leveraging a structured reasoning instruction, dynamic retrieval guidance template, and iterative verification generation. Experimental results demonstrate that LLMs augmented by ModSolAgent achieve an 83.3% success rate on real-world finite element simulation tasks, effectively meeting most Abaqus scripting requirements while significantly outperforming baseline models. To further enhance accessibility, we construct the AbqInstruct dataset by distilling knowledge from ModSolAgent to fine-tune the lightweight open-source models. Experiments show that fine-tuning on AbqInstruct leads to substantial performance improvements, with the lightweight model achieving proficiency across most finite element modeling and solving tasks. This work establishes a paradigm for integrating LLMs with specialized engineering software and providing novel insights for other structured, domain-specific code generation scenarios in industrial applications.
Zidi Li, Hong-Wei Ge, Guozhi Tang, Yuxuan Liu 0015
IEEE Trans. Ind. Informatics3
2025 ParGo: Bridging Vision-Language with Partial and Global Views
abstract
This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous works that rely on global attention-based projectors, our ParGo bridges the representation gap between the separately pre-trained vision encoders and the LLMs by integrating global and partial views, which alleviates the overemphasis on prominent regions. To facilitate the effective training of ParGo, we collect a large-scale detail-captioned image-text dataset named ParGoCap-1M-PT, consisting of 1 million images paired with high-quality captions. Extensive experiments on several MLLM benchmarks demonstrate the effectiveness of our ParGo, highlighting its superiority in aligning vision and language modalities. Compared to conventional Q-Former projector, our ParGo achieves an improvement of 259.96 in MME benchmark. Furthermore, our experiments reveal that ParGo significantly outperforms other projectors, particularly in tasks that emphasize detail perception ability.
An-Lan Wang, Bin Shan, Kun-Yu Lin, Guozhi Tang, Jingqun Tang, Wei-Shi Zheng 0001
AAAI6
2025 OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
abstract
Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10,000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1,500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https://github.com/Yuliang-Liu/MultimodalOCR.
Zhebin Kuang, Jiajun Song, Mingxin Huang, Linghao Zhu, Qidi Luo, Xinyu Wang 0010, Hao Lu 0003, Guozhi Tang, Bin Shan, Chunhui Lin, Binghong Wu, Hao Feng 0009, Hao Liu 0003, Can Huang 0002, Jingqun Tang, Wei Chen 0088, Xiang Bai
NeurIPS12
2025 Bi-VLDoc: bidirectional vision-language modeling for visually-rich document understanding
Chuwei Luo, Guozhi Tang, Qi Zheng 0002, Cong Yao, Yang Xue 0001, Luo Si
Int. J. Document Anal. Recognit.2
2025 Spiking Trans-YOLO: A range-adaptive energy-efficient bridge between YOLO and Transformer
Yushi Huo, Hong-Wei Ge, Guozhi Tang, Shengxuan Gao
Neurocomputing3
2025 Multiscale Recovery Diffusion Model With Unsupervised Learning for Video Anomaly Detection System
abstract
The rapid development of intelligent industry and smart city increases the number of surveillance devices, greatly enhancing the need for unsupervised automatic anomaly detection in real-time video surveillance, which uses raw data without laborious manual annotations. Existing video anomaly detection (VAD) methods encounter limitations when utilizing pretext tasks, such as reconstruction or prediction to identify abnormal events, as these tasks are not completely consistent and complementary with the essential objective of anomaly detection. Motivated by recent advances in diffusion models, we propose a multiscale recovery diffusion model, which relies on the proposed novel and effective pretext task named recovery to introduce the notion of generation speed. It utilizes critical step-by-step generation of diffusion probabilistic models in unsupervised anomaly detection scenarios. By incorporating a proposed multiscale spatial-temporal subtraction module, our model captures more detailed appearance and motion information of foreground objects without relying on other high-level pretrained models. Furthermore, an innovative push–pull loss further extends the disparity between normal and abnormal events through pseudolabels. We validate our model on five established benchmarks: UCSD Ped1, UCSD Ped2, CUHK Avenue, ShanghaiTech, and UCF-Crime, achieving frame-level area under the curves of 86.01%, 99.23%, 92.35%, 82.49%, and 74.79%, respectively, surpassing other state-of-the-art unsupervised VAD methods.
Hong-Wei Ge, Yuxuan Liu 0015, Guozhi Tang
IEEE Trans. Ind. Informatics4
2025 Multi-Memory Streams: A Paradigm for Online Video Super-Resolution in Complex Exposure Scenes
abstract
Existing online video super-resolution methods utilize implicit memories of previous frames to provide reference information, which have a single memory stream path and are highly dependent on the continuous memory stream. However, video capture in real-world scenes is typically affected by abnormal exposures resulting in sudden changes of lightness thus interrupting the memory stream, while long-term memories suffer from memory vanishing problems during transmission. To address this problem, we propose a novel multi-memory streams based online video super-resolution paradigm that adaptively corrects for abnormal exposures and creates multi-memory streams to accurately converge long-term memories. Specifically, we first propose an exposure detection-correction module, which utilizes optical flow overfitting property and temporal lightness information to detect and correct abnormal exposures to avoid interruption of memory streams. In addition, we propose a dynamic-static decoupled alignment strategy, which can adaptively select the alignment method based on pixel displacement, thus accurately aggregating past long-term memories to create multiple memory streams. Further, we propose an adaptive memory fusion module to mine complementary information between multiple memory streams to solve the memory vanishing problem. Extensive experimental results show that our method outperforms existing video super-resolution methods on complex exposure datasets. We also conduct detailed ablation experiments to analyze and validate our contributions.
Guozhi Tang, Hong-Wei Ge, Chunguo Wu
IEEE Trans. Multim.1
2024 Occluded Person Reidentification via a Universal Framework With Difference Consistency Guidance Learning
abstract
Occluded person reidentification (Re-ID) aims at learning discriminative identity features to match person images under the interference of occlusion situations in video surveillance of the visual Internet of Things (VIoT). Currently, occluded person Re-ID methods have made impressive improvements in nonperson occlusion situations. However, in real-world scenarios, the target person is commonly occluded by other nontarget persons, and the fine-grained differences between the persons make the model difficult to distinguish the discriminative identity features. To this end, we propose a difference consistency guidance (DCG) learning to enlarge the fine-grained differences by the guidance of the coarse-grained differences, which can distinguish the discriminative identity features in nontarget person occlusion situations. Then, DCG reduces the identity feature representation of the nonperson occlusion instances, which further improves the ability of the model in nonperson occlusion situations and improves the guidance ability of coarse-grained difference. Moreover, DCG can enhance the robustness of the model in the unsupervised occluded person Re-ID task and further improve the universal applicability of the model. Extensive experiment results under the supervised and unsupervised settings demonstrate the DCG outperforms the state-of-the-art methods in experiments conducted on the occluded person Re-ID benchmarks.
Yuxuan Liu 0015, Hong-Wei Ge, Guozhi Tang
IEEE Internet Things J.3
2024 Progressive reconstruction-decoupled face super-resolution framework with controllable knowledge guidance
Guozhi Tang, Hong-Wei Ge, Enxuan Gu, Yaqing Hou, Mingde Zhao 0002
Knowl. Based Syst.1
2024 Localization and saturation of degradation space for weakly-supervised real-world super-resolution
Guozhi Tang, Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu
Knowl. Based Syst.1
2024 A Virtual-Sensor Construction Network Based on Physical Imaging for Image Super-Resolution
abstract
Image imaging in the real world is based on physical imaging mechanisms. Existing super-resolution methods mainly focus on designing complex network structures to extract and fuse image features more effectively, but ignore the guiding role of physical imaging mechanisms for model design, and cannot mine features from a physical perspective. Inspired by the mechanism of physical imaging, we propose a novel network architecture called Virtual-Sensor Construction network (VSCNet) to simulate the sensor array inside the camera. Specifically, VSCNet first generates different splitting directions to distribute photons to construct virtual sensors, and then performs a multi-stage adaptive fine-tuning operation to fine-tune the number of photons on the virtual sensors to increase the photosensitive area and eliminate photon cross-talk, and finally converts the obtained photon distributions into RGB images. These operations can naturally be regarded as the virtual expansion of the camera's sensor array in the feature space, which makes our VSCNet bridge the physical space and feature space, and uses their complementarity to mine more effective features to improve performance. Extensive experiments on various datasets show that the proposed VSCNet achieves state-of-the-art performance with fewer parameters. Moreover, we perform experiments to validate the connection between the proposed VSCNet and the physical imaging mechanism. The implementation code is available at https://github.com/GZ-T/VSCNet.
Guozhi Tang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou, Mingde Zhao 0002
IEEE Trans. Image Process.1
2021 Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
abstract
Visual Information Extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing works decoupled this problem into several independent sub-tasks of text spotting (text detection and recognition) and information extraction, which completely ignored the high correlation among them during optimization. In this paper, we propose a robust Visual Information Extraction System (VIES) towards real-world scenarios, which is an unified end-to-end trainable framework for simultaneous text detection, recognition and information extraction by taking a single document image as input and outputting the structured information. Specifically, the information extraction branch collects abundant visual and semantic representations from text spotting for multimodal feature fusion and conversely, provides higher-level semantic clues to contribute to the optimization of text spotting. Moreover, regarding the shortage of public benchmarks, we construct a fully-annotated dataset called EPHOIE (https://github.com/HCIILAB/EPHOIE), which is the first Chinese benchmark for both text spotting and visual information extraction. EPHOIE consists of 1,494 images of examination paper head with complex layouts and background, including a total of 15,771 Chinese handwritten or printed text instances. Compared with the state-of-the-art methods, our VIES shows significant superior performance on the EPHOIE dataset and achieves a 9.01% F-score gain on the widely used SROIE dataset under the end-to-end scenario.
Chongyu Liu, Guozhi Tang, Jiaxin Zhang 0003, Shuaitao Zhang, Qianying Wang 0002, Yaqiang Wu, Mingxiang Cai
AAAI4
2021 Improving Machine Understanding of Human Intent in Charts
Sihang Wu, Canyu Xie, Guozhi Tang, Qianying Liao, Jiapeng Wang 0003, Bangdong Chen, Xinfeng Chang, Kai Ding 0009, Yichao Huang
ICDAR (3)4
2021 MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction
abstract
Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify each kind of semantics by introducing multimodal features, such as font, color, layout. But simply introducing multimodal features can't work well when faced with numeric semantic categories or some ambiguous texts. To address this issue, in this paper we propose a novel key-value matching model based on a graph neural network for VIE (MatchVIE). Through key-value matching based on relevancy evaluation, the proposed MatchVIE can bypass the recognitions to various semantics, and simply focuses on the strong relevancy between entities. Besides, we introduce a simple but effective operation, Num2Vec, to tackle the instability of encoded values, which helps model converge more smoothly. Comprehensive experiments demonstrate that the proposed MatchVIE can significantly outperform previous methods. Notably, to the best of our knowledge, MatchVIE may be the first attempt to tackle the VIE task by modeling the relevancy between keys and values and it is a good complement to the existing methods.
Guozhi Tang, Lele Xie, Jingdong Chen, Qianying Wang 0002, Yaqiang Wu
IJCAI1
2021 Tag, Copy or Predict: A Unified Weakly-Supervised Learning Framework for Visual Information Extraction using Sequences
abstract
Visual information extraction (VIE) has attracted increasing attention in recent years. The existing methods usually first organized optical character recognition (OCR) results in plain texts and then utilized token-level category annotations as supervision to train a sequence tagging model. However, it expends great annotation costs and may be exposed to label confusion, the OCR errors will also significantly affect the final performance. In this paper, we propose a unified weakly-supervised learning framework called TCPNet (Tag, Copy or Predict Network), which introduces 1) an efficient encoder to simultaneously model the semantic and layout information in 2D OCR results, 2) a weakly-supervised training method that utilizes only sequence-level supervision; and 3) a flexible and switchable decoder which contains two inference modes: one (Copy or Predict Mode) is to output key information sequences of different categories by copying a token from the input or predicting one in each time step, and the other (Tag Mode) is to directly tag the input sequence in a single forward pass. Our method shows new state-of-the-art performance on several public benchmarks, which fully proves its effectiveness.
Jiapeng Wang 0003, Guozhi Tang, Weihong Ma, Kai Ding 0009, Yichao Huang
IJCAI3
2021 Skeleton-based action recognition using sparse spatio-temporal GCN with edge effective resistance
Tasweer Ahmad, Luojun Lin, Guozhi Tang
Neurocomputing4