Zeqing Wang

dblp:322/3313 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Jump-teaching: Combating Sample Selection Bias via Temporal Disagreement
abstract
Sample selection is a straightforward technique to combat noisy labels, aiming to prevent mislabeled samples from degrading the robustness of neural networks. However, existing methods mitigate compounding selection bias either by leveraging dual-network disagreement or additional forward propagations, leading to multiplied training overhead. To address this challenge, we introduce Jump-teaching, an efficient sample selection framework for debiased model update and simplified selection criterion. Based on a key observation that a neural network exhibits significant disagreement across different training iterations, Jump-teaching proposes a jump-manner model update strategy to enable self-correction of selection bias by harnessing temporal disagreement, eliminating the need for multi-network or multi-round training. Furthermore, we employ a sample-wise selection criterion building on the intra variance of a decomposed single loss for a fine-grained selection without relying on batch-wise ranking or dataset-wise modeling. Extensive experiments demonstrate that Jump-teaching outperforms state-of-the-art counterparts while achieving a nearly overhead-free selection procedure, which boosts training speed by up to 4.47× and reduces peak memory footprint by 54%.
Kangye Ji, Zeqing Wang, Qichang Zhang, Bohu Huang
AAAI3
2026 SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage Efficiency
abstract
Segment Anything Model (SAM) struggles in open-world scenarios with diverse domains. In such settings, naive fine-tuning with a well-designed learning module is inadequate and often causes catastrophic forgetting issue when learning incrementally. To address this issue, we propose a novel continual learning (CL) method for SAM, termed SAMCL. Rather than relying on a fixed learning module, our method decomposes incremental knowledge into separate modules and trains a selector to choose the appropriate one during inference. However, this intuitive design introduces two key challenges: ensuring effective module learning and selection, and managing storage as tasks accumulate. To tackle these, we introduce two components: AugModule and Module Selector. AugModule reduces the storage of the popular LoRA learning module by sharing parameters across layers while maintaining accuracy. It also employs heatmaps—generated from point prompts—to further enhance domain adaptation with minimal additional cost. Module Selector leverages the observation that SAM’s embeddings can effectively distinguish domains, enabling high selection accuracy by training on low-consumed embeddings instead of raw images. Experiments show that SAMCL outperforms state-of-the-art methods, achieving only 0.19% forgetting and at least 2.5% gain on unseen domains. Each AugModule requires just 0.233 MB, reducing storage by at least 24.3% over other fine-tuning approaches. The buffer storage for Module Selector is further reduced by up to 256x.
Zeqing Wang, Kangye Ji
AAAI1
2026 Minute-Long Videos with Dual Parallelisms
abstract
Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address this, we propose a novel distributed inference strategy, termed DualParal. The core idea is that, instead of generating an entire video on a single GPU, we parallelize computation by partitioning both video frames and model layers across multiple GPUs. However, a naive parallel implementation is not feasible. Because all frames need to share the same noise level, they can't be processed independently. Instead, every step must wait for all others to finish, which cancels out the speed benefits of parallel processing. We overcome this obstacle with a block-wise denoising scheme. Namely, we segment the video into sequential blocks, each with a different noise level. As a result, we process them in a pipeline across the GPUs. Each GPU, holding a subset of the model layers, processes a specific block of frames and passes the results to the next GPU, enabling asynchronous computation and communication. To further optimize performance, we incorporate two key enhancements. Firstly, each GPU uses a feature cache technique to reduce the overhead of smooth transitions by reusing only features involved in cross-frame computation from the prior block, minimizing inter-GPU communication and redundant computation. Secondly, we employ a coordinated noise initialization strategy, ensuring globally consistent temporal dynamics by sharing initial noise patterns across GPUs. Together, these enable fast, artifact-free, and infinitely long video generation. Applied to the latest diffusion transformer video generator, our method efficiently produces 1,025-frame videos with up to 6.54x lower latency and 1.48x lower memory cost on 8xRTX 4090 GPUs.
Zeqing Wang, Xingyi Yang, Zhenxiong Tan, Yuecong Xu, Xinchao Wang
AAAI1
2026 Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
abstract
Lirong Gao, Zeqing Wang, Yuyan Cai, Jiayi Deng, Yanmei Gu, Yiming Zhang, Jia Zhou, Yanfei Zhang, Junbo Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lirong Gao, Zeqing Wang, Yuyan Cai, Jiayi Deng, Yanmei Gu, Junbo Zhao 0001
ACL (1)2
2026 Toward Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
abstract
Recent methods to enhance Vision-Language Models (VLMs) for Visual Question Answering (VQA) have focused on strengthening their inference capabilities, enabling them to tackle VQA tasks independently rather than merely as aids to Large Language Models (LLMs). However, these approaches often ignore the rich commonsense knowledge inside the given VQA image sampled from the real world, limiting the full potential of VLMs. Inspired by the human top-down reasoning process, i.e., systematically exploring relevant issues to derive a comprehensive answer, this work introduces a novel, explainable multi-agent collaboration framework by leveraging the expansive knowledge of LLMs to enhance the capabilities of VLMs themselves. Our framework comprises three agents, i.e.,Responder,Seeker, andIntegrator, to collaboratively answer the given VQA question by seeking its relevant issues and generating the final answer in such a top-down reasoning process. The VLM-basedResponderagent generates the answer candidates for the question and responds to other relevant issues. TheSeekeragent, primarily based on LLM, identifies relevant issues related to the question to inform theResponderagent and constructs a Multi-View Knowledge Base (MVKB) for the given visual scene by leveraging the build-in world knowledge of LLM. TheIntegratoragent combines knowledge from theSeekeragent and theResponderagent to produce the final VQA answer. Extensive and comprehensive evaluations on diverse VQA datasets with a variety of VLMs demonstrate the superior performance and interpretability of our framework over the baseline method, e.g., 5.7% improvement on VQA-RAD and 5.2% on Winoground in the zero-shot setting without extra training cost.
Zeqing Wang, Wentao Wan 0001, Qiqing Lao, Runmeng Chen, Minjie Lang, Xiao Wang 0002, Feng Gao 0014, Keze Wang, Liang Lin 0004
IEEE Trans. Multim.1
2025 LightCL: Compact Continual Learning with Low Memory Footprint For Edge Device
abstract
Continual learning (CL) is a technique that enables neural networks to constantly adapt to their dynamic surroundings. Despite being overlooked for a long time, this technology can considerably address the customized needs of users in edge devices. Actually, most CL methods require huge resource consumption by the training behavior to acquire generalizability among all tasks for delaying forgetting regardless of edge scenarios. Therefore, this paper proposes a compact algorithm called LightCL, which evaluates and compresses the redundancy of already generalized components in structures of the neural network. Specifically, we consider two factors of generalizability, learning plasticity and memory stability, and design metrics of both to quantitatively assess generalizability of neural networks during CL. This evaluation shows that generalizability of different layers in a neural network exhibits a significant variation. Thus, we Maintain Generalizability by freezing generalized parts without the resource-intensive training process and Memorize Feature Patterns by stabilizing feature extracting of previous tasks to enhance generalizability for less-generalized parts with a little extra memory, which is far less than the reduction by freezing. Experiments illustrate that LightCL outperforms other state-of-the-art methods and reduces at most 6.16× memory footprint. We also verify the effectiveness of LightCL on the edge device.
Zeqing Wang, Kangye Ji, Bohu Huang
ASP-DAC1
2025 Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body
abstract
Recent improvements in visual synthesis have significantly enhanced the depiction of generated human photos, which are pivotal due to their wide applicability and demand. Nonetheless, the existing text-to-image or text-to-video models often generate low-quality human photos that might differ considerably from real-world body structures, referred to as "abnormal human bodies". Such abnormalities, typically deemed unacceptable, pose considerable challenges in the detection and repair of them within human photos. These challenges require precise abnormality recognition capabilities, which entail pinpointing both the location and the abnormality type. Intuitively, Visual Language Models (VLMs) that have obtained remarkable performance on various visual tasks are quite suitable for this task. However, their performance on abnormality detection in human photos is quite poor. Hence, it is quite important to highlight this task for the research community. In this paper, we first introduce a simple yet challenging task, i.e., Fine-grained Human-body Abnormality Detection (FHAD), and construct two high-quality datasets for evaluation. Then, we propose a meticulous framework, named HumanCalibrator, which identifies and repairs abnormalities in human body structures while preserving the other content. Experiments indicate that our HumanCalibrator achieves high accuracy in abnormality detection and accomplishes an increase in visual comparisons while preserving the other visual content.
Zeqing Wang, Qingyang Ma, Wentao Wan 0001, Keze Wang, Yonghong Tian 0001
CVPR1
2025 Tracking-Aware Deformation Field Estimation for Non-rigid 3D Reconstruction in Robotic Surgeries
abstract
Minimally invasive procedures have been advanced rapidly by the robotic laparoscopic surgery. The latter greatly assists surgeons in sophisticated and precise operations with reduced invasiveness. Nevertheless, it is still safety critical to be aware of even the least tissue deformation during instrument-tissue interactions, especially in 3D space. To address this, recent works rely on NeRF to render 2D videos from different perspectives and eliminate occlusions. However, most of the methods fail to predict the accurate 3D shapes and associated deformation estimates robustly. Differently, we propose Tracking-Aware Deformation Field (TADF), a novel framework which reconstructs the 3D mesh along with the 3D tissue deformation simultaneously. It first tracks the key points of soft tissue by a foundation vision model, providing an accurate 2D deformation field. Then, the 2D deformation field is smoothly incorporated with a neural implicit reconstruction network to obtain tissue deformation in the 3D space. Finally, we experimentally demonstrate that the proposed method provides more accurate deformation estimation compared with other 3D neural reconstruction methods in two public datasets. Our demo is available at https://kasumigaoka-utaha.github.io/TADF-web/. Our code is available at https://github.com/Zing110/TADF.
Zeqing Wang, Yutong Ban
IROS1
2025 ACOC-MT: More Effective Handling of Real-World Noisy Labels in Remote Sensing Semantic Segmentation
abstract
In remote sensing semantic segmentation, imperfect labels are prevalent due to the complexities of data acquisition and annotation processes. Although recent approaches to noisy label correction in remote sensing segmentation have shown promising results, challenges remain in accuracy and generalizability, given the incomplete consideration of the complex mixing involved. Nearly all methods should address three key challenges to identify and correct noisy labels: when to select labels, which labels to select, and how to handle the selected labels. We propose a novel label correction framework, Adaptive Consistency-Guided Object-Level Label Correction with Mean Teacher (ACOC-MT), which effectively addresses all three of these challenges. The ACOC-MT framework determines when to conduct label correction through the Stable Early Learning Detection module, selects noisy labels using the Consistency Threshold Mask module, and corrects the labels through the Object Label Correction module. To validate against real-world noisy labels rather than simulated ones, we constructed three real-world noisy-labeled datasets, CPCTC-N, BBD250-N, and CFD-N, from three different remote sensing scenarios. Extensive experiments on these three datasets demonstrate the efficacy and superior performance of our approach.
Zeqing Wang, Yixiang Li, Zhaoming Wu, Zhengchao Chen
IEEE Trans. Geosci. Remote. Sens.1
2024 Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation
abstract
Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, often resulting in unrealistic facial animations. To the best of our knowledge, this work makes the first attempt to explore the coupled information between the speaking style and the semantic content in facial motions. Specifically, we introduce an innovative speaking style disentanglement method, which enables arbitrary-subject speaking style encoding and leads to a more realistic synthesis of speech-driven facial animations. Subsequently, we propose a novel framework called Mimic to learn disentangled representations of the speaking style and content from facial motions by building two latent spaces for style and content, respectively. Moreover, to facilitate disentangled representation learning, we introduce four well-designed constraints: an auxiliary style classifier, an auxiliary inverse classifier, a content contrastive loss, and a pair of latent cycle losses, which can effectively contribute to the construction of the identity-related style space and semantic-related content space. Extensive qualitative and quantitative experiments conducted on three publicly available datasets demonstrate that our approach outperforms state-of-the-art methods and is capable of capturing diverse speaking styles for speech-driven 3D facial animation. The source code and supplementary video are publicly available at: https://zeqing-wang.github.io/Mimic/
Zeqing Wang, Keze Wang, Tianshui Chen, Haifeng Zeng, Wenxiong Kang
AAAI2
2023 E3ID: An efficient end to end person search model
Yanchun Liang 0001, Zeqing Wang, Xiaosong Han
Pattern Recognit. Lett.4