VLDB 2026 Research / reviewers in the wild / expert
Zeren Zhang
dblp:322/1107
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Head-Aware KV Cache Compression for Efficient Visual Autoregressive ModelingabstractVisual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n4) to O(Bn2). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster in- ference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B. Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo 0002, Zeren Zhang, Danping Zou, Weiyao Lin |
AAAI | 5 |
| 2025 | AIR: Unifying Individual and Collective Exploration in Cooperative Multi-Agent Reinforcement LearningabstractExploration in cooperative multi-agent reinforcement learning (MARL) remains challenging for value-based agents due to the absence of an explicit policy. Existing approaches include individual exploration based on uncertainty towards the system and collective exploration through behavioral diversity among agents. However, the introduction of additional structures often leads to reduced training efficiency and infeasible integration of these methods. In this paper, we propose Adaptive exploration via Identity Recognition~(AIR), which consists of two adversarial components: a classifier that recognizes agent identities from their trajectories, and an action selector that adaptively adjusts the mode and degree of exploration. We theoretically prove that AIR can facilitate both individual and collective exploration during training, and experiments also demonstrate the efficiency and effectiveness of AIR across various tasks. Guangchong Zhou, Zeren Zhang |
AAAI | 2 |
| 2025 | TopSUMseg: A Topology-Aware Swin Transformer-Mamba Framework for 3D Seismic Fault Image SegmentationabstractSeismic fault image segmentation is crucial for interpreting subsurface geological structures, supporting geologists in resource exploration and structural analysis. However, current deep learning models struggle with single-architecture limitations and the distinctive characteristics of seismic faults, which are distinguished by elongated structures with uneven spatial distributions. To address these challenges, we propose TopSUMseg, a novel Topology-Aware Swin Transformer-Mamba framework for 3D seismic fault image segmentation. Our framework combines Swin Transformer’s local feature extraction with Mamba’s efficient sequence modeling, and boosts 3D spatial modeling in Mamba with a newly designed Global-Local Attention module (GLA). Additionally, we design a Topology-Aware Structural Constraint (TASC) to align predictions with ground-truth structures in the feature space, promoting the modeling of complex fault geometries. Experiments on Thebe, the largest public seismic dataset, demonstrate that TopSUMseg achieves state-of-the-art performance with OIS and ODS scores of 0.879 and 0.875, respectively. Trained entirely from scratch, TopSUMseg nonetheless achieves superior performance compared to extensively pre-trained counterparts. In addition, TopSUMseg maintains a significantly lower parameter count while achieving a favorable trade-off between segmentation performance and time complexity, making it a practical and generalizable solution for real-world seismic fault interpretation. Ran Chen 0002, Jingyang Deng, Zeren Zhang, Ruohua Shi, Jinwen Ma |
ECAI | 3 |
| 2025 | Reframing Multimodal Complex Document Layout Understanding: A Layout-Aware Multi-Source Reasoning Decision FrameworkabstractMultimodal large language models (MLLMs) have achieved significant progress in document understanding. However, complex layout reasoning, characterized by concise answers and cross-page integration, remains a challenge. Unlike conventional semantics-oriented tasks, this task demands accurate visual perception of fine-grained structural elements and logical reasoning across multi-page documents. Existing approaches primarily focus on information extraction and semantic understanding, limiting the capacity of fine-tuned autoregressive models to capture short-answer reasoning signals and generalize to complex layout structures. To address this, we propose the Layout-Aware Multi-Source Reasoning Decision Framework (LAMRD), which reframes complex layout reasoning as a decision-making task over multi-source reasoning paths. In the reasoning path construction stage, LAMRD generates layout-aware reasoning paths by integrating internal visual cues and external knowledge from three complementary perspectives: Visual Structural Awareness (VSA), Logical Reasoning Paths (LRP), and External Knowledge Augmentation (EKA). In the reasoning path decision stage, we employ Group Relative Policy Optimization (GRPO) to train a decision model that produces the final answer based on these paths. We conduct comprehensive evaluations using Qwen2.5-VL-7B-Instruct on the CEP-7K dataset, covering layout structure understanding, information extraction, and logical association. Experimental results demonstrate that LAMRD outperforms advanced MLLMs in accuracy, validating its effectiveness for complex document layout understanding. Ran Chen 0002, Jingyang Deng, Zeren Zhang, Xuefei Tong, Jinwen Ma, Qinghui Shi, Yuanjun Li |
ECAI | 4 |
| 2025 | Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference AlignmentabstractIn handling complex, domain-specific tasks, particularly in the context of state-owned assets and enterprises (SOAEs), general LLMs suffer from the knowledge gap due to insufficient exposure to domain-specific corpora, and the value disagreement, as they are aligned with universal values rather than domain-specific ones. To tackle these challenges, we propose a novel training strategy tailored for the SOAEs domain. This strategy includes a improved domain-adaptive pretraining (DAP) phase with a replay mechanism to mitigate catastrophic forgetting. Following DAP, we utilize a selective portion of domain-specific data for supervised fine-tuning (SFT), and innovatively integrate low-quality data with the remaining SFT data to curate tailored preference datasets, leveraging the Kahneman-Tversky Optimization technique to align our LLMs. Our proposed approach effectively utilizes the data that is often discarded in conventional training procedures, highlighting the substantial improvements in model performance and the importance of training methodologies for domain-specific tasks. Jingyang Deng, Zeren Zhang, Jo-Ku Cheng, Jinwen Ma |
ICASSP | 2 |
| 2025 | Diagram Formalization Enhanced Multi-Modal Geometry Problem SolverabstractMathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems, which require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geometry problem-solving than LLMs that only process text. This limitation is further amplified by the lack of effective methods for representing geometric relationships. To address these issues, we introduce the Diagram Formalization Enhanced Geometry Problem Solver (DFE-GPS), a new framework that integrates visual features, geometric formal language, and natural language representations. Specifically, we propose a novel synthetic data approach and construct a large-scale geometric dataset, SynthGeo228K, annotated with formal and natural language captions, designed to enhance the vision encoder to understand geometric structures better. Our framework improves MLLMs’ ability to process geometric diagrams and extends their application to open-ended tasks on the formalgeo7k dataset. Zeren Zhang, Jo-Ku Cheng, Jingyang Deng, Jinwen Ma, Ziran Qin, Tuo Leng |
ICASSP | 1 |
| 2025 | SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent SpaceabstractCombining face-swapping with lip synchronization offers a cost-effective solution for generating customized talking faces. However, directly cascading existing models can introduce significant interference and reduce video clarity due to limited interaction space in the low-level RGB domain. To solve this, we propose SwapTalk, a unified framework that performs face-swapping and lip synchronization within the same latent VQ-embedding space, known for its editability and fidelity. We enhance generalization to unseen identities with identity loss in the face-swapping module and improve synchronization quality with expert discriminator supervision. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experiments on HDTF show that SwapTalk outperforms existing methods in video quality, lip synchronization accuracy, face-swapping fidelity, and identity consistency. Zeren Zhang, Haibo Qin, Jo-Ku Cheng, Yitao Duan, Jinwen Ma |
ICASSP | 1 |
| 2025 | Unveiling Decision Intention for Cooperative Multi-Agent Reinforcement Learning
Zeren Zhang, Zhiwei Xu 0005, Guangchong Zhou, Dapeng Li 0001, Bin Zhang 0052 |
AAMAS | 1 |
| 2025 | GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions
Jo-Ku Cheng, Zeren Zhang, Ran Chen 0002, Jingyang Deng, Ziran Qin, Jinwen Ma |
ACM Multimedia | 2 |
| 2024 | Decentralized Extension for Centralized Multi-Agent Reinforcement Learning via Online Distillation
Zeren Zhang, Bin Zhang 0052, Guangchong Zhou, Dapeng Li 0001, Zhiwei Xu 0005 |
ICONIP (3) | 1 |
| 2023 | Consensus Learning for Cooperative Multi-Agent Reinforcement LearningabstractAlmost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and choose actions given local observations during execution. Inspired by viewpoint invariance and contrastive learning, we propose consensus learning for cooperative multi-agent reinforcement learning in this study. Although based on local observations, different agents can infer the same consensus in discrete spaces without communication. We feed the inferred one-hot consensus to the network of agents as an explicit input in a decentralized way, thereby fostering their cooperative spirit. With minor model modifications, our suggested framework can be extended to a variety of multi-agent reinforcement learning algorithms. Moreover, we carry out these variants on some fully cooperative tasks and get convincing results. Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Zeren Zhang, Guangchong Zhou, Hao Chen 0103 |
AAAI | 4 |
| 2023 | PCSalmix: Gradient Saliency-Based Mix Augmentation for Point Cloud ClassificationabstractPoint cloud classification has sparked many researchers’ interest for its cornerstone role in 3D applications. Inheriting the CutMix series augmentation that performs well in 2D images, PointCutMix and RSMix are proposed to generate new samples for 3D point clouds, by replacing partial points of one cloud with those of another. However, the selection of mixed regions is all built on randomness, ignoring the significance of point clouds’ saliency. To address this deficiency, we propose PCSalMix: a novel Saliency-based Mix augmentation for Point Cloud classification. The gradient of classification network on inputs is a natural tool to locate the saliency. Based on this discovery, we extract points with larger gradient values to make more representative samples. Afterward, the soft labels are weighted more accurately by accumulated gradients rather than count ratios of points. The experimental results verify the outperformance of our method on ModelNet40 and ModelNet10 benchmarks in terms of accuracy and robustness against adversarial attacks. Zeren Zhang, Jinwen Ma |
ICASSP | 2 |
| 2023 | Mastering Complex Coordination Through Attention-Based Dynamic Graph
Guangchong Zhou, Zhiwei Xu 0005, Zeren Zhang |
ICONIP (1) | 3 |
| 2023 | SORA: Improving Multi-agent Cooperation with a Soft Role Assignment Mechanism
Guangchong Zhou, Zhiwei Xu 0005, Zeren Zhang |
ICONIP (1) | 3 |
| 2023 | Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARLabstractValue decomposition methods have gained popularity in the field of cooperative multi-agent reinforcement learning. However, almost all existing methods follow the principle of Individual Global Max (IGM) or its variants, which limits their problem-solving capabilities. To address this, we propose a dual self-awareness value decomposition framework, inspired by the notion of dual self-awareness in psychology, that entirely rejects the IGM premise. Each agent consists of an ego policy for action selection and an alter ego value function to solve the credit assignment problem. The value function factorization can ignore the IGM assumption by utilizing an explicit search procedure. On the basis of the above, we also suggest a novel anti-ego exploration mechanism to avoid the algorithm becoming stuck in a local optimum. As the first fully IGM-free value decomposition method, our proposed framework achieves desirable performance in various cooperative tasks. Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Guangchong Zhou, Zeren Zhang |
NeurIPS | 5 |
| 2023 | Overcoming Catastrophic Forgetting for Fine-Tuning Pre-trained GANs
Zeren Zhang, Xingjian Li 0002, Tianyang Wang 0004, Jinwen Ma, Haoyi Xiong, Cheng-Zhong Xu 0001 |
ECML/PKDD (5) | 1 |
| 2022 | SG-Net: Semantic Guided Network for Image Dehazing
Xiangyang Guo, Zeren Zhang, Jinwen Ma |
ACCV (3) | 3 |