EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Zhang 0006
dblp:44/8352-6
· DBLP profile ↗
32ranked-venue papers
4as first author
32since 2021 · last 2026
0009-0006-8819-3741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D CoCa: Contrastive Learners are 3D Captionersabstract3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address these challenges, we propose 3D CoCa, a novel unified framework that seamlessly combines contrastive vision-language learning with 3D caption generation into a single architecture. We design a frozen CLIP vision-language backbone to provide rich semantic priors, a spatially-aware 3D scene encoder to capture geometric context, and a multi-modal decoder to generate descriptive captions. Unlike the prior two-stage methods that rely on explicit object proposals, 3D CoCa jointly optimizes contrastive and captioning objectives in a shared feature space, eliminating the need for external detectors or handcrafted proposals. This joint training paradigm yields stronger spatial reasoning and richer semantic grounding by aligning 3D and textual representations. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that 3D CoCa significantly outperforms current state-of-the-arts by 10.2% and 5.76% in [email protected], respectively. Code will be available at https://github.com/AIGeeksGroup/3DCoCa. Zeyu Zhang 0006, Yemin Wang, Hao Tang 0005 |
3DV | 2 |
| 2026 | Revisiting Depth Representations for Feed-Forward 3D Gaussian SplattingabstractDepth maps are widely used in feed-forward 3D Gaussian Splatting (3DGS) pipelines by unprojecting them into 3D point clouds for novel view synthesis. This approach offers advantages such as efficient training, the use of known camera poses, and accurate geometry estimation. However, depth discontinuities, which are particularly problematic at the boundaries of the reconstructed geometry, often lead to fragmented or sparse point clouds, degrading rendering quality-a well-known limitation of depth-based representations. To tackle this issue, we introduce PM-Loss, a novel regularization loss based on a pointmap predicted by a pre-trained transformer. Although the pointmap itself may be less accurate than the depth map, it provides a powerful prior for geometric coherence and structural completeness, especially at the very edges where depth prediction falters. With the improved depth map, our method significantly improves the feed-forward 3DGS across various architectures and scenes, delivering consistently better rendering results. Project page: https://aimuofa.github.io/PMLoss Duochao Shi, Weijie Wang 0014, Donny Y. Chen, Zeyu Zhang 0006, Jiawang Bian, Bohan Zhuang |
3DV | 4 |
| 2026 | V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision TransformerabstractVision Transformer (ViT) has become one of the cornerstones of the computer vision field, demonstrating exceptional performance. However, its inherent high computational complexity and inference latency still pose significant obstacles for deployment in resource-constrained environments. Token pruning, by removing less informative tokens, offers an effective strategy to reduce computational overhead. However, existing pruning methods largely rely on static or local token importance scores. This myopic approach fundamentally overlooks the sequential dependency of pruning decisions and fails to capture the interaction effects between pruning decisions across layers, often neglecting the global interactions between mask variables. To address this limitation, we propose V-Pruner, a fast and globally-informed token pruning framework for Vision Transformer. V-Pruner first leverages Fisher information to perform an initial assessment of token importance, providing a principled initial prior for pruning decisions. Building on this, V-Pruner introduces a Reinforcement Learning (RL) Proximal Policy Optimization (PPO) algorithm, refining token pruning into a global sequential decision process. The algorithm combines a composite reward signal that incorporates both model performance and computational cost to guide policy exploration, effectively evaluating the long-term impact of different pruning decision combinations on global model performance. Extensive experiments on ViT-L, DeiT-B, DeiT-S, and DeiT-T demonstrate that V-Pruner achieves a better balance between accuracy, GFLOPs, inference speed, and training time, surpassing existing mainstream ViT pruning algorithms in overall performance. Guangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang 0005, Renda Han, Chuangxin Zhao, Zeyu Zhang 0006 |
AAAI | 7 |
| 2026 | MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-TrainingabstractVision-and-language pretraining (VLP) in medicine leverages contrastive learning on image-text pairs, often enhanced with masked modeling.However, existing methods face two challenges: difficulty reconstructing key pathological features due to limited data, and reliance on either paired or image-only datasets without combining both.To address this, we propose MMCLIP (Masked Medical Contrastive Language-Image Pre-training), which introduces two modules: AttMIM, masking image features highly correlated with text to improve reconstruction of fine medical details, and EntMLM, masking key medical entities in text and reconstructing them using visual cues.Furthermore, MMCLIP incorporates unpaired data through disease-kind prompts, achieving state-of-the-art performance in zero-shot and fine-tuning across five benchmarks.Code Biao Wu 0006, Yutong Xie 0001, Zeyu Zhang 0006, Vu Minh Hieu Phan, Qi Chen 0014, Ling Chen 0006, Qi Wu 0001 |
ACL (1) | 3 |
| 2026 | A Unified Graph Clustering NetworkabstractClustering is a fundamental task in graph data mining, including both node-level and graph-level clustering. While the former has been extensively explored to capture local structures and features, the latter has gained attention for its ability to capture global relationships and high-level abstractions. However, existing methods often address these two tasks in isolation, which not only wastes computational resources but also fails to fully leverage the knowledge from both levels to improve each other, hindering consistent performance improvement. To this end, we propose a novel Unified Graph Clustering Network called UGCN, which employs both local and global graph information to address node- and graph-level clustering collaboratively. In detail, we design a dual-branch projector that performs joint learning at both node and graph levels. The first branch extracts node-level features and projects them into distinct cluster layers, where the derived prototypes are used to refine graph attributes and highlight clustering-friendly substructures. In parallel, the second branch captures subgraph embeddings and aggregates them into discriminative graph-level representations. we align the two branches through joint contrastive objectives to establish a bidirectional interaction: refined prototypes guide subgraph and graph-level clustering, while graph-level pseudo-labels provide feedback to enhance node-level clustering. Extensive experimental results across seven datasets demonstrate that our method significantly outperforms existing state-of-the-art approaches. Renda Han, Xiaobao Wang, Longbiao Wang, Wenxin Zhang 0005, Ronghao Fu, Kaiming Wang, Zeyu Zhang 0006, Kuntharrgyal Khysru |
WWW | 7 |
| 2026 | DOEI: Dual optimization of embedding information for attention-enhanced class activation maps
Zeyu Zhang 0006, Huazhang Wang, Shimin Wen, Daji Ergu, Ying Cai 0002, Yang Zhao 0019 |
Neurocomputing | 2 |
| 2026 | Federated graph-level clustering network with adaptive knowledge compensation
Renda Han, Guangzhen Yao, Wenxin Zhang 0005, Ronghao Fu, Zeyu Zhang 0006 |
Neural Networks | 7 |
| 2026 | Attribute-incomplete graph anomaly detection network
Renda Han, Xiaobao Wang, Guangzhen Yao, Wenxin Zhang 0005, Ronghao Fu, Dayu Hu, Zeyu Zhang 0006, Kaiming Wang |
Pattern Recognit. | 9 |
| 2026 | Exploring dynamic interpretable brain networks via hierarchical graph transformer
Rundong Xue, Shaoyi Du, Xiangmin Han, Jingxi Feng, Zeyu Zhang 0006, Wei Zeng 0003, Yue Gao 0002 |
Pattern Recognit. | 6 |
| 2025 | ProjectedEx: Enhancing Generation in Explainable AI for Prostate CancerabstractProstate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy. Recent advancements in explainable AI and representation learning have significantly improved prostate cancer diagnosis by enabling automated and precise lesion classification. However, existing explainable AI methods, particularly those based on frameworks like generative adversarial networks (GANs), are predominantly developed for natural image generation, and their application to medical imaging often leads to suboptimal performance due to the unique characteristics and complexity of medical image. To address these challenges, our paper introduces three key contributions. First, we propose ProjectedEx, a generative framework that provides interpretable, multi-attribute explanations, effectively linking medical image features to classifier decisions. Second, we enhance the encoder module by incorporating feature pyramids, which enables multiscale feedback to refine the latent space and improves the quality of generated explanations. Additionally, we conduct comprehensive experiments on both the generator and classifier, demonstrating the clinical relevance and effectiveness of ProjectedEx in enhancing interpretability and supporting the adoption of AI in medical settings. Code will be released at https://github.com/Richardqiyi/ProjectedEx. Xuyin Qi, Zeyu Zhang 0006, Aaron Berliano Handoko, Huazhan Zheng, Mingxi Chen, Ta Duc Huy, Vu Minh Hieu Phan, Linqi Cheng, Zhibin Liao, Yang Zhao 0019, Minh-Son To |
CBMS | 2 |
| 2025 | OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadabstractAlthough foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks. Code Luyao Tang, Chaoqi Chen, Zeyu Zhang 0006, Yue Huang 0001, Kun Zhang 0001 |
CVPR | 4 |
| 2025 | PedDet: Adaptive Spectral Optimization for Multimodal Pedestrian DetectionabstractPedestrian detection in intelligent transportation systems has made significant progress but faces two critical challenges: (1) insufficient fusion of complementary information between visible and infrared spectra, particularly in complex scenarios, and (2) sensitivity to illumination changes, such as low-light or overexposed conditions, leading to degraded performance. To address these issues, we propose PedDet, an adaptive spectral optimization complementarity framework which specifically enhanced and optimized for multispectral pedestrian detection. PedDet introduces the Multi-scale Spectral Feature Perception Module (MSFPM) to adaptively fuse visible and infrared features, enhancing robustness and flexibility in feature extraction. Additionally, the Illumination Robustness Feature Decoupling Module (IRFDM) improves detection stability under varying lighting by decoupling pedestrian and background features. We further design a contrastive alignment to enhance intermodal feature discrimination. Experiments on LLVIP and MSDS datasets demonstrate that PedDet achieves state-of-the-art performance, improving the mAP by 6.6 % with superior detection accuracy even in low-light conditions, marking a significant step forward for road safety. Zeyu Zhang 0006, Wenxin Zhang 0005, Zirui Song, Xiuying Chen, Yang Zhao 0019 |
ECAI | 2 |
| 2025 | Efficient Learning with Sine-Activated Low-Rank MatricesabstractLow-rank decomposition has emerged as a vital tool for enhancing parameter efficiency in neural network architectures, gaining traction across diverse applications in machine learning. These techniques significantly lower the number of parameters, striking a balance between compactness and performance. However, a common challenge has been the compromise between parameter efficiency and the accuracy of the model, where reduced parameters often lead to diminished accuracy compared to their full-rank counterparts. In this work, we propose a novel theoretical framework that integrates a sinusoidal function within the low-rank decomposition process. This approach not only preserves the benefits of the parameter efficiency characteristic of low-rank methods but also increases the decomposition's rank, thereby enhancing model performance. Our method proves to be a plug in enhancement for existing low-rank models, as evidenced by its successful application in Vision Transformers (ViT), Large Language Models (LLMs), Neural Radiance Fields (NeRF) and 3D shape modelling. Yiping Ji, Hemanth Saratchandran, Cameron Gordon, Zeyu Zhang 0006, Simon Lucey |
ICLR | 4 |
| 2025 | MSDet: Receptive Field Enhanced Multiscale Detection for Tiny Pulmonary NoduleabstractPulmonary nodules are critical for early lung cancer diagnosis, but traditional CT imaging methods suffer from low detection rates and poor localization. Small nodule detection is challenging due to subtle differences in density and issues like occlusion. Existing methods such as FPN, with its fixed feature fusion and limited receptive field, struggle to effectively overcome these issues. To address these challenges, our paper proposed three key contributions: Firstly, we proposed MSDet, a multiscale attention and receptive field network for detecting tiny pulmonary nodules. Secondly, we proposed the extended receptive domain (ERD) strategy to capture richer contextual information and reduce false positives caused by nodule occlusion. We also proposed the position channel attention mechanism (PCAM) to optimize feature learning and reduce multiscale detection errors, and designed the tiny object detection block (TODB) to enhance the detection of tiny nodules. Experiments on the LUNA16 dataset show an 8.8% improvement in mAP over YOLOv8, achieving state-of-the-art performance. The code is available at https://github.com/CaiGuoHui123/MSDet. Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang 0006, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhibin Liao, Yang Zhao 0019, Ying Cai 0002 |
ICME | 4 |
| 2025 | Multi-Relation Graph-Kernel Strengthen Network for Graph-Level ClusteringabstractGraph-level clustering is a fundamental task of data mining, aiming at dividing unlabeled graphs into distinct groups. However, existing deep methods that are limited by pooling have difficulty extracting diverse and complex graph structure features, while traditional graph kernel methods rely on exhaustive substructure search, unable to adaptively handle multi-relational data. This limitation hampers producing robust and representative graph-level embeddings. To address this issue, we propose a novel Multi-Relation Graph-Kernel Strengthen Network for Graph-Level Clustering (MGSN), which integrates Multi-Relation Modeling (MRM) with graph kernel to fully employ their respective advantages. Specifically, MGSN constructs multi-relation graphs to capture diverse semantic relationships between nodes and graphs, which employ graph kernel methods to extract graph affinity, enriching the representation space. Moreover, a Relation-aware Embedding Strengthening Strategy (RESS) is designed, which adaptively aligns multi-relation information across views while strengthening graph-level features through a progressive fusion process. Extensive experiments on multiple benchmark datasets demonstrate the superiority of MGSN over state-of-the-art methods. The results highlight its ability to leverage multi-relation structures and graph kernel features, establishing a new paradigm for robust graph-level clustering. Renda Han, Guangzhen Yao, Wenxin Zhang 0005, Yu Li 0047, Wen Xin, Huajie Lei, Zeyu Zhang 0006, Chengze Du 0001, Yahe Tian |
IJCNN | 9 |
| 2025 | JTFM: Joint Time-Frequency Method For Long-term Time Series ForecastingabstractLong-term Time Series Forecasting (LTSF) is an important task with extensive applications across diverse domains. While contemporary methodologies have achieved notable results through the integration of time and frequency domain features, significant challenges persist. Current approaches frequently disregard the information degradation inherent in Fast fourier transform (FFT) and inverse Fast fourier transform (IFFT) operations, substantially compromising predictive accuracy. Furthermore, conventional weighting mechanisms demonstrate limitations in their capacity to capture the intricate relationships between temporal and frequency representations, leading to suboptimal feature fusion and consequent information loss. To address these limitations, we present the Joint Time-Frequency Method (JTFM), a novel framework that simultaneously extracts sequence features from both temporal and frequency domains, thereby transcending single-domain constraints and enhancing feature comprehensiveness. Additionally, we introduce the Dynamic Harmonic Accumulation Weighting Mechanism (DHAWM), which surpasses traditional weighting approaches by dynamically modulating the relative contributions of temporal and frequency domain features based on sequence-specific characteristics. This adaptive mechanism strengthens the model’s feature representation capabilities and enhances forecasting precision. Empirical validation on eight real-world datasets demonstrates the JTFM’s superior performance compared to state-of-the-art baseline methods, establishing its efficacy in long-term time series forecasting applications. Yu Li 0047, Wenxin Zhang 0005, Renda Han, Guangzhen Yao, Zeyu Zhang 0006, Cuicui Luo |
IJCNN | 5 |
| 2025 | Applying Large Language Models with Active Learning and Ensemble Learning for Sentiment Recognition in Children's ReadingabstractSentiment recognition in children’s reading plays a crucial role in supporting their mental health and overall development. However, existing datasets often suffer from issues such as limited sample sizes, cross-linguistic ambiguities, and subjective human annotations, hindering the accurate modeling of complex emotional expressions in real-world contexts. These challenges result in suboptimal performance and limited generalization in sentiment recognition models. To address these issues, we leverage the extensive domain knowledge of Large Language Models (LLMs) to improve sentiment recognition performance. First, to mitigate data scarcity, we design a pretraining strategy that utilizes publicly available sentiment recognition datasets in both Chinese and English, enhancing LLMs’ ability to adapt to sentiment recognition tasks. Second, we propose a trustworthy ensemble strategy that integrates the outputs of LLMs across different languages. This strategy employs an uncertainty-aware weighted fusion mechanism to handle semantic ambiguities introduced by linguistic variations. Third, we develop an active learning framework for LLMs that iteratively resolves label inconsistencies, improving data quality and model performance. To validate the proposed methods, we construct a high-quality sentiment recognition dataset for children’s reading and incorporate the publicly available 40-Thai-Children dataset. Extensive experiments show that our method achieves an F1 score of 96.71% and accuracy of 97.42% on the collected dataset, and 86.31% F1 score and 87.02% accuracy on the 40-Thai-Children dataset. Furthermore, the active learning framework demonstrates its potential for generalization, improving accuracy from 82.05% to 86.32% in cross-domain testing by reducing data bias. The collected dataset and code are available at https://anonymous.4open.science/r/Children-s-Reading-Sentiment-Recognition-Dataset-745D. Hongtao Mao, Jincai Yang, Xusheng Yang, Zhanghao Qin, Zeyu Zhang 0006 |
IJCNN | 5 |
| 2025 | RL-Pruner: Retraining-Free Global Exploration Pruning Method Based on Reinforcement LearningabstractLarge language models (LLMs) have achieved significant success in complex tasks across various domains, but these achievements come with high computational costs and long inference delays. Pruning, as an effective optimization technique, simplifies model structures by removing redundant components, thereby improving model generalization and operational efficiency. Although existing pruning retraining-free algorithms perform excellently in pruning time, these algorithms often focus on local optimal solutions in encoder-based language models, lacking comprehensive exploration of global optimal solutions, which may affect the overall model performance. To address this issue, we propose a novel retraining-free structured pruning algorithm, named RL-Pruner. The algorithm consists of two main stages: the Mask Rearrangement Based on Asynchronous Advantage Actor-Critic (MA3C) stage and the BiConjugate Gradient Solver for Mask Tuning (BGMT) stage. It aims to explore the intra-layer interactions of mask variables and efficiently find the global optimal solution without requiring retraining. We evaluate this method using BERTBASEand DistilBERT models on the GLUE and SQuAD benchmark tests. Experimental results show that RL-Pruner significantly improves accuracy on the SQuAD1.1benchmark. Under a 60% FLOPs constraint, compared with existing pruning retraining-free algorithms, the F1 score increases by 4.25%. Guangzhen Yao, Wenxin Zhang 0005, Xaioyu Deng, Chengze Du 0001, Renda Han, Zhanghao Qin, Yu Li 0047, Bobin Xie, Haiming Peng, Sandong Zhu, Zezhou Wang, Zeyu Zhang 0006 |
IJCNN | 15 |
| 2025 | Dual-channel Heterophilic Message Passing for Graph Fraud DetectionabstractFraudulent activities have significantly increased across various domains, such as e-commerce, online review platforms, and social networks, making fraud detection a critical task. Spatial Graph Neural Networks (GNNs) have been successfully applied to fraud detection tasks due to their strong inductive learning capabilities. However, existing spatial GNN-based methods often enhance the graph structure by excluding heterophilic neighbors during message passing to align with the homophilic bias of GNNs. Unfortunately, this approach can disrupt the original graph topology and increase uncertainty in predictions. To address these limitations, this paper proposes a novel framework, Dual-channel Heterophilic Message Passing (DHMP), for fraud detection. DHMP leverages a heterophily separation module to divide the graph into homophilic and heterophilic subgraphs, mitigating the low-pass inductive bias of traditional GNNs. It then applies shared weights to capture signals at different frequencies independently and incorporates a customized sampling strategy for training. This allows nodes to adaptively balance the contributions of various signals based on their labels. Extensive experiments on three real-world datasets demonstrate that DHMP outperforms existing methods, highlighting the importance of separating signals with different frequencies for improved fraud detection. The code is available at https://github.com/shaieesss/DHMP. Wenxin Zhang 0005, Jingxing Zhong, Guangzhen Yao, Renda Han, Xiaojian Lin, Zeyu Zhang 0006, Cuicui Luo |
IJCNN | 7 |
| 2025 | Adaptive Embedding for Long-Range High-Order Dependencies via Time-Varying Transformer on fMRI
Rundong Xue, Xiangmin Han, Zeyu Zhang 0006, Shaoyi Du, Yue Gao 0002 |
MICCAI (12) | 4 |
| 2025 | DHGFormer: Dynamic Hierarchical Graph Transformer for Disorder Brain Disease Diagnosis
Rundong Xue, Zeyu Zhang 0006, Xiangmin Han, Yue Gao 0002, Shaoyi Du |
MICCAI (12) | 3 |
| 2025 | Unified Medical Image Segmentation with State Space Modeling SnakeabstractUnified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-based approaches, lacking object-level anatomical insight and inter-organ relational modeling, struggle with morphological complexity and feature conflicts, limiting their efficacy in UMIS. We propose Mamba Snake, a novel deep snake framework enhanced by state space modeling for UMIS. Mamba Snake frames multi-contour evolution as a hierarchical state space atlas, effectively modeling macroscopic inter-organ topological relationships and microscopic contour refinements. We introduce a snake-specific vision state space module, the Mamba Evolution Block (MEB), which leverages effective spatiotemporal information aggregation for adaptive refinement of complex morphologies. Energy map shape priors further ensures robust long-range contour evolution in heterogeneous data. Additionally, a dual-classification synergy mechanism is incorporated to concurrently optimize detection and segmentation, mitigating under-segmentation of microstructures in UMIS. Extensive evaluations across five clinical datasets reveal Mamba Snake's superior performance. Ruicheng Zhang, Haowei Guo, Kanghui Tian, Mingliang Yan, Zeyu Zhang 0006 |
ACM Multimedia | 6 |
| 2025 | MARL-MambaContour: Unleashing Multi-Agent Deep Reinforcement Learning for Active Contour Optimization in Medical Image SegmentationabstractWe introduce MARL-MambaContour, the first contour-based medical image segmentation framework based on Multi-Agent Reinforcement Learning (MARL). Our approach reframes segmentation as a multi-agent cooperation task focused on generating topologically consistent object-level contours, addressing the limitations of traditional pixel-based methods which could lack topological constraints and holistic structural awareness of anatomical regions. Each contour point is modeled as an autonomous agent that iteratively adjusts its position to align precisely with the target boundary, enabling adaptation to blurred edges and intricate morphologies common in medical images. This iterative adjustment process is optimized by a contour-specific Soft Actor-Critic (SAC) algorithm, further enhanced with the Entropy Regularization Adjustment Mechanism (ERAM) which dynamically balances agent exploration with contour smoothness. Furthermore, the framework incorporates a Mamba-based policy network featuring a novel Bidirectional Cross-attention Hidden-state Fusion Mechanism (BCHFM). This mechanism mitigates potential memory confusion limitations associated with long-range modeling in state space models, thereby facilitating more accurate inter-agent information exchange and informed decision-making. Extensive experiments on five diverse medical imaging datasets demonstrate the state-of-the-art performance of MARL-MambaContour, highlighting its potential as an accurate and robust clinical application. Ruicheng Zhang, Zeyu Zhang 0006, Jinai Li, Hoi Fan Au, Haowei Guo, Puxin Yan |
ACM Multimedia | 3 |
| 2025 | You Can Generate It Again: Data-to-Text Generation with Verification and Correction PromptingabstractSmall language models like T5 excel in generating high-quality text for data-to-text tasks, offering adaptability and cost-efficiency compared to Large Language Models (LLMs). However, they frequently miss keywords, which is considered one of the most severe and common errors in this task. In this work, we explore the potential of using feedback systems to enhance semantic fidelity in smaller language models for data-to-text generation tasks, through our Verification and Correction Prompting (VCP) approach. In the inference stage, our approach involves a multi-step process, including generation, verification, and regeneration stages. During the verification stage, we implement a simple rule to check for the presence of every keyword in the prediction. Recognizing that this rule can be inaccurate, we have developed a carefully designed training procedure, which enabling the model to incorporate feedback from the error-correcting prompt effectively, despite its potential inaccuracies. The VCP approach effectively reduces the Semantic Error Rate (SER) while maintaining the text’s quality. Xuan Ren, Zeyu Zhang 0006, Lingqiao Liu |
MMAsia | 2 |
| 2025 | Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve AnomaliesabstractZirui Song, Guangxian Ouyang, Meng Fang, Hongbin Na, Zijing Shi, Zhenhao Chen, Fu Yujie, Zeyu Zhang, Shiyu Jiang, Miao Fang, Ling Chen, Xiuying Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zirui Song, Guangxian Ouyang, Hongbin Na, Zijing Shi, Zhenhao Chen, Yujie Fu, Zeyu Zhang 0006, Miao Fang 0001, Ling Chen 0006, Xiuying Chen |
NAACL (Long Papers) | 8 |
| 2025 | FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video DiffusionabstractDiffusion generative models have become the standard for producing high-quality, coherent video content, yet their slow inference speeds and high computational demands hinder practical deployment. Although both quantization and sparsity can independently accelerate inference while maintaining generation quality, naively combining these techniques in existing training-free approaches leads to significant performance degradation, as they fail to achieve proper joint optimization.
We introduce FPSAttention, a novel training-aware co-design of FP8 quantization and Sparsity for video generation, with a focus on the 3D bi-directional attention mechanism. Our approach features three key innovations: 1) A unified 3D tile-wise granularity that simultaneously supports both quantization and sparsity. 2) A denoising step-aware strategy that adapts to the noise schedule, addressing the strong correlation between quantization/sparsity errors and denoising steps. 3) A native, hardware-friendly kernel that leverages FlashAttention and is implemented with optimized Hopper architecture features, enabling highly efficient execution.
Trained on Wan2.1's 1.3B and 14B models and evaluated on the vBench benchmark, FPSAttention achieves a 7.09$\times$ kernel speedup for attention operations and a 4.96$\times$ end-to-end speedup for video generation compared to the BF16 baseline at 720p resolution—without sacrificing generation quality. Akide Liu, Zeyu Zhang 0006, Zhexin Li, Xuehai Bai, Yuanjie Xing, Yizeng Han, Jiasheng Tang, Jichao Wu, Mingyang Yang, Yuanyu He, Fan Wang 0019, Gholamreza Haffari, Bohan Zhuang |
NeurIPS | 2 |
| 2025 | ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGSabstractFeed-forward 3D Gaussian Splatting (3DGS) models have recently emerged as a promising solution for novel view synthesis, enabling one-pass inference without the need for per-scene 3DGS optimization. However, their scalability is fundamentally constrained by the limited capacity of their encoders, leading to degraded performance or excessive memory consumption as the number of input views increases. In this work, we analyze feed-forward 3DGS frameworks through the lens of the Information Bottleneck principle and introduce ZPressor, a lightweight architecture-agnostic module that enables efficient compression of multi-view inputs into a compact latent state $Z$ that retains essential scene information while discarding redundancy. Concretely, ZPressor enables existing feed-forward 3DGS models to scale to over 100 input views at 480P resolution on an 80GB GPU, by partitioning the views into anchor and support sets and using cross attention to compress the information from the support views into anchor views, forming the compressed latent state $Z$. We show that integrating ZPressor into several state-of-the-art feed-forward 3DGS models consistently improves performance under moderate input views and enhances robustness under dense view settings on two large-scale benchmarks DL3DV-10K and RealEstate10K. Weijie Wang 0014, Donny Y. Chen, Zeyu Zhang 0006, Duochao Shi, Akide Liu, Bohan Zhuang |
NeurIPS | 3 |
| 2025 | FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationabstractDiffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net inductive bias. To address these challenges, we propose **FlashMo**, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design. We further introduce *MotionSiT*, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over $\mathrm{SO}(3)$ manifolds, enabling principled generation of joint rotations. Extensive experiments on the large-scale MotionHub V2 dataset and standard benchmarks including HumanML3D and KIT-ML demonstrate that our method significantly outperforms previous approaches in motion quality, efficiency, and scalability. Compared to the state-of-the-art 1-step distillation baseline, FlashMo reduces **12.9%** inference time and FID by **34.1%**. Project website: https://steve-zeyu-zhang.github.io/FlashMo. Zeyu Zhang 0006, Danning Li, Dong Gong, Ian D. Reid 0001, Richard I. Hartley |
NeurIPS | 1 |
| 2025 | Medical artificial intelligence for early detection of lung cancer: A survey
Guohui Cai, Ying Cai 0002, Zeyu Zhang 0006, Yuanzhouhan Cao, Daji Ergu, Zhibin Liao, Yang Zhao 0019 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | MedDet: Generative Adversarial Distillation for Efficient Cervical Disc Herniation DetectionabstractCervical disc herniation (CDH) is a prevalent musculoskeletal disorder that significantly impacts health and requires labor-intensive analysis from experts. Despite advancements in automated detection of medical imaging, two significant challenges hinder the real-world application of these methods. First, the computational complexity and resource demands present a significant gap for real-time application. Second, noise in MRI reduces the effectiveness of existing methods by distorting feature extraction. To address these challenges, we propose three key contributions: Firstly, we introduced MedDet, which leverages the multi-teacher single-student knowledge distillation for model compression and efficiency, meanwhile integrating generative adversarial training to enhance performance. Additionally, we customize the second-order nmODE to improve the model’s resistance to noise in MRI. Lastly, we conducted comprehensive experiments on the CDH-1848 dataset, achieving up to a 5% improvement in mAP compared to previous methods. Our approach also delivers over 5 times faster inference speed, with approximately 67.8% reduction in parameters and 36.9% reduction in FLOPs compared to the teacher model. These advancements significantly enhance the performance and efficiency of automated CDH detection, demonstrating promising potential for future application in clinical practice. Zeyu Zhang 0006, Nengmin Yi, Shengbo Tan, Ying Cai 0002, Yi Yang 0001, Lei Xu 0001, Qingtai Li, Daji Ergu, Yang Zhao 0019 |
BIBM | 1 |
| 2024 | Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
Zeyu Zhang 0006, Biao Wu 0006, Shiya Huang, Wenbo Zhang 0009, Ling Chen 0006, Yang Zhao 0019 |
BMVC | 1 |
| 2024 | Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang 0006, Akide Liu, Ian D. Reid 0001, Richard I. Hartley, Bohan Zhuang, Hao Tang 0005 |
ECCV (1) | 1 |