EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Ma 0002
dblp:170/8990-2
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-1846-8889ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DWTSG: Parameter-Efficient Fine-Tuning of Large Pre-trained Models via Discrete Wavelet Transform and Subband GuidanceabstractFully fine-tuning large pre-trained models for each downstream task is impractical due to prohibitive memory, computation, and storage costs. Although parameter-efficient fine-tuning (PEFT) methods address this issue, leading methods like LoRA still exhibit linear scaling of trainable parameters with hidden size. Recent studies have explored PEFT in the frequency domain to reduce computational costs by employing fast Fourier transform and discrete cosine transform with sparse frequency selection. These methods rely on global frequency representations that lack spatial locality and disperse energy across the domain. As a result, sparse coefficient selection struggles to preserve fine-grained structural information and often introduces artifacts such as ringing near boundaries. To address these limitations, we propose DWTSG, a novel PEFT framework based on discrete wavelet transform (DWT) and subband guidance. DWTSG decomposes pre-trained weights into four wavelet subbands that jointly encode global context and local details. It fine-tunes only the most informative coefficients in each subband through an energy-based selection strategy that prioritizes coefficients based on their individual importance and interactions. Finally, inverse DWT reconstructs the updated weights, enabling efficient and precise adaptation. Extensive experiments on natural language understanding, commonsense reasoning, and image classification demonstrate that DWTSG outperforms existing PEFT methods, achieving superior performance and higher parameter efficiency. Chengwei Sun, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Ran Ran 0001, Jie Zou 0001, Yang Yang 0002 |
AAAI | 4 |
| 2026 | CARD: Non-Uniform Quantization of Visual Semantic Unit for Generative Recommendation
Yibiao Wei, Jie Zou 0001, Xiao Ao, Weikang Guo, Zeyu Ma 0002, Yang Yang 0002 |
SIGIR | 6 |
| 2026 | ORCA: Object Recognition and Comprehension for Archiving Marine SpeciesabstractMarine visual understanding is essential for monitoring and protecting marine ecosystems, enabling automatic and scalable biological surveys. However, progress is hindered by limited training data and the lack of a systematic task formulation that aligns domain-specific marine challenges with well-defined computer vision tasks, thereby limiting effective model application. To address this gap, we present ORCA, a multi-modal benchmark for marine research comprising 14,647 images from 478 species, with 42,217 bounding box annotations and 22,321 expert-verified instance captions. The dataset provides fine-grained visual and textual annotations that capture morphology-oriented attributes across diverse marine species. To catalyze methodological advances, we evaluate 18 state-of-the-art models on three tasks: object detection (closed-set and open-vocabulary), instance captioning, and visual grounding. Results highlight key challenges, including species diversity, morphological overlap, and specialized domain demands, underscoring the difficulty of marine understanding. ORCA thus establishes a comprehensive benchmark to advance research in marine domain. Yuk-Kwan Wong, Haixin Liang, Zeyu Ma 0002, Yiwei Chen 0003, Ziqiang Zheng, Rinaldi Gotama, Pascal Sebastian, Lauren D. Sparks, Sai-Kit Yeung |
WACV | 3 |
| 2026 | GASE: Generalized adaptive static enhancement for temporal sentence grounding
Ran Ran 0001, Kaiwen Shen, Jiwei Wei, Ruikun Chai, Shiyuan He, Zeyu Ma 0002, Malu Zhang, Yang Yang 0002 |
Knowl. Based Syst. | 6 |
| 2025 | KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal Grounding
Ran Ran 0001, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
ICCV | 4 |
| 2025 | SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeaturesabstractGenerating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatched mouth movements due to unstable Gaussian initialization and incomplete speech features. To overcome these limitations, we introduce SyncGaussian, a 3DGS-based framework that ensures stable head poses, enhanced lip sync, and realistic appearances with real-time rendering. SyncGaussian employs a stable head Gaussian initialization strategy to mitigate head jitter by optimizing commonly used rough head pose parameters. To enhance lip sync, we propose a sync-enhanced encoder that leverages audio-to-text and audio-to-visual speech features. Guided by a tailored cosine similarity loss function, the encoder integrates discriminative speech features through a multi-level sync adaptation mechanism, enabling the learning of an adaptive speech feature space. Extensive experiments demonstrate that SyncGaussian outperforms state-of-the-art methods in image quality, dynamic motion, and lip sync, with the potential for real-time applications. Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
IJCAI | 4 |
| 2025 | Bipolar Self-attention for Spiking TransformersabstractHarnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores.
To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation.
Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers. Shuai Wang 0058, Malu Zhang, Dehao Zhang, Yimeng Shan, Jieyuan Zhang, Yichen Xiao, Honglin Cao, Zeyu Ma 0002, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 10 |
| 2025 | Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingabstractThe explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. This mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms. Dehao Zhang, Malu Zhang, Shuai Wang 0058, Wenjie Wei, Zeyu Ma 0002, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 6 |
| 2025 | SGCDiff: Sketch-Guided Cross-modal Diffusion Model for 3D shape completion
Zhenjiang Du, Zhitao Liu, Zeyu Ma 0002, Ning Xie 0003, Yang Yang 0002 |
Neurocomputing | 5 |
| 2025 | Text-guided dynamic mouth motion capturing for person-generic talking face generation
Jiwei Wei, Ruiqi Yuan, Ruikun Chai, Shiyuan He, Zeyu Ma 0002, Yang Yang 0002 |
Knowl. Based Syst. | 6 |
| 2024 | Instance-Dictionary Learning for Open-World Object Detection in Autonomous Driving ScenariosabstractThis paper addresses an important and valuable open-world object detection (OWOD) in autonomous driving scenarios, which aims to detect objects under bothdomain-agnosticandcategory-agnosticsettings simultaneously. Existing OWOD algorithms mainly focus on the detection of pre-defined object categories under various conditions (domain-agnostic) or instead perform zero-shot object detection (category-agnostic), separately. The knowledge gap between seen and unseen object categories poses challenges for models optimized with supervision from the only seen object categories. The domain difference across different scenarios also causes further challenges in aligning observations with different appearances. To address these two challenges simultaneously, we propose our Instance Dictionary Learning (IDL for short) for more robust and accurate OWOD performance. We first design a pre-training procedure to build up the mappings between region features and category semantic embeddings by introducing instance contrastive learning. The joint vision-semantic space is formulated through the more detailed instance-level “Dictionary”, which expresses the region-category correspondences and helps link the seen and unseen object categories. The domain discrimination is further designed for extracting the domain invariance feature representations in the further training procedure seamlessly. The proposed IDL could detect the unseen categories from unseen domains without any bounding box annotations while there is no obvious performance drop on detecting seen categories meanwhile. Comprehensive experiments have been conducted and our method could achieve a new state-of-the-art OWOD performance over previous algorithms. Zeyu Ma 0002, Ziqiang Zheng, Jiwei Wei, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Set of Diverse Queries With Uncertainty Regularization for Composed Image RetrievalabstractComposed image retrieval aims to search a target image by concurrently understanding the composed inputs with a reference image and the complementary modification text. It aims to find a shared latent space where the representation of the composed inputs is close to the desired target image. Most previous methods capture the one-to-one correspondence between the composed inputs and target image, which encodes the composed inputs and the target image into single points in the feature space. However, the one-to-one correspondence cannot effectively handle this task due to the inherent ambiguity problem arising from the various semantic meanings and data uncertainty. Specifically, the composed inputs and target image always exhibit various semantic meanings, affecting the retrieval results. Moreover, given the composed inputs (resp. target image), there are multiple target images (resp. composed inputs) that equally make sense. In this paper, we propose a novel method termed Set of Diverse Queries with Uncertainty Regularization (SDQUR) to solve such inherent ambiguity problem. First, we utilize diverse queries to adaptively aggregate the composed inputs and target image into multiple deterministic embeddings that capture different semantic meanings in the triplet affecting the retrieval process. It can exploit the deterministic many-to-many correspondence within each triple through these set-based queries. Moreover, we provide an uncertainty regularization module to encode the composed inputs and target image into gaussian distribution. Multiple potential positive candidates are sampled from the distribution for probabilistic many-to-many correspondence. Through the complementary deterministic and probabilistic many-to-many correspondence manner, we achieve consistent improvements on the standard FashionIQ, CIRR, and Shoes benchmarks, surpassing the state-of-the-art methods by a large margin. Yahui Xu, Jiwei Wei, Yi Bin, Yang Yang 0002, Zeyu Ma 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Open-Scenario Domain Adaptive Object Detection in Autonomous DrivingabstractExisting domain adaptive object detection algorithms (DAOD) have demonstrated their effectiveness in discriminating and localizing objects across scenarios. However, these algorithms typically assume a single source and target domain for adaptation, which is not representative of the more complex data distributions in practice. To address this issue, we propose a novel Open-Scenario Domain Adaptive Object Detection (OSDA), which leverages multiple source and target domains for more practical and effective domain adaptation. We are the first to increase the granularity of the background category by building the foundation model using contrastive vision-language pre-training in an open-scenario setting for better distinguishing foreground and background, which is under-explored in previous studies. The performance gains by introducing the pre-training have been observed and have validated the model's ability to detect objects across domains. To further fine-tune the model for domain-specific object detection, we propose a hierarchical feature alignment strategy to obtain a better common feature space among the various source and target domains. In the case of multi-source domains, the cross-reconstruction framework is introduced for learning more domain invariances. The proposed method is able to alleviate knowledge forgetting without any additional computational costs. Extensive experiments across different scenarios demonstrate the effectiveness of the proposed model. Zeyu Ma 0002, Ziqiang Zheng, Jiwei Wei, Xiaoyong Wei, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 1 |
| 2022 | Rethinking Open-World Object Detection in Autonomous Driving ScenariosabstractExisting object detection models have been demonstrated to successfully discriminate and localize the predefined object categories under the seen or similar situations. However, the open-world object detection as required by autonomous driving perception systems refers to recognizing unseen objects under various scenarios. On the one hand, the knowledge gap between seen and unseen object categories poses extreme challenges for models trained with supervision only from the seen object categories. On the other hand, the domain differences across different scenarios also cause an additional urge to take the domain gap into consideration by aligning the sample or label distribution. Aimed at resolving these two challenges simultaneously, we firstly design a pre-training model to formulate the mappings between visual images and semantic embeddings from the extra annotations as guidance to link the seen and unseen object categories through a self-supervised manner. Within this formulation, the domain adaptation is then utilized for extracting the domain-agnostic feature representations and alleviating the misdetection of unseen objects caused by the domain appearance changes. As a result, the more realistic and practical open-world object detection problem is visited and resolved by our novel formulation, which could detect the unseen categories from unseen domains without any bounding box annotations while there is no obvious performance drop in detecting the seen categories. We are the first to formulate a unified model for open-world task and establish a new state-of-the-art performance for this challenge. Zeyu Ma 0002, Yang Yang 0002, Guoqing Wang 0001, Xing Xu 0001, Heng Tao Shen |
ACM Multimedia | 1 |
| 2021 | Disentangled Representation Learning and Enhancement Network for Single Image De-RainingabstractIn this paper, we present a disentangled representation learning and enhancement network (DRLE-Net) to address the challenging single image de-raining problems, i.e., raindrop and rain streak removal. Specifically, the DRLE-Net is formulated as a multi-task learning framework, and an elegant knowledge transfer strategy is designed to train the encoder of DRLE-Net to embed a rainy image into two separated latent spaces representing the task (clean image reconstruction in this paper) relevant and irrelevant variations respectively, such that only the essential task-relevant factors will be used by the decoder of DRLE-Net to generate high-quality de-raining results. Furthermore, visual attention information is modeled and fed into the disentangled representation learning network to enhance the task-relevant factor learning. To facilitate the optimization of the hierarchical network, a new adversarial loss formulation is proposed and used together with the reconstruction loss to train the proposed DRLE-Net. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and DRLE-Net is demonstrated to produce significantly better results than state-of-the-art models. Guoqing Wang 0001, Changming Sun, Xing Xu 0001, Jingjing Li 0001, Zheng Wang 0044, Zeyu Ma 0002 |
ACM Multimedia | 6 |