EDBT 2026 Demo / reviewers in the wild / expert
Ziyu Ma
dblp:152/6687
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 21% Video understanding and tracking · 19% Vision and language · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution · ACL (1) 2026 |
Machine learning › Reinforcement learning
LLM agent training |
1.0 | 1 | 2026 | CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution · ACL (1) 2026 |
Computer vision › Vision and language
multimodal in-context learning |
1.0 | 1 | 2026 | Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning · AAAI 2026 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
1.0 | 1 | 2026 | Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
task vector |
1.0 | 1 | 2026 | Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning · AAAI 2026 |
Computer vision › Video understanding and tracking
long video understanding |
0.9 | 1 | 2025 | DrVideo: Document Retrieval Based Long Video Understanding · CVPR 2025 |
Information retrieval
document retrieval |
0.9 | 1 | 2025 | DrVideo: Document Retrieval Based Long Video Understanding · CVPR 2025 |
Computer vision › Face, body and person analysis › gaze analysis
eye contact detection |
0.6 | 1 | 2022 | TA-CNN: A Unified Network for Human Behavior Analysis in Multi-Person Conversations · ACM Multimedia 2022 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis
human behavior analysis |
0.6 | 1 | 2022 | TA-CNN: A Unified Network for Human Behavior Analysis in Multi-Person Conversations · ACM Multimedia 2022 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.7chain-of-thought · 1.7agent-based retrieval · 1.7uncertainty estimation · 1.0task synthesis · 1.0sensitivity-aware insertion · 1.0reinforcement learning · 1.0forgetting signal · 1.0activation clustering · 1.0convolutional neural network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context LearningabstractLarge Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and high inference cost. To address these challenges, task-vector-based methods have been explored by inserting compact representations of many-shot in-context demonstrations into model activations. However, existing task-vector-based methods either overlook the importance of where to insert task vectors or struggle to determine suitable values for each location. To this end, we propose a novel Sensitivity-aware Task Vector insertion framework (STV) to figure out where and what to insert. Our key insight is that activation deltas across query-context pairs exhibit consistent structural patterns, providing a reliable cue for insertion. Based on the identified sensitive-aware locations, we construct a pre-clustered activation bank for each location by clustering the activation values, and then apply reinforcement learning to choose the most suitable one to insert. We evaluate STV across a range of multimodal models (e.g., Qwen-VL, Idefics-2) and tasks (e.g., VizWiz, OK-VQA), demonstrating its effectiveness and showing consistent improvements over previous task-vector-based methods with strong generalization. Ziyu Ma, Chenhui Gou, Yiming Hu, Bohan Zhuang, Jianfei Cai 0001 |
AAAI | 1 |
| 2026 | CoEvolve: Training LLM Agents via Agent-Data Mutual EvolutionabstractReinforcement learning for LLM agents is typically conducted on a static data distribution, which fails to adapt to the agent's evolving behavior and leads to poor coverage of complex environment interactions.To address these challenges, we propose CoEvolve, an agentdata mutual evolution framework that enables LLM agents to improve through closed-loop, interaction-driven training.Specifically, CoEvolve extracts feedback signals such as forgetting and uncertainty from rollout trajectories to identify failure-prone interaction patterns, and utilizes them to guide LLM-based task synthesis.The synthesized tasks are validated through environment interaction and utilized to update the data distribution, enabling joint adaptation of the agent and its data.Extensive experiments on AppWorld and BFCL across Qwen2.5-7B,Qwen3-4B, and Qwen3-30B-A3B demonstrate consistent and significant improvements over strong base models, yielding absolute gains of 19.43%, 15.58%, and 18.14%, respectively. Shidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu, Xiangxiang Chu |
ACL (1) | 2 |
| 2025 | DrVideo: Document Retrieval Based Long Video UnderstandingabstractMost of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing long-range reasoning. Thus, we propose DrVideo, a document-retrieval-based system designed for long video understanding. Our key idea is to convert the long-video understanding problem into a long-document understanding task so as to effectively leverage the power of large language models. Specifically, DrVideo first transforms a long video into a coarse text-based long document to initially retrieve key frames and then updates the documents with the augmented key frame information. It then employs an agent-based iterative loop to continuously search for missing information and augment the document until sufficient question-related information is gathered for making the final predictions in a chain-of-thought manner. Extensive experiments on long video benchmarks confirm the effectiveness of our method. DrVideo significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema benchmark (3 minutes), MovieChat-1K benchmark (10 minutes), and the long split of Video-MME benchmark (average of 44 minutes). Code is available at https://github.com/Upper9527/DrVideo. Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 0001, Shutao Li 0001, Seyed Hamid Rezatofighi, Jianfei Cai 0001 |
CVPR | 1 |
| 2025 | A Pix2pix-Based Generative Channel Estimator for CF-MIMO SystemsabstractIn cell-free massive multiple-input multiple-output (CF-MIMO) systems, pilot compression improves spectral efficiency but often destroys the spatial structural features of the channel matrix, significantly degrading the performance of traditional estimators. To address this issue, we propose a generative adversarial network (GAN)-based method named least squares based Pix2pix (LS-P2P), which formulates the channel estimation process as a conditional image-to-image generation task. The proposed approach takes the coarse least squares (LS) estimation under compressed pilots as input and trains a Pix2pix generator with a compound loss consisting of L1reconstruction loss and a structure-aware structural similarity index measure (SSIM) loss to recover the missing channel structure. Experimental results demonstrate that under low pilot and low signal-to-noise ratio (SNR) conditions, the proposed method reduces the normalized mean squared error (NMSE) by approximately 58.9% compared with the traditional LS method. Moreover, compared to the two-stage GAN-based estimator (TwoStage-G) architecture, the proposed framework exhibits superior performance, highlighting the structural modeling and recovery capability of generative adversarial networks under sparse pilot constraints. Ziyu Ma, Chunyu Pan, Zhonghao Luo |
VTC2025-Fall | 1 |
| 2025 | Cascade Fusion and Correlation Enhancement for Knowledge DistillationabstractKnowledge distillation (KD) improves the performance of a compact student network by transferring learned knowledge from a cumbersome teacher network. In the existing approaches, the multiscale feature knowledge is transferred via densely connected paths, which increases the optimization difficulty. Moreover, correlations among the labels are neglected despite their capability to enhance the intraclass similarity of samples. To solve these issues, we propose cascade fusion and correlation enhancement for KD (CC-KD). The multiscale feature knowledge is transferred via much simpler paths, which are constructed by fusing features of different scales with cross-scale attention (CSA) in a cascade manner, thereby reducing the optimization difficulty. On the other hand, the relational knowledge of teacher logits is further enhanced by correlations of the corresponding labels, so that the student can produce more similar logits for the samples in the same category. Extensive experimental results on five public datasets (i.e., CIFAR100/10, ImageNet, RAF-DB, and FERPlus) indicate superior performance of the proposed method over several state-of-the-arts (SOTAs). More specifically, our method obtains an accuracy of 71.70% on ImageNet and achieves a new record of 90.20% on RAF-DB with fewer calculations and parameters. Bin Sun 0001, Zuxiang Long, Ziyu Ma, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Two-Stage Selective Fusion Framework for Joint Intent Detection and Slot FillingabstractSpoken language understanding (SLU) is the core of the speech-centric human-robot interaction system, which mainly involves intent detection and slot filling. The recent SLU research focuses on the joint modeling of the two tasks due to their correlation. Furthermore, the slot information consists of slot position and slot type. Although the slot types are semantically related to the intent, the slot positions of the same intent may vary a lot in different utterances due to the diversity of spoken language. Thus, the conventional one-stage slot filling task may introduce unrelated information for slot position prediction in the slot-intent interaction of the joint modeling. Therefore, we propose a novel two-stage selective fusion framework for joint intent detection and slot filling. Unlike the previous one-stage framework, the proposed framework decomposes the slot filling into two stages, i.e., the slot proposal and slot classification. The slot proposal network consisting of BERT and bidirectional long short-term memory (Bi-LSTM)-conditional random field (CRF) predicts the slot positions. Instead of the tokenwise fusion in the existing methods, the slot-intent feature fusion is only performed in the slot classification. A selective fusion mechanism is designed to facilitate the slot-intent interaction within each slot candidate for more accurate slot-type classification. Experiments on five standard benchmarks (i.e., ATIS, SNIPS, MixATIS, MixSNIPS, and DSTC4) show that the proposed framework achieves the best performance in comparison with several state-of-the-art methods. Ziyu Ma, Bin Sun 0001, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | TA-CNN: A Unified Network for Human Behavior Analysis in Multi-Person ConversationsabstractHuman behavior analysis in multi-person conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on single-person behavior analysis, therefore, can hardly be generalized in real-world application scenarios. Fortunately, the MultiMediate'22 Challenge provides various video clips of multi-party conversations. In this paper, we present a unified network named TA-CNN for both sub-challenges. Our TA-CNN can not only model the spatio-temporal dependencies for eye contact detection, but also capture the group-level discriminative features for multi-label next speaker prediction. We empirically evaluate the performance of our method on the officially provided datasets. Our method achieves the state-of-the-art result of 0.7261 for eye contact detection in terms of accuracy and the UAR of 0.5965 for next speaker prediction on the corresponding test sets. Fuyan Ma, Ziyu Ma, Bin Sun 0001, Shutao Li 0001 |
ACM Multimedia | 2 |
| 2022 | Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation
Bin Li 0083, Yixuan Weng, Ziyu Ma, Bin Sun 0001, Shutao Li 0001 |
NLPCC (2) | 3 |