EDBT 2026 Demo / reviewers in the wild / expert
Zhangling Duan
dblp:252/7815
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-3246-8022ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationabstractThe Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting cross-modal salient anchors, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a Mutual Event Agreement Evaluation module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a Cross-modal Salient Anchor Identification module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an Anchor-based Temporal Propagation module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance. Jinxing Zhou, Yanghao Zhou, Yuxin Mao, Zhangling Duan, Dan Guo 0001 |
AAAI | 5 |
| 2026 | Modeling Stage-wise Evolution of User Interests for News RecommendationabstractPersonalized news recommendation is highly time-sensitive, as user interests are often driven by emerging events, trending topics, and shifting real-world contexts. These dynamics make it essential to model not only users' long-term preferences, which reflect stable reading habits and high-order collaborative patterns, but also their short-term, context-dependent interests that change rapidly over time. However, most existing approaches rely on a single static interaction graph, which struggles to capture both long-term preference patterns and short-term interest changes as user behavior evolves. To address this challenge, we propose a unified framework that learns user preferences from both global and local temporal perspectives. A global preference modeling component captures long-term collaborative signals from the overall interaction graph, while a local preference modeling component partitions historical interactions into stage-wise temporal subgraphs to represent short-term dynamics. Within this module, an LSTM branch models the progressive evolution of recent interests, and a self-attention branch captures long-range temporal dependencies. Extensive experiments on two large-scale real-world datasets show that our approach consistently outperforms strong baselines and delivers fresher and more relevant recommendations across diverse user behaviors and temporal settings. Zhiyong Cheng 0001, Yike Jin, Huilin Chen 0002, Zhangling Duan, Meng Wang 0001 |
WWW | 5 |
| 2026 | From Social Media to Psychological Scale: An Adaptive Framework with Two-Hop Retrieval for Depression ScreeningabstractDepressive disorders represent a major global public health challenge. As an increasing number of individuals share their emotional experiences and concerns on social media, researchers have shown growing interest in leveraging such data for early depression screening. However, most existing methods rely on a fixed model and a singular reasoning paradigm, which constrains their adaptability to depression detection. The limited availability of mental health-related data and variability in training data distributions across different LLMs hinder their consistent and comprehensive understanding of diverse psychological symptoms. In this paper, we propose AdaDepression, a framework that enables explainable depression screening through a two-hop retrieval algorithm to identify symptom-relevant posts and a two-stage adaptive routing mechanism for selecting appropriate reasoning strategies and LLMs. Specifically, we first collect representative posts from the training dataset to capture the real-world symptom expressions, and then utilize these posts to retrieve symptom-relevant posts from the user's posting history. Subsequently, we employ the Mixture of Routers (MoR), which integrates the Mixture of Experts (MoE) into the routing mechanism to select the optimal reasoning strategies and LLMs in a cascaded manner. Finally, we complete the standardized psychological questionnaire using the selected LLMs and reasoning strategies. Experimental results on the Reddit-based benchmarks demonstrate the effectiveness of the proposed method, outperforming existing studies on various metrics. Our code is released at https://github.com/MindIntLab-HFUT/AdaDepression. Yangyang Xu 0002, Jinpeng Hu, Peipei Song, Zhangling Duan, Xun Yang 0001 |
WWW | 4 |
| 2026 | Towards personalized long-term learning modeling in knowledge tracing
Shanshan Wang 0008, Jianqi Qiu, Jiaxin Pang, Xun Yang 0001, Ke Xu 0011, Zhangling Duan, Yuanhong Zhong, Xingyi Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Local-Global Feature Fusion for Enhancing 3D Human Pose EstimationabstractBased on its excellent capability to extract temporal features, transformer has been widely used in monocular 3D human pose estimation. However, due to its global perspective, it performs inadequately in extracting spatial features, which hinders breakthroughs in performance. In this paper, we propose a local-global feature fusion method based on GCN and transformer for 3D human pose estimation. Our method integrates GCN with multiscale transformer to extract local spatiotemporal features of poses. These are then integrated with the global spatiotemporal features extracted by vanilla transformer to reconstruct 3D human poses accurately. In addition, we introduce a hierarchical feature fusion method to better capturing the underlying 3D pose structure. It blends deep abstract features with shallow raw features. We evaluate our model on the Human3.6M and MPI-INF-3DHP datasets, and experimental results demonstrate that our approach outperforms existing state-of-the-art methods. We achieve advanced performance on both datasets with errors of 37.7mm and 16.4mm under MPJPE, respectively. The code and model are available at https://github.com/ygx7/LG3DPose. Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008, Zhangling Duan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support ConversationabstractThe development of Emotional Support Conversation (ESC) systems is critical for delivering mental health support tailored to the needs of help-seekers.Recent advances in large language models (LLMs) have contributed to progress in this domain, while most existing studies focus on generating responses directly and overlook the integration of domain-specific reasoning and expert interaction.Therefore, in this paper, we propose a training-free Multi-Agent collaboration framework for ESC (Mul-tiAgentESC).The framework is designed to emulate the human-like process of providing emotional support through three stages: dialogue analysis, strategy deliberation, and response generation.At each stage, a multi-agent system is employed to iteratively enhance information understanding and reasoning, simulating real-world decision-making processes by incorporating diverse interactions among these expert agents.Additionally, we introduce a novel response-centered approach to handle the one-to-many problem on strategy selection, where multiple valid strategies are initially employed to generate diverse responses, followed by the selection of the optimal response through multi-agent collaboration.Experiments on the ESConv dataset reveal that our proposed framework excels at providing emotional support as well as diversifying support strategy selection 1 . Yangyang Xu 0002, Jinpeng Hu, Zhuoer Zhao, Zhangling Duan, Xiao Sun 0003, Xun Yang 0001 |
EMNLP | 4 |
| 2025 | Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language ModelsabstractKnowledge editing aims to efficiently and cost-effectively correct inaccuracies and update outdated information. Recently, there has been growing interest in extending knowledge editing from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs), which integrate both textual and visual information, introducing additional editing complexities. Existing multimodal knowledge editing works primarily focus on text-oriented, coarse-grained scenarios, failing to address the unique challenges posed by multimodal contexts. In this paper, we propose a visual-oriented, fine-grained multimodal knowledge editing task that targets precise editing in images with multiple interacting entities. We introduce the Fine-Grained Visual Knowledge Editing (FGVEdit) benchmark to evaluate this task. Moreover, we propose a Multimodal Scope Classifier-based Knowledge Editor (MSCKE) framework. MSCKE leverages a multimodal scope classifier that integrates both visual and textual information to accurately identify and update knowledge related to specific entities within images. This approach ensures precise editing while preserving irrelevant information, overcoming the limitations of traditional text-only editing methods. Extensive experiments on the FGVEdit benchmark demonstrate that MSCKE outperforms existing methods, showcasing its effectiveness in solving the complex challenges of multimodal knowledge editing. Leijiang Gu, Xun Yang 0001, Zhangling Duan, Zenglin Shi, Meng Wang 0001 |
ICCV | 4 |
| 2025 | Gradient-based Causal Feature SelectionabstractCausal feature selection leverages causal discovery techniques to identify critical features associated with a target variable using observational data. Traditional methodologies primarily rely on constraint-based or score-based techniques, which are fraught with limitations. For example, conditional independence tests often yield unreliable results in the presence of noise and complex data generation processes, while the computational complexity of learning directed acyclic graphs increases exponentially with the number of variables involved. In light of recent advancements in deep learning, gradient-based methods have shown promise for global causal discovery. However, significant challenges arise when focusing on the identification of local causal features, particularly in defining the local causal constraint space to achieve both minimality and completeness. To address these issues, we introduce a novel gradient-based causal feature selection method (GCFS) that leverages an AutoEncoder to simultaneously model the target variable alongside other variables, thereby capturing of causal associations within a divide-and-conquer framework. Additionally, our approach incorporates a mask pruning strategy that transforms the search process into the minimization of a non-cyclic local reconstruction loss objective function. This function is then effectively optimized using a gradient-based method to accurately identify the causal features related to the target variable. Experimental results substantiate that GCFS surpasses existing methodologies across both synthetic and real datasets. Zhaolong Ling, Mengxiang Guo, Debo Cheng, Peng Zhou 0008, Zhangling Duan |
IJCAI | 7 |
| 2025 | Hierarchical Matrix-Contrastive Bilateral Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) seeks to understand human sentiment by leveraging the correlations across multimodal data. Current approaches often employ contrastive learning and text-centric fusion methods to explore the sentiment mapping space, improve the ability to extract and integrate multimodal features, and capture modality correlations. However, these methods typically depend on complex sampling strategies to select predefined positive and negative samples and perform unidirectional fusion of other modalities aligned with the text. This process overlooks the collaborative information that could be shared between modalities, leading to a loss of valuable insights. To address these limitations, we propose Hierarchical MAtrix-Contrastive BiLateral FusiOn (HALO), which integrates two key components: Matrix-Aware Contrastive Learning (MACL) and Hierarchical Bilateral Fusion (HBF). Specifically, MACL uses two supervisory signals to sample positive and negative pairs within the same batch and assigns different weights according to the difficulty of samples, thereby enhancing the cross-modal discrimination ability of the model. In addition, HBF introduces a bilateral fusion method by guiding vision and audio fusion with text, while using vision and audio information to enhance the overall expressive ability of text. Extensive experiments on datasets MOSI and MOSEI demonstrate the effectiveness and superiority of HALO. Chaoxing Tang, Anyang Tong, Fei Wang 0073, Zhangling Duan |
ICMR | 4 |
| 2025 | Confidence-aware iterative training for cross-lingual entity alignment
Zhangling Duan, Zhaolong Ling, Yun Yang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Fair streaming feature selection
Zhangling Duan, Zhaolong Ling, Jingye Yang |
Neurocomputing | 1 |
| 2025 | Partial multi-label learning with label and classifier correlations
Ke Wang 0047, Yahu Guan, Yunyu Xie, Zhangling Duan, Dong Liang 0009 |
Inf. Sci. | 6 |
| 2025 | Video Corpus Moment Retrieval With Query-Specific Context Learning and Progressive LocalizationabstractVideo corpus moment retrieval (VCMR) aims to retrieve a moment from a large corpus of untrimmed videos corresponding to a given language query. However, existing methods often fall short due to their reliance on simple cross-modal attention mechanisms and one-stop localization, which fail to handle the complex multimodal information and large search space effectively. To address these challenges, we propose a novel VCMR method with Query-specific Context Learning and Progressive Localization (QCLPL). First, we construct query-specific multimodal contexts that capture complementary and consistent semantics across subtitles and frames, ensuring informative and efficient context building. We further introduce a semantic contrastive loss to refine these multimodal contexts, filtering out query-irrelevant information. Additionally, we introduce a progressive localization strategy that transforms the moment localization task into a two-stage process. By classifying frames into foreground and background regions, we present a simplified binary classification problem before boundary prediction, constrained by a region-aware loss. This progressive approach leverages region priors to improve subsequent moment localization. Extensive experiments on the TVR and DiDeMo datasets demonstrate that our method significantly outperforms existing approaches, setting a new state of the art for VCMR. Peipei Song, Zhangling Duan, Shuo Wang 0008, Xiaojun Chang, Xun Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Entropy-minimization Mean Teacher for Source-Free Domain Adaptive Object Detection
Xing Wei 0002, Ting Bai 0006, Zhangling Duan, Yang Lu 0015 |
ICONIP (1) | 3 |
| 2020 | Class-unbalanced domain adaptation for object detection via dynamic weighting mechanismabstractThe state-of-the-art object detection frameworks often suffer from a performance decline when the feature distribution differences between the source (training) domain and the target (testing) domain are existed. To alleviate this problem, recent works proposed various domain adaptation methods to improve object detection frameworks. Existing methods only consider the feature discrepancy between the source and target domains and ignores unbalanced in class space under cross-domain settings, which can lead to serious negative transfer problems. To address this issue, we propose a novel domain adaptation for object detection to reduce class space discrepancy between domains. Specifically, the weighted mechanism is used to increase the weight of public categories between domains to promote positive transfer and reduce the weight of non-public categories to retard the impact of negative transfer. Moreover, the model reduces domain feature distribution discrepancy by adding domain classifiers and employing adversarial training methods. The results of our experiments on several datasets demonstrate that our model can effectively solve the problem of performance degradation caused by the discrepancy in class space and significantly improve the detection accuracy in each domain. Xing Wei 0002, Shaofan Liu, Changguang Wang, Yaoci Xiang, Xuanyuan Qiao, Zhangling Duan, Yang Lu 0015 |
3DV | 6 |
| 2020 | Incremental learning based multi-domain adaptation for object detection
Xing Wei 0002, Shaofan Liu, Yaoci Xiang, Zhangling Duan, Yang Lu 0015 |
Knowl. Based Syst. | 4 |
| 2020 | Simulated annealing-based reprogramming scheme of wireless sensor nodes
Zhangling Duan, Xing Wei 0002, Jianghong Han, Yang Lu 0015, Lei Shi 0011 |
Wirel. Networks | 1 |