EDBT 2026 Demo / reviewers in the wild / expert
Duzhen Zhang
dblp:235/0398
· DBLP profile ↗
28ranked-venue papers
13as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 11 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bring Your Dreams to Life: Continual Text-to-Video CustomizationabstractCustomized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting and concept neglect when continuously learning new concepts, including subjects and motions. To resolve the above challenges, we develop a novel Continual Customized Video Diffusion (CCVD) model, which can continuously learn new concepts to generate videos across various text-to-video generation tasks by tackling forgetting and concept neglect. To address catastrophic forgetting, we introduce a concept-specific attribute retention module and a task-aware concept aggregation strategy. They can capture the unique characteristics and identities of old concepts during training, while combining all subject and motion adapters of old concepts based on their relevance during testing. Besides, to tackle concept neglect, we develop a controllable conditional synthesis to enhance regional features and align video contexts with user conditions, by incorporating layer-specific region attention-guided noise estimation. Extensive experimental comparisons demonstrate that our CCVD outperforms existing CTVG models. Jiahua Dong 0001, Wenqi Liang, Zongyan Han, Meng Cao 0002, Duzhen Zhang, Hanbin Zhao, Zhi Han, Salman Khan 0001, Fahad Shahbaz Khan |
AAAI | 6 |
| 2026 | TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction TasksabstractWhile Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previous works predominantly focusing on text and thus overlooking critical visual information, such as molecular structures. Current approaches that directly adopt standard VLMs for chemical tasks suffer from two primary issues: (i) computational inefficiency of processing entire chemical images with non-informative backgrounds. (ii) a narrow scope on molecular-level tasks that restricts progress in chemical reasoning. In this work, we propose TinyChemVL, an efficient and powerful chemical VLM that leverages visual token reduction and reaction-level tasks to improve model efficiency and reasoning capacity. Also, we propose ChemRxn-V, a reaction-level benchmark for assessing vision-based reaction recognition and prediction tasks. Directly predicting reaction products from molecular images poses a non-trivial challenge, as it requires models to integrate both recognition and reasoning capacities. Our results demonstrate that, with only 4B parameters, TinyChemVL achieves superior performance on both molecular and reaction tasks, while also demonstrating faster inference and training speeds compared to existing models. Notably, TinyChemVL outperforms ChemVLM while utilizing only 1/16th of the visual tokens. This work builds efficient yet powerful VLMs for chemical domains by co-designing model architecture and task complexity. Xuanle Zhao, Shuxin Zeng, Xinyuan Cai, Duzhen Zhang, Xiuyi Chen, Bo Xu 0002 |
AAAI | 5 |
| 2026 | AutoReproduce: Automatic AI Experiment Reproduction with Paper LineageabstractXuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanle Zhao, Zilin Sang, Qi Shi 0002, Wei-Lun Zhao, Shuo Wang 0013, Duzhen Zhang, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 7 |
| 2026 | From System 1 to System 2: A Survey of Reasoning Large Language ModelsabstractAchieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, trace the evolution of various reasoning models, and examine the core methods that enable advanced reasoning behind them. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time GitHub Repository to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field. Duzhen Zhang, Zhongzhi Li, Jiaxin Zhang 0024, Zengyan Liu, Junhao Zheng, Xiuyi Chen, Jiahua Dong 0001, Zhijiang Guo, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Lifelong Learning of Large Language Model Based Agents: A RoadmapabstractLifelong learning, also known as continual or incremental learning, is a crucial component for advancing Artificial General Intelligence (AGI) by enabling systems to continuously adapt in dynamic environments. While large language models (LLMs) have demonstrated impressive capabilities in natural language processing, existing LLM agents are typically designed for static systems and lack the ability to adapt over time in response to new challenges. This survey is the first to systematically summarize the potential techniques for incorporating lifelong learning into LLM-based agents. We categorize the core components of these agents into three modules: the perception module for multimodal input integration, the memory module for storing and retrieving evolving knowledge, and the action module for grounded interactions with the dynamic environment. We highlight how these pillars collectively enable continuous adaptation, mitigate catastrophic forgetting, and improve long-term performance. This survey provides a roadmap for researchers and practitioners working to develop lifelong learning capabilities in LLM agents, offering insights into emerging trends, evaluation metrics, and application scenarios. Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu 0001, Qianli Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Enhancing Multimodal Continual Instruction Tuning with BranchLoRAabstractMultimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes. Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai |
ACL (1) | 1 |
| 2025 | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren 0006, Chenxing Li, Duzhen Zhang, Yujie Chen 0006, Manjie Xu, Ruibo Fu, Shan Yang 0001, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2025 | Information-theoretic complementary prompts for improved continual text classification
Duzhen Zhang, Yong Ren 0006, Chenxing Li, Dong Yu 0001, Tielin Zhang |
Neural Networks | 1 |
| 2024 | AC-CAM: Affinity-Aware Contrast CAM for Weakly-Supervised Semantic Segmentation on MRI Brain TumorabstractApplying the latest visual transformer (ViT) to Weakly-Supervised Semantic Segmentation (WSSS) can compensate for the local perception limitations of CNN, but it also brings about the over-smoothing problem, that is, the final patch labels tend to be uniform. To overcome this challenge, we present an Affinity-Aware Contrast Class Activation Maps (AC-CAM) framework aimed at enhancing WSSS for MRI Brain Tumor analysis by exploiting only image-level labels. We propose two main components: the Affinity-Aware Token Contrast Module (ATCM) and the Affinity-Aware Refine Module (ARM). ATCM utilizes semantic affinities from attention maps to improve the contrast between patch tokens, effectively reducing the over-smoothing tendency of Vision Transformers (ViT). ARM refines the pseudo labels further, incorporating RGB and affinity information to capture the intricate details of the target objects. Our approach capitalizes on the global feature capturing capabilities of ViT, producing more accurate pseudo-labels for WSSS. The framework is optimized through a composite loss function that ensures the consistency of representations for positive token pairs and discriminability for negative ones. Experiments show that our method achieves state-of-the-art performance on the BraTS 2021 dataset. Jingqian Wu, Changwei Wang 0001, Duzhen Zhang, Rongtao Xu |
BIBM | 5 |
| 2024 | Prompt-guided Precise Audio Editing with Diffusion ModelsabstractAudio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as **PPAE**, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks. Manjie Xu, Chenxing Li, Duzhen Zhang, Dan Su 0002, Dong Yu 0001 |
ICML | 3 |
| 2024 | DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic SegmentationabstractThe complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks. Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
ICRA | 3 |
| 2024 | How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?abstractCustom diffusion models (CDMs) have attracted widespread attention due to their astonishing generative ability for personalized concepts. However, most existing CDMs unreasonably assume that personalized concepts are fixed and cannot change over time. Moreover, they heavily suffer from catastrophic forgetting and concept neglect on old personalized concepts when continually learning a series of new concepts. To address these challenges, we propose a novel Concept-Incremental text-to-image Diffusion Model (CIDM), which can resolve catastrophic forgetting and concept neglect to learn new customization tasks in a concept-incremental manner. Specifically, to surmount the catastrophic forgetting of old concepts, we develop a concept consolidation loss and an elastic weight aggregation module. They can explore task-specific and task-shared knowledge during training, and aggregate all low-rank weights of old concepts based on their contributions during inference. Moreover, in order to address concept neglect, we devise a context-controllable synthesis strategy that leverages expressive region features and noise estimation to control the contexts of generated images according to user conditions. Experiments validate that our CIDM surpasses existing custom diffusion models. The source codes are available at https://github.com/JiahuaDong/CIFC. Jiahua Dong 0001, Wenqi Liang, Hongliu Li, Duzhen Zhang, Meng Cao 0002, Henghui Ding, Salman Khan 0001, Fahad Shahbaz Khan |
NeurIPS | 4 |
| 2024 | Image Denoise Model Based on Structural Re-Parameterized Uniformer Transformer and UNetabstractImage denoising is a fundamental problem in computer vision (CV). Vision Transformer (ViT) is an improvement in CV after convolutional neural networks (CNNs). Recently, it has been demonstrated that the RepVGG network of the structural re-parameterization performs well in image tasks. Experiments indicate that the ViT-based uniformer transformer network has successfully balanced local and global information. As image denoising tasks require more local and global information about the image, a novel image denoising model named structural Re-parameterization Uniformer Transformer-UNet (Rep-UUNet) is proposed in this paper. The model is structural and re-parameterized the structure of the Uniformer Transformer network and uses the UNet skip connections to reconstruct the output image. To complete image downstream tasks, the RepVGG model which utilizes the image local information is used. Indicators such as PSNR, SSIM and others are used to assess the performance of image denoising. Experimental results demonstrate that our Rep-UUNet network model outperforms other five models. Zhengwei Lu, Duzhen Zhang, Tao Wang 0160, Hengchang Jiang, Yingkai Pu |
Int. J. Comput. Intell. Appl. | 2 |
| 2024 | Structure Aware Multi-Graph Network for Multi-Modal Emotion Recognition in ConversationsabstractMulti-Modal Emotion Recognition in Conversations (MMERC) is an increasingly active research field that leverages multi-modal signals to understand the feelings behind each utterance. Modeling contextual interactions and multi-modal fusion lie at the heart of this field, with graph-based models recently being widely used for MMERC to capture global multi-modal contextual information. However, these models generally mix all modality representations in a single graph, and utterances in each modality are fully connected, potentially ignoring three problems: 1) the heterogeneity of the multi-modal context, 2) the redundancy of contextual information, and 3) over-smoothing of the graph networks. To address these problems, we propose a Structure Aware Multi-Graph Network (SAMGN) for MMERC. Specifically, we construct multiple modality-specific graphs to model the heterogeneity of the multi-modal context. Instead of fully connecting the utterances in each modality, we design a structure learning module that determines whether edges exist between the utterances. This module reduces redundancy by forcing each utterance to focus on the contextual ones that contribute to its emotion recognition, acting like a message propagating reducer to alleviate over-smoothing. Then, we develop the SAMGN via Dual-Stream Propagation (DSP), which contains two propagation streams, i.e., intra- and inter-modal, performed in parallel to aggregate the heterogeneous modality information from multi-graphs. DSP also contains a gating unit that adaptively integrates the co-occurrence information from the above two propagations for emotion recognition. Experiments on two popular MMERC datasets demonstrate that SAMGN achieves new State-Of-The-Art (SOTA) results. Duzhen Zhang, Jianlong Chang, Xiuyi Chen, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech RecognitionabstractThe spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different scales of neuronal dynamics is necessary. Here we introduce four types of neuronal dynamics to post-process the sequential patterns generated from the spiking transformer to get the complex dynamic neuron improved spiking transformer neural network (DyTr-SNN). We found that the DyTr-SNN could handle the non-toy automatic speech recognition task well, representing a lower phoneme error rate, lower computational cost, and higher robustness. These results indicate that the further cooperation of SNNs and neural dynamics at the neuron and network scales might have much in store for the future, especially on the ASR tasks. Tielin Zhang, Minglun Han, Yi Wang 0077, Duzhen Zhang, Bo Xu 0002 |
AAAI | 5 |
| 2023 | DualGATs: Dual Graph Attention Networks for Emotion Recognition in ConversationsabstractCapturing complex contextual dependencies plays a vital role in Emotion Recognition in Conversations (ERC).Previous studies have predominantly focused on speaker-aware context modeling, overlooking the discourse structure of the conversation.In this paper, we introduce Dual Graph ATtention networks (Dual-GATs) to concurrently consider the complementary aspects of discourse structure and speaker-aware context, aiming for more precise ERC.Specifically, we devise a Discourseaware GAT (DisGAT) module to incorporate discourse structural information by analyzing the discourse dependencies between utterances.Additionally, we develop a Speakeraware GAT (SpkGAT) module to incorporate speaker-aware contextual information by considering the speaker dependencies between utterances.Furthermore, we design an interaction module that facilitates the integration of the DisGAT and SpkGAT modules, enabling the effective interchange of relevant information between the two modules.We extensively evaluate our method on four datasets, and experimental results demonstrate that our proposed DualGATs surpass state-of-the-art baselines on the majority of the datasets.1 Duzhen Zhang, Xiuyi Chen |
ACL (1) | 1 |
| 2023 | Task Relation Distillation and Prototypical Pseudo Label for Incremental Named Entity RecognitionabstractIncremental Named Entity Recognition (INER) involves the sequential learning of new entity types without accessing the training data of previously learned types. However, INER faces the challenge of catastrophic forgetting specific for incremental learning, further aggravated by background shift (i.e., old and future entity types are labeled as the non-entity type in the current task). To address these challenges, we propose a method called task Relation Distillation and Prototypical pseudo label (RDP) for INER. Specifically, to tackle catastrophic forgetting, we introduce a task relation distillation scheme that serves two purposes: 1) ensuring inter-task semantic consistency across different incremental learning tasks by minimizing inter-task relation distillation loss, and 2) enhancing the model's prediction confidence by minimizing intra-task self-entropy loss. Simultaneously, to mitigate background shift, we develop a prototypical pseudo label strategy that distinguishes old entity types from the current non-entity type using the old model. This strategy generates high-quality pseudo labels by measuring the distances between token embeddings and type-wise prototypes. We conducted extensive experiments on ten INER settings of three benchmark datasets (i.e., CoNLL2003, I2B2, and OntoNotes5). The results demonstrate that our method achieves significant improvements over the previous state-of-the-art methods, with an average increase of 6.08% in Micro F1 score and 7.71% in Macro F1 score. Duzhen Zhang, Hongliu Li, Wei Cong, Rongtao Xu, Jiahua Dong 0001, Xiuyi Chen |
CIKM | 1 |
| 2023 | Federated Incremental Semantic SegmentationabstractFederated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fixed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive new categories incrementally while have no memory storage to access old classes. Moreover, new clients collecting novel classes may join in the global training of FSS, which further exacerbates catastrophic forgetting. To surmount the above challenges, we propose a Forgetting-Balanced Learning (FBL) model to address heterogeneous forgetting on old classes from both intra-client and interclient aspects. Specifically, under the guidance of pseudo labels generated via adaptive class-balanced pseudo labeling, we develop a forgetting-balanced semantic compensation loss and a forgetting-balanced relation consistency loss to rectify intra-client heterogeneous forgetting of old categories with background shift. It performs balanced gradient propagation and relation consistency distillation within local clients. Moreover, to tackle heterogeneous forgetting from inter-client aspect, we propose a task transition monitor. It can identify new classes under privacy protection and store the latest old global model for relation distillation. Qualitative experiments reveal large improvement of our model against comparison methods. The code is available at https://github.com/JiahuaDong/FISS. Jiahua Dong 0001, Duzhen Zhang, Yang Cong, Wei Cong, Henghui Ding, Dengxin Dai |
CVPR | 2 |
| 2023 | Continual Named Entity Recognition without Catastrophic ForgettingabstractContinual Named Entity Recognition (CNER) is a burgeoning area, which involves updating an existing model by incorporating new entity types sequentially.Nevertheless, continual learning approaches are often severely afflicted by catastrophic forgetting.This issue is intensified in CNER due to the consolidation of old entity types from previous steps into the non-entity type at each step, leading to what is known as the semantic shift problem of the non-entity type.In this paper, we introduce a pooled feature distillation loss that skillfully navigates the trade-off between retaining knowledge of old entity types and acquiring new ones, thereby more effectively mitigating the problem of catastrophic forgetting.Additionally, we develop a confidence-based pseudo-labeling for the non-entity type, i.e., predicting entity types using the old model to handle the semantic shift of the non-entity type.Following the pseudo-labeling process, we suggest an adaptive re-weighting type-balanced learning strategy to handle the issue of biased type distribution.We carried out comprehensive experiments on ten CNER settings using three different datasets.The results illustrate that our method significantly outperforms prior state-of-the-art approaches, registering an average improvement of 6.3% and 8.0% in Micro and Macro F1 scores, respectively.1 * Equal contributions.† The corresponding author is Dr. Duzhen Zhang, Wei Cong, Jiahua Dong 0001, Yahan Yu, Xiuyi Chen, Yonggang Zhang 0003, Zhen Fang 0001 |
EMNLP | 1 |
| 2023 | Crucial Semantic Classifier-based Adversarial Learning for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA), which aims to explore the transferrable features from a well-labeled source domain to a related unlabeled target domain, has been widely progressed. Nevertheless, as one of the mainstream, existing adversarial-based methods neglect to filter the irrelevant semantic knowledge, hindering adaptation performance improvement. Besides, they require an additional domain discriminator that strives extractor to generate confused representations, but discrete designing may cause model collapse. To tackle the above issues, we propose Crucial Semantic Classifier-based Adversarial Learning (CSCAL), which pays more attention to crucial semantic knowledge transferring and leverages the classifier to implicitly play the role of domain discriminator without extra network designing. CSCAL effectively mitigates distribution shifts between the source and target domains from both intra- and inter-class perspectives. Specifically, in intra-class-wise alignment, a Paired-Level Discrepancy (PLD) is designed to transfer crucial semantic knowledge. Additionally, based on classifier predictions, a Nuclear Norm-based Discrepancy (NND) is formed that considers inter-class-wise information and improves the adaptation performance. Moreover, CSCAL can be effortlessly merged into different UDA methods as a regularizer and dramatically promote their performance. Yajun Gao, Hongliu Li, Ating Yin, Duzhen Zhang, Xiuyi Chen |
IJCNN | 5 |
| 2023 | ODE-based Recurrent Model-free Reinforcement Learning for POMDPsabstractNeural ordinary differential equations (ODEs) are widely recognized as the standard for modeling physical mechanisms, which help to perform approximate inference in unknown physical or biological environments. In partially observable (PO) environments, how to infer unseen information from raw observations puzzled the agents. By using a recurrent policy with a compact context, context-based reinforcement learning provides a flexible way to extract unobservable information from historical transitions. To help the agent extract more dynamics-related information, we present a novel ODE-based recurrent model combines with model-free reinforcement learning (RL) framework to solve partially observable Markov decision processes (POMDPs). We experimentally demonstrate the efficacy of our methods across various PO continuous control and meta-RL tasks. Furthermore, our experiments illustrate that our method is robust against irregular observations, owing to the ability of ODEs to model irregularly-sampled time series. Xuanle Zhao, Duzhen Zhang, Liyuan Han, Tielin Zhang, Bo Xu 0002 |
NeurIPS | 2 |
| 2023 | Decomposing Logits Distillation for Incremental Named Entity RecognitionabstractIncremental Named Entity Recognition (INER) aims to continually train a model with new data, recognizing emerging entity types without forgetting previously learned ones. Prior INER methods have shown that Logits Distillation (LD), which involves preserving predicted logits via knowledge distillation, effectively alleviates this challenging issue. In this paper, we discover that a predicted logit can be decomposed into two terms that measure the likelihood of an input token belonging to a specific entity type or not. However, the traditional LD only preserves the sum of these two terms without considering the change in each component. To explicitly constrain each term, we propose a novel Decomposing Logits Distillation (DLD) method, enhancing the model's ability to retain old knowledge and mitigate catastrophic forgetting. Moreover, DLD is model-agnostic and easy to implement. Extensive experiments show that DLD consistently improves the performance of state-of-the-art INER methods across ten INER settings in three datasets. Duzhen Zhang, Yahan Yu, Xiuyi Chen |
SIGIR | 1 |
| 2022 | Multi-Sacle Dynamic Coding Improved Spiking Actor Network for Reinforcement LearningabstractWith the help of deep neural networks (DNNs), deep reinforcement learning (DRL) has achieved great success on many complex tasks, from games to robotic control. Compared to DNNs with partial brain-inspired structures and functions, spiking neural networks (SNNs) consider more biological features, including spiking neurons with complex dynamics and learning paradigms with biologically plausible plasticity principles. Inspired by the efficient computation of cell assembly in the biological brain, whereby memory-based coding is much more complex than readout, we propose a multiscale dynamic coding improved spiking actor network (MDC-SAN) for reinforcement learning to achieve effective decision-making. The population coding at the network scale is integrated with the dynamic neurons coding (containing 2nd-order neuronal dynamics) at the neuron scale towards a powerful spatial-temporal state representation. Extensive experimental results show that our MDC-SAN performs better than its counterpart deep actor network (based on DNNs) on four continuous control tasks from OpenAI gym. We think this is a significant attempt to improve SNNs from the perspective of efficient coding towards effective decision-making, just like that in biological networks. Duzhen Zhang, Tielin Zhang, Shuncheng Jia, Bo Xu 0002 |
AAAI | 1 |
| 2022 | TSAM: A Two-Stream Attention Model for Causal Emotion EntailmentabstractCausal Emotion Entailment (CEE) aims to discover the potential causes behind an emotion in a conversational utterance. Previous works formalize CEE as independent utterance pair classification problems, with emotion and speaker information neglected. From a new perspective, this paper considers CEE in a joint framework. We classify multiple utterances synchronously to capture the correlations between utterances in a global view and propose a Two-Stream Attention Model (TSAM) to effectively model the speaker’s emotional influences in the conversational history. Specifically, the TSAM comprises three modules: Emotion Attention Network (EAN), Speaker Attention Network (SAN), and interaction module. The EAN and SAN incorporate emotion and speaker information in parallel, and the subsequent interaction module effectively interchanges relevant information between the EAN and SAN via a mutual BiAffine transformation. Extensive experimental results demonstrate that our model achieves new State-Of-The-Art (SOTA) performance and outperforms baselines remarkably. Duzhen Zhang, Fandong Meng, Xiuyi Chen, Jie Zhou 0016 |
COLING | 1 |
| 2022 | Recent Advances and New Frontiers in Spiking Neural NetworksabstractIn recent years, spiking neural networks (SNNs) have received extensive attention in brain-inspired intelligence due to their rich spatially-temporal dynamics, various encoding methods, and event-driven characteristics that naturally fit the neuromorphic hardware. With the development of SNNs, brain-inspired intelligence, an emerging research field inspired by brain science achievements and aiming at artificial general intelligence, is becoming hot. This paper reviews recent advances and discusses new frontiers in SNNs from five major research topics, including essential elements (i.e., spiking neuron models, encoding methods, and topology structures), neuromorphic datasets, optimization algorithms, software, and hardware frameworks. We hope our survey can help researchers understand SNNs better and inspire new works to advance this field. Duzhen Zhang, Tielin Zhang, Shuncheng Jia, Bo Xu 0002 |
IJCAI | 1 |
| 2022 | Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual DialogabstractVisual dialog requires models to give reasonable answers according to a series of coherent questions and related visual concepts in images. However, most current work either focuses on attention-based fusion or pre-training on large-scale image-text pairs, ignoring the critical role of explicit vision-language alignment in visual dialog. To remedy this defect, we propose a novel unsupervised and pseudo-supervised vision-language alignment approach for visual dialog (AlignVD). Firstly, AlginVD utilizes the visual and dialog encoder to represent images and dialogs. Then, it explicitly aligns visual concepts with textual semantics via unsupervised and pseudo-supervised vision-language alignment (UVLA and PVLA). Specifically, UVLA utilizes a graph autoencoder, while PVLA uses dialog-guided visual grounding to conduct alignment. Finally, based on the aligned visual and textual representations, AlignVD gives a reasonable answer to the question via the cross-modal decoder. Extensive experiments on two large-scale visual dialog datasets have demonstrated the effectiveness of vision-language alignment, and our proposed AlignVD achieves new state-of-the-art results. In addition, our single model has won first place on the visual dialog challenge leaderboard with a NDCG metric of 78.70, surpassing the previous best ensemble model by about 1 point. Duzhen Zhang, Xiuyi Chen, Jing Shi 0003, Bo Xu 0002 |
ACM Multimedia | 2 |
| 2020 | Knowledge Aware Emotion Recognition in Textual Conversations via Multi-Task Incremental TransformerabstractEmotion recognition in textual conversations (ERTC) plays an important role in a wide range of applications, such as opinion mining, recommender systems, and so on.ERTC, however, is a challenging task.For one thing, speakers often rely on the context and commonsense knowledge to express emotions; for another, most utterances contain neutral emotion in conversations, as a result, the confusion between a few non-neutral utterances and much more neutral ones restrains the emotion recognition performance.In this paper, we propose a novel Knowledge Aware Incremental Transformer with Multi-task Learning (KAITML) to address these challenges.Firstly, we devise a dual-level graph attention mechanism to leverage commonsense knowledge, which augments the semantic information of the utterance.Then we apply the Incremental Transformer to encode multi-turn contextual utterances.Moreover, we are the first to introduce multi-task learning to alleviate the aforementioned confusion and thus further improve the emotion recognition performance.Extensive experimental results show that our KAITML model outperforms the state-of-the-art models across five benchmark datasets. Duzhen Zhang, Xiuyi Chen, Bo Xu 0002 |
COLING | 1 |
| 2019 | Top-Down Saliency Detection Based on Deep-Learned FeaturesabstractHow to localize objects in images accurately and efficiently is a challenging problem in computer vision. In this paper, a novel top–down fine-grained salient object detection method based on deep-learned features is proposed, which can detect the same object in input image as the query image. The query image and its three subsample images are used as top–down cues to guide saliency detection. We ameliorate convolutional neural network (CNN) using the fast VGG network (VGG-f) pre-trained on ImageNet and re-trained on the Pascal VOC 2012 dataset. Experiment on the FiFA dataset demonstrates that proposed method can localize the saliency region and find the specific object (e.g., human face) as the query. Experiments on the David1 and Face1 sequences conclusively prove that the proposed algorithm is able to effectively deal with many challenging factors including illumination change, shape deformation, scale change and partial occlusion. Duzhen Zhang, Zakir Ali |
Int. J. Comput. Intell. Appl. | 1 |