EDBT 2026 Demo / reviewers in the wild / expert
Tianjiao Wan
dblp:348/9607
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-2423-4982ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental LearningabstractExemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques. Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Dynamic Confidence Variance for Generalized Coreset in Active LearningabstractActive Learning (AL) aims to reduce data annotation costs by selecting the most informative samples from an unlabeled data pool. Traditional AL methods often rely on a single snapshot to identify uncertain or representative samples, often overlooking the poor generalization of a single model. Recent AL studies have attempted to address this issue by tracking a broader range of training dynamics for data selection, typically using averaging or accumulating manner. However, both our theoretical and experimental analyses reveal that these methods obscure the variability inherent in the training process, potentially prioritizing hard-to-learn samples that result in poor generalization. In this paper, we propose a novel AL method termed as Dynamic Confidence Variance (DCoV), that seamlessly integrates variability with the training dynamic to effectively identify a well-generalized Coreset. DCoV leverages the variance of the model’s prediction confidence throughout the training process for active sampling and model training. Our theoretical analysis demonstrates that DCoV provides a lower bound on the population risk of the model learned from selected labeled subset, spanning the entire training process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art AL methods on various balanced and imbalanced benchmark datasets across various modalities. Tianjiao Wan, Zijian Gao, Xudong Gong, Bo Ding 0001, Huaimin Wang 0001, Kele Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | JI2S: Joint Influence-Aware Instruction Data Selection for Efficient Fine-TuningabstractInstruction tuning (IT) improves large language models (LLMs) by aligning their outputs with human instructions, but its success depends critically on training data quality, and datasets such as Alpaca often contain noisy or suboptimal examples that undermine fine-tuning.Prior selection strategies score samples using general-purpose LLMs (e.g., GPT), leveraging their strong language understanding yet introducing inherent biases that misalign with the target model's behavior and yield unstable downstream performance.Influence-based methods address this by estimating each example's marginal contribution to overall performance, but they typically assume additive contributions and therefore overlook higher-order interactions among samples.To overcome these limitations, we propose JI 2 S, a novel framework that jointly models both marginal and combinatorial influences within sample groups.Applying JI 2 S to select the top 1,000 most influential examples from Alpaca, we fine-tune LLaMA2-7B, Mistral-7B, and LLaMA2-13B and evaluate them on Open LLM Benchmarks, MT-Bench, and GPT-4-judged pairwise comparisons.Our experiments show that JI 2 S consistently outperforms full-dataset training and strong baselines, highlighting the value of capturing joint influence for high-quality instruction fine-tuning.We provide our code in this GitHub repository. Jingyu Wei, Bo Liu 0014, Tianjiao Wan, Baoyun Peng, Xingkong Ma, Mengmeng Guo |
EMNLP | 3 |
| 2025 | Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-ModelingabstractDeep learning-based bioacoustic audio analysis holds immense potential across various applications. However, existing studies in bioacoustics often focus on a limited number of species, potentially hindering the transferability of models across different species. Furthermore, the manual annotation of bioacoustic data is both costly and labor-intensive. To address these challenges, self-supervised learning on large-scale bioacoustic audio data presents a promising solution. In this paper, we introduce GPM-BT (General Pre-training Model for Bioacoustic Tasks), a self-supervised, Transformer-based model pre-trained on approximately 1.2 million unannotated bioacoustic audio samples. We evaluate the scalability and effectiveness of this pre-training approach through comprehensive experiments across a broad range of classification and detection tasks. Our results demonstrate that pre-training on large-scale bioacoustic data significantly enhances model performance, improving both generalization and robustness. Notably, GPM-BT achieves state-of-the-art performance on the BEANS benchmark and secures first place in the Few-shot Bioacoustic Event Detection task at the IEEE DCASE 2024 Challenge1. To further advance research in bioacoustics, we have open-sourced our models and code2. Xuyao Deng, Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
ICASSP | 2 |
| 2025 | Complementary Learning System Theory-based Active Learning for Audio ClassificationabstractDeep learning has significantly advanced the audio classification, achieving remarkable results. However, these successes often rely on extensive manual annotation of audio, a labor-intensive and costly process. Active Learning (AL) presents a promising solution by minimizing the required amount of annotation through the iterative selection of the most informative audio samples. Current AL methods for audio classification typically depend solely on the latest model checkpoint, overlooking the dynamics of the entire training process. The Complementary Learning Systems (CLS) theory posits that the interplay between short-term and long-term memory systems can effectively measure sample uncertainty, offering a means to capture training dynamics. In this work, we introduce a novel AL framework for audio classification, termed CLS-AL, which addresses the limitations of existing methods by simultaneously maintaining both short-term and long-term memory models. This dual-memory approach allows for a more comprehensive consideration of training dynamics. The divergence in predictions between these memory models provides a new metric for evaluating the uncertainty of unlabeled samples, enhancing the effectiveness of the AL sample selection process. We demonstrate the effectiveness and generalizability of CLS-AL through extensive experiments on a diverse set of audio datasets, showing that CLS-AL obviously outperforms existing state-of-the-art methods. Hui Geng, Zijian Gao, Tianjiao Wan, Kele Xu |
ICASSP | 3 |
| 2025 | Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint UnderstandingabstractEmotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address this challenge, we propose the Text-guided Multimodal Emotion-Intent Joint Recognition method. By leveraging the text modality to guide the fusion process, it effectively reduces the noise introduced by other modalities. To strengthen the text modality’s guiding role, we use large language models (LLMs) for multi-turn targeted data augmentation and oversampling strategies to address data imbalance. Our approach achieved first place in Track 1 (English) of the ICASSP 2025 MEIJU Challenge, demonstrating its effectiveness in practical applications. Yu Zhang 0133, Bin Chen 0006, Hongfei Ye, Zijian Gao, Tianjiao Wan, Long Lan, Kele Xu |
ICASSP | 5 |
| 2025 | Segment Anything for Visual Bird Sound DenoisingabstractCurrent audio denoising methods perform well with synthetic noise but struggle with complex natural noise, especially for bird sounds, which contain natural environmental sounds such as wind and rain, making it challenging to extract clean bird sounds. This issue becomes more pronounced with short and faint bird sounds, where existing methods are less effective. In this paper, we introduceBudSAM, a novel audio denoising model that incorporates theSegment Anything Model (SAM), originally designed for image segmentation task, into the field of visual bird sound denoising. By treating audio denoising as a segmentation task, BudSAM utilizes SAM's powerful segmentation capabilities and we incorporates BCE and Dice losses to enhance the model's ability to segment weak signals, effectively isolating the clean bird sounds that are often masked by background noise. Our method is evaluated on the BirdSoundsDenoising dataset, achieving a 4.0% improvement in IoU and a 0.77 dB increase in SDR compared to state-of-the-art methods. To the best knowledge of the authors, BudSAM marks the first attempt which employs SAM in audio denoising task, offering a promising direction for future research and real-world bird sound processing tasks. Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
IEEE Signal Process. Lett. | 2 |
| 2025 | One-Step Multi-View Clustering With Diverse RepresentationabstractMulti-View clustering has attracted broad attention due to its capacity to utilize consistent and complementary information among views. Although tremendous progress has been made recently, most existing methods undergo high complexity, preventing them from being applied to large-scale tasks. Multi-View clustering via matrix factorization is a representative to address this issue. However, most of them map the data matrices into a fixed dimension, limiting the model's expressiveness. Moreover, a range of methods suffers from a two-step process, i.e., multimodal learning and the subsequent k-means, inevitably causing a suboptimal clustering result. In light of this, we propose a one-step multi-view clustering with diverse representation (OMVCDR) method, which incorporates multi-view learning and k-means into a unified framework. Specifically, we first project original data matrices into various latent spaces to attain comprehensive information and auto-weight them in a self-supervised manner. Then, we directly use the information matrices under diverse dimensions to obtain consensus discrete clustering labels. The unified work of representation learning and clustering boosts the quality of the final results. Furthermore, we develop an efficient optimization algorithm with proven convergence to solve the resultant problem. Comprehensive experiments on various datasets demonstrate the promising clustering performance of our proposed method. The code is publicly available at https://github.com/wanxinhang/OMVCDR. Xinhang Wan, Jiyuan Liu 0003, Xinbiao Gan, Xinwang Liu 0002, Siwei Wang 0001, Yi Wen 0001, Tianjiao Wan, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Temporal Inconsistency-Based Active LearningabstractDeep supervised learning has demonstrated strong capabilities; however, such progress relies on massive and expensive data annotation. Active Learning (AL) has been introduced to selectively annotate samples, thus reducing the human labeling effort. Previous AL research has focused on employing recently trained models to design sampling strategies, based on uncertainty or representativeness. Drawing inspiration from the issue of model forgetting, we propose a novel AL framework called Temporal Inconsistency-Based Active Learning (TIR-AL). In this framework, multiple snapshots of the models across consecutive cycles are jointly utilized to select samples with higher temporal inconsistency, by computing the proposed self-weighted nuclear norm metric. Furthermore, we introduce a consistency regularization term to mitigate the issue of forgetting. Together, these components make full use of the potential of data and facilitate effective interaction within the AL loop. To demonstrate the efficacy of TIR-AL, we conducted a set of experiments illustrating how our approach outperforms state-of-the-art methods without incurring any additional training costs. Tianjiao Wan, Yutao Dou, Kele Xu, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001 |
ICASSP | 1 |
| 2024 | Decouple then Classify: A Dynamic Multi-view Labeling Strategy with Shared and Specific InformationabstractSample labeling is the most primary and fundamental step of semi-supervised learning. In literature, most existing methods randomly label samples with a given ratio, but achieve unpromising and unstable results due to the randomness, especially in multi-view settings. To address this issue, we propose a Dynamic Multi-view Labeling Strategy with Shared and Specific Information. To be brief, by building two classifiers with existing labels to utilize decoupled shared and specific information, we select the samples of low classification confidence and label them in high priorities. The newly generated labels are also integrated to update the classifiers adaptively. The two processes are executed alternatively until a satisfying classification performance. To validate the effectiveness of the proposed method, we conduct extensive experiments on popular benchmarks, achieving promising performance. The code is publicly available at https://github.com/wanxinhang/ICML2024_decouple_then_classify. Xinhang Wan, Jiyuan Liu 0003, Xinwang Liu 0002, Yi Wen 0001, Hao Yu 0017, Siwei Wang 0001, Shengju Yu, Tianjiao Wan, Jun Wang 0118, En Zhu |
ICML | 8 |
| 2024 | Higher-Order Vision-Language Alignment for Social Media PredictionabstractThe prediction task of social media popularity aims to automatically forecast the future popularity of the posts by leveraging vast amounts of social media data. This data encompasses diverse visual and textual content, including photos, categories, custom tags, temporal information, and geographical data. Existing methods have explored multiple feature types to enhance popularity prediction. Despite their success, visual and textual features-both crucial pieces of information-are often simply concatenated after extraction, ignoring the divergence between these two feature spaces. In this paper, we propose a method to project visual and language information into an aligned semantic representation, thereby uncovering intricate associations between these two modalities. Specifically, we leverage the BLIP-2 model to understand and generate visual description text that encapsulates the content of photos. Semantic embeddings are then extracted from all available visual and textual information. Additionally, we deeply exploit user-related behavior and characteristic information to extract features, uncovering hidden clues for post popularity prediction. Leveraging these improvements, we conduct extensive experiments to demonstrate the effectiveness of our proposed method. Mingsheng Tu, Tianjiao Wan, Qisheng Xu, Xinhao Jiang, Kele Xu, Cheng Yang 0004 |
ACM Multimedia | 2 |
| 2024 | Tracing Training Progress: Dynamic Influence Based Selection for Active LearningabstractActive learning (AL) aims to select highly informative data points from an unlabeled dataset for annotation, mitigating the need for extensive human labeling effort. However, classical AL methods heavily rely on human expertise to design the sampling strategy, inducing limited scalability and generalizability. Many efforts have sought to address this limitation by directly connecting sample selection with model performance improvement, typically through influence function. Nevertheless, these approaches often ignore the dynamic nature of model behavior during training optimization, despite empirical evidence highlights the importance of dynamic influence to track the sample contribution. This oversight can lead to suboptimal selection, hindering the generalizability of model. In this study, we explore the dynamic influence based data selection strategy by tracing the impact of unlabeled instances on model performance throughout the training process. Our theoretical analyses suggest that selecting samples with higher projected gradients along the accumulated optimization direction at each checkpoint leads to improved performance. Furthermore, to capture a wider range of training dynamics without incurring excessive computational or memory costs, we introduce an additional dynamic loss term designed to encapsulate more generalized training progress information. These insights are integrated into a universal and task-agnostic AL framework termed Dynamic Influence Scoring for Active Learning (DISAL). Comprehensive experiments across various tasks have demonstrated that DISAL significantly surpasses existing state-of-the-art AL methods, demonstrating its ability to facilitate more efficient and effective learning in different domains. Tianjiao Wan, Kele Xu, Long Lan, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001 |
ACM Multimedia | 1 |
| 2023 | Complementary Learning System Based Intrinsic Reward in Reinforcement LearningabstractDeep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learning systems utilizing short-term and long-term memories. Inspired by the fact that humans evaluate curiosity by comparing current observations with historical information, we propose a novel intrinsic reward, namely CLS-IR, which aims to address the problems caused by sparse extrinsic rewards. Specifically, we train a self-supervised predictive model with short-term and long-term memories via exponential moving averages. We employ the information gain between the two memories as the intrinsic reward, which does not incur additional training costs but leads to better exploration. To investigate the effectiveness of CLS-IR, we conduct extensive experimental evaluations; the results demonstrate that CLS-IR can achieve state-of-the-art performance on Atari games and DeepMind Control Suite. Zijian Gao, Kele Xu, Hongda Jia, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
ICASSP | 4 |
| 2023 | Bi-level Multi-Agent Actor-Critic Methods with ransformersabstractRecently, deep multi-agent reinforcement learning methods have witnessed great progress, including multi-agent actor-critic methods. However, it’s worth noticing there is a performance gap between multi-agent actor-critic methods and state-of-the-art value-based methods. In this paper, we investigate the causes and attribute inferior performance to issues of contribution-mismatch and indiscriminate guidance. To overcome these problems, we introduce a novel bi-level multi-agent actorcritic reinforcement learning approach with transformers, called BMT. Specifically, we propose a simple but efficient bi-level optimization mechanism to learn both global critic and agentspecific critic, thus jointly guiding the policy update. In addition, we adopt the transformer-based model as the policy network to decouple complicated relationships and generate flexible policy. BMT is also general enough to be plugged into any actor-critic multi-agent reinforcement learning approach, such as MAPPO, and equips it with strong expression. On multiple benchmarks including multi-agent particle environments and a challenging set of StarCraft II micromanagement tasks, large-scale empirical experiments demonstrate that BMT-based multi-agent reinforcement learning methods achieve superior performance over both state-of-the-art actor-critic and value-based approaches. Tianjiao Wan, Haibo Mi, Zijian Gao, Yuanzhao Zhai, Bo Ding 0001 |
JCC | 1 |
| 2023 | VTQAGen: BART-based Generative Model For Visual Text Question AnsweringabstractVisual Text Question Answering (VTQA) is a challenging task that requires answering questions pertaining to visual content by combining image understanding and language comprehension. The main objective is to develop models that can accurately provide relevant answers based on complementary information from both images and text, as well as the semantic meaning of the question. Despite ongoing efforts, the VTQA task presents several challenges, including multimedia alignment, multi-step cross-media reasoning, and handling open-ended questions. This paper introduces a novel generative framework called VTQAGen, which leverages a Multi- modal Attention Layer to combine image-text pairs and question inputs, as well as a BART-based model for reasoning and entity extraction from both images and text. The framework incorporates a step-based ensemble method to enhance model performance and generalization ability. VTQAGen utilizes an encoder-decoder generative model based on BART. Faster R-CNN is employed to extract visual regions of interest, while BART's encoder is modified to handle multi-modal interaction. The decoder stage utilizes the shift-predict approach and introduces step-based logits fusion to improve stability and accuracy. In the experiments, the proposed VTQAGen demonstrates superior performance on the testing set, securing second place in the ACM Multimedia Visual Text Question Answer Challenge. Haoru Chen, Tianjiao Wan, Zhimin Lin, Kele Xu, Jin Wang 0006, Huaimin Wang 0001 |
ACM Multimedia | 2 |
| 2022 | Uncertainty Estimation based Intrinsic Reward For Efficient Reinforcement LearningabstractFor reinforcement learning, the extrinsic reward is a core factor for the learning process which however can be very sparse or completely missing. In response, researchers have proposed the idea of intrinsic reward, such as encouraging the agent to visit novel states through prediction error. However, the deep prediction model can provide over-confident and miscalibrated predictions. To mitigate the impact of inaccurate prediction, previous research applied deep ensembles and achieved superior results, despite the increased computation and storage space. In this paper, inspired by the uncertainty estimation, we leverage Monte Carlo Dropout to generate intrinsic reward from the perspective of uncertainty estimation with the goal to decrease the demands for computing resources while retaining superior performance. Utilizing the simple yet effective approach, we conduct extensive experiments across a variety of benchmark environments. The experimental results suggest that our method provides a competitive performance in final score and is faster in running speed, while requiring much fewer computing resources and storage space. Tianjiao Wan, Peichang Shi, Bo Ding 0001, Zijian Gao |
JCC | 2 |