Tianle Li

dblp:242/0053 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Expert-Inspired Multi-Agent Coordination for Multi-Objective Molecular Optimization
abstract
Multi-objective molecular optimization is a fundamental yet inherently challenging task in drug discovery, as it requires simultaneously optimizing multiple, often conflicting, molecular properties. Although recent deep learning methods have shown promise, they often lack objective-specific specialization and dynamic coordination, making them ineffective in handling competing objectives and difficult to scale in complex, high-dimensional molecular design tasks. Inspired by the division of labor among domain experts in medicinal chemistry, we propose MAMO, a multi-agent framework for molecular design that simulates expert collaboration. Each agent specializes in optimizing a single objective, and their interactions are orchestrated by a central scheduling module that dynamically reallocates tasks based on evaluation feedback. This coordination mechanism enables interpretable and goal-conditioned optimization while adaptively balancing conflicting objectives. Extensive experiments on benchmark datasets demonstrate that MAMO consistently achieves superior performance in both objective quality and Pareto diversity, particularly in scenarios with strong inter-objective conflict. Our results highlight the potential of multi-agent coordination strategies for scalable and conflict-aware molecular design.
Daojian Zeng, Tianle Li, Jiacai Yi, Lincheng Jiang, Tengfei Ma 0002, Xiangxiang Zeng
AAAI2
2026 BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
abstract
Encoder-free multimodal large language models (MLLMs) eliminate the need for a well-trained vision encoder by directly processing image tokens before the language model. While this approach reduces computational overhead and model complexity, it often requires large amounts of training data to effectively capture the visual knowledge typically encoded by vision models like CLIP. The absence of a vision encoder implies that the model is likely to rely on substantial data to learn the necessary visual-semantic alignments. In this work, we present BREEN, a data-efficient encoder-free multimodal architecture that mitigates this issue. BREEN leverages a learnable query and image experts to achieve comparable performance with significantly less training data. The learnable query, positioned between image and text tokens, is supervised by the output of a pretrained CLIP model to distill visual knowledge, bridging the gap between visual and textual modalities. Additionally, the image expert processes image tokens and learnable queries independently, improving efficiency and reducing interference with the LLM’s textual capabilities. BREEN achieves comparable performance to prior encoder-free state-of-the-art models like Mono-InternVL, using only 13 million text-image pairs in training—about one percent of the data required by existing methods. Our work highlights a promising direction for data-efficient encoder-free multimodal learning, offering an alternative to traditional encoder-based approaches.
Tianle Li, Yongming Rao, Winston Hu
WACV1
2026 Counterfactual Debiasing Heterogeneous Ability-Induced Exercise Indices Estimation for Cognitive Diagnosis
abstract
Cognitive Diagnosis is a critical task in computer-assisted education, aimed at assessing students' mastery of knowledge concepts and analyzing exercise indices. In fact, this direction has received a lot of research attention in the past few decades. However, the inherent heterogeneity in students' abilities introduces significant challenges to accurate exercise indices estimation, resulting biases that lead to inaccurate diagnostics within student groups and undermining the generalizability of exercise indices across diverse groups. To address these challenges, we propose a Counterfactual Adaptive-Debiasing Framework (CADF) for Cognitive Diagnosis, which employs a causal graph to model the intricate relationships among key variables influencing student performance and knowledge mastery. Specifically, by introducing exercise adjustment factors, we capture both the intrinsic attributes of exercises and their dynamic adaptability to individual students. Then, to disentangle the direct and indirect effects of these factors, we adopt a counterfactual inference approach to answer the critical question:How would the diagnostic feedback from a cognitive diagnosis model change if it were only directly influenced by exercise adjustment factors?This allows CADF to retain the beneficial indirect effects while neutralizing the direct effects that introduce bias, thereby achieving debiased exercise indices estimation. Finally, Extensive experiments on three real-world datasets demonstrate that CADF significantly reduces bias in exercise indices estimation and enhances the accuracy of diagnostic feedback.
Haiping Ma, Tianle Li, Changqian Wang, Siyu Song, Limiao Zhang, Xingyi Zhang 0001
IEEE Trans. Knowl. Data Eng.2
2025 How to Evaluate Reward Models for RLHF
abstract
We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-standard approach is to run a full RLHF training pipeline and directly probe downstream LLM performance. However, this process is prohibitively expensive. To address this, we build a predictive model of downstream LLM performance by evaluating the reward model on proxy tasks. These proxy tasks consist of a large-scale human preference and a verifiable correctness preference dataset, in which we measure 12 metrics across 12 domains. To investigate which reward model metrics are most correlated to gold-standard RLHF outcomes, we launch an end-to-end RLHF experiment on a large-scale crowd-sourced human preference platform to view real reward model downstream performance as ground truth. Ultimately, we compile our data and findings into Preference Proxy Evaluations (PPE), the first reward model benchmark explicitly linked to post-RLHF real-world human preference performance, which we opensource for public use and further development at https://github.com/lmarena/PPE.
Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, Anastasios Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph Gonzalez 0001, Ion Stoica
ICLR2
2025 AutoEval Done Right: Using Synthetic Data for Model Evaluation
abstract
The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data can be used to decrease the number of human annotations required for this purpose in a process called autoevaluation. We suggest efficient and statistically principled algorithms for this purpose that improve sample efficiency while remaining unbiased.
Pierre Boyeau, Anastasios Angelopoulos, Tianle Li, Nir Yosef, Jitendra Malik, Michael I. Jordan
ICML3
2025 Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations
abstract
Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and prompt-specific variations in model performance. To address this, we propose Prompt-to-Leaderboard (P2L), a method that produces leaderboards specific to a prompt or set of prompts. The core idea is to train an LLM taking natural language prompts as input to output a vector of Bradley-Terry coefficients which are then used to predict the human preference vote. The resulting prompt-dependent leaderboards allow for unsupervised task-specific evaluation, optimal routing of queries to models, personalization, and automated evaluation of model strengths and weaknesses. Data from Chatbot Arena suggest that P2L better captures the nuanced landscape of language model performance than the averaged leaderboard. Furthermore, our findings suggest that P2L’s ability to produce prompt-specific evaluations follows a power law scaling similar to that observed in LLMs themselves. In January 2025, the router we trained based on this methodology achieved the #1 spot on the Chatbot Arena leaderboard. Our code is available at this GitHub link: https://github.com/lmarena/p2l.
Evan Frick, Connor Chen, Joseph Tennyson, Tianle Li, Wei-Lin Chiang, Anastasios Angelopoulos, Ion Stoica
ICML4
2025 From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline
abstract
The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned benchmarks is expensive and time-consuming. To address this, we introduce BenchBuilder, an automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop. We apply BenchBuilder to datasets such as Chatbot Arena and WildChat-1M, extracting challenging prompts and utilizing LLM-as-a-Judge for automatic model evaluation. To validate benchmark quality, we propose new metrics to measure a benchmark’s alignment with human preferences and ability to separate models. We release Arena-Hard-Auto, a benchmark consisting 500 challenging prompts curated by BenchBuilder. Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. Our work sets a new framework for the scalable curation of automated benchmarks from extensive data.
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu 0002, Banghua Zhu, Joseph Gonzalez 0001, Ion Stoica
ICML1
2025 FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios
abstract
Federated Learning (FL) enables decentralized model training while preserving data privacy. Despite its benefits, FL faces challenges with non-identically distributed (non-IID) data, especially in long-tailed scenarios with imbalanced class samples. Momentum-based FL methods, often used to accelerate FL convergence, struggle with these distributions, resulting in biased models and making FL hard to converge. To understand this challenge, we conduct extensive investigations into this phenomenon, accompanied by a layer-wise analysis of neural network behavior. Based on these insights, we propose FedWCM, a method that dynamically adjusts momentum using global and per-round data to correct directional biases introduced by long-tailed distributions. Extensive experiments show that FedWCM resolves non-convergence issues and outperforms existing methods, enhancing FL’s efficiency and effectiveness in handling client heterogeneity and data imbalance.
Tianle Li, Yongzhi Huang 0002, Linshan Jiang, Qipeng Xie, Chang Liu 0093, Wenfeng Du, Lu Wang 0002, Kaishun Wu
ICPP1
2025 FedWMSAM: Fast and Flat Federated Learning via Weighted Momentum and Sharpness-Aware Minimization
abstract
In federated learning (FL), models must \emph{converge quickly} under tight communication budgets while \emph{generalizing} across non-IID client distributions. These twin requirements have naturally led to two widely used techniques: client/server \emph{momentum} to accelerate progress, and \emph{sharpness-aware minimization} (SAM) to prefer flat solutions. However, simply combining momentum and SAM leaves two structural issues unresolved in non-IID FL. We identify and formalize two failure modes: \emph{local–global curvature misalignment} (local SAM directions need not reflect the global loss geometry) and \emph{momentum-echo oscillation} (late-stage instability caused by accumulated momentum). To our knowledge, these failure modes have not been jointly articulated and addressed in the FL literature. We propose \textbf{FedWMSAM} to address both failure modes. First, we construct a momentum-guided global perturbation from server-aggregated momentum to align clients' SAM directions with the global descent geometry, enabling a \emph{single-backprop} SAM approximation that preserves efficiency. Second, we couple momentum and SAM via a cosine-similarity adaptive rule, yielding an early-momentum, late-SAM two-phase training schedule. We provide a non-IID convergence bound that \emph{explicitly models the perturbation-induced variance} $\sigma_\rho^2=\sigma^2+(L\rho)^2$ and its dependence on $(S,K,R,N)$ on the theory side. We conduct extensive experiments on multiple datasets and model architectures, and the results validate the effectiveness, adaptability, and robustness of our method, demonstrating its superiority in addressing the optimization challenges of Federated Learning. Our code is available at \url{https://github.com/Li-Tian-Le/NeurlPS_FedWMSAM}.
Tianle Li, Yongzhi Huang 0002, Linshan Jiang, Chang Liu 0093, Qipeng Xie, Wenfeng Du, Lu Wang 0002, Kaishun Wu
NeurIPS1
2025 On the Robustness of Transformers against Context Hijacking for Linear Classification
abstract
Transformer-based Large Language Models (LLMs) have demonstrated powerful in-context learning capabilities. However, their predictions can be disrupted by factually correct context, a phenomenon known as context hijacking, revealing a significant robustness issue. To understand this phenomenon theoretically, we explore an in-context linear classification problem based on recent advances in linear transformers. In our setup, context tokens are designed as factually correct query-answer pairs, where the queries are similar to the final query but have opposite labels. Then, we develop a general theoretical analysis on the robustness of the linear transformers, which is formulated as a function of the model depth, training context lengths, and number of hijacking context tokens. A key finding is that a well-trained deeper transformer can achieve higher robustness, which aligns with empirical observations. We show that this improvement arises because deeper layers enable more fine-grained optimization steps, effectively mitigating interference from context hijacking. This is also well supported by our numerical and real-world experiments. Our findings provide theoretical insights into the benefits of deeper architectures and contribute to enhancing the understanding of transformer architectures.
Tianle Li, Xingwu Chen, Yuan Cao 0006, Difan Zou
NeurIPS1
2025 Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
abstract
While large language models (LLMs) demonstrate strong reasoning capabilities utilizing reinforcement learning (RL) with verifiable reward, whether large vision-language models (VLMs) can directly inherit such capabilities through similar post-training strategies remains underexplored. In this work, we conduct a systematic compositional probing study to evaluate whether current VLMs trained with RL or other post-training strategies can compose capabilities across modalities or tasks under out-of-distribution conditions. We design a suite of diagnostic tasks that train models on unimodal tasks or isolated reasoning skills, and evaluate them on multimodal, compositional variants requiring skill integration. Through comparisons between supervised fine-tuning (SFT) and RL-trained models, we identify three key findings: (1) RL-trained models consistently outperform SFT on compositional generalization, demonstrating better integration of learned skills; (2) although VLMs achieve strong performance on individual tasks, they struggle to generalize compositionally under cross-modal and cross-task scenarios, revealing a significant gap in current training strategies; (3) enforcing models to explicitly describe visual content before reasoning (e.g., caption-before-thinking), along with rewarding progressive vision-to-text grounding, yields notable gains. It highlights two essential ingredients for improving compositionality in VLMs: visual-to-text alignment and accurate visual grounding. Our findings shed light on the current limitations of RL-based reasoning VLM training and provide actionable insights toward building models that reason compositionally across modalities and tasks.
Tianle Li, Yongming Rao
NeurIPS1
2025 Optimizing Spectrum Sharing in Low-Altitude Intelligent Networks: A Model-Based Reinforcement Learning Approach
abstract
The Low-Altitude Intelligent Network (LAIN) leverages cutting-edge Artificial Intelligence (AI) technologies to enable fast and reliable communications for Uncrewed Aerial Vehicles (UAVs), supporting the rapid growth of various low-altitude economy applications. However, as UAV deployments increase, the scarcity of spectrum resources has emerged as a critical bottleneck that severely limits the quality of service in LAIN. While Reinforcement Learning (RL) has shown promise in optimizing spectrum sharing and improving spectral efficiency of UAVs, existing RL-based methods suffer from high training costs and require extensive interactions with the environment to converge to near-optimal solutions. To address these challenges, we propose a novel model-based RL framework for dynamic spectrum sharing of UAVs in LAIN. Our approach employs a world model to learn compact state representations and reconstruct environment dynamics, significantly enhancing the sample efficiency in complex scenarios. Extensive simulations demonstrate that our approach drastically reduces the number of environmental interactions required for UAVs to achieve near-optimal spectrum allocations, lowering training costs while maintaining high performance.
Tianle Li, Peixi Peng
TrustCom1
2024 ImagenHub: Standardizing the evaluation of conditional image generation models
abstract
Recently, a myriad of conditional image generation and editing models have been developed to serve different downstream tasks, including text-to-image generation, text-guided image editing, subject-driven image generation, control-guided image generation, etc. However, we observe huge inconsistencies in experimental conditions: datasets, inference, and evaluation metrics -- render fair comparisons difficult. This paper proposes ImagenHub, which is a one-stop library to standardize the inference and evaluation of all the conditional image generation models. Firstly, we define seven prominent tasks and curate high-quality evaluation datasets for them. Secondly, we built a unified inference pipeline to ensure fair comparison. Thirdly, we design two human evaluation scores, i.e. Semantic Consistency and Perceptual Quality, along with comprehensive guidelines to evaluate generated images. We train expert raters to evaluate the model outputs based on the proposed metrics. Our human evaluation achieves a high inter-worker agreement of Krippendorff’s alpha on 76\% models with a value higher than 0.4. We comprehensively evaluated a total of around 30 models and observed three key takeaways: (1) the existing models’ performance is generally unsatisfying except for Text-guided Image Generation and Subject-driven Image Generation, with 74\% models achieving an overall score lower than 0.5. (2) we examined the claims from published papers and found 83\% of them hold with a few exceptions. (3) None of the existing automatic metrics has a Spearman's correlation higher than 0.2 except subject-driven image generation. Moving forward, we will continue our efforts to evaluate newly published models and update our leaderboard to keep track of the progress in conditional image generation.
Max Ku, Tianle Li, Kai Zhang 0033, Wenwen Zhuang, Wenhu Chen
ICLR2
2024 LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
abstract
Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng 0007, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang 0001, Zhuohan Li 0001, Zi Lin, Eric P. Xing, Joseph Gonzalez 0001, Ion Stoica, Hao Zhang 0025
ICLR4
2024 Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
abstract
Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing. The platform has been operational for several months, amassing over 240K votes. This paper describes the platform, analyzes the data we have collected so far, and explains the tried-and-true statistical methods we are using for efficient and accurate evaluation and ranking of models. We confirm that the crowdsourced questions are sufficiently diverse and discriminating and that the crowd-sourced human votes are in good agreement with those of expert raters. These analyses collectively establish a robust foundation for the credibility of Chatbot Arena. Because of its unique value and openness, Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. The platform is publicly available at https://chat.lmsys.org.
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng 0007, Anastasios Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang 0025, Michael I. Jordan, Joseph Gonzalez 0001, Ion Stoica
ICML5
2024 GenAI Arena: An Open Evaluation Platform for Generative Models
abstract
Generative AI has made remarkable strides to revolutionize fields such as image and video generation. These advancements are driven by innovative algorithms, architecture, and data. However, the rapid proliferation of generative models has highlighted a critical gap: the absence of trustworthy evaluation metrics. Current automatic assessments such as FID, CLIP, FVD, etc often fail to capture the nuanced quality and user satisfaction associated with generative outputs. This paper proposes an open platform GenAI-Arena to evaluate different image and video generative models, where users can actively participate in evaluating these models. By leveraging collective user feedback and votes, GenAI-Arena aims to provide a more democratic and accurate measure of model performance. It covers three tasks of text-to-image generation, text-to-video generation, and image editing respectively. Currently, we cover a total of 35 open-source generative models. GenAI-Arena has been operating for seven months, amassing over 9000 votes from the community. We describe our platform, analyze the data, and explain the statistical methods for ranking the models. To further promote the research in building model-based evaluation metrics, we release a cleaned version of our preference data for the three tasks, namely GenAI-Bench. We prompt the existing multi-modal models like Gemini, and GPT-4o to mimic human voting. We compute the accuracy by comparing the model voting with the human voting to understand their judging abilities. Our results show existing multimodal models are still lagging in assessing the generated visual content, even the best model GPT-4o only achieves an average accuracy of $49.19\%$ across the three generative tasks. Open-source MLLMs perform even worse due to the lack of instruction-following and reasoning ability in complex vision scenarios.
Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, Wenhu Chen
NeurIPS3
2024 MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
abstract
In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates part of the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16\% to 33\% compared to MMLU, but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5\% in MMLU to just 2\% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is more discriminative benchmark to better track progress in the field.
Yubo Wang 0019, Xueguang Ma, Ge Zhang 0009, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang 0068, Alex Zhuang, Rongqi Fan, Xiang Yue, Wenhu Chen
NeurIPS11
2023 Few-shot In-context Learning on Knowledge Base Question Answering
abstract
Question answering over knowledge bases is considered a difficult problem due to the challenge of generalizing to a wide variety of possible natural language questions.Additionally, the heterogeneity of knowledge base schema items between different knowledge bases often necessitates specialized training for different knowledge base question-answering (KBQA) datasets.To handle questions over diverse KBQA datasets with a unified trainingfree framework, we propose KB-BINDER, which for the first time enables few-shot incontext learning over KBQA tasks.Firstly, KB-BINDER leverages large language models like Codex to generate logical forms as the draft for a specific question by imitating a few demonstrations.Secondly, KB-BINDER grounds on the knowledge base to bind the generated draft to an executable one with BM25 score matching.The experimental results on four public heterogeneous KBQA datasets show that KB-BINDER can achieve a strong performance with only a few in-context demonstrations.Especially on GraphQA and 3-hop MetaQA, KB-BINDER can even outperform the state-of-the-art trained models.On GrailQA and WebQSP, our model is also on par with other fully-trained models.We believe KB-BINDER can serve as an important baseline for future research.Our code is available at
Tianle Li, Xueguang Ma, Alex Zhuang, Yu Gu 0016, Yu Su 0001, Wenhu Chen
ACL (1)1
2023 Rethinking graph anomaly detection: A self-supervised Group Discrimination paradigm with Structure-Aware
abstract
Structural anomalies are the core problem in graph anomaly detection. However, the current mainstream self-supervised graph anomaly detection models do not directly model structural anomalies and their expensive time consumption limits the efficiency of graph anomaly detection. For this reason, we rethink graph anomaly detection and propose a self-supervised Group Discrimination paradigm with Structure-Aware (GDSA). Our model can be explicitly aware of the graph topology changes by multi-view structure disturbance. Moreover, GDSA transforms graph anomaly detection into discriminating the scalar summaries of positive and negative group nodes. The results of extensive experiments on four benchmark datasets show that GDSA outperforms current state-of-the-art methods, with the most significant AUC performance improvement of 28.7%. Notably, in scalability testing on a large-scale dataset, the training time and testing time of GDSA are 1181.0× and 5064.7× faster than the baseline, respectively, with 61.9% savings in memory usage.
Junyi Yan, Enguang Zuo, Chen Chen 0078, Tianle Li, Xiaoyi Lv
ICME6
2023 A Masked Attention Network with Query Sparsity Measurement for Time Series Anomaly Detection
abstract
Time series aomaly detection has been widely studied in recent years. Previous research focuses on point-wise features and pairwise associations for feature learning or designed anomaly scores based on prior knowledge. However, these methods cannot fully learn the intricate abnormal dynamic information and can only identify a limited class of anomalies. We propose a Masked Attention Network with Query Sparsity Measurement (MAN-QSM) to address the above challenges. This model uses two kinds of prior knowledge to fully exploit the differences between normal and abnormal points from two perspectives: pairwise association and sequence-level information. We designs the anomaly mask mechanism to collaborate with the training strategy to amplify the difference between normal and abnormal points. In experiments, we compare the model with classical methods, reconstruction-based models, autoregressive-based models, and state-of-the-art models, and the MAN-QSM achieves state-of-the-art results on SMD, PSM, and MSL datasets with an average of 16% reduction in error rate.
Enguang Zuo, Chen Chen 0078, Junyi Yan, Tianle Li, Xiaoyi Lv
ICME6