Weidong Tang

dblp:76/8363 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 50% Trustworthy machine learning · 13% Efficient and distributed learning · 11%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 74% Computing education · 26%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.922026
MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs · ACL (1) 2026
Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.012026
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
1.012026
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs · ACL (1) 2026
Computer vision › Vision and language
visual reasoning
1.012026
MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs · ACL (1) 2026
Machine learning › Representation and self-supervised learning
cognitive alignment
0.912025
Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
EA-Vit: Efficient Adaptation for Elastic Vision Transformer · ICCV 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
EA-Vit: Efficient Adaptation for Elastic Vision Transformer · ICCV 2025
Computational social science and digital humanities
social media analysis
0.912025
Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies · ACM Multimedia 2025
Machine learning › Trustworthy machine learning
robustness
0.312026
MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs · ACL (1) 2026
Computing education
AI education
0.312026
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models · AAAI 2026
Computer vision › Vision and language
multimodal understanding
0.312025
Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies · ACM Multimedia 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.312025
EA-Vit: Efficient Adaptation for Elastic Vision Transformer · ICCV 2025

Methods — techniques the papers use, named apart from their topics

knowledge-point reference-augmented generation · 2.0dynamic evaluation · 2.0optimal transport · 1.7multi-agent framework · 1.7diagnostic benchmark · 1.0causal intervention · 1.0benchmark construction · 1.0knowledge distillation · 0.9curriculum learning · 0.9NSGA-II · 0.9
YearPublicationVenuePosition
2026 MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education.
Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang
AAAI10
2026 MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs
abstract
Multimodal Large Language Models typically assume linguistic context invariably enhances visual understanding.We study this assumption in semantic adversarial scenarios, specifically magic tricks, where narration deliberately diverges from physical reality.We introduce MagicBench, a diagnostic benchmark of 402 videos for evaluating MLLMs under hierarchical linguistic interference, together with a Physical Constraint Set (PCS) protocol for assessing adherence to physical laws.Evaluation uncovers a Semantic Dependency Paradox: (1) Semantic anchoring: Entity nouns act as anchors aiding localization, paradoxically boosting performance despite false predicates.(2) Visual Agency Loss: In semantic vacuums, multimodal performance collapses 12.4% (p < 0.01) below the vision-only capability probe.This gap persists under symmetric prompting, suggesting a form of functional perception suppression in which autonomous visual search may be under-utilized in multimodal settings without linguistic triggers.Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck.Our findings suggest MLLMs function as language-guided passive observers, advocating for perceptuallyindependent architectures that decouple sensory agency from linguistic dominance.
Tang Da Huang, Weidong Tang, Wen Qi Xu, Xianpeng Guo
ACL (1)2
2026 GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs
abstract
Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Xinyan Wan, Zhiyuan Liang, Yang You 0001, Wangbo Zhao
ACL (1)1
2026 JCD: Just-in-Time Category Discovery Beyond Post-Hoc Alignment
Weidong Tang, Feifan Zhang
ICIC (18)1
2026 Superpixel-based scribble-level supervision diffusion for the interactive segmentation of multi-configuration chips
Weidong Tang, Shengfeng Chen, Yuanqiang Luo
Eng. Appl. Artif. Intell.3
2026 SLNALog: A Log Anomaly Detection Scheme Based on Swift Layer Normalization Attention Mechanism for Next-Generation Power Communication Networks
abstract
Log anomaly detection is a critical first line of defense for securing next-generation power communication networks against malicious attacks.serves as the initial line of defense for safeguarding the security of the next-generation power communication networks, which can protect it from attackers invasion damage. However, in industrial settingsin the industrial Internet domain, limitedthe scarcity of computational resources on edge devices result in longin devices leads to prolonged inference times for anomaly detection models, hindering the timely detection of anomalous log activities.impeding the prompt identification of unusual activities logged within these devices. To address these challenges, we propose SLNALog, an anomaly detection workflow centered around a Swift Layer Normalization Attention module.the aforementioned issues, the SLNALog anomaly detection workflow has been proposed. Its core comprises a Swift Layer Normalization Attention module. This module leveragesis based on linear attention to optimizeand optimizes the key-value interactions found in traditional attention mechanisms, thereby reducing the computational complexity of the detection process. This optimization reduces the complexity of log anomaly detection. As a result, the model’s receptive field for log data is expanded, and the efficiency of log anomaly detection is improved. ExperimentalThe experimental results on the HDFS and BGL datasets demonstrate the superiority of our approach.indicate that it achieves higher accuracy. SLNALog achieves higher accuracy, with F1-scores increasing by 0.08 and 0.04, respectively, while reducing detection time by 5.7% and 28.3%. The model provides an effective solution to enhance the cyber security of smart grids. Furthermore, the workflow incorporates an LLM-based log template analysis module and an Adapter-based model tuning module to enhance the model’s generalization in real-world scenarios. The proposed model provides an effective solution for enhancing the cybersecurity of smart grids.
Weidong Tang, Lintao Tan
IEEE Trans. Netw. Serv. Manag.2
2025 EA-Vit: Efficient Adaptation for Elastic Vision Transformer
abstract
Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT.
Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036
ICCV5
2025 Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies
abstract
In the digital age, brand meaning is increasingly shaped through user participation and content sharing on social media platforms. However, significant perceptual gaps often exist between official brand narratives and consumer interpretations. These multimodal and cognitively nuanced gaps are challenging to detect and model using traditional analytical methods. To address this, we propose a multi-agent framework that metaphorically models perception as an optical process-propagation, interference, and measurement---termed OPIM. We construct a novel dual-perspective dataset from representative social media platforms, integrating text and image content from both user-generated and official brand communications. We evaluate brand perception along six psychological dimensions. Experiments across 15 brands demonstrate that our framework effectively captures key perception gaps, particularly in sincerity, professionalism, and attractiveness. In contrast, materialism and sophistication exhibit higher alignment between brand messaging and consumer perception. Our framework enhances the cognitive alignment and multimodal interpretability of large language models, offering actionable insights for brand strategy and bridging computational modeling with human-centric understanding. The dataset will be available at https://github.com/htgan-ai/OPIM.
Haotian Gan, Yudong Li 0001, Weidong Tang
ACM Multimedia4
2025 Inference-Time Scaling for Visual AutoRegressive Modeling by Searching Representative Samples
Weidong Tang, Xinyan Wan
PRCV (4)1
2025 BHRAM: a knowledge graph embedding model based on bidirectional and heterogeneous relational attention mechanism
Wanqiu Li, Yuanbin Mo, Weidong Tang, Zhilin Zeng
Appl. Intell.4
2025 LDM-KGC: A low-dimensional knowledge graph completion model based on multi-head attention mechanism
Bingjie Qiu, Weidong Tang, Bicheng Liang, Danyang Cui, Haisheng Luo
Neurocomputing3
2024 A multi-level thresholding image segmentation method using hybrid Arithmetic Optimization and Harris Hawks Optimizer algorithms
Yanfeng Xue, Weidong Tang, Taybeh Salehnia
Expert Syst. Appl.4
2024 Robotic Process Automation Efficiency for Mobile App Testing: An Empirical Investigation
abstract
In today’s rapidly evolving software environment, the graphical user interface (GUI) plays a crucial role in providing intuitive, user-friendly interaction. However, traditional intrusive GUI testing methods often face challenges such as interrupting user workflows, requiring significant manual effort and insufficient test scenario coverage. Non-intrusive testing methods, such as Robotic Process Automation (RPA), offer a solution to validate GUI functionality without modifying the application’s code or affecting the user experience. RPA systems automate repetitive tasks by simulating user interactions, becoming valuable tools in GUI testing. However, challenges like limited computational resources, time constraints, or restricted exploration capabilities may limit the efficiency of individual RPA agents, thus restricting coverage and effectiveness. To address this issue, this study explores the performance of a single RPA agent versus an RPA cluster under different testing conditions, using three popular testing methods: Monkey, Stoat and Q-testing. Experimental results indicate that an RPA cluster outperforms a single RPA in GUI coverage and error detection, making a significant contribution to the field of non-intrusive GUI exploration testing. The findings of this study provide directions for future research to ensure the delivery of high-quality mobile applications.
Xiang Wang 0026, Weidong Tang, Jinhui Zhang 0002, Shaolei Wang, Peng Wang 0131
Int. J. Softw. Eng. Knowl. Eng.4
2023 A novel UAV path planning approach: Heuristic crossing search and rescue optimization algorithm
Wenjuan Zhou, Weidong Qin, Weidong Tang
Expert Syst. Appl.4
2021 An Optimized k-means Algorithm Based on Information Entropy
abstract
Abstract Clustering is a widely used technique in data mining applications and various pattern recognition applications, in which data objects are divided into groups. K-means algorithm is one of the most classical clustering algorithms. In this algorithm, the initial clustering centers are randomly selected, this results in unstable clustering results. To solve this problem, an optimized algorithm to select the initial centers is proposed. In the proposed algorithm, dispersion degree is defined, which is based on entropy. In the algorithm, all the objects are firstly grouped into a big cluster, and the object that has the maximum dispersion degree and the object that has the minimum dispersion degree are selected as the initial clustering centers from the initial big cluster. And then other objects in the biggest cluster are partitioned to the initial clusters to which the objects are nearest. The partition process will be repeated until the cluster number is equal to the specified value k. Finally, the partitioned k clusters and their cluster centers are applied to k-means algorithm as initial clusters and centers. Several experiments are conducted on real data sets to evaluate the proposed algorithm. The proposed algorithm is compared with traditional k-means algorithm and max-min distance clustering algorithm, and experimental results show that the improved k-means algorithm is stable in selecting initial clustering, because it can select unique initial clustering centers. The optimized algorithm’s effectiveness and feasibility are also verified by experiments, and the algorithm can reduce the times of iterations and has more stable clustering results and higher accuracy.
Beixian Zhang, Weidong Tang, Gangqiang Zhang
Comput. J.4
2020 Parameters Selection of Twin Support Vector Regression Based on Cloud Particle Swarm Optimization
Xiuxi Wei, Huajuan Huang, Weidong Tang
ICIC (3)3