EDBT 2026 Demo / reviewers in the wild / expert
Wen Shen 0002
dblp:55/8186-2
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-4210-5447ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual EnhancementabstractLarge Vision-Language Models (LVLMs) have achieved remarkable performance on diverse vision-language tasks.However, LVLMs still suffer from hallucinations, generating text that contradicts the visual input.Existing research has primarily focused on mitigating object hallucinations, but often overlooks more complex relation hallucinations, particularly action relations involving interactions between objects.In this study, we empirically observe that the primary cause of action-relation hallucinations in LVLMs is the insufficient attention allocated to visual information.Thus, we propose a framework to locate action-relevant image regions and enhance the LVLM's attention to those regions.Specifically, we define the Action-Relation Sensitivity (ARS) score to identify attention heads that are most sensitive to actionrelation changes, thereby localizing actionrelevant image regions that contain key visual cues.Then, we propose the Relation-aware Visual Enhancement (RVE) method to enhance the LVLM's attention to these action-relevant image regions.Extensive experiments demonstrate that, compared to existing baselines, our method achieves superior performance in mitigating action-relation hallucinations with negligible additional inference cost.Furthermore, it effectively generalizes to spatial-relation hallucinations and object hallucinations.The code is available at https://github.com/ LandHqzx/ARS-RVE. Zhenxin Qin, Qingzhuo Wang, Ruiyang Qin, Zhihua Wei 0001, Wen Shen 0002 |
ACL (1) | 6 |
| 2025 | A Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via InteractionsabstractRecently, many self-supervised pre-training methods have been proposed to improve the performance of deep neural networks (DNNs) for 3D point clouds processing. However, the common mechanism underlying the effectiveness of different pre-training methods remains unclear. In this paper, we use game-theoretic interactions as a unified approach to explore the common mechanism of pre-training methods. Specifically, we decompose the output score of a DNN into the sum of numerous effects of interactions, with each interaction representing a distinct 3D substructure of the input point cloud. Based on the decomposed interactions, we draw the following conclusions. (1) The common mechanism across different pre-training methods is that they enhance the strength of high-order interactions encoded by DNNs, which represent complex and global 3D structures, while reducing the strength of low-order interactions, which represent simple and local 3D structures. (2) Sufficient pre-training and adequate fine-tuning data for downstream tasks further reinforce the mechanism described above. (3) Pre-training methods carry a potential risk of reducing the transferability of features encoded by DNNs. Inspired by the observed common mechanism, we propose a new method to directly enhance the strength of high-order interactions and reduce the strength of low-order interactions encoded by DNNs, improving performance without the need for pre-training on large-scale datasets. Experiments show that our method achieves performance comparable to traditional pre-training methods. Jian Ruan, Fanghao Wu, Yuchi Chen, Zhihua Wei 0001, Wen Shen 0002 |
CVPR | 6 |
| 2025 | Leveraging Debiased Cross-Modal Attention Maps and Code-Based Reasoning for Zero-Shot Referring Expression Comprehension
Wen Shen 0002, Zhihua Wei 0001, Hongyun Zhang 0001 |
ICCV | 2 |
| 2025 | StyleCSR: Style-controllable Sequential Recommendation with Prompt TuningabstractControllability is a key factor in building user-trustworthy recommendation systems. However, most existing studies lack consideration for controllability, relying solely on users’ historical interaction data to continuously recommend items that align with their past preferences. This approach may isolate users from the outside world, leading to the "filter bubble" phenomenon. As a result, users may still feel dissatisfied, even when the recommendation accuracy is high. In this work, we propose a user-friendly, style-controllable sequential recommendation framework with prompt tuning (StyleCSR). This framework enables users to provide instructions for adjusting item features, such as popularity and similarity, to control the distribution of recommendation results. Specifically, we map user instructions into prompts and fuse them with user historical behavior sequences. To enhance the controllability of recommendation results while maintaining high recommendation performance, we design an interest alignment module to retain users’ original interests and an instruction discrimination module to emphasize the role of user instructions. We conduct extensive experiments on multiple datasets and various types of pre-trained sequential recommendation models. The results validate that our StyleCSR can be applied to different types of pre-trained models, significantly enhancing the controllability of recommendation results while maintaining high recommendation performance. Leilei Wen, Qi Shen 0001, Shixuan Zhu, Zhihua Wei 0001, Wen Shen 0002 |
IJCNN | 5 |
| 2025 | Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsabstractIn recent years, large language models (LLMs) have made significant advancements in arithmetic reasoning.
However, the internal mechanism of how LLMs solve arithmetic problems remains unclear.
In this paper, we propose explaining arithmetic reasoning in LLMs using game-theoretic interactions.
Specifically, we disentangle the output score of the LLM into numerous interactions between the input words.
We quantify different types of interactions encoded by LLMs during forward propagation to explore the internal mechanism of LLMs for solving arithmetic problems.
We find that (1) the internal mechanism of LLMs for solving simple one-operator arithmetic problems is their capability to encode operand-operator interactions and high-order interactions from input samples.
Additionally, we find that LLMs with weak one-operator arithmetic capabilities focus more on background interactions.
(2) The internal mechanism of LLMs for solving relatively complex two-operator arithmetic problems is their capability to encode operator interactions and operand interactions from input samples.
(3) We explain the task-specific nature of the LoRA method from the perspective of interactions. Leilei Wen, Liwei Zheng, Zhihua Wei 0001, Wen Shen 0002 |
NeurIPS | 6 |
| 2025 | Infusing Multi-Hop Medical Knowledge Into Smaller Language Models for Biomedical Question AnsweringabstractMedQA-USMLE is a challenging biomedical question answering (BQA) task, as its questions typically involve multi-hop reasoning. To solve this task, BQA systems should possess not only extensive medical professional knowledge but also strong medical reasoning capabilities. While state-of-the-art larger language models, such as Med-PaLM 2, have overcome this challenge, smaller language models (SLMs) still struggle with it. To bridge this gap, we introduces a multi-hop medical knowledge infusion (MHMKI) procedure to endow SLMs with medical reasoning capabilities. Specifically, we categorize MedQA-USMLE questions into distinct reasoning types, then tailor pre-training instances for each type of questions using the semi-structured information and hyperlinks of Wikipedia articles. To enable SLMs to efficiently capture the multi-hop knowledge contained in these instances, we design a reasoning chain masked language model to further pre-train BERT models. Moreover, we convert the pre-training instances into a composite question answering dataset for intermediate fine-tuning of GPT models. We evaluate MHMKI on six SLMs across five datasets spanning three BQA tasks. The results demonstrate that MHMKI consistently improves SLMs' performance, particularly on tasks requiring substantial medical reasoning. For instance, the accuracy of MedQA-USMLE shows a significant increase of 5.3% on average. Jing Chen 0043, Zhihua Wei 0001, Wen Shen 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Clarifying the Behavior and the Difficulty of Adversarial TrainingabstractAdversarial training is usually difficult to optimize. This paper provides conceptual and analytic insights into the difficulty of adversarial training via a simple theoretical study, where we derive an approximate dynamics of a recursive multi-step attack in a simple setting. Despite the simplicity of our theory, it still reveals verifiable predictions about various phenomena in adversarial training under real-world settings. First, compared to vanilla training, adversarial training is more likely to boost the influence of input samples with large gradient norms in an exponential manner. Besides, adversarial training also strengthens the influence of the Hessian matrix of the loss w.r.t. network parameters, which is more likely to make network parameters oscillate and boosts the difficulty of adversarial training. Xu Cheng 0005, Hao Zhang 0063, Wen Shen 0002, Quanshi Zhang |
AAAI | 4 |
| 2024 | Batch Normalization Is Blind to the First and Second Derivatives of the LossabstractWe prove that when we do the Taylor series expansion of the loss function, the BN operation will block the influence of the first-order term and most influence of the second-order term of the loss. We also find that such a problem is caused by the standardization phase of the BN operation. We believe that proving the blocking of certain loss terms provides an analytic perspective for potential detects of a deep model with BN operations, although the blocking problem is not fully equivalent to significant damages in all tasks on benchmark datasets. Experiments show that the BN operation significantly affects feature representations in specific tasks. Zhanpeng Zhou, Wen Shen 0002, Huixin Chen, Ling Tang 0002, Yuefeng Chen, Quanshi Zhang |
AAAI | 2 |
| 2024 | Explaining Generalization Power of a DNN Using Interactive ConceptsabstractThis paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well explained by a small number of interactions between input variables. In this way, to some extent, we can consider such interactions as interactive concepts encoded by the DNN. Therefore, in this paper, we derive an analytic explanation of inconsistency of concepts of different complexities. This may shed new lights on using the generalization power of concepts to explain the generalization power of the entire DNN. Besides, we discover that the DNN with stronger generalization power usually learns simple concepts more quickly and encodes fewer complex concepts. We also discover the detouring dynamics of learning complex concepts, which explains both the high learning difficulty and the low generalization power of complex concepts. The code will be released when the paper is accepted. Huilin Zhou, Hao Zhang 0063, Huiqi Deng, Dongrui Liu, Wen Shen 0002, Shih-Han Chan, Quanshi Zhang |
AAAI | 5 |
| 2024 | Where We Have Arrived in Proving the Emergence of Sparse Interaction Primitives in DNNsabstractThis study aims to prove the emergence of symbolic concepts (or more precisely, sparse primitive inference patterns) in well-trained deep neural networks (DNNs). Specifically, we prove the following three conditions for the emergence. (i) The high-order derivatives of the network output with respect to the input variables are all zero. (ii) The DNN can be used on occluded samples, and when the input sample is less occluded, the DNN will yield higher confidence. (iii) The confidence of the DNN does not significantly degrade on occluded samples. These conditions are quite common, and we prove that under these conditions, the DNN will only encode a relatively small number of sparse interactions between input variables. Moreover, we can consider such interactions as symbolic primitive inference patterns encoded by a DNN, because we show that inference scores of the DNN on an exponentially large number of randomly masked samples can always be well mimicked by numerical effects of just a few interactions. Qihan Ren, Jiayang Gao, Wen Shen 0002, Quanshi Zhang |
ICLR | 3 |
| 2024 | Interpretable Rotation-Equivariant Quaternion Neural Networks for 3D Point Cloud ProcessingabstractThis study proposes a set of generic rules to revise existing neural networks for 3D point cloud processing to rotation-equivariant quaternion neural networks (REQNNs), in order to make feature representations of neural networks to be rotation-equivariant and permutation-invariant. Rotation equivariance of features means that the feature computed on a rotated input point cloud is the same as applying the same rotation transformation to the feature computed on the original input point cloud. We find that the rotation-equivariance of features is naturally satisfied, if a neural network uses quaternion features. Interestingly, we prove that such a network revision also makes gradients of features in the REQNN to be rotation-equivariant w.r.t. inputs, and the training of the REQNN to be rotation-invariant w.r.t. inputs. Besides, permutation-invariance examines whether the intermediate-layer features are invariant, when we reorder input points. We also evaluate the stability of knowledge representations of REQNNs, and the robustness of REQNNs to adversarial rotation attacks. Experiments have shown that REQNNs outperform traditional neural networks in both terms of classification accuracy and robustness on rotated testing samples. Wen Shen 0002, Zhihua Wei 0001, Qihan Ren, Shikun Huang, Quanshi Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Defects of Convolutional Decoder Networks in Frequency RepresentationabstractIn this paper, we prove the representation defects of a cascaded convolutional decoder network, considering the capacity of representing different frequency components of an input sample. We conduct the discrete Fourier transform on each channel of the feature map in an intermediate layer of the decoder network. Then, we extend the 2D circular convolution theorem to represent the forward and backward propagations through convolutional layers in the frequency domain. Based on this, we prove three defects in representing feature spectrums. First, we prove that the convolution operation, the zero-padding operation, and a set of other settings all make a convolutional decoder network more likely to weaken high-frequency components. Second, we prove that the upsampling operation generates a feature spectrum, in which strong signals repetitively appear at certain frequencies. Third, we prove that if the frequency components in the input sample and frequency components in the target output for regression have a small shift, then the decoder usually cannot be effectively learned. Ling Tang 0002, Wen Shen 0002, Zhanpeng Zhou, Yuefeng Chen, Quanshi Zhang |
ICML | 2 |
| 2021 | Verifiability and Predictability: Interpreting Utilities of Network Architectures for Point Cloud ProcessingabstractIn this paper, we diagnose deep neural networks for 3D point cloud processing to explore utilities of different intermediate-layer network architectures. We propose a number of hypotheses on the effects of specific intermediate-layer network architectures on the representation capacity of DNNs. In order to prove the hypotheses, we design five metrics to diagnose various types of DNNs from the following perspectives, information discarding, information concentration, rotation robustness, adversarial robustness, and neighborhood inconsistency. We conduct comparative studies based on such metrics to verify the hypotheses. We further use the verified hypotheses to revise intermediate-layer architectures of existing DNNs and improve their utilities. Experiments demonstrate the effectiveness of our method. The code will be released when this paper is accepted. Wen Shen 0002, Zhihua Wei 0001, Shikun Huang, Panyue Chen, Quanshi Zhang |
CVPR | 1 |
| 2021 | Interpretable Compositional Convolutional Neural NetworksabstractThis paper proposes a method to modify a traditional convolutional neural network (CNN) into an interpretable compositional CNN, in order to learn filters that encode meaningful visual patterns in intermediate convolutional layers. In a compositional CNN, each filter is supposed to consistently represent a specific compositional object part or image region with a clear meaning. The compositional CNN learns from image labels for classification without any annotations of parts or regions for supervision. Our method can be broadly applied to different types of CNNs. Experiments have demonstrated the effectiveness of our method. The code will be released when the paper is accepted. Wen Shen 0002, Zhihua Wei 0001, Shikun Huang, Quanshi Zhang |
IJCAI | 1 |
| 2021 | Interpreting Representation Quality of DNNs for 3D Point Cloud ProcessingabstractIn this paper, we evaluate the quality of knowledge representations encoded in deep neural networks (DNNs) for 3D point cloud processing. We propose a method to disentangle the overall model vulnerability into the sensitivity to the rotation, the translation, the scale, and local 3D structures. Besides, we also propose metrics to evaluate the spatial smoothness of encoding 3D structures, and the representation complexity of the DNN. Based on such analysis, experiments expose representation problems with classic DNNs, and explain the utility of the adversarial training. The code will be released when this paper is accepted. Wen Shen 0002, Qihan Ren, Dongrui Liu, Quanshi Zhang |
NeurIPS | 1 |
| 2020 | 3D-Rotation-Equivariant Quaternion Neural Networks
Wen Shen 0002, Shikun Huang, Zhihua Wei 0001, Quanshi Zhang |
ECCV (20) | 1 |
| 2020 | Improved general attribute reduction algorithms
Baizhen Li, Zhihua Wei 0001, Duoqian Miao 0001, Nan Zhang 0041, Wen Shen 0002, Hongyun Zhang 0001 |
Inf. Sci. | 5 |
| 2020 | Three-way decisions based blocking reduction models in hierarchical classification
Wen Shen 0002, Zhihua Wei 0001, Qianwen Li, Hongyun Zhang 0001, Duoqian Miao 0001 |
Inf. Sci. | 1 |
| 2019 | A Dustbin Category Based Feedback Incremental Learning Strategy for Hierarchical Image Classification
Wen Shen 0002, Qianwen Li, Zhihua Wei 0001 |
PRCV (1) | 2 |
| 2019 | A self-adaptive cascade ConvNets model based on label relation mining
Zhihua Wei 0001, Wen Shen 0002, Cairong Zhao, Duoqian Miao 0001 |
Neurocomputing | 2 |
| 2015 | Multiple Granular Analysis of TCM Data with Applications on Diagnosis of Hepatitis BabstractThe objectiveness of Traditional Chinese Medicine (TCM) limits its further development and generalization. Big data provide the golden opportunity for TCM quantization. The main purpose of this paper is to build a bridge between data analysis and clinical experience and provide experimental support for TCM experience. Taking the Hepatitis B disease data as experimental subject, we propose a framework for mining latent relations between features and disease categories based on Granular Computing theory. That is, kmeans clustering and correlation analysis is adopted to analyze the intra-relationship of disease stages and relationship between clinical symptoms and stages respectively. Algorithm based on Latent Dirichlet Allocation model is proposed to mining the mapping relationships of the three layers: clinical symptoms, stages of Hepatitis B and their middle layer syndromes. Experimental results indicate that the results of the data analysis are consistent with the clinical experience of TCM. It is proved that the diagnose based on syndromes is the scientific results of manually mining large amount of history data. Our study is a useful attempt that using data mining techniques to make TCM quantifiable and objective. Wen Shen 0002, Zhihua Wei 0001, Yunyi Li |
SMC | 1 |