Mengxue Zhang

dblp:195/8218 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Theoretical Convergence Analysis and Initialization Comparisons of Deep Soft-Thresholding Networks
abstract
Soft-thresholding (ST) has been widely used in deep neural networks. Its fundamental network structure is a deep soft-thresholding fully connected network (ST-FCN). However, training deep ST-FCN to achieve convergence remains time-consuming or even encounters gradient explosion, in part because the convergence behavior is not fully understood. To address this issue, this article proves the relationship between the convergence of deep ST-FCN and the values of network weights and biases. Theoretical analysis shows that, as the number of network layers approaches infinity, deep ST-FCN converges when the network weights tend to an identity matrix, while the biases tend to zero. Following this guidance, we initialize the network weights as the identity matrix, compare it with other representative initialization methods (Gaussian, He, LeCun, Xavier, and Uniform), and quantify their effects on network convergence. Extensive results on a synthetic spectrum dataset and real-world datasets (MNIST and CIFAR-10) demonstrate that initializing the weights to the identity matrix and the bias to zero leads to fast and stable convergence. These conclusions are further supported by additional experiments and statistical analysis on deeper ST networks (with more than ten layers) and other representative architectures (DenseNet-161, ResNet-152, and VGG-19), and more challenging benchmarks (CIFAR-100, STL-10, and Tiny ImageNet). This work provides a theoretical foundation for understanding the convergence of ST neural networks. Furthermore, convergence theory analysis for deep recurrent neural networks (RNNs) with ST is deduced.
Chunyan Xiong, Mengxue Zhang, Qingrui Cai, Zhong Chen 0005, Xiaobo Qu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
abstract
Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded harmfulness levels, and coverage of jailbreak-style attacks. We introduce CARES (Clinical Adversarial Robustness and Evaluation of Safety), a benchmark for evaluating LLM safety in healthcare. CARES includes over 18,000 prompts spanning eight medical safety principles, four harm levels, and four prompting styles—direct, indirect, obfuscated, and role-play—to simulate both malicious and benign use cases. We propose a three-way response evaluation protocol (Accept, Caution, Refuse) and a fine-grained Safety Score metric to assess model behavior. Our analysis reveals that many state-of-the-art LLMs remain vulnerable to jailbreaks that subtly rephrase harmful prompts, while also over-refusing safe but atypically phrased queries. Finally, we propose a mitigation strategy using a lightweight classifier to detect jailbreak attempts and steer models toward safer behavior via reminder-based conditioning. CARES provides a rigorous framework for testing and improving medical LLM safety under adversarial and ambiguous conditions.
Mengxue Zhang, Eric Hanchen Jiang, Qingcheng Zeng, Chen-Hsiang Yu
NeurIPS3
2024 DEMAE: Diffusion-Enhanced Masked Autoencoder for Hyperspectral Image Classification With Few Labeled Samples
abstract
Unlike other deep learning (DL) models, Transformer has the ability to extract long-range dependency features from hyperspectral image (HSI) data. Masked autoencoder (MAE), which is based on Transformer architecture, employs a “mask-reconstruction” strategy for training, allowing the model to be effective for downstream tasks. However, existing MAE-based methods only apply spectral or spatial masking to HSI and reconstruct them for feature learning, which is too simplistic and insufficient for the model to learn robust features. Additionally, the issue of lacking labeled samples in HSI and the primary objective of MAE to reduce the reliance on labeled samples are often overlooked. To address these issues, we are inspired by diffusion-based representation learning and propose diffusion-enhanced MAE (DEMAE) for HSI classification with few labeled samples. First, an asymmetric encoder–decoder framework is constructed as the backbone by stacking both conditional and standard Transformer blocks. Second, we devise an auxiliary task aimed at simultaneous denoising and reconstruction, facilitating heuristic feature learning from HSI data. Third, the encoder of DEMAE is isolated for training with few labeled samples. Finally, the encoder is used for classification, and a novel signal-to-noise ratio enhanced (SNR-Enhanced) loss function is introduced to regularize the model training process. The performance of DEMAE is evaluated on four benchmark datasets, demonstrating its superiority in classification accuracy and mapping capabilities on unlabeled areas compared to existing state-of-the-art methods with few labeled samples. The source code will be available online athttps://github.com/ZhaohuiXue/DEMAE.
Zhaohui Xue, Xiangyu Nie, Hao Wu 0082, Mengxue Zhang, Hongjun Su
IEEE Trans. Geosci. Remote. Sens.6
2024 SPFormer: Self-Pooling Transformer for Few-Shot Hyperspectral Image Classification
abstract
Transformer has shown great potential in extracting global features, and it can achieve better classification performance with a large number of training samples compared with other deep learning (DL) models. However, most of the existing Transformer-based models for hyperspectral image (HSI) classification simply use the multihead self-attention and channel multilayer perceptron (MLP) modules that contain many parameters to learn, resulting in poor performance in a few-shot learning scenario. To overcome the above issue, a lightweight self-pooling Transformer (SPFormer) is proposed for few-shot HSI classification. First, a one-layer autoencoder based on self-supervised learning is built to reduce the dimensionality of HSI. Second, two parameter-free modules, channel shuffle for multihead self-pooling with sparse mapping (CSSM-MHSP) and central token mixer (CTM), are proposed for mapping spectral features to higher dimensions and promoting information interaction between pixels, respectively. Third, a lightweight channel embedding is designed to extract deep spectral features. Finally, a fully connected layer is used for classification. The classification performance of the proposed method is evaluated on four benchmark datasets, showing its superiority in classification accuracy, generalization performance, and model complexity compared to existing state-of-the-art methods with limited training samples.
Zhaohui Xue, Tianzhi Zhu, Mengxue Zhang
IEEE Trans. Geosci. Remote. Sens.6
2023 Interpretable Math Word Problem Solution Generation via Step-by-step Planning
abstract
Solutions to math word problems (MWPs) with step-by-step explanations are valuable, especially in education, to help students better comprehend problem-solving strategies.Most existing approaches only focus on obtaining the final correct answer.A few recent approaches leverage intermediate solution steps to improve final answer correctness but often cannot generate coherent steps with a clear solution strategy.Contrary to existing work, we focus on improving the correctness and coherence of the intermediate solutions steps.We propose a step-by-step planning approach for intermediate solution generation, which strategically plans the generation of the next solution step based on the MWP and the previous solution steps.Our approach first plans the next step by predicting the necessary math operation needed to proceed, given history steps, then generates the next step, token-by-token, by prompting a language model with the predicted math operation.Experiments on the GSM8K dataset demonstrate that our approach improves the accuracy and interpretability of the solution on both automatic metrics and human evaluation.
Mengxue Zhang, Zichao Wang 0001, Zhichao Yang 0001, Weiqi Feng, Andrew S. Lan
ACL (1)1
2023 Algebra Error Classification with Large Language Models
Hunter McNichols, Mengxue Zhang, Andrew S. Lan
AIED2
2023 Modeling and Analyzing Scorer Preferences in Short-Answer Math Questions
Mengxue Zhang, Neil T. Heffernan, Andrew S. Lan
EDM1
2023 DSR-GCN: Differentiated-Scale Restricted Graph Convolutional Network for Few-Shot Hyperspectral Image Classification
abstract
Graph convolution networks (GCNs) have shown great potentials for few-shot hyperspectral image (HSI) classification. Mainstream GCNs construct graph according to single scale segmentation, which usually ignores subtle adjacency relation between small regions, leading to unreliable initial local graph. To overcome the above issue, we propose a differentiated-scale restricted GCN (DSR-GCN) for HSI classification. Firstly, we propose a differentiated-scale graph construction method considering both the subtle and relative wider range spectral-spatial relation. Secondly, restricted fusion loss is designed to restrict the fusion of features extracted with differentiated-scale GCN branches. Finally, we design a lightweight spatial-spectral siamese network to remedy local pixel-level features. The proposed DSR-GCN can better model spatial structure with a reliable and refined graph, and it can capture more discriminate features in few-shot learning (FSL) scenario. Extensive experiments conducted on four benchmark data sets demonstrate that DSR-GCN outperforms the other deep learning methods in terms of classification accuracy and generalization performance, with the improvements in terms of OA around 6.20%~23.41% (Indian Pines), 4.45%~ 16.48% (University of Pavia), 4.25%~11.85% (Salinas), and 2.0%~17.23% (University of Houston) under 5 labeled samples per class.
Zhaohui Xue, Mengxue Zhang
IEEE Trans. Geosci. Remote. Sens.3
2023 Multistage Relation Network With Dual-Metric for Few-Shot Hyperspectral Image Classification
abstract
Recently, few-shot learning (FSL) has exhibited great potentials in hyperspectral image (HSI) classification due to its promising performance under few training samples. Although existing FSL methods have achieved great success, some limitations can still be witnessed. On the one hand, current methods mainly rely on the single metric to identify, which cannot effectively represent the class distribution with few labeled samples. On the other hand, existing methods usually only use the last deep feature of feature extractor, which may lead to the under-utilization of scarce labeled samples. To overcome the above issues, a novel multistage relation network with dual-metric (DM-MRN) is proposed for few-shot HSI classification. Firstly, a sample recombination strategy is designed to increase the variety of classification tasks in training period. Secondly, an embedding module is employed to extract deep features of the input image patches. Thirdly, we propose two relation modules: image-to-class (I2C) block and image-to-image (I2I) block. I2C block is designed to compute I2C-level relation score between second-order features, and I2I block is conceived to generate I2I-level relation score between first-order features. Finally, DM-MRN is constructed by integrating one embedding module, two I2C blocks, and one I2I block. In addition, an adaptive weighting strategy is designed to fuse the obtained relation scores, and classification can be achieved by assigning each query sample to the class with the highest value of the fused relation score. Extensive experiments carried out on five popular HSI data sets demonstrate that the proposed method outperforms other traditional and advanced models under few training samples in terms of classification accuracy and generalization performance, i.e., the performance improvement in terms of OA is around 0.30%-27.98% under 10 labeled samples per class.
Zhaohui Xue, Qiuping Lan, Mengxue Zhang
IEEE Trans. Geosci. Remote. Sens.5
2022 Automatic Short Math Answer Grading via In-context Meta-learning
Mengxue Zhang, Sami Baral, Neil T. Heffernan, Andrew S. Lan
EDM1
2022 Incremental Dictionary Learning-Driven Tensor Low-Rank and Sparse Representation for Hyperspectral Image Classification
abstract
Low-rank and sparse representation (LRSR) has gained popularity in hyperspectral image (HSI) classification. However, existing LRSR models usually treat HSI as a two-dimensional matrix, which may destroy the original 3D intrinsic structure of HSI. Moreover, the dictionary consisting of only training samples lacks completeness and may be suboptimal for representation. To overcome the above issues, we propose an incremental dictionary learning-driven tensor low-rank and sparse representation (TLRSR-IDL) model for HSI classification. First, we represent HSI as a third-order tensor to retain its original 3D intrinsic structure by using the TLRSR model, which also combines both sparsity and low rankness to maintain global and local data structures. Second, we design an optimal reconstruction within regularized neighborhood (ORRN) method to exploit spectral-spatial information by avoiding the interference of heterogeneous samples in the neighborhood. Finally, an incremental dictionary learning (IDL) scheme is designed to iteratively introduce augmented samples into the dictionary, and the final classification map is produced by feeding back the last round of the incremental dictionary into the TLRSR-IDL model. The main innovative contribution lies in that the proposed IDL scheme can leverage supervised and unsupervised information, which greatly enhances traditional LRSR and TLRSR models. Experimental results based on three popular hyperspectral datasets demonstrate that the proposed method outperforms other related counterparts in terms of classification accuracy and generalization performance, with OA improvements of 0.97%-16.83%, 1.25%-6.89%, and 0.85%-6.67% for Indian Pines, Pavia University, and Salinas, respectively.
Zhaohui Xue, Xiangyu Nie, Mengxue Zhang
IEEE Trans. Geosci. Remote. Sens.3
2021 Scientific Formula Retrieval via Tree Embeddings
abstract
Exploiting the ever-growing corpus of scientific content calls for new ways and means to effectively organize, search, and retrieve scientific formulae. We propose a new data-driven framework for retrieving similar scientific formulae via learned formula representations based on tree embeddings. FORTE (for FOrmula Representation learning via Tree Embeddings) leverages operator tree representations of symbolic scientific formulae (such as math equations) to explicitly capture their inherent structural and semantic properties. FORTE employs i) a tree encoder that encodes the formula’s operator tree into an embedding vector and ii) a tree decoder that directly generates a formula’s operator tree from the embedding vector. We also develop a novel tree beam search algorithm that improves the quality of the decoded operator trees. We demonstrate that FORTE (sometimes significantly) outperforms various baseline methods on formula reconstruction and retrieval using a real-world dataset comprising 770k scientific formulae collected on-line.
Zichao Wang 0001, Mengxue Zhang, Richard G. Baraniuk, Andrew S. Lan
IEEE BigData2
2021 Math Operation Embeddings for Open-ended Solution Analysis and Feedback
Mengxue Zhang, Zichao Wang 0001, Richard G. Baraniuk, Andrew S. Lan
EDM1
2021 Attention-Based Second-Order Pooling Network for Hyperspectral Image Classification
abstract
Deep learning (DL) has exhibited huge potentials for hyperspectral image (HSI) classification due to its powerful nonlinear modeling and end-to-end optimization characteristics. Although the superior performance of DL-based methods has been witnessed, some limitations can still be found. On the one hand, existing DL frameworks usually resorted to first-order statistical features, whereas they rarely considered second-order or higher order statistical features. On the other hand, the optimization of complex hyperparameters (e.g., the layer number and convolutional kernel size) is time-consuming and a very tough task, making the designed DL framework unexplainable. To overcome these challenges, we propose a novel attention-based second-order pooling network (A-SPN). First, a first-order feature operator is designed to model the spectral–spatial information of HSI. Second, an attention-based second-order pooling (A-SOP) operator is designed to model discriminative and representative features. Finally, a fully connected layer with softmax loss is used for classification. The proposed framework can obtain second-order statistical features in an end-to-end manner. In addition, A-SPN is free of complex hyperparameters tuning, making it more explainable and easily equipped for classification tasks. Experimental results based on three common hyperspectral data sets demonstrate that A-SPN outperforms other traditional and state-of-the-art DL-based HSI classification methods in terms of generalization performance with limited training samples, classification accuracy, convergence rate, and computational complexity.
Zhaohui Xue, Mengxue Zhang, Peijun Du
IEEE Trans. Geosci. Remote. Sens.2
2020 Matching Questions and Answers in Dialogues from Online Forums
abstract
Matching question-answer relations between two turns in conversations is not only the first step in analyzing dialogue structures, but also valuable for training dialogue systems.This paper presents a QA matching model considering both distance information and dialogue history by two simultaneous attention mechanisms called mutual attention.Given scores computed by the trained model between each non-question turn with its candidate questions, a greedy matching strategy is used for final predictions.Because existing dialogue datasets such as the Ubuntu dataset are not suitable for the QA matching task, we further create a dataset with 1,000 labeled dialogues and demonstrate that our proposed model outperforms the state-of-the-art and other strong baselines, particularly for matching long-distance QA pairs.
Qi Jia 0003, Mengxue Zhang, Shengyao Zhang, Kenny Q. Zhu
ECAI2
2020 Evaluating the Performance of Reinforcement Learning Algorithms
abstract
Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results are often inconsistent and difficult to replicate. In this work, we argue that the inconsistency of performance stems from the use of flawed evaluation metrics. Taking a step towards ensuring that reported results are consistent, we propose a new comprehensive evaluation methodology for reinforcement learning algorithms that produces reliable measurements of performance both on a single environment and when aggregated across environments. We demonstrate this method by evaluating a broad class of reinforcement learning algorithms on standard benchmark tasks.
Scott M. Jordan, Yash Chandak, Mengxue Zhang, Philip S. Thomas
ICML4
2019 Automatic discovery of adverse reactions through Chinese social media
Mengxue Zhang, Meizhuo Zhang, Chen Ge, Quanyang Liu, Jiemin Wang, Kenny Q. Zhu
Data Min. Knowl. Discov.1
2018 Anomaly Explanation Using Metadata
abstract
Anomaly detection is the well-studied task of identifying when data is atypical in some way with respect to its source. In this work, by contrast, we are interested in finding possible descriptions of what may be causing anomalies. We propose a new task, attaching semantics drawn from metadata to a portion of the anomalous examples from some data source. Such a partial description of the anomalous data in terms of the meta-data is useful both because it may help to explain what causes the identified anomalies, and also because it may help to identify the truly unusual examples that defy such simple categorization. This is especially significant when the data set is too large for a human analyst to inspect the anomalies manually. The challenge is that anomalies are, by definition, relatively rare, and so we are seeking to learn a precise characterization of a rare event. We examine algorithms for this task in a webcam domain, generating human-understandable explanations for a pixellevel characterization of anomalies. We find that using a recently proposed algorithm that prioritizes precision over recall, it is possible to attach good descriptions to a moderate fraction of the anomalies in webcam data so long as the data set is fairly large.
Joshua Arfin, Mengxue Zhang, Tushar Mathew, Robert Pless, Brendan Juba
WACV3
2017 An Improved Algorithm for Learning to Perform Exception-Tolerant Abduction
abstract
Inference from an observed or hypothesized condition to a plausible cause or explanation for this condition is known as abduction. For many tasks, the acquisition of the necessary knowledge by machine learning has been widely found to be highly effective. However, the semantics of learned knowledge are weaker than the usual classical semantics, and this necessitates new formulations of many tasks. We focus on a recently introduced formulation of the abductive inference task that is thus adapted to the semantics of machine learning. A key problem is that we cannot expect that our causes or explanations will be perfect, and they must tolerate some error due to the world being more complicated than our formalization allows. This is a version of the qualification problem, and in machine learning, this is known as agnostic learning. In the work by Juba that introduced the task of learning to make abductive inferences, an algorithm is given for producing k-DNF explanations that tolerates such exceptions: if the best possible k-DNF explanation fails to justify the condition with probability ε, then the algorithm is promised to find a k-DNF explanation that fails to justify the condition with probability at most O(nkε), where n is the number of propositional attributes used to describe the domain. Here, we present an improved algorithm for this task. When the best k- DNF fails with probability ε, our algorithm finds a k-DNF that fails with probability at most O ̃(nk/2ε) (i.e., suppressing logarithmic factors in n and 1/ε). We also examine the empirical advantage of this new algorithm over the previous algorithm in two test domains, one of explaining conditions generated by a “noisy” k-DNF rule, and another of explaining conditions that are actually generated by a linear threshold rule.
Mengxue Zhang, Tushar Mathew, Brendan Juba
AAAI1