Huiwen Dong

dblp:273/5262 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-0426-4102ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 67% Representation and self-supervised learning · 33%
Theoretical computer science
1 paper
Algorithms and data structures · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
high-dimensional data analysis
0.912025
On Finding Hubs in High Dimensions with Sampling · AAAI 2025
Algorithms and data structures › similarity search › nearest neighbor search
hubness
0.912025
On Finding Hubs in High Dimensions with Sampling · AAAI 2025
Algorithms and data structures › similarity search
nearest neighbor search
0.912025
On Finding Hubs in High Dimensions with Sampling · AAAI 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
ImageNet Pre-training Also Transfers Non-robustness · AAAI 2023
Machine learning › Representation and self-supervised learning
pre-training
0.712023
ImageNet Pre-training Also Transfers Non-robustness · AAAI 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
ImageNet Pre-training Also Transfers Non-robustness · AAAI 2023

Methods — techniques the papers use, named apart from their topics

sampling · 1.7approximate kNN indexes · 1.7robust pre-training · 0.7fine-tuning · 0.7
YearPublicationVenuePosition
2026 A Gradient-Based Causal Discovery Framework With Applications to Complex Industrial Processes
abstract
With the rapid development of deep learning, a wide range of neural network-based causal discovery frameworks have emerged. Although these methods have achieved significant advancements, they still face several limitations when deployed in real-world industrial processes. Many existing models follow the component-wise modeling design, where an individual model must be built for each variable. This leads to significant computational overhead, especially in high-dimensional industrial systems. Moreover, imposing sparsity constraints on the first-layer weights of neural networks to discover causal relationships limits their ability to capture complex and nonlinear interactions among variables. To address these challenges, we propose a novel lightweight causal discovery framework, termed gradient-based causal discovery (GCD). Different from conventional component-wise models, GCD only employs a single multilayer perceptron for time-series prediction and leverages$\ell _{1}$regularization on the neural network’s input–output gradient to infer causal relationships. Numerical simulations on the Lorenz-96 and CausalTime show that GCD consistently achieves the state-of-the-art performance. Moreover, evaluations across four industrial processes, including Tennessee-Eastman, ultra-processed food, debutanizer, and gas turbine power generation, demonstrate that GCD significantly reduces computational overhead while maintaining high causal discovery accuracy, highlighting its applicability to complex industrial processes.
Meiliang Liu, Huiwen Dong, Xiaoxiao Yang, Yunfang Xu, Mingbao Yang, Zhengye Si, Zhiwen Zhao
IEEE Trans. Ind. Informatics2
2025 On Finding Hubs in High Dimensions with Sampling
abstract
Hubs are a few points that frequently appear in the k-nearest neighbors (kNN) of many other points in a high-dimensional data set. The hubs' effects, called the hubness phenomenon, degrade the performance of kNN based models in high dimensions. We present SamHub, a simple sampling approach to efficiently identify hubs with theoretical guarantees. Apart from previous works based on approximate kNN indexes, SamHub is generic and applicable to any distance measure with negligible additional memory footprint. Empirically, by sampling only 10% of points, SamHub runs significantly faster and offers higher accuracy than existing hub detection methods on many real-world data sets with dot product, L1, L2, and dynamic time warping distances. Our ablation studies of SamHub on improving kNN-based classification show potential for other high-dimensional data analysis tasks.
Huiwen Dong, Linghan Zeng, Zhiwen Zhao, Francesco Silvestri 0001, Ninh Pham
AAAI1
2023 ImageNet Pre-training Also Transfers Non-robustness
abstract
ImageNet pre-training has enabled state-of-the-art results on many tasks. In spite of its recognized contribution to generalization, we observed in this study that ImageNet pre-training also transfers adversarial non-robustness from pre-trained model into fine-tuned model in the downstream classification tasks. We first conducted experiments on various datasets and network backbones to uncover the adversarial non-robustness in fine-tuned model. Further analysis was conducted on examining the learned knowledge of fine-tuned model and standard model, and revealed that the reason leading to the non-robustness is the non-robust features transferred from ImageNet pre-trained model. Finally, we analyzed the preference for feature learning of the pre-trained model, explored the factors influencing robustness, and introduced a simple robust ImageNet pre-training solution. Our code is available at https://github.com/jiamingzhang94/ImageNet-Pretraining-transfers-non-robustness.
Jiaming Zhang 0006, Jitao Sang 0001, Qi Yi, Yunfan Yang, Huiwen Dong, Jian Yu 0001
AAAI5
2023 Chain-Based Outlier Detection for Complex Data Scenarios
abstract
Outlier detection is a challenging problem due to the complexity of real-life data. Specifically, an effective outlier detection method should be able to handle (1) different types of outliers: local outliers, global outliers, and cluster outliers; (2) a lack of prior knowledge of outlier number or percentage; (3) neighboring clusters with different densities; and (4) real-time detection. Unfortunately, few algorithms can tackle all these challenges simultaneously. In this paper, we propose a chain-based theory to address these issues, where a minority of data points will be considered outliers if their distances from normal data points change abruptly. We present the parallel chaining method based on this theory. Experiments demonstrate that the proposed method exhibits comparable performance to other state-of-the-art methods in the synthetic datasets and outperform other methods in the real-life datasets.
Huiwen Dong, Qing-Guo Wang
IEEE Big Data1