Hang Zhang 0003

dblp:49/6156-3 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0401-4066ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SCoNE: Spherical Consistent Neighborhoods Ensemble for Effective and Efficient Multi-View Anomaly Detection
abstract
The core problem in multi-view anomaly detection is to represent local neighborhoods of normal instances consistently across all views. Recent approaches consider a representation of local neighborhood in each view independently, and then capture the consistent neighbors across all views via a learning process. They suffer from two key issues. First, there is no guarantee that they can capture consistent neighbors well, especially when the same neighbors are in regions of varied densities in different views, resulting in inferior detection accuracy. Second, the learning process has a high computational cost of O(N^2), rendering them inapplicable for large datasets. To address these issues, we propose a novel method termed Spherical Consistent Neighborhoods Ensemble (SCoNE). It has two unique features: (a) the consistent neighborhoods are represented with multi-view instances directly, requiring no intermediate representations as used in existing approaches; and (b) the neighborhoods have data-dependent properties, which lead to large neighborhoods in sparse regions and small neighborhoods in dense regions. The data-dependent properties enable local neighborhoods in different views to be represented well as consistent neighborhoods, without learning. This leads to O(N) time complexity. Empirical evaluations show that SCoNE has superior detection accuracy and runs orders-of-magnitude faster in large datasets than existing approaches.
Hang Zhang 0003, Ye Zhu 0002, Kai Ming Ting
AAAI2
2026 Kernel-bounded clustering: Achieving the objective of spectral clustering without eigendecomposition
Hang Zhang 0003, Kai Ming Ting, Ye Zhu 0002
Artif. Intell.1
2025 Machine Unlearning for Random Forest via Method of Images
Hang Zhang 0003, Kai Ming Ting
ECML/PKDD (5)1
2025 Towards a Robust Persistence Diagram via Data-dependent Kernel
abstract
Topological Data Analysis (TDA) is used to extract topological features such as rings from point clouds. Recent works have identified that existing methods, which construct persistence diagrams in TDA, are not robust to noise and varied densities in a point cloud. This causes these methods to obtain incorrect topological features. We analyze the necessary properties of an approach that can address these two issues, and propose a new filter function for TDA based on a new data-dependent kernel that possesses these properties. Our empirical evaluation reveals that (i) the proposed kernel provides a better mean for UMAP dimensionality reduction (ii) the proposed filter function can significantly improve the performance of Topological Point Cloud Clustering (iii) the proposed filter function is a more effective way of constructing Persistence Diagram for t-SNE visualization and SVM classification than three existing methods of TDA, In addition, we explore the proposed filter’s performance on a more complex deformation named Riemannian stretching. Our proposed filter equipped with Sample Fermat distance outperforms all the other filters when noise and Riemannian stretching coexist. Code is available at https://github.com/IsolationKernel/Codes/tree/main/Lambda-kernel.
Hang Zhang 0003, Kaifeng Zhang 0002, Kai Ming Ting, Ye Zhu 0002
J. Artif. Intell. Res.1
2025 A new filter for deformation-invariant persistence diagram
Kaifeng Zhang 0002, Hang Zhang 0003, Kai Ming Ting, Tianrun Liang
Mach. Learn.2
2024 Local Subsequence-Based Distribution for Time Series Clustering
Lei Gong 0001, Hang Zhang 0003, Zongyou Liu, Kai Ming Ting, Yang Cao 0019, Ye Zhu 0002
PAKDD (1)2
2024 A new distributional treatment for time series anomaly detection
Kai Ming Ting, Zongyou Liu, Lei Gong 0001, Hang Zhang 0003, Ye Zhu 0002
VLDB J.4
2023 Towards a Persistence Diagram that is Robust to Noise and Varied Densities
abstract
Recent works have identified that existing methods, which construct persistence diagrams in Topological Data Analysis (TDA), are not robust to noise and varied densities in a point cloud. We analyze the necessary properties of an approach that can address these two issues, and propose a new filter function for TDA based on a new data-dependent kernel which possesses these properties. Our empirical evaluation reveals that the proposed filter function provides a better means for t-SNE visualization and SVM classification than three existing methods of TDA.
Hang Zhang 0003, Kaifeng Zhang 0002, Kai Ming Ting, Ye Zhu 0002
ICML1
2023 Isolation Kernel Estimators
Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Hang Zhang 0003, Ye Zhu 0002
Knowl. Inf. Syst.4
2022 A New Distributional Treatment for Time Series and An Anomaly Detection Investigation
abstract
Time series is traditionally treated with two main approaches, i.e., the time domain approach and the frequency domain approach. These approaches must rely on a sliding window so that time-shift versions of a periodic subsequence can be measured to be similar. Coupled with the use of a root point-to-point measure, existing methods often have quadratic time complexity. We offer the third R domain approach. It begins with an insight that subsequences in a periodic time series can be treated as sets of independent and identically distributed (iid) points generated from an unknown distribution in R. This R domain treatment enables two new possibilities: (a) the similarity between two subsequences can be computed using a distributional measure such as Wasserstein distance (WD), kernel mean embedding or Isolation Distributional kernel (IDK); and (b) these distributional measures become non-sliding-window-based. Together, they offer an alternative that has more effective similarity measurements and runs significantly faster than the point-to-point and sliding-window-based measures. Our empirical evaluation shows that IDK and WD are effective distributional measures for time series; and IDK-based detectors have better detection accuracy than existing sliding-window-based detectors, and they run faster with linear time complexity.
Kai Ming Ting, Zongyou Liu, Hang Zhang 0003, Ye Zhu 0002
Proc. VLDB Endow.3
2021 Isolation Kernel Density Estimation
abstract
This paper shows that adaptive kernel density estimator (KDE) can be derived effectively from Isolation Kernel. Existing adaptive KDEs often employ a data independent kernel such as Gaussian kernel. Therefore, it requires an additional means to adapt its bandwidth locally in a given dataset. Because Isolation Kernel is a data dependent kernel which is derived directly from data, no additional adaptive operation is required. The resultant estimator called IKDE is the only KDE that is fast and adaptive. Existing KDEs are either fast but non-adaptive or adaptive but slow. In addition, using IKDE for anomaly detection, we identify two advantages of IKDE over LOF (Local Outlier Factor), contributing to significantly faster runtime.
Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Hang Zhang 0003
ICDM4
2012 VDoc+: a virtual document based approach for matching large ontologies using MapReduce
abstract
Many ontologies have been published on the Semantic Web, to be shared to describe resources. Among them, large ontologies of real-world areas have the scalability problem in presenting semantic technologies such as ontology matching (OM). This either suffers from too long run time or has strong hypotheses on the running environment. To deal with this issue, we propose a three-stage MapReduce-based approach V-Doc + for matching large ontologies, based on the MapReduce framework and virtual document technique. Specifically, two MapReduce processes are performed in the first stage to extract the textual descriptions of named entities (classes, properties, and instances) and blank nodes, respectively. In the second stage, the extracted descriptions are exchanged with neighbors in Resource Description Framework (RDF) graphs to construct virtual documents. This extraction process also benefits from the MapReduce-based implementation. A word-weight-based partitioning method is proposed in the third stage to conduct parallel similarity calculation using the term frequency-inverse document frequency (TF-IDF) model. Experimental results on two large-scale real datasets and the benchmark testbed from Ontology Alignment Evaluation Initiative (OAEI) are reported, showing that the proposed approach significantly reduces the run time with minor loss in precision and recall.
Hang Zhang 0003, Wei Hu 0007, Yuzhong Qu
J. Zhejiang Univ. Sci. C1
2011 How Matchable Are Four Thousand Ontologies on the Semantic Web
Wei Hu 0007, Hang Zhang 0003, Yuzhong Qu
ESWC (1)3