Li Qian 0001

dblp:25/6847-1 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0001-8951-1832ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 DynoGraph: Dynamic Graph Construction for Nonlinear Dimensionality Reduction
abstract
Most well-known graph-based dimensionality re-duction algorithms, such as t-SNE and UMAP, use a two-step approach: first to construct a graph out of the high-dimensional data and then to embed the graph into the low-dimensional space. The main challenges of these algorithms include how to construct a good graph and how to maintain the similarity structure of the high-dimensional data in the low-dimensional space. This study proposes DynoGraph, a novel algorithm called Dynamic Graph Construction for Nonlinear Dimensionality Reduction, to address these two challenges. First, we develop an adaptive neighborhood graph construction method that accurately captures the intrinsic geometry of the high-dimensional data. Second, for the first time, we introduce a dynamic graph modification process during dimensionality reduction, ensuring that the data structure in the low-dimensional space faithfully reflects the high-dimensional data. For vertex pairs that are connected by edges in high-dimensional space exhibit far apart in low-dimensional space, additional edges are inserted to strengthen the connection between them. Conversely, for vertex pairs that are not connected in high-dimensional space exhibit close together in the low-dimensional space, edges are deleted to reduce the connection between them. These adjustments help to update their positions in subsequent embeddings, aligning them toward the high-dimensional data. Extensive experiments have demonstrated the superiority of DynoGraph against various comparative algorithms in tasks such as visualization, classification and clustering.
Li Qian 0001, Claudia Plant, Yalan Qin, Christian Böhm 0001
ICDM1
2024 ADOD: Adaptive Density Outlier Detection
abstract
Outlier detection plays a dual role in data analysis: cleansing data to optimize the performance of downstream tasks and identifying potentially rare valuable events or patterns. Proximity-based methods, which are independent of data distribution assumptions, are plagued by parameter selection and performance challenges when handling data with varying densities. This study proposed a novel unsupervised algorithm named Adaptive Density Outlier Detection (ADOD) to address these challenges. The core innovation of ADOD involves two main aspects: adaptive neighborhood boundaries and density consistency scoring. First, instead of relying on a predefined fixed radius, ADOD employs perplexity to calculate the local scale of each data point. It then dynamically adjusts the neighborhood boundaries according to this scale to adapt to data with varying densities. Second, ADOD estimates local density using a mutual neighbor graph and combines the density differences between data points and their neighbors to compute outlier scores, effectively distinguishing outliers that significantly deviate from their surroundings. This study evaluated ADOD on one synthetic and 32 real datasets, and compared it with 14 classical and state-of-the-art algorithms from different categories. Extensive experimental results demonstrated the superior performance of ADOD, achieving the highest average accuracy across ROC, P@N, and AP metrics. This study promotes the development of outlier detection techniques and expands their potential for real-time applications.
Li Qian 0001, Xin Sun 0003, Wengang Guo, Christian Böhm 0001
ICDM1
2024 Fast Elastic-Net Multi-view Clustering: A Geometric Interpretation Perspective
abstract
Multi-view clustering methods have been extensively explored in the last decades. This kind of methods is built on the assumption that the data are sampled from multiple subspaces with low dimension and each group fits into one of these subspaces. The quadratic or cubic computation complexity produced by these methods is inevitable, resulting in the difficulty for clustering multi-view datasets with large scales. Some efforts have been presented to select key anchors beforehand to capture the data distributions in different views. Despite significant progress, these methods pay few attentions to deriving provably scalable and correct method for finding the optimal shared anchor graph from the geometric interpretation perspective. They also ignore to give a well balance between the connectedness and subspace preserving properties of the shared anchor graph. In this paper, we propose the Fast Elastic- Net Multi-view Clustering (FENMC) from a geometric interpretation perspective. We provide the geometric analysis in determining the optimal shared anchor graph based on the introduced elastic-net regularizer for fast multi-view clustering, where the elastic-net regularizer is built on the mixture of L_2 and L_1 norms. We also give a theoretical justification for the balance between the connectedness and subspace preserving properties of the shared anchor graph for multi-view clustering. Our experiments on different datasets show that the proposed method not only obtains the satisfied clustering performance, but also deals with large-scale datasets with high efficiency.
Yalan Qin, Li Qian 0001
ACM Multimedia2
2021 Density-Based Clustering for Adaptive Density Variation
abstract
Cluster analysis plays a crucial role in data mining and knowledge discovery. Although many researchers have investigated clustering algorithms over the past few decades, most of the well-known algorithms have shortcomings when dealing with clusters of arbitrary shapes and varying sizes and in the presence of noise and outliers. Density-based methods partially solve these issues but fail to discover clusters with varying densities. In this paper, we propose a novel Density-Based clustering algorithm for Adaptive Density Variation (DBADV), which is based on the classic clustering algorithm DBSCAN. To address the problem of density variation, we define the local density information, which not only reflects the individual property of each object but also describes the density distribution of clusters, and finds the adaptive search range of each object by collecting information from its neighbors. Moreover, we design a new metric to obtain the mutual nearest neighbors of each object to better detect the objects around the boundaries between clusters. We show the effectiveness of our method in extensive experiments on synthetic and realworld data sets, which demonstrate that the performance of the proposed algorithm DBADV is superior to other competitive clustering algorithms.
Li Qian 0001, Claudia Plant, Christian Böhm 0001
ICDM1