Hao Cheng 0001

dblp:09/5158-1 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2012
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 50% Information retrieval · 32% Indexing and storage engines · 18%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
similarity search
0.122008
Bounded Approximation: A New Criterion for Dimensionality Reduction Approximation in Similarity Search · IEEE Trans. Knowl. Data Eng. 2008
A non-linear dimensionality-reduction technique for fast similarity search in large databases · SIGMOD Conference 2006
Data mining
clustering
0.112008
Constrained locally weighted clustering · Proc. VLDB Endow. 2008
Data mining › clustering
constrained clustering
0.112008
Constrained locally weighted clustering · Proc. VLDB Endow. 2008
Data mining
dimensionality reduction
0.112008
Bounded Approximation: A New Criterion for Dimensionality Reduction Approximation in Similarity Search · IEEE Trans. Knowl. Data Eng. 2008
Information retrieval › similarity search
high-dimensional similarity search
0.112008
Bounded Approximation: A New Criterion for Dimensionality Reduction Approximation in Similarity Search · IEEE Trans. Knowl. Data Eng. 2008
Data mining › clustering › high-dimensional clustering
subspace clustering
0.112008
Constrained locally weighted clustering · Proc. VLDB Endow. 2008
Indexing and storage engines › multidimensional indexing
dimensionality reduction for indexing
0.112006
A non-linear dimensionality-reduction technique for fast similarity search in large databases · SIGMOD Conference 2006
Indexing and storage engines
multidimensional indexing
0.112006
A non-linear dimensionality-reduction technique for fast similarity search in large databases · SIGMOD Conference 2006
Data mining › clustering
instance-level constraints
0.012008
Constrained locally weighted clustering · Proc. VLDB Endow. 2008

Methods — techniques the papers use, named apart from their topics

spherical range search · 0.1rectangular range search · 0.1pairwise constraints · 0.1nonlinear transformation · 0.1locally weighted clustering · 0.1dimensionality reduction · 0.1
YearPublicationVenuePosition
2012 A Multi-Directional Search technique for image annotation propagation
Ning Yu 0001, Kien A. Hua, Hao Cheng 0001
J. Vis. Commun. Image Represent.3
2010 An automatic feature generation approach to multiple instance learning and its applications to image databases
Hao Cheng 0001, Kien A. Hua, Ning Yu 0001
Multim. Tools Appl.1
2009 SubSpace Projection: A unified framework for a class of partition-based dimension reduction techniques
Hao Cheng 0001, Khanh Vu, Kien A. Hua
Inf. Sci.1
2008 Boost image clustering with user query log
abstract
Image clustering is to derive a salient grouping of images such that similar ones are placed in the same cluster, which is useful in many applications. In this paper, we propose a constrained clustering algorithm, which leverages the collected user query log to guide the clustering process. Our method models a set of images as a graph and randomly contracts two vertices into a meta vertex iteratively with regarding to their similarity until the desired number of image groups has been reached. The experimental results demonstrate the superiority of our proposal.
Hao Cheng 0001, Kien A. Hua, Ning Yu 0001
ICME1
2008 Dynamic Directional Navigation in Content-Based Image Retrieval
abstract
Nowadays’ image retrieval techniques focus on modifying a query contour based on the relative correlations of relevant images. The spatial relationship between the relevant images and all the other images in database (i.e. the directional information) is not explored and utilized well. From the aspect of feature space, when a user selects relevant images, the direction to the potential relevant images is implicitly expressed. In this paper, we propose a Dynamic Directional Navigation (DDN) system to explore the directional information and navigate in the data space to find similar images. Multiple queries will be used when the relevant images point to different directions. The experiment results indicate that our technique can address the semantic gap well in Content-Based Image Retrieval (CBIR) and achieves better precision and recall than the previous techniques.
Ning Yu 0001, Kien A. Hua, Hao Cheng 0001
ICME3
2008 Constrained locally weighted clustering
abstract
Data clustering is a difficult problem due to the complex and heterogeneous natures of multidimensional data. To improve clustering accuracy, we propose a scheme to capture the local correlation structures: associate each cluster with an independent weighting vector and embed it in the subspace spanned by an adaptive combination of the dimensions. Our clustering algorithm takes advantage of the known pairwise instance-level constraints. The data points in the constraint set are divided into groups through inference; and each group is assigned to the feasible cluster which minimizes the sum of squared distances between all the points in the group and the corresponding centroid. Our theoretical analysis shows that the probability of points being assigned to the correct clusters is much higher by the new algorithm, compared to the conventional methods. This is confirmed by our experimental results, indicating that our design indeed produces clusters which are closer to the ground truth than clusters created by the current state-of-the-art algorithms.
Hao Cheng 0001, Kien A. Hua, Khanh Vu
Proc. VLDB Endow.1
2008 Bounded Approximation: A New Criterion for Dimensionality Reduction Approximation in Similarity Search
abstract
We examine the problem of efficient distance-based similarity search over high-dimensional data. We show that a promising approach to this problem is to reduce dimensions and allow fast approximation. Conventional reduction approaches, however, entail a significant shortcoming: The approximation volume extends across the dataspace, which causes overestimation of retrieval sets and impairs performance. This paper focuses on a new criterion for dimensionality reduction methods: bounded approximation. We show that this requirement can be accomplished by a novel nonlinear transformation scheme that extracts two important parameters from the data. We devise two approximation formulations, namely, rectangular and spherical range search, each corresponding to a closed volume around the original search sphere. We discuss in detail how we can derive tight bounds for the parameters and prove further results, as well as highlight insights into the problems and our proposed solutions. To demonstrate the benefits of the new criterion, we study the effects of (un)boundedness on approximation performance, including selectivity, error toleration, and efficiency. Extensive experiments confirm the superiority of this technique over recent state-of-the-art schemes.
Khanh Vu, Kien A. Hua, Hao Cheng 0001, Sheau-Dong Lang
IEEE Trans. Knowl. Data Eng.3
2007 Local and Global Structures Preserving Projection
abstract
In this paper, we propose Local and Global Structures Preserving Projection (LGSPP), which is to find a small set of projection directions so as to properly preserve the local and global structures for a given set of data. Specifically, for each point in the dataset, its local neighborhood is extracted as well as a set of sampled points far away from this point, which characterize the global structure. The embedding minimizes the distances of the points in each local neighborhood while dispersing them far apart from their corresponding remote points. In this way, the local-global relationships between data points are well kept.
Hao Cheng 0001, Kien A. Hua, Khanh Vu
ICTAI (2)1
2006 A non-linear dimensionality-reduction technique for fast similarity search in large databases
abstract
To enable efficient similarity search in large databases, many indexing techniques use a linear transformation scheme to reduce dimensions and allow fast approximation. In this reduction approach the approximation is unbounded, so that the approximation volume extends across the dataspace. This causes over-estimation of retrieval sets and impairs performance.This paper presents a non-linear transformation scheme that extracts two important parameters specifying the data. We prove that these parameters correspond to a bounded volume around the search sphere, irrespective of dimensionality. We use a special workspace-mapping mechanism to derive tight bounds for the parameters and to prove further results, as well as highlighting insights into the problems and our proposed solutions. We formulate a measure that lower-bounds the Euclidean distance, and discuss the implementation of the technique upon a popular index structure. Extensive experiments confirm the superiority of this technique over recent state-of-the-art schemes.
Khanh Vu, Kien A. Hua, Hao Cheng 0001, Sheau-Dong Lang
SIGMOD Conference3