Hao Huang 0001

dblp:04/5616-1 · DBLP profile ↗
← Back
43ranked-venue papers in the field
12as first author
17since 2021 · last 2026
0000-0002-3777-1488ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 17 (5 first)Data Mining & Knowledge Discovery · 13 (4 first)Information Retrieval & Web Search · 8 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Multi-task Inference of Diffusion Networks
abstract
Inferring the underlying structures of diffusion networks based on observed diffusion results is a fundamental problem in network analysis. Traditional approaches typically address this problem by inferring each diffusion network in isolation, relying on the assumption that sufficient observation data is available for each individual inference task. However, in many real-world scenarios, it is common to observe diffusion processes occur across multiple networks with similar structures, while the amount of observable data collected on each network is often limited. In this work, we study how to infer multiple similar diffusion networks jointly with limited observation data for each network. To this end, we propose a novel iterative strategy which in turn updates the inference results for all diffusion networks by exploiting the similarity between the networks, and theoretically guarantee the monotonicity and convergence of the iterative process. Extensive experiments on both synthetic and real-world networks demonstrate that our method not only achieves superior inference accuracy compared to existing techniques, but also maintains high computational efficiency.
Ting Gan, Kudereti Kuerban, Qian Yan 0001, Zhigao Zheng 0001, Hao Huang 0001
WWW6
2026 Text-attributed Graph Condensation via Text Selection and Attribute Matching
abstract
Text-Attributed Graph (TAG) is an important type of graph structured data, where each node has a text description. TAG models usually train a Graph Neural Network (GNN) and language model jointly, which leads to high space and time consumption, especially on large datasets. To mitigate this, we propose TAGSAM, a condensation method that compresses TAGs while preserving training accuracy. TAGSAM comes with two key designs, i.e., subgraph text Selection and Attribute similarity Matching, which compress the text description and graph topology of TAG, respectively. For the texts, subgraph text selection selects and merges representative text chunks from multiple related text descriptions by maximizing mutual information. For the graph topology, popular condensation methods based on Matching Training Trajectories (MTT) suffer from high variance, which hinders accuracy. Our attribute similarity matching mitigates this issue by aligning stable similarity matrices. We evaluate TAGSAM against six state-of-the-art baselines, where it showcases superior performance. For the same compressed size, TAGSAM improves upon the best-performing baseline by an average of 4.9% in accuracy. Furthermore, it maintains competitive training accuracy even when the TAG is condensed to just 1% size. Our code is available at https://github.com/SundayVHan/TAGSAM
Haowei Han, Yuxiang Wang 0013, Guojia Wan, Hao Wang 0013, Shanshan Feng 0001, Hao Huang 0001, Jiawei Jiang 0001, Xiao Yan 0002
WWW6
2026 DeepUL: Deep Unlearning via Model Sparsity
Zhigao Zheng 0001, Yaowen Kuang, Tao Wang 0037, Yahong Chen, Shihong Yao, Hao Huang 0001
WWW7
2025 VF-FD: Feature Deduplication for Vertical Federated Learning
Xiao Yan 0002, Yuanyuan Zhu 0001, Hao Huang 0001, Qinbo Zhang, Guojia Wan, Jiawei Jiang 0001
DASFAA (4)5
2025 InC: A Vertical Federated Learning Framework with Multiple Noisy Labels
Xiao Yan 0002, Xiaokai Zhou, Hao Wang 0013, Hao Huang 0001, Jiawei Jiang 0001
DASFAA (5)6
2025 Hounding Data Diversity: Towards Participant Selection in Vertical Federated Learning
abstract
Due to the rising concerns on privacy protection, how to build machine learning models from distributed databases with privacy guarantees has gained more popularity. Vertical federated learning (VFL) trains machine learning models in a privacy-preserving way when the data features are scattered over distributed databases. We study the participant selection problem (PSP) for VFL, which chooses a given number of participants to conduct training while maximizing model accuracy. Compared to training with all participants, PSP can filter out hitch-riders that contribute marginally to model quality and reduce training time by involving fewer participants. To achieve good model accuracy, we formulate PSP as choosing a set of participants that maximizes the likelihood of the data samples. Then, utilizing the k-nearest neighbors (KNN) classifier as the proxy model, we express the likelihood as a function of the selected participants and prove that the function is sub modular. The submodular property is favorable as it can account for the feature diversity among the participants and allows to greedily select the participant with the maximum gain in each step. However, the selection process requires finding the top-k neighbors of a data sample as the basic operation, which is expensive in VFL setting as it involves encrypted communication. As such, we adapt the Fagin's algorithm, a famous top-k query algorithm, to reduce the amount of encrypted communication. We deploy our solution VFPS-SM across five distributed nodes and conduct experiments with 10 datasets and 3 models to evaluate its performance. The results show that VFPS-SM can reduce the end-to-end running time by up to$35\times$, selection time$365\times$and improve model accuracy by 6.0% compared with state-of-the-art baselines.
Xiaokai Zhou, Xiao Yan 0002, Fangcheng Fu, Hao Huang 0001, Quanqing Xu, Chuanhui Yang, Bo Du 0001, Tieyun Qian, Jiawei Jiang 0001
ICDE5
2025 Secure spatial remote sensing image matching
Hao Huang 0001, Hao Wang 0013, Chuang Hu, Jiawei Jiang 0001
GeoInformatica3
2025 Diffusion pattern mining
Qian Yan 0001, Yulan Yang, Ting Gan, Hao Huang 0001
Knowl. Inf. Syst.5
2025 Online Billboard Auction With Social Welfare Maximization
abstract
Outdoor billboard advertising has proven effective for commercial promotions, attracting potential customers, and boosting product sales. Auction serves as a popular method for leasing billboard usage rights, enabling a seller to rent billboards to winning users for predefined periods according to their bids. An effective auction algorithm is of great significance to maximize the efficiency of the billboard ecosystem. In contrast to a rich literature on Internet advertising auctions, well-crafted algorithms tailored for outdoor billboard auctions remain rare. In this work, we investigate the problem of outdoor billboard auctions, in the practical setting where bids are received and processed on the fly. Our goal is to maximize social welfare, namely the total benefits of auction participants, including the billboard service provider and the bidding users. To this end, we first formulate the billboard social welfare maximization problem into an Integer Linear Problem (ILP), and then reformulate the ILP into a compact form with a reduced size of constraints (at the cost of involving exponentially many primal variables), based on which we derive the dual problem. Furthermore, we design a dual oracle to handle the exponentially many dual constraints, avoiding exhaustive enumeration. We present a primal-dual online algorithm with an incentive-compatible pricing mechanism. Theoretical analysis proves the individual rationality, incentive compatibility, and computational efficiency of our online algorithm. Extensive experimental results show that the online algorithm is both effective and efficient, and achieves a good competitive ratio.
Hao Huang 0001, Mengqi Shan, Zhigao Zheng 0001, Ting Gan, Jiawei Jiang 0001, Zongpeng Li
IEEE Trans. Knowl. Data Eng.1
2025 Detecting and Analyzing Motifs in Large-Scale Online Transaction Networks
abstract
Motif detection is a graph algorithm that detects certain local structures in a graph. Although network motif has been studied in graph analytics, e.g., social network and biological network, it is yet unclear whether network motif is useful for analyzingonline transaction networkthat is generated in applications such as instant messaging and e-commerce. In an online transaction network, each vertex represents a user’s account and each edge represents a money transaction between two users. In this work, we try to analyze online transaction networks with network motifs. We design motif-based vertex embedding that integrates motif counts and centrality measurements. Furthermore, we design a distributed framework to detect motifs in large-scale online transaction networks. Our framework obtains the edge directions using a bi-directional tagging method and avoids redundant detection with a reduced view of neighboring vertices. We implement the proposed framework under the parameter server architecture. In the evaluation, we analyze different kinds of online transaction networks w.r.t the distribution of motifs and evaluate the effectiveness of motif-based embedding in downstream graph analytical tasks. The experimental results also show that our proposed motif detection framework can efficiently handle large-scale graphs.
Jiawei Jiang 0001, Hao Huang 0001, Zhigao Zheng 0001, Fangcheng Fu, Xiaosen Li, Bin Cui 0001
IEEE Trans. Knowl. Data Eng.2
2024 VFDV-IM: An Efficient and Securely Vertical Federated Data Valuation
Xiaokai Zhou, Xiao Yan 0002, Hao Huang 0001, Quanqing Xu, Qinbo Zhang, Yen Jerome, Zhaohui Cai, Jiawei Jiang 0001
DASFAA (1)4
2024 Self-Supervised Learning for Graph Dataset Condensation
abstract
Graph dataset condensation (GDC) reduces a dataset with many graphs into a smaller dataset with fewer graphs while maintaining model training accuracy. GDC saves the storage cost and hence accelerates training. Although several GDC methods have been proposed, they are all supervised and require massive labels for the graphs, while graph labels can be scarce in many practical scenarios. To fill this gap, we propose a self-supervised graph dataset condensation method called SGDC, which does not require label information. Our initial design starts with the classical bilevel optimization paradigm for dataset condensation and incorporates contrastive learning techniques. But such a solution yields poor accuracy due to the biased gradient estimation caused by data augmentation. To solve this problem, we introduce representation matching, which conducts training by aligning the representations produced by the condensed graphs with the target representations generated by a pre-trained SSL model. This design eliminates the need for data augmentation and avoids biased gradient. We further propose a graph attention kernel, which not only improves accuracy but also reduces running time when combined with self-supervised kernel ridge regression (KRR). To simplify SGDC and make it more robust, we adopt a adjacency matrix reusing approach, which reuses the topology of the original graphs for the condensed graphs instead of repeatedly learning topology during training. Our evaluations on seven graph datasets find that SGDC improves model accuracy by up to 9.7% compared with 5 state-of-the-art baselines, even if they use label information. Moreover, SGDC is significantly more efficient than the baselines.
Yuxiang Wang 0013, Xiao Yan 0002, Shiyu Jin, Hao Huang 0001, Quanqing Xu, Qingchen Zhang 0001, Bo Du 0001, Jiawei Jiang 0001
KDD4
2024 Incremental Maximal Clique Enumeration for Hybrid Edge Changes in Large Dynamic Graphs
abstract
Incremental maximal clique enumeration (IMCE), which maintains maximal cliques in dynamic graphs, is a fundamental problem in graph analysis. A maximal clique has a solid descriptive power of dense structures in graphs. Real-world graph data is often large and dynamic. Studies on IMCE face significant challenges in the efficiency of incremental batch computation and hybrid edge changes. Moreover, with growing graph sizes, new requirements occur on indexing global maximal cliques and obtaining maximal cliques under specific vertex scope constraints. This work presents a new data structure SOMEi to maintain intermediate maximal cliques during construction. SOMEi serves as a space-efficient index to retrieve scope-constrained maximal cliques on the fly. Based on SOMEi, we design a procedure-oriented IMCE algorithm to deal with hybrid edge changes within a unified algorithm framework. In particular, the algorithm is able to process a large batch of edge changes and significantly improve the average processing time of a single edge change through an efficient pruning strategy. Experimental results on real and synthetic graph data demonstrate that the proposed algorithm outperforms all the baselines and achieves good efficiency through pruning.
Ting Yu 0004, Ting Jiang 0006, Mohamed Jaward Bah, Chen Zhao 0019, Hao Huang 0001, Mengchi Liu, Shuigeng Zhou, Zhao Li 0007, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.5
2023 Multi-aspect Diffusion Network Inference
abstract
To learn influence relationships between nodes in a diffusion network, most existing approaches resort to precise timestamps of historical node infections. The target network is customarily assumed as an one-aspect diffusion network, with homogeneous influence relationships. Nonetheless, tracing node infection timestamps is often infeasible due to high cost, and the type of influence relationships may be heterogeneous because of the diversity of propagation media. In this work, we study how to infer a multi-aspect diffusion network with heterogeneous influence relationships, using only node infection statuses that are more readily accessible in practice. Equipped with a probabilistic generative model, we iteratively conduct a posteriori, quantitative analysis on historical diffusion results of the network, and infer the structure and strengths of homogeneous influence relationships in each aspect. Extensive experiments on both synthetic and real-world networks are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Keqi Han, Beicheng Xu, Ting Gan
WWW1
2021 Metric Learning via Penalized Optimization
abstract
Metric learning aims to project original data into a new space, where data points can be classified more accurately using kNN or similar types of classification algorithms. To avoid trivial learning results such as indistinguishably projecting the data onto a line, many existing approaches formulate metric learning as a constrained optimization problem, like finding a metric that minimizes the distance between data points from the same class, with a constraint of ensuring a certain separation for data points from different classes, and then they approximate the optimal solution to the constrained optimization in an iterative way. In order to improve the classification accuracy as much as possible, we try to find a metric that is able to minimize the intra-class distance and maximize the inter-class distance simultaneously. Towards this, we formulate metric learning as a penalized optimization problem, and provide design guideline, paradigms with a general formula, as well as two representative instantiations for the penalty term. In addition, we provide an analytical solution for the penalized optimization, with which costly computation can be avoid, and more importantly, there is no need to worry about the convergence rates or approximation ratios any more. Extensive experiments on real-world data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Yanan Peng, Ting Gan, Weiping Tu, Ruiting Zhou, Sai Wu
KDD1
2021 An Improved KNN-Based Efficient Log Anomaly Detection Method with Automatically Labeled Samples
abstract
Logs that record system abnormal states (anomaly logs) can be regarded as outliers, and the k-Nearest Neighbor (kNN) algorithm has relatively high accuracy in outlier detection methods. Therefore, we use the kNN algorithm to detect anomalies in the log data. However, there are some problems when using the kNN algorithm to detect anomalies, three of which are: excessive vector dimension leads to inefficient kNN algorithm, unlabeled log data cannot support the kNN algorithm, and the imbalance of the number of log data distorts the classification decision of kNN algorithm. In order to solve these three problems, we propose an efficient log anomaly detection method based on an improved kNN algorithm with an automatically labeled sample set. This method first proposes a log parsing method based on N-gram and frequent pattern mining (FPM) method, which reduces the dimension of the log vector converted with Term frequency.Inverse Document Frequency (TF-IDF) technology. Then we use clustering and self-training method to get labeled log data sample set from historical logs automatically. Finally, we improve the kNN algorithm using average weighting technology, which improves the accuracy of the kNN algorithm on unbalanced samples. The method in this article is validated on six log datasets with different types.
Bingming Wang, Lu Wang 0014, Qingshan Li, Yishi Zhao, Jianga Shang, Hao Huang 0001, Guoli Cheng, Jiangyi Geng
ACM Trans. Knowl. Discov. Data7
2021 Statistical Inference of Diffusion Networks
abstract
To infer structures in diffusion networks, existing approaches mostly need to know not only the final infection statuses of network nodes, but also the exact times when infections occur. In contrast, in many real-world settings, such as disease propagation, monitoring exact infection times is often infeasible due to a high cost. We investigate the problem of how to learn diffusion network structures based on only the final infection statuses of nodes. Instead of utilizing sequences of timestamps to determine potential parent-child influence relationships between nodes, we propose to find influence relationships with high statistical significance. To this end, we design a probabilistic generative model of the final infection statuses to quantitatively measure the likelihood of potential structures of the objective diffusion network, taking into account network complexity. Based on this model, we can infer an appropriate number of most probable parent nodes for each node in the network. Furthermore, to reduce redundant inference computations, we are able to preclude insignificant candidate parent nodes from being considered during inferencing, if their infections have little correlation with the infections of the corresponding child nodes. Extensive experiments on both synthetic and real-world networks offer evidence that the proposed approach is effective and efficient.
Hao Huang 0001, Qian Yan 0001, Lu Chen 0001, Yunjun Gao, Christian S. Jensen
IEEE Trans. Knowl. Data Eng.1
2020 From Code to Natural Language: Type-Aware Sketch-Based Seq2Seq Learning
Yuhang Deng, Hao Huang 0001, Xu Chen 0042, Zuopeng Liu, Sai Wu, Jifeng Xuan, Zongpeng Li
DASFAA (1)2
2020 Statistical Estimation of Diffusion Network Topologies
abstract
Reconstructing the topology of a diffusion network based on observed diffusion results is an open challenge in data mining. Existing approaches mostly assume that the observed diffusion results are available and consist of not only the final infection statuses of nodes, but also the exact timestamps that pinpoint when infections occur. Nonetheless, the exact infection timestamps are often unavailable in practice, due to a high cost and uncertainties in the monitoring of node infections. In this work, we investigate the problem of how to infer the topology of a diffusion network from only the final infection statuses of nodes. To this end, we propose a new scoring criterion for diffusion network reconstruction, which is able to estimate the likelihood of potential topologies of the objective diffusion network based on infection status results with a relatively low statistical error. As the proposed scoring criterion is decomposable, our problem is transformed into finding for each node in the network a set of most probable parent nodes that maximizes the value of a local score. Furthermore, to eliminate redundant computations during the search of most probable parent nodes, we identify insignificant candidate parent nodes by checking whether their infections have negative or extremely low positive correlations with the infections of a corresponding child node, and exclude them from the search space. Extensive experiments on both synthetic and real-world networks are conducted, and the results verify the effectiveness and efficiency of our approach.
Keqi Han, Hao Huang 0001, Yunjun Gao
ICDE5
2020 LERI: Local Exploration for Rare-Category Identification
abstract
To identify the data examples of rare categories that form small compact clusters in large data sets, existing approaches mostly require enough labeled data examples as a training set to learn a classifier, assuming that the rare-category clusters are spherical or nearly spherical. Nonetheless, a large enough training set is usually difficult to obtain in practice, and rare categories in many real-world applications often form small compact clusters with arbitrary shapes. In this paper, we investigate how to identify all data examples of a rare category with an arbitrary shape based on only one seed (i.e., a labeled rare-category data example). Instead of finding a compact and spherical local region around the seed, we locally explore the data set from the seed by continuously searching and visiting the k-nearest neighbors of each newly visited data example. The local exploration connects the data examples in the objective rare category by the relationship of k-nearest neighbors, and meanwhile, suspected external data examples are filtered out if they are not close enough to any visited data example. Experimental results on both synthetic and real-world data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Qian Yan 0001, Wei Lu 0015, Huaizhong Lin, Yunjun Gao, Lei Chen 0002
IEEE Trans. Knowl. Data Eng.1
2018 Efficient and Scalable Mining of Frequent Subgraphs Using Distributed Graph Processing Systems
Hao Huang 0001, Wei Lu 0015, Zhe Peng, Xiaoyong Du 0001
DASFAA (1)2
2018 Evaluation of local community metrics: from an experimental perspective
Lianhang Ma, Kevin Chiew, Hao Huang 0001, Qinming He
J. Intell. Inf. Syst.3
2018 Mining frequent subgraphs from tremendous amount of small graphs using MapReduce
Zhe Peng, Wei Lu 0015, Hao Huang 0001, Xiaoyong Du 0001, Feng Zhao 0009, Anthony K. H. Tung
Knowl. Inf. Syst.4
2018 MSQL+: a Plugin Toolkit for Similarity Search under Metric Spaces in Distributed Relational Database Systems
abstract
Similarity search is a primitive operation in various database applications. Thus far, a large number of access methods have been proposed to accelerate the similarity query processing. Nonetheless, these methods mostly focus on developing standalone systems by proposing new indices. Given the fact that existing RDBMS merely support traditional indices, it is of great necessity and practical importance to develop a standard RDBMS built-in index based approach to speeding up the query processing. In this demonstration, we introduce MSQL+, a plugin toolkit that enable users to answer similarity queries in metric spaces simply using standard SQL statements. This toolkit can help existing RDBMS to effectively and efficiently handle with big data due to the following three advantages. First, MSQL+ enables users to find similar objects by submitting SELECT-FROM-WHERE statements so that it can be easily integrated into existing RDBMS. Second, MSQL+ works in a more general data space. Objects of any type can be indexed by B + -trees and the query processing can be boosted by using index seeks, as long as the similarity function is metric. Third, MSQL+ supports the parallelization of both pre-processing and query processing in distributed RDBMS.
Wei Lu 0015, Xinyi Zhang 0002, Zhiyu Shui, Zhe Peng, Xiao Zhang 0001, Xiaoyong Du 0001, Hao Huang 0001, Anqun Pan, Haixiang Li
Proc. VLDB Endow.7
2017 Group-Level Influence Maximization with Budget Constraint
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Wei Lu 0015, Qinming He
DASFAA (1)2
2017 False data separation for data security in smart grids
Hao Huang 0001, Qian Yan 0001, Wei Lu 0015, Zhenguang Liu, Zongpeng Li
Knowl. Inf. Syst.1
2016 Fast Rare Category Detection Using Nearest Centroid Neighborhood
Hao Huang 0001, Yunjun Gao, Tieyun Qian, Liang Hong 0001, Zhiyong Peng 0001
APWeb (1)2
2016 Modeling for Noisy Labels of Crowd Workers
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Chen Ying, Qingyang Hu, Tieyun Qian, Qinming He
APWeb (2)2
2016 Mining Arbitrary Shaped Clusters and Outputting a High Quality Dendrogram
Hao Huang 0001, Shuangke Wu, Yunjun Gao, Wei Lu 0015, Qinming He
DEXA (1)1
2016 A formalized framework for incorporating expert labels in crowdsourcing environment
Qingyang Hu, Qinming He, Hao Huang 0001, Kevin Chiew, Zhenguang Liu
J. Intell. Inf. Syst.3
2015 Rare Category Exploration on Linear Time Complexity
Zhenguang Liu, Hao Huang 0001, Qinming He, Kevin Chiew, Yunjun Gao
DASFAA (2)2
2014 Towards effective and efficient mining of arbitrary shaped clusters
abstract
Mining arbitrary shaped clusters in large data sets is an open challenge in data mining. Various approaches to this problem have been proposed with high time complexity. To save computational cost, some algorithms try to shrink a data set size to a smaller amount of representative data examples. However, their user-defined shrinking ratios may significantly affect the clustering performance. In this paper, we present CLASP an effective and efficient algorithm for mining arbitrary shaped clusters. It automatically shrinks the size of a data set while effectively preserving the shape information of clusters in the data set with representative data examples. Then, it adjusts the positions of these representative data examples to enhance their intrinsic relationship and make the cluster structures more clear and distinct for clustering. Finally, it performs agglomerative clustering to identify the cluster structures with the help of a mutual k-nearest neighbors-based similarity metric called Pk. Extensive experiments on both synthetic and real data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Yunjun Gao, Kevin Chiew, Lei Chen 0002, Qinming He
ICDE1
2014 Learning from Crowds under Experts' Supervision
Qingyang Hu, Qinming He, Hao Huang 0001, Kevin Chiew, Zhenguang Liu
PAKDD (1)3
2014 Rare Category Detection on O(dN) Time Complexity
Zhenguang Liu, Hao Huang 0001, Qinming He, Kevin Chiew, Lianhang Ma
PAKDD (2)2
2014 Recovering Missing Labels of Crowdsourcing Workers
abstract
Data sets collected from crowdsourcing platforms are well known for their cheap costs. But cheap costs may lead to low quality, i.e., labels may be incorrect or missing. Most of the existing work focuses on modeling the labeling errors of crowd workers, but missing labels can also cause problems when modeling the data. In this paper, we present an algorithm to predict the missing labels of crowd workers, in which we adopt thoughts from semi-supervised learning and utilize the particular consistency between crowd workers. We also define the consistency between workers by crowd labels and develop an algorithm to learn them from the data automatically. Experiments on both benchmark and real data show that our algorithm outperforms traditional semi-supervised learning algorithms in predicting missing labels, and the recovered crowd labels are capable of predicting the ground truth and reflecting real properties of crowd workers.
Qingyang Hu, Kevin Chiew, Hao Huang 0001, Qinming He
SDM3
2014 Toward seed-insensitive solutions to local community detection
Lianhang Ma, Hao Huang 0001, Qinming He, Kevin Chiew, Zhenguang Liu
J. Intell. Inf. Syst.2
2014 Mining regional co-location patterns with kNNG
Feng Qian 0006, Kevin Chiew, Qinming He, Hao Huang 0001
J. Intell. Inf. Syst.4
2013 GMAC: A Seed-Insensitive Approach to Local Community Detection
Lianhang Ma, Hao Huang 0001, Qinming He, Kevin Chiew, Jianan Wu, Yanzhe Che
DaWaK2
2013 Discovery of Regional Co-location Patterns with k-Nearest Neighbor Graph
Feng Qian 0006, Kevin Chiew, Qinming He, Hao Huang 0001, Lianhang Ma
PAKDD (1)4
2013 Commodity query by snapping
abstract
Commodity information such as prices and public reviews is always the concern of consumers. Helping them conveniently acquire these information as an instant reference is often of practical significance for their purchase activities. Nowadays, Web 2.0, linked data clouds, and the pervasiveness of smart hand held devices have created opportunities for this demand, i.e., users could just snap a photo of any commodity that is of interest at anytime and anywhere, and retrieve the relevant information via their Internet-linked mobile devices. Nonetheless, compared with the traditional keyword-based information retrieval, extracting the hidden information related to the commodities in photos is a much more complicated and challenging task, involving techniques such as pattern recognition, knowledge base construction, semantic comprehension, and statistic deduction. In this paper, we propose a framework to address this issue by leveraging on various techniques, and evaluate the effectiveness and efficiency of this framework with experiments on a prototype.
Hao Huang 0001, Yunjun Gao, Kevin Chiew, Qinming He, Lu Chen 0001
SIGIR1
2013 Browse with a social web directory
abstract
Browse with either web directories or social bookmarks is an important complementation to search by keywords in web information retrieval. To improve users' browse experiences and facilitate the web directory construction, in this paper, we propose a novel browse system called Social Web Directory (SWD for short) by integrating web directories and social bookmarks. In SWD, (1) web pages are automatically categorized to a hierarchical structure to be retrieved efficiently, and (2) the popular web pages, hottest tags, and expert users in each category are ranked to help users find information more conveniently. Extensive experimental results demonstrate the effectiveness of our SWD system.
Hao Huang 0001, Yunjun Gao, Lu Chen 0001, Kevin Chiew, Qinming He
SIGIR1
2013 CLOVER: a faster prior-free approach to rare-category detection
Hao Huang 0001, Qinming He, Kevin Chiew, Feng Qian 0006, Lianhang Ma
Knowl. Inf. Syst.1
2011 RADAR: Rare Category Detection via Computation of Boundary Degree
Hao Huang 0001, Qinming He, Jiangfeng He, Lianhang Ma
PAKDD (2)1