VLDB 2026 Research / reviewers in the wild / expert
Liang Hong 0001
dblp:92/399-1
· DBLP profile ↗
26ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-1466-9843ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Centric Dual-Layer Adversarial Learning for Fraud Detection under Label Delay and Concept DriftabstractBanks often uncover new fraud cases by tracing connections from previously confirmed ones. In practice, this process faces two intertwined challenges: audits delay label confirmation, and fraud groups continually evolve their behaviors. These challenges manifest as label delay (LD) and concept drift (CD), which jointly degrade fraud detection. Errors introduced by delayed labels can propagate along network links, further amplifying the impact of concept drift. We propose a Data-Centric Dual-Layer Adversarial Framework (DDAF) that jointly addresses LD and CD by improving data robustness in a backbone-agnostic way. DDAF introduces a counterfactual attribution module to disentangle and quantify the individual and interaction effects of LD and CD. Guided by such attributions, we construct adversarial training graphs at two levels: node-level feature perturbations to mitigate delay-induced bias, and community-level structural perturbations to adapt to evolving group behavior. This dual-layer design coordinates correction across node and community scales, reducing error accumulation over time. On public datasets, DDAF improves AUC by 3.4 points and F1 by 3.5 points over strong baselines. In a large commercial bank deployment, DDAF identified 1,449 previously unknown fraud accounts, and expanded risk coverage in daily operations. Mingxuan Shen, Liang Hong 0001, Qingying Xu |
KDD (1) | 2 |
| 2026 | Hyper-Relational Knowledge Graph Representation Learning Based on Multi-Granularity Semantic Aware Message PassingabstractReal-world networks containing hyper-relational facts can be represented by Hyper-relational Knowledge Graphs (HKGs), which contain rich semantics compared to hypergraphs. In an HKG, multiple entities with a shared property playing different roles form a hyperedge. To mitigate information loss in HKG representation learning, it is important to integrate entity-level roles and hyperedge-level properties. However, such multi-grained semantics are often inadequately learned, as conventional message passing only operates along one granularity, i.e, hyperedges. We propagate information between entities and hyperedges to aggregate entity-level roles at the hyperedge level for message passing. In this paper, we propose an HKG representation learning method based on Multi-granularity Semantic aware Message Passing (HyperSMP), preserving both semantic and structural information. Specifically, we design a fine-grained aggregation layer in HyperSMP to aggregate different roles of entities in hyperedge embeddings using an entity-level attention mechanism. Based on hyperedge embeddings, we propose a granularity-aligned propagation layer that recursively propagates information from hyperedges to entities and properties, capturing message passing paths in and out of hyperedges, respectively. As a result, multi-grained semantics are learned through unification at the hyperedge level to adapt to the message passing mechanism, and the high-order structure is captured through entity information exchange channeled via hyperedge embeddings. Further, HyperSMP discovers multi-hop associations among entities by stacking the above layers. Experimental results on public datasets show that HyperSMP outperforms state-of-the-art methods by at least 9.66% in F1 score. Qingying Xu, Liang Hong 0001, Mingxuan Shen, Aoyuan Jiang |
KDD (1) | 2 |
| 2025 | Disclosing Actual Controller based on Equity Knowledge Graph LearningabstractDisclosing Actual Controllers (ACs) of a company has been the basis for financial risk governance. A shareholder in a winning stable coalition, where members make consistent decisions and win in votes, is considered an AC. However, existing methods fail to discover stable coalitions due to the ignorance of various relations other than the shareholding relation among shareholders, such as kinship, subsidiary and so on. Moreover, the above relations form a large-scale equity network, which brings challenges for efficiently identifying winning stable coalitions. Qingying Xu, Liang Hong 0001, Mingxuan Shen, Baokun Yi |
KDD (1) | 2 |
| 2024 | KGRED: Knowledge-graph-based rule discovery for weakly supervised data labeling
Liang Hong 0001 |
Inf. Process. Manag. | 2 |
| 2023 | RoRED: Bootstrapping labeling rule discovery for robust relation extraction
Liang Hong 0001, Haoshuai Xu |
Inf. Sci. | 2 |
| 2023 | TrajMesa: A Distributed NoSQL-Based Trajectory Data Management SystemabstractWith the development of positioning technology, a large number of trajectories have been generated, which are very useful for many urban applications. However, it is challenging to manage trajectory data for its spatio-temporal dynamics and high-volume properties. Existing trajectory data management frameworks suffer from efficiency or scalability problem, and only support limited trajectory query types. This paper takes the first attempt to build a holistic distributed NoSQL trajectory storage engine, named TrajMesa, based on GeoMesa, an open-source indexing toolkit for spatio-temporal data. TrajMesa can manage a prohibitively large number of trajectories, and support plenty of query types efficiently. Specifically, we first design a novel trajectory storage schema, which reduces the storage size tremendously. We then devise a novel indexing key schema for time ranges, based on which ID temporal query can be supported efficiently. To reduce the amount of retrieved trajectory data for a spatial range query, we innovatively propose a position code to indicate the spatial location of trajectories accurately. We also propose a bunch of pruning strategies for similarity query and k-NN query in the NoSQL environment. Extensive experiments are conducted using two real datasets and one synthetic dataset, verifying the powerful query efficiency and scalability of TrajMesa. Huajun He, Rubin Wang, Sijie Ruan, Tianfu He, Jie Bao 0003, Junbo Zhang 0004, Liang Hong 0001, Yu Zheng 0004 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2023 | Temporal Graph CubeabstractData warehouse and OLAP (Online Analytical Processing) are effective tools for decision support on traditional relational data and static multidimensional network data. However, many real-world multidimensional networks are often modeled as temporal multidimensional networks, where the edges in the network are associated with temporal information. Such temporal multidimensional networks typically cannot be handled by traditional data warehouse and OLAP techniques. To fill this gap, we propose a novel data warehouse model, named$\mathsf {Temporal{ }\; Graph{ }\; Cube}$, to support OLAP queries on temporal multidimensional networks. Through supporting OLAP queries in any time range, users can obtain summarized information of the network in the time range of interest, which cannot be derived by using traditional static graph OLAP techniques. We propose a segment-tree based indexing technique to speed up the OLAP queries, and also develop an index-updating technique to maintain the index when the temporal multidimensional network evolves over time. In addition, we also propose a novel concept called$\mathsf {similarity{ }\; of{ }\; snapshots}$which shows a strong correlation with the efficiency of indexing technique and can provide a good reference on the necessity of building the index. The results of extensive experiments on two large real-world datasets demonstrate the effectiveness and efficiency of the proposed method. Guoren Wang, Yue Zeng 0004, Rong-Hua Li 0001, Hongchao Qin, Xuanhua Shi, Yubin Xia, Xuequn Shang 0001, Liang Hong 0001 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2022 | Evaluating Knowledge Graph Accuracy Powered by Optimized Human-machine CollaborationabstractEstimating the accuracy of an automatically constructed knowledge graph (KG) becomes a challenging task as the KG often contains a large number of entities and triples. Generally, two major components information extraction (IE) and entity linking (EL) are involved in KG construction. However, the existing approaches just focus on evaluating the triple accuracy that indicates the IE quality, completely ignoring the entity accuracy. Motivated by the fact that the major advance of machines is the strong computing power while humans are skilled in correctness verification, we propose an efficient interactive method to reduce the overall cost for evaluating the KG quality, which produces accuracy estimates with a statistical guarantee for both triples and entities. Instead of annotating triples and entities separately, we design a general annotation cost that blends triples and entities generated from the identical source text. During human verification, the machine can pre-compute and infer triples to be annotated in the next round by speculating human feedback. The human-machine collaborative mechanism is optimized by formulating an order selection problem of triples which is NP-hard. Thus, a Monte Carlo Tree Search is proposed to guide the annotation process by finding an approximate solution. Extensive experiments demonstrate that our method takes less annotation cost while yielding higher accuracy estimation quality compared to the state-of-the-art approaches. Yifan Qi, Weiguo Zheng, Liang Hong 0001, Lei Zou 0001 |
KDD | 3 |
| 2021 | Distributed Spatio-Temporal k Nearest Neighbors JoinabstractThe rapid development of positioning technology produces an extremely large volume of spatio-temporal data with various geometry types such as point, line string, polygon, or a mixed combination of them. As one of the most basic but time-consuming operations, k nearest neighbors join (kNN join) has attracted much attention. However, most existing works for kNN join either ignore temporal information or consider point data only. Rubin Wang, Junwen Liu, Zisheng Yu, Huajun He, Tianfu He, Sijie Ruan, Jie Bao 0003, Chao Chen 0004, Fuqiang Gu, Liang Hong 0001, Yu Zheng 0004 |
SIGSPATIAL/GIS | 11 |
| 2020 | Discovering Real-Time Reachable Area Using Trajectory Connections
Jie Bao 0003, Huajun He, Sijie Ruan, Tianfu He, Liang Hong 0001, Zhongyuan Jiang, Yu Zheng 0004 |
DASFAA (2) | 6 |
| 2020 | Efficient Path Query Processing Over Massive Trajectories on the CloudabstractA path query aims to find trajectories passing a given sequence of connected road segments within a time period. It is very useful in many urban applications: 1) traffic modeling, 2) frequent path mining, 3) intersection coordination, and 4) traffic anomaly detection. Existing solutions for path query processing are implemented based on single machines, which are not efficient for the following tasks: 1) indexing large-scale historical data; 2) handling real-time trajectory updates; and 3) processing concurrent path queries from urban data mining applications. In this paper, we design and implement a cloud-based path query processing framework based on Microsoft Azure. We modify existing suffix tree structure to index trajectories using Azure Table. The proposed system consists of two main parts: 1) back-end processing, which performs pre-processing (i.e., parsing and map-matching) and index building tasks with a distributed computing platform (i.e., Storm) used to efficiently handle massive real-time trajectory updates; and 2) query processing, which answers path queries using Azure Storm to improve efficiency and overcome I/O bottleneck. Extensive experiments are performed based on the real-time taxi trajectories from Guiyang City, the capital of Guizhou Province, China to confirm the system efficiency. We also demonstrate a real deployed traffic analysis system based on our query processing framework. Sijie Ruan, Jie Bao 0003, Yingcai Wu, Liang Hong 0001, Yu Zheng 0004 |
IEEE Trans. Big Data | 6 |
| 2018 | Time and Location Aware Points of Interest Recommendation in Location-Based Social Networks
Tieyun Qian, Liang Hong 0001, Zhenni You |
J. Comput. Sci. Technol. | 3 |
| 2017 | SoC-constrained team formation with self-organizing mechanism in social networks
Yuling Shi, Zhiyong Peng 0001, Liang Hong 0001 |
Knowl. Based Syst. | 3 |
| 2017 | A cloud-based taxi trace mining framework for smart cityabstractSummary As a well‐known field of big data applications, smart city takes advantage of massive data analysis to achieve efficient management and sustainable development in the current worldwide urbanization process. An important problem in smart city is how to discover frequent trajectory sequence pattern and cluster trajectory. To solve this problem, this paper proposes a cloud‐based taxi trajectory pattern mining and trajectory clustering framework for smart city. Our work mainly includes (1) preprocessing raw Global Positioning System trace by calling the Baidu API Geocoding; (2) proposing a distributed trajectory pattern mining (DTPM) algorithm based onSpark; and (3) proposing a distributed trajectory clustering (DTC) algorithm based onSpark. The proposed DTPM algorithm and DTC algorithm can overcome the high input/output overhead and communication overhead by adopting in‐memory computation. In addition, the proposed DTPM algorithm can avoid generating redundant local trajectory patterns to significantly improve the overall performance. The proposed DTC algorithm can enhance the performance of trajectory similarity computation by transforming the trajectory similarity calculation into AND and OR operators. Experimental results indicate that DTPM algorithm and DTC algorithm can significantly improve the overall performance and scalability of trajectory pattern mining and trajectory clustering on massive taxi trace data. Copyright © 2016 John Wiley & Sons, Ltd. Jin Liu 0016, Xiao Yu 0008, Zheng Xu 0001, Kim-Kwang Raymond Choo, Liang Hong 0001, Xiaohui Cui |
Softw. Pract. Exp. | 5 |
| 2016 | Fast Rare Category Detection Using Nearest Centroid Neighborhood
Hao Huang 0001, Yunjun Gao, Tieyun Qian, Liang Hong 0001, Zhiyong Peng 0001 |
APWeb (1) | 5 |
| 2016 | Developer Recommendation with Awareness of Accuracy and CostabstractAs the scale and complexity of software products increase, software maintenance on bug resolution has become a challenging work.In the process of software implementation, developers often use bug reports, source code and change history to help solve bugs.However, hundreds of bug reports are being submitted every day.It is time-consuming and effortless for developers to review all the bug reports.To facilitate the assignment of bug reports, existing developer recommendation systems typically recommend the developer who has the fullest potential.However, bug reports are highly varied; time that the developers may spend fixing them is also important.To address the problem of developer recommendation, we propose a developer recommendation system with awareness of accuracy and cost (DRAC).This recommendation system is based on modern portfolio theory by striking a balance between accuracy and cost (time).We evaluate our approach with experiments on data collected from Bugzilla 1 . Jin Liu 0016, Yiqiuzi Tian, Liang Hong 0001, Xu Chen 0017, Zhou Xu 0003 |
SEKE | 3 |
| 2016 | Online Subgraph Skyline Analysis over Knowledge GraphsabstractSubgraph search is very useful in many real-world applications. However, users may be overwhelmed by the masses of matches. In this paper, we propose a subgraph skyline analysis problem, denoted as S2A, to support more complicated analysis over graph data. Specifically, given a large graph G and a query graph q, we want to find all the subgraphs g in G, such that g is graph isomorphic to q and not dominated by any other subgraphs. In order to improve the efficiency, we devise a hybrid feature encoding incorporating both structural and numeric features based on a partitioning strategy, and discuss how to optimize the space partitioning. We also present a skylayer index to facilitate the dynamic subgraph skyline computation. Moreover, an attribute cluster-based method is proposed to deal with the curse of dimensionality. Extensive experiments over real datasets confirm the effectiveness and efficiency of our algorithm. Weiguo Zheng, Xiang Lian 0001, Lei Zou 0001, Liang Hong 0001, Dongyan Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Detecting urban black holes based on human mobility dataabstractMany types of human mobility data, such as flows of taxicabs, card swiping data of subways, bike trip data and Call Details Records (CDR), can be modeled by a Spatio-Temporal Graph (STG). STG is a directed graph in which vertices and edges are associated with spatio-temporal properties (e.g. the traffic flow on a road and the geospatial location of an intersection). In this paper, we instantly detect interesting phenomena, entitled black holes and volcanos, from an STG. Specifically, a black hole is a subgraph (of an STG) that has the overall inflow greater than the overall outflow by a threshold, while a volcano is a subgraph with the overall outflow greater than the overall inflow by a threshold (detecting volcanos from an STG is proved to be equivalent to the detection of black holes). The online detection of black holes/volcanos can timely reflect anomalous events, such as disasters, catastrophic accidents, and therefore help keep public safety. The patterns of black holes/volcanos and the relations between them reveal human mobility patterns in a city, thus help formulate a better city planning or improve a system's operation efficiency. Based on a well-designed STG index, we propose a two-step black hole detection algorithm: The first step identifies a set of candidate grid cells to start from; the second step expands an initial edge in a candidate cell to a black hole and prunes other candidate cells after a black hole is detected. Then, we adapt this detection algorithm to a continuous black hole detection scenario. We evaluate our method based on Beijing taxicab data and the bike trip data in New York, finding urban anomalies and human mobility patterns. Liang Hong 0001, Yu Zheng 0004, Duncan Yung, Jingbo Shang, Lei Zou 0001 |
SIGSPATIAL/GIS | 1 |
| 2015 | Context-Aware Recommendation Using Role-Based Trust NetworkabstractRecommender systems have been studied comprehensively in both academic and industrial fields over the past decade. As user interests can be affected by context at any time and any place in mobile scenarios, rich context information becomes more and more important for personalized context-aware recommendations. Although existing context-aware recommender systems can make context-aware recommendations to some extent, they suffer several inherent weaknesses: (1) Users’ context-aware interests are not modeled realistically, which reduces the recommendation quality; (2) Current context-aware recommender systems ignore trust relations among users. Trust relations are actually context-aware and associated with certain aspects (i.e., categories of items) in mobile scenarios. In this article, we define a term role to model common context-aware interests among a group of users. We propose an efficient role mining algorithm to mine roles from a “user-context-behavior” matrix, and a role-based trust model to calculate context-aware trust value between two users. During online recommendation, given a user u in a context c , an efficient weighted set similarity query (WSSQ) algorithm is designed to build u ’s role-based trust network in context c . Finally, we make recommendations to u based on u ’s role-based trust network by considering both context-aware roles and trust relations. Extensive experiments demonstrate that our recommendation approach outperforms the state-of-the-art methods in both effectiveness and efficiency. Liang Hong 0001, Lei Zou 0001, Jian Wang 0018, Jilei Tian |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | Subgraph Matching with Set Similarity in a Large Graph DatabaseabstractIn real-world graphs such as social networks, Semantic Web and biological networks, each vertex usually contains rich information, which can be modeled by a set of tokens or elements. In this paper, we study a subgraph matching with set similarity (SMS2) query over a large graph database, which retrieves subgraphs that are structurally isomorphic to the query graph, and meanwhile satisfy the condition of vertex pair matching with the (dynamic) weighted set similarity. To efficiently process the SMS2query, this paper designs a novel lattice-based index for data graph, and lightweight signatures for both query vertices and data vertices. Based on the index and signatures, we propose an efficient two-phase pruning strategy including set similarity pruning and structure-based pruning, which exploits the unique features of both (dynamic) weighted set similarity and graph topology. We also propose an efficient dominating-set-based subgraph matching algorithm guided by a dominating set selection algorithm to achieve better query performance. Extensive experiments on both real and synthetic datasets demonstrate that our method outperforms state-of-the-art methods by an order of magnitude. Liang Hong 0001, Lei Zou 0001, Xiang Lian 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Retargeting Semantically-Rich PhotosabstractSemantically-rich photos contain a rich variety of semantic objects (e.g., pedestrians and bicycles). Retargeting these photos is a challenging task since each semantic object has fixed geometric characteristics. Shrinking these objects simultaneously during retargeting is prone to distortion. In this paper, we propose to retarget semantically-rich photos by detecting photo semantics from image tags, which are predicted by a multi-label SVM. The key technique is a generative model termed latent stability discovery (LSD). It can robustly localize various semantic objects in a photo by making use of the predicted noisy image tags. Based on LSD, a feature fusion algorithm is proposed to detect salient regions at both the low-level and high-level. These salient regions are linked into a path sequentially to simulate human visual perception . Finally, we learn the prior distribution of such paths from aesthetically pleasing training photos. The prior enforces the path of a retargeted photo to be maximally similar to those from the training photos. In the experiment, we collect 217 1600 ×1200 photos, each containing over seven salient objects. Comprehensive user studies demonstrate the competitiveness of our method. Meng Wang 0001, Liqiang Nie, Liang Hong 0001, Yong Rui, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | Efficient Subgraph Skyline Search Over Large GraphsabstractSubgraph search is very useful in many real-world applications. However, users may be overwhelmed by the masses of matches. In this paper, we propose subgraph skyline search problem, denoted as S3, to support more complicated analysis over graph data. Specifically, given a large graph G and a query graph q, we want to find all the subgraphs g in G, such that g is graph isomorphic to q and not dominated by any other subgraphs. In order to improve the efficiency, we devise a hybrid feature encoding incorporating both structural and numeric features. Moreover, we present some optimizations based on partitioning strategy. We also propose a skylayer index to facilitate the dynamic subgraph skyline computation. Extensive experiments over real dataset confirm the effectiveness and efficiency of our algorithm. Weiguo Zheng, Lei Zou 0001, Xiang Lian 0001, Liang Hong 0001, Dongyan Zhao 0001 |
CIKM | 4 |
| 2014 | Novel Community Recommendation Based on a User-Community Total Relation
Zhiyong Peng 0001, Liang Hong 0001, Haiping Peng |
DASFAA (2) | 3 |
| 2014 | An adaptive backoff algorithm for multi-channel CSMA in wireless sensor networks
Yantao Li 0001, Gang Zhou 0002, Liang Hong 0001 |
Neural Comput. Appl. | 4 |
| 2012 | A Latent Topic Based Collaborative Filtering Recommendation Algorithm for Web CommunitiesabstractProviding personalized high quality community recommendation for Web community members has become increasingly important. Traditional collaborative filtering methods based on explicit topic associations cannot solve the information sparsity problem. The recommendation methods based on latent topic association results in inaccurate results. To solve the above problems, we propose a collaborative Web community recommendation algorithm based on latent topic. Our algorithm generates the latent link between communities and members using latent topic associations to overcome the sparsity problem. Our algorithm also reduces inaccurate results by combining similar members' behaviors and interests. The experiment indicates that our recommendation algorithm has higher recommendation accuracy than traditional methods. Zhiyong Peng 0001, Liang Hong 0001, Dawen Jia |
WISA | 3 |
| 2009 | Event-Based Location Dependent Data Services in Mobile WSNsabstractMobile sensors are widely deployed in Wireless Sensor Networks (WSNs) to satisfy emerging application requirements. Specifically, processing location dependent queries in mobile WSNs is still a challenging problem due to sensor mobility. We present an Event-based Location Dependent Query (ELDQ) model that continuously aggregate data in specific areas around mobile sensors of interests to provide event-based location dependent data services to the users. ELDQs generalize several typical query types and are important in many applications. However, existing approaches are incapable of efficiently answering ELDQs. In this paper, we propose a set of techniques to process ELDQs while optimizing system performance. Cost analysis and simulation results indicate that our techniques greatly reduce the cost of processing ELDQs while achieving relatively high accuracy and short response time. Liang Hong 0001, Yafeng Wu, Sang Hyuk Son, Yansheng Lu |
RTCSA | 1 |