EDBT 2026 Demo / reviewers in the wild / expert
Xia Chen 0004
dblp:06/1899-4
· DBLP profile ↗
12ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-8223-5641ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Adaptive Label Propagation for Group Anomaly Detection in Large-Scale NetworksabstractThis paper concentrates on group anomalies in general large-scale networks. Existing algorithms on group anomalies mainly focus on homogeneous or bipartite networks, and thus are difficult to apply to heterogeneous networks directly. Moreover, these algorithms follow the non-overlapping hypothesis of groups implicitly, which is improper in many scenarios. For example, fraud users in Alibaba E-commerce platform may join more than one organization at the same time. In this paper, we introduce a novel algorithm calledAdaptive Label Propagation (ALP)to solve these problems. ALP is designed based on label propagation (LP) frameworks, for the reason that LP-based frameworks are simple in thought and easy to scale. ALP is able to find overlapping groups by label propagation with belonging coefficients, and can be applied to heterogeneous networks for its design of adaptive neighbor weighting. Assigning different weights to neighbors in label propagation is a challenging task. Inspired by the combinatorial multi-armed bandit mechanism, ALP views the neighbors of each node as arms to be selected, and iteratively updates their weights by evaluating their expected rewards in following iterations. Experiments are conducted on four real-world networks (including two bipartite ones and two heterogeneous ones). The results show that LP-based methods are effective for detecting group anomalies, and the comparison results with several state-of-the-art label propagation based community detection methods show the effectiveness of the proposed method. Zhao Li 0007, Xia Chen 0004, Junshuai Song, Jun Gao 0003 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | CMAL: Cost-Effective Multi-Label Active Learning by Querying SubexamplesabstractMulti-label active learning (MAL) aims to learn an accurate multi-label classifier by selecting which examples (or example-label pairs) will be annotated and reducing query effort. MAL is a more complicated and expensive process than single-label active learning, due to one example can be associated with a set of non-exclusive labels and the annotator has to scrutinize the whole example and label space to provide correct annotations. Instead of scrutinizing the whole example for annotation, we may just examine some of its subexamples with respect to a label for annotation. In this way, we can not only save the annotation cost but also speedup the annotation process. Given this observation, we introduce CMAL, a two-stage Cost-effective MAL strategy (CMAL) by querying subexamples. CMAL first selects the most informative example-label pairs by leveraging uncertainty, label correlation and label space sparsity. Specifically, the uncertainty of a label to an example can be reduced if its correlated labels already annotated to the example, and its uncertainty can be reduced also if more examples annotated to this label. Next, CMAL greedily queries the most probable positive subexample-label pairs of the selected example-label pair. In addition, we propose rCMAL to account for the representative of examples to more reliably select example-label pairs in the first stage. Extensive experiments on multi-label datasets from diverse domains show that our proposed CMAL and rCMAL can better save the query cost than state-of-the-art MAL methods. The contribution of leveraging label correlation, label sparsity, and representative for saving cost is also confirmed. Guoxian Yu, Xia Chen 0004, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001, Xiangliang Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | From Community Search to Community Understanding: A Multimodal Community Query EngineabstractIn this demo, we present an online multi-modal community query engine (MQE) on Alibaba's billion-scale heterogeneous network. MQE has two distinct features in comparison with existing community query engines. Firstly, MQE supports multimodal community search on heterogeneous graphs with keyword and image queries. Secondly, to facilitate community understanding in real business scenarios, MQE generates natural language descriptions for the retrieved community in combination with other useful demographic information. The distinct features of MQE benefit many downstream applications in Alibaba's e-commerce platform like recommendation. Our experiments confirm the effectiveness and efficiency of MQE on graphs with billions of edges. Zhao Li 0007, Pengcheng Zou, Xia Chen 0004, Shichang Hu, Peng Zhang 0001, Yumou Zhang, Bingsheng He, Yuchen Li 0001 |
CIKM | 3 |
| 2020 | Link Inference via Heterogeneous Multi-view Graph Neural Networks
Yuying Xing, Zhao Li 0007, Pengrui Hui, Xia Chen 0004, Guoxian Yu |
DASFAA (1) | 5 |
| 2020 | Strong Statistical Correlation Revealed by Quantum Entanglement for Supervised LearningabstractIn supervised learning, the generative approach is an important one, which obtains the generative model by learning the joint probability between features and categories. In quantum mechanics, Quantum Entanglement (QE) can provide a statistical correlation between subsystems (or attributes) that is stronger than what classical systems are able to produce. It inspires us to use entangled systems (states) to characterize this strong statistical correlation between features and categories, that is, to use the joint probability derived from QE to model the correlation. Based on the separability of the density matrix of entangled systems, this paper formally clarifies the manifestation of the strong statistical correlation revealed by QE, and implements a classification algorithm (called ECA) to verify the feasibility and superiority of the correlation in specific tasks. Since QE arises from the measurement process of entangled systems, the core of ECA is quantum measurement operations. In this paper, we use the GHZ [25] and W [22] states to prepare the entangled system and use a fully connected network layer to learn the measurement operator. It can also be understood as replacing the output layer of the Multi-Layer Perceptron (MLP) with a quantum measurement operation. The experimental results show that ECA is superior to most representative classification algorithms in multiple evaluation metrics. Junwei Zhang 0009, Yuexian Hou, Zhao Li 0007, Xia Chen 0004 |
ECAI | 5 |
| 2019 | Density Matrix Based Preference Evolution Networks for E-Commerce Recommendation
Zhao Li 0007, Xuming Pan, Donghui Ding, Xia Chen 0004, Yuexian Hou |
DASFAA (2) | 5 |
| 2019 | ActiveHNE: Active Heterogeneous Network EmbeddingabstractHeterogeneous network embedding (HNE) is a challenging task due to the diverse node types and/or diverse relationships between nodes. Existing HNE methods are typically unsupervised. To maximize the profit of utilizing the rare and valuable supervised information in HNEs, we develop a novel Active Heterogeneous Network Embedding (ActiveHNE) framework, which includes two components: Discriminative Heterogeneous Network Embedding (DHNE) and Active Query in Heterogeneous Networks (AQHN).In DHNE, we introduce a novel semi-supervised heterogeneous network embedding method based on graph convolutional neural network. In AQHN, we first introduce three active selection strategies based on uncertainty and representativeness, and then derive a batch selection method that assembles these strategies using a multi-armed bandit mechanism. ActiveHNE aims at improving the performance of HNE by feeding the most valuable supervision obtained by AQHN into DHNE. Experiments on public datasets demonstrate the effectiveness of ActiveHNE and its advantage on reducing the query cost. Xia Chen 0004, Guoxian Yu, Jun Wang 0035, Carlotta Domeniconi, Zhao Li 0007, Xiangliang Zhang 0001 |
IJCAI | 1 |
| 2019 | SHOAL: Large-scale Hierarchical Taxonomy via Graph-based Query Coalition in E-commerceabstractE-commerce taxonomy plays an essential role in online retail business. Existing taxonomy of e-commerce platforms organizes items into an ontology structure. However, the ontology-driven approach is subject to costly manual maintenance and often does not capture user's search intention, particularly when user searches by her personalized needs rather than a universal definition of the items. Observing that search queries can effectively express user's intention, we present a novel large-Scale Hierarchical taxOnomy via grAph based query coaLition ( SHOAL ) to bridge the gap between item taxonomy and user search intention. SHOAL organizes hundreds of millions of items into a hierarchical topic structure . Each topic that consists of a cluster of items denotes a conceptual shopping scenario, and is tagged with easy-to-interpret descriptions extracted from search queries. Furthermore, SHOAL establishes correlation between categories of ontology-driven taxonomy, and offers opportunities for explainable recommendation. The feedback from domain experts shows that SHOAL achieves a precision of 98% in terms of placing items into the right topics, and the result of an online A/B test demonstrates that SHOAL boosts the Click Through Rate (CTR) by 5%. SHOAL has been deployed in Alibaba and supports millions of searches for online shopping per day. Zhao Li 0007, Xia Chen 0004, Xuming Pan, Pengcheng Zou, Yuchen Li 0001, Guoxian Yu |
Proc. VLDB Endow. | 2 |
| 2018 | Cost Effective Multi-label Active Learning via Querying SubexamplesabstractMulti-label active learning addresses the scarce labeled example problem by querying the most valuable unlabeled examples, or example-label pairs, to achieve a better performance with limited query cost. Current multi-label active learning methods require the scrutiny of the whole example in order to obtain its annotation. In contrast, one can find positive evidence with respect to a label by examining specific patterns (i.e., subexample), rather than the whole example, thus making the annotation process more efficient. Based on this observation, we propose a novel two-stage cost effective multi-label active learning framework, called CMAL. In the first stage, a novel example-label pair selection strategy is introduced. Our strategy leverages label correlation and label space sparsity of multi-label examples to select the most uncertain example-label pairs. Specifically, the unknown relevant label of an example can be inferred from the correlated labels that are already assigned to the example, thus reducing the uncertainty of the unknown label. In addition, the larger the number of relevant examples of a particular label, the smaller the uncertainty of the label is. In the second stage, CMAL queries the most plausible positive subexample-label pairs of the selected example-label pairs. Comprehensive experiments on multi-label datasets collected from different domains demonstrate the effectiveness of our proposed approach on cost effective queries. We also show that leveraging label correlation and label sparsity contribute to saving costs. Xia Chen 0004, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001 |
ICDM | 1 |
| 2018 | Feature-Induced Partial Multi-label LearningabstractCurrent efforts on multi-label learning generally assume that the given labels of training instances are noise-free. However, obtaining noise-free labels is quite difficult and often impractical, and the presence of noisy labels may compromise the performance of multi-label learning. Partial multi-label learning (PML) addresses the scenario in which each instance is annotated with a set of candidate labels, of which only a subset corresponds to the ground-truth. The PML problem is more challenging than partial-label learning, since the latter assumes that only one label is valid and may ignore the correlation among candidate labels. To tackle the PML challenge, we introduce a feature induced PML approach called fPML, which simultaneously estimates noisy labels and trains multi-label classifiers. In particular, fPML simultaneously factorizes the observed instance-label association matrix and the instance-feature matrix into low-rank matrices to achieve coherent low-rank matrices from the label and the feature spaces, and a low-rank label correlation matrix as well. The low-rank approximation of the instance-label association matrix is leveraged to estimate the association confidence. To predict the labels of unlabeled instances, fPML learns a matrix that maps the instances to labels based on the estimated association confidence. An empirical study on public multi-label datasets with injected noisy labels, and on archived proteomic datasets, shows that fPML can more accurately identify noisy labels than related solutions, and consequently can achieve better performance on predicting labels of instances than competitive methods. Guoxian Yu, Xia Chen 0004, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001, Xindong Wu 0001 |
ICDM | 2 |
| 2018 | Matrix Factorization for Identifying Noisy Labels of Multi-label Instances
Xia Chen 0004, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001 |
PRICAI | 1 |
| 2017 | Semi-supervised Multi-label Linear Discriminant Analysis
Yanming Yu, Guoxian Yu, Xia Chen 0004, Yazhou Ren 0001 |
ICONIP (1) | 3 |