EDBT 2026 Demo / reviewers in the wild / expert
Jian Yu 0001
dblp:52/5812-1
· DBLP profile ↗
39ranked-venue papers in the field
3as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 14Knowledge Engineering, Semantic Web & Information Systems · 11 (3 first)Information Retrieval & Web Search · 7Database Systems & Data Management · 4Other / Interdisciplinary · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning OOD Robust Neural Operator with Risk-Averse Stochastic OptimizationabstractExisting work in physical-informed machine learning (PIML) has shown that data-driven learning of solution operators can provide a fast approximate alternative to classical numerical ordinary/partial differential equations (ODEs/PDEs) solvers. Of these, Neural Operators (NOs) have emerged as particularly promising. However, a key challenge in the field of NOs lies in developing methods that can effectively handle out-of-distribution (OOD) forecasting problems. Such problems involve the ability to adaptively learn from observations of the same dynamical system governed by ODEs/PDEs, where the underlying parameters are unknown and vary across instances. These tasks further require precise predictions even when faced with initial conditions and PDEs/ODEs parameters outside the training distribution. In this study, we consider the problem of training models in a risk-reverse manner. We introduce a risk-aware framework aimed at enhancing the OOD robustness of NOs by stochastically optimizing the conditional value-at-risk (CVAR) of a loss distribution. Through experiments on different distinct OOD tasks, our approach demonstrates a significant performance improvement over existing advanced NOs. Huafeng Liu 0001, Yiran Fu, Jingyue Shi, Liping Jing, Jian Yu 0001 |
KDD (2) | 5 |
| 2025 | Towards Robust Recommendation: A Review and an Adversarial Robustness Evaluation LibraryabstractRecently, recommender system has achieved significant success. However, due to the openness of recommender systems, they remain vulnerable to malicious attacks. Additionally, natural noise in training data and issues such as data sparsity can also degrade the performance of recommender systems. Therefore, enhancing the robustness of recommender systems has become an increasingly important research topic. In this survey, we provide a comprehensive overview of the robustness of recommender systems. Based on our investigation, we categorize the robustness of recommender systems into adversarial robustness and non-adversarial robustness. In the adversarial robustness, we introduce the fundamental principles and classical methods of recommender system adversarial attacks and defenses. In the non-adversarial robustness, we analyze nonadversarial robustness from the perspectives of data sparsity, natural noise, and data imbalance. Additionally, we summarize commonly used datasets and evaluation metrics for evaluating the robustness of recommender systems. Finally, we also discuss the current challenges in the field of recommender system robustness and potential future research directions. Additionally, to facilitate fair and efficient evaluation of attack and defense methods in adversarial robustness, we propose an adversarial robustness evaluation library–ShillingREC, and we conduct evaluations of basic attack models and recommendation models. ShillingREC project is released at https://github.com/ chengleileilei/ShillingREC. Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Deep fair clustering with multi-level decorrelation
Xiang Wang 0023, Liping Jing, Huafeng Liu 0001, Jian Yu 0001, Weifeng Geng, Gencheng Ye |
Inf. Sci. | 4 |
| 2024 | Attributed graph clustering under the contrastive mechanism with cluster-preserving augmentation
Yimei Zheng, Caiyan Jia, Jian Yu 0001 |
Inf. Sci. | 3 |
| 2024 | Structure-Driven Representation Learning for Deep ClusteringabstractAs an important branch of unsupervised learning methods, clustering makes a wide contribution in the area of data mining. It is well known that capturing the group-discriminative properties of each sample for clustering is crucial. Among them, deep clustering delivers promising results due to the strong representational power of neural networks. However, most of them adopt sample-level learning strategies, and the standalone data point barely captures its holistic cluster’s context and may undergo sub-optimal cluster assignment. To tackle this issue, we propose a Structure-driven Representation Learning (SRL) method by introducing latent structure information into the representation learning process at both the local and global levels. Specifically, a local-structure-driven sample representation strategy is proposed to approximate the estimation of data distribution, which models the neighborhood distribution of samples with potential structure information and exploits statistical dependencies between them to improve cluster consistency. A global-structure-driven cluster representation strategy is designed, where the context of each cluster is sufficiently encoded according to its samples (exemplar-theory) and corresponding prototype (prototype-theory). In this case, each cluster can only be related to its most similar samples, and different clusters are separated as much as possible. These two models are seamlessly combined into a joint optimization problem, which can be efficiently solved. Experiments on six widely-used datasets demonstrate the superiority of SRL over state-of-the-art clustering methods. Xiang Wang 0023, Liping Jing, Huafeng Liu 0001, Jian Yu 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Learning Hierarchical Preferences for Recommendation With Mixture Intention Neural Stochastic ProcessesabstractUser preferences behind users' decision-making processes are highly diverse and may range from lower-level concepts with more specific intentions and higher-level concepts with more general intentions. In this case, user preferences tend to be expressed hierarchically. However, learning such intentions with different levels from user behaviors is challenging, and remains largely neglected by the existing literature. Meanwhile, user behavior data tends to be sparse because of the limited user response and the vast combinations of users and items, which results in cold-start problems with unclear user intentions. In this paper, we propose a mixture intention neural stochastic process (MINSP), a new view of the stochastic processes family using a general meta-learning mechanism and mixture strategy for robust recommendation with hierarchical preferences modeling. By considering the recommendation process for each user as a stochastic process, MINSP defines distributions over functions and is capable of rapid adaptation to different users. To capture the user's intention on different levels, an iterative additive algorithm is proposed that minimizes the approximation error by backfitting the residuals of previous approximations. In this case, the induced tree intention hierarchies serve as an aggregated structured representation of the whole preference, summarizing the gist for convenient navigation and better generalization. Furthermore, we theoretically analyze the generalization error bound of the proposed MINSP to guarantee the model performance. Empirical results show that our approach can achieve substantial improvement over the state-of-the-art baselines in terms of recommendation performance, and obtain an interpretable hierarchical intention structure. Huafeng Liu 0001, Liping Jing, Jian Yu 0001, Michael Kwok-Po Ng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Not All Tasks Are Equal: A Parameter-Efficient Task Reweighting Method for Few-Shot Learning
Xin Liu 0086, Yilin Lyu, Liping Jing, Tieyong Zeng, Jian Yu 0001 |
ECML/PKDD (2) | 5 |
| 2023 | vMF Loss: Exploring a Scattered Intra-class Hypersphere for Few-Shot Learning
Xin Liu 0086, Shijing Wang, Kairui Zhou, Yilin Lyu, Liping Jing, Tieyong Zeng, Jian Yu 0001 |
ECML/PKDD (2) | 8 |
| 2023 | Contrastive Learning with Cluster-Preserving Augmentation for Attributed Graph Clustering
Yimei Zheng, Caiyan Jia, Jian Yu 0001 |
ECML/PKDD (1) | 3 |
| 2023 | Knowledge Graph-Enhanced Sampling for Conversational Recommendation SystemabstractThe traditional recommendation systems mainly use offline user data to train offline models, and then recommend items for online users, thus suffering from the unreliable estimation of user preferences based on sparse and noisy historical data. Conversational Recommendation System(CRS) uses the interactive form of the dialogue systems to solve the intrinsic problems of traditional recommendation systems. However, due to the lack of contextual information modeling, the existing CRS models are unable to deal with the exploitation and exploration(E&E) problem well, resulting in the heavy burden on users. To address the aforementioned issue, this work proposes a contextual information enhancement model tailored for CRS, called Knowledge Graph-enhanced Sampling(KGenSam). KGenSam integrates the dynamic graph of user interaction data with the external knowledge into one heterogeneous Knowledge Graph(KG) as the contextual information environment. Then, two samplers are designed to enhance knowledge by sampling fuzzy samples with high uncertainty for obtaining user preferences and reliable negative samples for updating recommender to achieve efficient acquisition of user preferences and model updating, and thus provide a powerful solution for CRS to deal with E&E problem. Experimental results on two real-world datasets demonstrate the superiority of KGenSam with significant improvements over state-of-the-art methods. Xiaowen Huang 0001, Lixi Zhu, Jitao Sang 0001, Jian Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Adaptive distribution calibration for few-shot learning via optimal transport
Xin Liu 0086, Kairui Zhou, Pengbo Yang, Liping Jing, Jian Yu 0001 |
Inf. Sci. | 5 |
| 2022 | Bayesian Additive Matrix Approximation for Social RecommendationabstractSocial relations between users have been proven to be a good type of auxiliary information to improve the recommendation performance. However, it is a challenging issue to sufficiently exploit the social relations and correctly determine the user preference from both social and rating information. In this article, we propose a unified Bayesian Additive Matrix Approximation model (BAMA), which takes advantage of rating preference and social network to provide high-quality recommendation. The basic idea of BAMA is to extract social influence from social networks, integrate them to Bayesian additive co-clustering for effectively determining the user clusters and item clusters, and provide an accurate rating prediction. In addition, an efficient algorithm with collapsed Gibbs Sampling is designed to inference the proposed model. A series of experiments were conducted on six real-world social datasets. The results demonstrate the superiority of the proposed BAMA by comparing with the state-of-the-art methods from three views, all users, cold-start users, and users with few social relations. With the aid of social information, furthermore, BAMA has ability to provide the explainable recommendation. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Pengyu Xu, Jian Yu 0001, Michael Kwok-Po Ng |
ACM Trans. Knowl. Discov. Data | 5 |
| 2022 | Learning to Learn a Cold-start Sequential RecommenderabstractThe cold-start recommendation is an urgent problem in contemporary online applications. It aims to provide users whose behaviors are literally sparse with as accurate recommendations as possible. Many data-driven algorithms, such as the widely used matrix factorization, underperform because of data sparseness. This work adopts the idea of meta-learning to solve the user’s cold-start recommendation problem. We propose a meta-learning-based cold-start sequential recommendation framework called metaCSR, including three main components: Diffusion Representer for learning better user/item embedding through information diffusion on the interaction graph; Sequential Recommender for capturing temporal dependencies of behavior sequences; and Meta Learner for extracting and propagating transferable knowledge of prior users and learning a good initialization for new users. metaCSR holds the ability to learn the common patterns from regular users’ behaviors and optimize the initialization so that the model can quickly adapt to new users after one or a few gradient updates to achieve optimal performance. The extensive quantitative experiments on three widely used datasets show the remarkable performance of metaCSR in dealing with the user cold-start problem. Meanwhile, a series of qualitative analysis demonstrates that the proposed metaCSR has good generalization. Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001, Changsheng Xu |
ACM Trans. Inf. Syst. | 3 |
| 2021 | Social Recommendation With Learning Personal and Social Latent FactorsabstractDue to leveraging social relationships between users as well as their past social behavior, social recommendation becomes a core component in recommendation systems. Most existing social recommendation methods only consider direct social relationships among users (e.g., explicit and observed social relations). Recently, researchers proved that indirect social relationships can be effective to improve the recommendation quality when users only have few social connections, because it can identify the user interesting group even though the users have no observed social connection. In the literature, separate two-stage methods are studied, but they cannot explicitly capture the natural relationship between indirect social relations and latent user/item factors. In this paper, the main contribution is to propose a new joint recommendation model taking advantage of the Indirect Social Relations detection and Matrix Factorization collaborative filtering on social network and rating behavior information, which is called as InSRMF. In our work, the user latent factors can simultaneously and seamlessly capture user's personal preferences and social group characteristics. To optimize the InSRMF model, we develop a parallel graph vertex programming algorithm for efficiently handling large scale social recommendation data. Experiments based on four real-world datasets (Ciao, Epinions, Douban and Yelp) are conducted to demonstrate the performance of the proposed model. The experimental results have shown that InSRMF has ability to mine the proper indirect social relations and improve the recommendation performance compared with the testing methods in the literature, especially on the users with few social neighbors, Near-cold-start Users, Pure-cold-start Users and Long-tail Items. Huafeng Liu 0001, Liping Jing, Jian Yu 0001, Michael Kwok-Po Ng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Deep Generative Recommendation with Maximizing Reciprocal Rank
Xiaoyi Sun, Huafeng Liu 0001, Liping Jing, Jian Yu 0001 |
KSEM (2) | 4 |
| 2020 | Deep Global and Local Generative Model for RecommendationabstractDeep generative model, especially variational auto-encoder (VAE), has been successfully employed by more and more recommendation systems. The reason is that it combines the flexibility of probabilistic generative model with the powerful non-linear feature representation ability of deep neural networks. The existing VAE-based recommendation models are usually proposed under global assumption by incorporating simple priors, e.g., a single Gaussian, to regularize the latent variables. This strategy, however, is ineffective when the user is simultaneously interested in different kinds of items, i.e., the user’s preference may be highly diverse. In this paper, thus, we propose a Deep Global and Local Generative Model for recommendation to consider both local and global structure among users (DGLGM) under the Wasserstein auto-encoder framework. Besides keeping the global structure like the existing model, DGLGM adopts a non-parametric Mixture Gaussian distribution with several components to capture the diversity of the users’ preferences. Each component is corresponding to one local structure and its optimal size can be determined via the automatic relevance determination technique. These two parts can be seamlessly integrated and enhance each other. The proposed DGLGM can be efficiently inferred by minimizing its penalized upper bound with the aid of local variational optimization technique. Meanwhile, we theoretically analyze its generalization error bounds to guarantee its performance in sparse feedback data with diversity. By comparing with the state-of-the-art methods, the experimental results demonstrate that DGLGM consistently benefits the recommendation system in top-N recommendation task. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Zhicheng Wu, Xiaoyi Sun, Jiaqi Wang 0006, Jian Yu 0001 |
WWW | 8 |
| 2020 | Traffic Flow Prediction via Spatial Temporal Graph Neural NetworkabstractTraffic flow analysis, prediction and management are keystones for building smart cities in the new era. With the help of deep neural networks and big traffic data, we can better understand the latent patterns hidden in the complex transportation networks. The dynamic of the traffic flow on one road not only depends on the sequential patterns in the temporal dimension but also relies on other roads in the spatial dimension. Although there are existing works on predicting the future traffic flow, the majority of them have certain limitations on modeling spatial and temporal dependencies. In this paper, we propose a novel spatial temporal graph neural network for traffic flow prediction, which can comprehensively capture spatial and temporal patterns. In particular, the framework offers a learnable positional attention mechanism to effectively aggregate information from adjacent roads. Meanwhile, it provides a sequential component to model the traffic flow dynamics which can exploit both local and global temporal dependencies. Experimental results on various real traffic datasets demonstrate the effectiveness of the proposed framework. Yao Ma 0001, Yiqi Wang 0001, Wei Jin 0009, Xin Wang 0035, Jiliang Tang, Caiyan Jia, Jian Yu 0001 |
WWW | 8 |
| 2019 | In2Rec: Influence-based Interpretable RecommendationabstractInterpretability of recommender systems has caused increasing attention due to its promotion of the effectiveness and persuasiveness of recommendation decision, and thus user satisfaction. Most existing methods, such as Matrix Factorization (MF), tend to be black-box machine learning models that lack interpretability and do not provide a straightforward explanation for their outputs. In this paper, we focus on probabilistic factorization model and further assume the absence of any auxiliary information, such as item content or user review. We propose an influence mechanism to evaluate the importance of the users' historical data, so that the most related users and items can be selected to explain each predicted rating. The proposed method is thus called Influencebased Interpretable Recommendation model (In2Rec). To further enhance the recommendation accuracy, we address the important issue of missing not at random, i.e., missing ratings are not independent from the observed and other unobserved ratings, because users tend to only interact what they like. In2Rec models the generative process for both observed and missing data, and integrates the influence mechanism in a Bayesian graphical model. A learning algorithm capitalizing on iterated condition modes is proposed to tackle the non-convex optimization problem pertaining to maximum a posteriori estimation for In2Rec. A series of experiments on four real-world datasets (Movielens 10M, Netflix, Epinions, and Yelp) have been conducted. By comparing with the state-of-the-art recommendation methods, the experimental results have shown that In2Rec can consistently benefit the recommendation system in both rating prediction and ranking estimation tasks, and friendly interpret the recommendation results with the aid of the proposed influence mechanism. Huafeng Liu 0001, Jingxuan Wen, Liping Jing, Jian Yu 0001, Xiangliang Zhang 0001, Min Zhang 0006 |
CIKM | 4 |
| 2019 | Comprehensive Event Storyline Generation from MicroblogsabstractMicroblogging data contains a wealth of information of trending events and has gained increased attention among users, organizations, and research scholars for social media mining in different disciplines. Event storyline generation is one typical task of social media mining, whose goal is to extract the development stages with associated description of events. Existing storyline generation methods either generate storyline with less integrity or fail to guarantee the coherence between the discovered stages. Secondly, there are no scientific method to evaluate the quality of the storyline. In this paper, we propose a comprehensive storyline generation framework to address the above disadvantages. Given Microblogging data related to the specified event, we first propose Hot-Word-Based stage detection algorithm to identify the potential stages of event, which can effectively avoid ignoring important stages and preventing inconsistent sequence between stages. Community detection algorithm is applied then to select representative data for each stage. Finally, we conduct graph optimization algorithm to generate the logically coherent storylines of the event. We also introduce a new evaluation metric, SLEU, to emphasize the importance of the integrity and coherence of the generated storyline. Extensive experiments on real-world Chinese microblogging data demonstrate the effectiveness of the proposed methods in each module and the overall framework. Wenjin Sun, Yuhang Wang 0007, Yuqi Gao, Zesong Li, Jitao Sang 0001, Jian Yu 0001 |
MMAsia | 6 |
| 2019 | Multi-source User Attribute Inference based on Hierarchical Auto-encoderabstractWith the rapid development of Online Social Networks (OSNs), it is crucial to construct users' portraits from their dynamic behaviors to address the increasing needs for customized information services. Previous work on user attribute inference mainly concentrated on developing advanced features/models or exploiting external information and knowledge but ignored the contradiction between dynamic behaviors and stable demographic attributes, which results in deviation of user understanding Xiangguo Ding, Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001 |
MMAsia | 6 |
| 2019 | Deep generative ranking for personalized recommendationabstractRecommender systems offer critical services in the age of mass information. Personalized ranking has been attractive both for content providers and customers due to its ability of creating a user-specific ranking on the item set. Although the powerful factor-analysis methods including latent factor models and deep neural network models have achieved promising results, they still suffer from the challenging issues, such as sparsity of recommendation data, uncertainty of optimization, and etc. To enhance the accuracy and generalization of recommender system, in this paper, we propose a deep generative ranking (DGR) model under the Wasserstein autoencoder framework. Specifically, DGR simultaneously generates the pointwise implicit feedback data (via a Beta-Bernoulli distribution) and creates the pairwise ranking list by sufficient exploiting both interacted and non-interacted items for each user. DGR can be efficiently inferred by minimizing its penalized evidence lower bound. Meanwhile, we theoretically analyze the generalization error bounds of DGR model to guarantee its performance in extremely sparse feedback data. A series of experiments on four large-scale datasets (Movielens (20M), Netflix, Epinions and Yelp in movie, product and business domains) have been conducted. By comparing with the state-of-the-art methods, the experimental results demonstrate that DGR consistently benefit the recommendation system in ranking estimation task, especially for the near-cold-start-users (with less than five interacted items). Huafeng Liu 0001, Jingxuan Wen, Liping Jing, Jian Yu 0001 |
RecSys | 4 |
| 2019 | Adaptive Local Low-rank Matrix Approximation for RecommendationabstractLow-rank matrix approximation (LRMA) has attracted more and more attention in the community of recommendation. Even though LRMA-based recommendation methods (including Global LRMA and Local LRMA) obtain promising results, they suffer from the complicated structure of the large-scale and sparse rating matrix, especially when the underlying system includes a large set of items with various types and a huge amount of users with diverse interests. Thus, they have to predefine the important parameters, such as the rank of the rating matrix and the number of submatrices. Moreover, most existing Local LRMA methods are usually designed in a two-phase separated framework and do not consider the missing mechanisms of rating matrix. In this article, a non-parametric unified Bayesian graphical model is proposed for A daptive Lo cal low-rank M atrix A pproximation ( ALoMA ). ALoMA has ability to simultaneously identify rating submatrices, determine the optimal rank for each submatrix, and learn the submatrix-specific user/item latent factors. Meanwhile, the missing mechanism is adopted to characterize the whole rating matrix. These four parts are seamlessly integrated and enhance each other in a unified framework. Specifically, the user-item rating matrix is adaptively divided into proper number of submatrices in ALoMA by exploiting the Chinese Restaurant Process. For each submatrix, by considering both global/local structure information and missing mechanisms, the latent user/item factors are identified in an optimal latent space by adopting automatic relevance determination technique. We theoretically analyze the model’s generalization error bounds and give an approximation guarantee. Furthermore, an efficient Gibbs sampling-based algorithm is designed to infer the proposed model. A series of experiments have been conducted on six real-world datasets ( Epinions , Douban , Dianping , Yelp , Movielens (10M), and Netflix ). The results demonstrate that ALoMA outperforms the state-of-the-art LRMA-based methods and can easily provide interpretable recommendation results. Huafeng Liu 0001, Liping Jing, Jian Yu 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2018 | Neural gaussian mixture model for review-based rating predictionabstractReview has been proven to be an important information in recommendation. Different from the overall user-item rating matrix, it can provide textual information that exhibits why a user likes an item or not. Recently, more and more researchers have paid attention on review-based rating prediction. There are two challenging issues: how to extract representative features to characterize users / items from reviews and how to leverage them for recommendation system. In this paper, we propose a Neural Gaussian Mixture Model (NGMM) for review-based rating prediction task. Among it, the review textual information is used to construct two parallel neural networks for users and items respectively, so that the users' preferences and items' properties can be sufficiently extracted and represented as two latent vectors. A shared layer is introduced on the top to couple these two networks together and model user-item rating based on the features learned from reviews. Specifically, each rating is modeled via a Gaussian mixture model, where each Gaussian component has zero variance, the mean described by the corresponding component in user's latent vector and the weight indicated by the corresponding component in item's latent vector. Extensive experiments are conducted on five real-world Amazon review datasets. The experimental results have demonstrated that our proposed NGMM model achieves the state-of-the-art performance in review-based rating prediction task. Dong Deng 0002, Liping Jing, Jian Yu 0001, Shaolong Sun, Haofei Zhou |
RecSys | 3 |
| 2017 | Deterministic annealing Gustafson-Kessel fuzzy clustering algorithm
Chaomurilige Wang, Jian Yu 0001, Miin-Shen Yang |
Inf. Sci. | 2 |
| 2014 | Weighting Exponent Selection of Fuzzy C-Means via Jacobian Matrix
Liping Jing, Dong Deng 0002, Jian Yu 0001 |
KSEM | 3 |
| 2014 | Constrained Least Squares Regression for Semi-Supervised Learning
Bo Liu 0050, Liping Jing, Jian Yu 0001, Jia Li 0002 |
PAKDD (2) | 3 |
| 2013 | Clustering construction on a multimodal probability model
Jian Yu 0001, Miin-Shen Yang, Pengwei Hao |
Inf. Sci. | 1 |
| 2013 | A novel semi-supervised learning framework with simultaneous text representing
Yan Zhu 0012, Jian Yu 0001, Liping Jing |
Knowl. Inf. Syst. | 2 |
| 2011 | Intrinsic Dimension Induced Similarity Measure for Clustering
Jian Yu 0001, Shu Gong |
ADMA (2) | 2 |
| 2011 | Text Clustering via Constrained Nonnegative Matrix FactorizationabstractSemi-supervised nonnegative matrix factorization (NMF)receives more and more attention in text mining field. The semi-supervised NMF methods can be divided into two types, one is based on the explicit category labels, the other is based on the pair wise constraints including must-link and cannot-link. As it is hard to obtain the category labels in some tasks, the latter one is more widely used in real applications. To date, all the constrained NMF methods treat the must-link and cannot-link constraints in a same way. However, these two kinds of constraints play different roles in NMF clustering. Thus a novel constrained NMF method is proposed in this paper. In the new method, must-link constraints are used to control the distance of the data in the compressed form, and cannot-ink constraints are used to control the encoding factor. Experimental results on real-world text data sets have shown the good performance of the proposed method. Yan Zhu 0012, Liping Jing, Jian Yu 0001 |
ICDM | 3 |
| 2011 | Data Clustering by Scaled Adjacency Matrix
Jian Yu 0001, Caiyan Jia |
KSEM | 1 |
| 2011 | High-Order Co-clustering Text Data on Semantics-Based Representation Model
Liping Jing, Jiali Yun, Jian Yu 0001, Joshua Zhexue Huang |
PAKDD (1) | 3 |
| 2011 | Unsupervised Feature Weighting Based on Local Feature Relatedness
Jiali Yun, Liping Jing, Jian Yu 0001, Houkuan Huang |
PAKDD (1) | 3 |
| 2010 | Affinity Propagation on Identifying Communities in Social and Biological Networks
Caiyan Jia, Yawen Jiang, Jian Yu 0001 |
KSEM | 3 |
| 2010 | Text Clustering via Term Semantic UnitsabstractHow best to represent text data is an important problem in text mining tasks including information retrieval, clustering, classification and etc.. In this paper, we proposed a compact document representation with term semantic units which are identified from the implicit and explicit semantic information. Among it, the implicit semantic information is extracted from syntactic content via statistical methods such as latent semantic indexing and information bottleneck. The explicit semantic information is mined from the external semantic resource (Wikipedia). The proposed compact representation model can map a document collection in a low-dimension space (term semantic units which are much less than the number of all unique terms). Experimental results on real data sets have shown that the compact representation efficiently improve the performance of text clustering. Liping Jing, Jiali Yun, Jian Yu 0001, Houkuan Huang |
Web Intelligence | 3 |
| 2009 | Convergence Analysis of Affinity Propagation
Jian Yu 0001, Caiyan Jia |
KSEM | 1 |
| 2009 | New Labeling Strategy for Semi-supervised Document Categorization
Yan Zhu 0012, Liping Jing, Jian Yu 0001 |
KSEM | 3 |
| 2007 | Clustering Ensembles Based on Normalized Edges
Yan Li 0009, Jian Yu 0001, Pengwei Hao, Zhulin Li |
PAKDD | 2 |
| 2007 | Learning Bayesian Networks with Combination of MRMR Criterion and EMI Method
Fengzhan Tian, Jian Yu 0001 |
PAKDD | 4 |