Jian Yin 0001

dblp:95/578-1 · DBLP profile ↗
← Back
71ranked-venue papers in the field
0as first author
51since 2021 · last 2026
ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 31Data Mining & Knowledge Discovery · 18Information Retrieval & Web Search · 13Knowledge Engineering, Semantic Web & Information Systems · 5Other / Interdisciplinary · 4
YearPublicationVenuePosition
2026 STMHTNet: A Spatio-Temporal Masked Hourglass Transformer Network for Traffic Flow Forecasting
Yixin Hong, Huaijie Zhu, Wei Liu 0061, Zixin Qin, Jianxing Yu, Jian Yin 0001
DASFAA (3)6
2026 SAGE-LLM: Spatially-Aware Generation and Explanation via Large Language Models for Imbalanced Spatial Data Classification
Wenhui Tu, Wei Liu 0061, Huaijie Zhu, Jianxing Yu, Jian Yin 0001
DASFAA (4)5
2026 Inductive Controlled Generation Based on Adaptive Templates for Answering Subjective Product Questions
Yian Yao, Jianxing Yu, Huaijie Zhu, Hanjiang Lai, Wei Liu 0061, Yanghui Rao, Jian Yin 0001
DASFAA (6)7
2026 Geography-Aware Large Language Models for Next POI Recommendation
Wei Liu 0061, Muzu Xie, Huaijie Zhu, Jianxing Yu, Jian Yin 0001, Wang-Chien Lee
ICDE6
2026 Trajectory-User Linking via Heterogeneous Preference Graph and Dual-Encoder Mutual Distillation
Zeming Tian, Zixin Qin, Huaijie Zhu, Ningning Cui, Jianxing Yu, Jian Yin 0001
ICDE7
2026 Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
Xujie Yuan, Shimin Di, Jielong Tang, Libin Zheng 0001, Jian Yin 0001
KDD (1)5
2026 CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
abstract
High-quality annotated datasets are crucial for advancing machine learning in medical image analysis. However, a critical gap exists: most datasets either offer a single, clean ground truth, which hides real-world expert disagreement, or they provide multiple annotations without a separate gold standard for objective evaluation. To bridge this gap, we introduce CytoCrowd, a new public benchmark for cytology analysis. The dataset features 446 high-resolution images, each with two key components: (1) raw, conflicting annotations from four independent pathologists, and (2) a separate, high-quality gold-standard ground truth established by a senior expert. This dual structure makes CytoCrowd a versatile resource. It serves as a benchmark for standard computer vision tasks, such as object detection and classification, using the ground truth. Simultaneously, it provides a realistic testbed for evaluating annotation aggregation algorithms that must resolve expert disagreements. We provide comprehensive baseline results for both tasks. Our experiments demonstrate the challenges presented by CytoCrowd and establish its value as a resource for developing the next generation of models for medical image analysis.
Yonghao Si, Xingyuan Zeng, Zhao Chen 0003, Libin Zheng 0001, Caleb Chen Cao, Lei Chen 0002, Jian Yin 0001
WWW7
2026 DA-RAG: Dynamic Attributed Community Search for Retrieval-Augmented Generation
abstract
Owing to their unprecedented comprehension capabilities, large language models (LLMs) have become indispensable components of modern web search engines. From a technical perspective, this integration represents retrieval-augmented generation (RAG), which enhances LLMs by grounding them in external knowledge base. A prevalent technical approach in this context is graph-based RAG (G-RAG). However, current G-RAG methodologies frequently underutilize graph topology, predominantly focusing on low-order structures or pre-computed static communities. This limitation affects their effectiveness in addressing dynamic and complex queries. Thus, we propose DA-RAG, which leverages attributed community search (ACS) to dynamically extract relevant subgraphs based on the queried question. DA-RAG captures high-order graph structures, allowing for the retrieval of self-complementary knowledge. Furthermore, DA-RAG is equipped with a chunk-layer oriented graph index, which facilitates efficient multi-granularity retrieval while significantly reducing both computational and economic costs. We evaluate DA-RAG on multiple datasets, demonstrating that it outperforms existing RAG methods by up to 40% in head-to-head comparisons across four metrics while reducing index construction time and token overhead by up to 37% and 41%, respectively.
Xingyuan Zeng, Zuohan Wu, Yue Wang 0012, Chen Zhang 0013, Quanming Yao, Libin Zheng 0001, Jian Yin 0001
WWW7
2026 Learning from easy to hard: Curriculum meta-learning for few-shot node classification
Qilong Yan, Weinan Guan, Yifei Xing 0001, Jingpu Duan, Jian Yin 0001
Inf. Sci.5
2026 VPLight: A Reinforcement Learning Approach for Traffic Signal Control With Pedestrian Dynamics
abstract
Traffic Signal Control plays a vital role in modern traffic management. However, most existing methods focus exclusively on vehicle flow, neglecting the critical role of pedestrians, leading to suboptimal performance in intersections with mixed vehicle-pedestrian traffic. Pedestrian behavior presents unique challenges due to its irregularity and flexibility, such as non-lane-based movements and uncertain crossing directions, which cannot be modeled by existing methods. To address this limitation, we propose VPLight, a comprehensive framework designed to manage bothVehicle andPedestrian dynamics in traffic signal control. Specifically, we first design the Pedestrian Feature Extractor to capture the spatiotemporal dynamics of pedestrian movement, offering a robust representation of their irregular patterns. Subsequently, to coordinate traffic signal control at multiple intersections, we develop a novel communication approach called V-Comm to enable effective integration among intersections. Extensive experiments show that VPLight outperforms state-of-the-art baselines with significant margins (up to +44.04%). Our results demonstrate that VPLight can remarkably address the challenges of mixed vehicle-pedestrian traffic control and enhance the overall traffic flow efficiency across the road network.
Xinyu Zhang 0019, Zuohan Wu, Chen Zhang 0013, Libin Zheng 0001, Peng Cheng 0003, Jian Yin 0001, Cyrus Shahabi
IEEE Trans. Knowl. Data Eng.6
2025 DiffSTRec: A Diffusion-Based Framework for Spatiotemporal Next POI Recommendation
Junchao Zeng, Wei Liu 0061, Jianxing Yu, Huaijie Zhu, Jian Yin 0001
WISA8
2025 Demand-Oriented Route Recommendation for Shared Mobility Services
Zhijia Chen, Chen Zhang 0013, Peng Cheng 0003, Libin Zheng 0001, Jian Yin 0001
DASFAA (5)6
2025 Asking Diversified Reasonable Questions with External Commonsense Knowledge to Infer Inconsistency for Multi-modal Clickbait Detection
Jianxing Yu, Shiqi Wang 0016, Huaijie Zhu, Libin Zheng 0001, Wenqing Chen, Jian Yin 0001
DASFAA (2)7
2025 Emotion-Based Conversational Recommendation by Inferring Implicit Users' Preferences from Their Subjective Claims
Xuanming Zhang, Yonghe Lu, Jianxing Yu, Huaijie Zhu, Wei Liu 0061, Wenqing Chen, Jian Yin 0001
DASFAA (5)7
2025 Causal-TSF: A Causal Intervention Approach to Mitigate Confounding Bias in Time Series Forecasting
abstract
Time series forecasting, aiming to learn models from historical data and predict future values in time series, is a fundamental research topic in machine learning. However, few efforts have been devoted to addressing the confounding effects in time series data, e.g., the historical data are affected by some hidden surrounding factors (i.e., confounders), leading to biased forecasting models for future data. This paper presents a causal intervention approach to eliminate the bias that is raised by some hidden confounders. By using a causal graph, we illustrate why hidden confounders can bring bias in time series forecasting and how to tackle it. We implement causal intervention by a deep architecture that consists of two modules, a Confounders Estimation module to estimate the hidden confounders and a Debiasing module to eliminate the confounding bias in the forecasting model via sampling on confounders. We conduct comprehensive evaluations on various time series datasets. The experiment results indicate that the proposed method can reduce the negative confounding effects in time series data, and it achieves superior gains over state-of-the-art baselines for time series forecasting.
Qinkang Gong, Yan Pan 0002, Hanjiang Lai, Rongbang Qiu, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.5
2024 Accelerating Training of Large Neural Models by Gradient-Based Growth Learning
Haowei Jiang, Jianxing Yu, Libin Zheng 0001, Huaijie Zhu, Wei Liu 0061, Jian Yin 0001
DASFAA (2)6
2024 Variational Kernel Density Estimation Recommendation Algorithm for Users with Diverse Activity Levels
Wei Liu 0061, Shangsong Liang, Huaijie Zhu, Leong Hou U, Jianxing Yu, Xiang Li 0067, Jian Yin 0001
DASFAA (2)7
2024 Cross-Domain-Aware Worker Selection with Training for Crowdsourced Annotation
abstract
Annotation through crowdsourcing draws incremental attention, which relies on an effective selection scheme given a pool of workers. Existing methods propose to select workers based on their performance on tasks with ground truth, while two important points are missed. 1) The historical performances of workers in other tasks. In real-world scenarios, workers need to solve a new task whose correlation with previous tasks is not well-known before the training, which is called cross-domain. 2) The dynamic worker performance as workers will learn from the ground truth. In this paper, we consider both factors in designing an allocation scheme named cross-domain-aware worker selection with training approach. Our approach proposes two estimation modules to both statistically analyze the cross-domain correlation and simulate the learning gain of workers dynamically. A framework with a theoretical analysis of the worker elimination process is given. To validate the effectiveness of our methods, we collect two novel real-world datasets and generate synthetic datasets. The experiment results show that our method outperforms the baselines on both real-world and synthetic datasets.
Yushi Sun, Jiachuan Wang, Peng Cheng 0003, Libin Zheng 0001, Lei Chen 0002, Jian Yin 0001
ICDE6
2024 Privacy-Preserving Traffic Flow Release with Consistency Constraints
abstract
Urban traffic flow data is useful in transport ap-plications, playing an important role in various tasks such as road planning, site selection, ad services, etc. However, traffic flow data is the composition of personal driving trajectories, which can reveal sensitive information such as home and work locations, leading to privacy issues. Thus publishing traffic flow data while not disclosing private information remains a challenge for urban managers. To address this challenge, we study the noisy publication of traffic flow data in this paper. The noise is added to the data with respect to the differential privacy paradigm, which ensures data safety but deteriorates its utility. On the other hand, we find that the inherent relations of the flow data inherited from the road network structure can be used to correct data without hurting the privacy property. Hence, we propose post-processing techniques, which exploit the data's inherent relations for corrections over the global and local differentially private traffic flow data, respectively. Extensive experiments on real data show that the proposed post-processing techniques improve the data utility by 29.7%-41.1% and 17.3%-48.6% subjecting to the global and local differential privacy paradigm, respectively.
Xiaoting Zhu, Libin Zheng 0001, Chen Zhang 0013, Peng Cheng 0003, Lei Chen 0002, Xuemin Lin 0001, Jian Yin 0001
ICDE8
2024 Next POI Recommendation Based on Time Slot Preferences and Bidirectional Transformation Modeling
Wei Liu 0061, Huaijie Zhu, Jianxing Yu, Jian Yin 0001
WISE (3)5
2024 Weighted Linear Regression with Optimized Gap for Learned Index
Libin Zheng 0001, Jian Yin 0001
WISE (4)3
2024 Category-Level Contrastive Learning for Unsupervised Hashing in Cross-Modal Retrieval
abstract
Abstract Unsupervised hashing for cross-modal retrieval has received much attention in the data mining area. Recent methods rely on image-text paired data to conduct unsupervised cross-modal hashing in batch samples. There are two main limitations for existing models: (1) learning of cross-modal representations is restricted to batches; (2) semantically similar samples may be wrongly treated as negative. In this paper, we propose a novel category-level contrastive learning for unsupervised cross-modal hashing, which alleviates the above problems and improves cross-modal query accuracy. To break the limitation of learning in small batches, a selected memory module is first proposed to take global relations into account. Then, we obtain pseudo labels through clustering and combine the labels with the Hadamard Matrix for category-centered learning. To reduce wrong negatives, we further propose a memory bank to store clusters of samples and construct negatives by selecting samples from different categories for contrastive learning. Extensive experiments show the significant superiority of our approach over the state-of-the-art models on MIRFLICKR-25K and NUS-WIDE datasets.
Mengying Xu, Linyin Luo, Hanjiang Lai, Jian Yin 0001
Data Sci. Eng.4
2024 VAE*: A Novel Variational Autoencoder via Revisiting Positive and Negative Samples for Top-N Recommendation
abstract
Due to the easy access, implicit feedback is often used for recommender systems. Compared with point-wise learning and pair-wise learning methods, list-wise rank learning methods have superior performance for top- \(N\) recommendation. Recent solutions, especially the list-wise methods, simply treat all interacted items of a user as equally important positives and annotate all no-interaction items of a user as negatives. For the list-wise approaches, we argue that this annotation scheme of implicit feedback is over-simplified due to the sparsity and missing fine-grained labels of the feedback data. To overcome this issue, we revisit the so-called positive and negative samples. First, considering the loss function of list-wise ranking, we analyze the impact of false positives and negatives theoretically. Second, based on the observation, we propose a self-adjusting credibility weight mechanism to re-weigh the positive samples and exploit the higher-order relation based on item–item matrix to sample the critical negative samples. In order to prevent the introduction of noise, we design a pruning strategy for critical negatives. Besides, to combine the reconstruction loss function for the positive samples and critical negative samples, we develop a simple yet effective VAEs framework with linear structure, which abandons the complex non-linear structure. Extensive experiments are conducted on six public real-world datasets. The results demonstrate that, our VAE* outperforms other VAE-based models by a large margin. Besides, we also verify the effect of denoising positives and exploring critical negatives by ablation study.
Wei Liu 0061, Leong Hou U, Shangsong Liang, Huaijie Zhu, Jianxing Yu, Jian Yin 0001
ACM Trans. Knowl. Discov. Data7
2024 Revisiting Few-Shot Learning From a Causal Perspective
abstract
Few-shot learning with$N$-way$K$-shot scheme is an open challenge in machine learning. Many metric-based approaches have been proposed to tackle this problem, e.g., the Matching Networks and CLIP-Adapter. Despite that these approaches have shown significant progress, the mechanism of why these methods succeed has not been well explored. In this paper, we try to interpret these metric-based few-shot learning methods via causal mechanism. We show that the existing approaches can be viewed as specific forms of front-door adjustment, which can alleviate the effect of spurious correlations and thus learn the causality. This causal interpretation could provide us a new perspective to better understand these existing metric-based methods. Further, based on this causal interpretation, we simply introduce two causal methods for metric-based few-shot learning, which considers not only the relationship between examples but also the diversity of representations. Experimental results demonstrate the superiority of our proposed methods in few-shot classification on various benchmark datasets. Code is available inhttps://github.com/lingl1024/causalFewShot.
Guoliang Lin, Yongheng Xu, Hanjiang Lai, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.4
2024 Anomaly Detection Under Contaminated Data With Contamination-Immune Bidirectional GANs
abstract
Anomaly detection aims to detect instances that deviate significantly from the majority. Due to the difficulties of collecting a large amount of anomalies in practice, existing methods generally assume the availability of a clean normal dataset and leverage it to detect anomalies by characterizing the normality of normal samples. However, for many application scenarios, collecting a normal dataset that is sufficiently clean is not easy. What is often observed is that a small amount of anomalies are often falsely mixed into the normal dataset, resulting in a contaminated dataset. Obviously, the contamination in the normal dataset could significantly compromise the model's ability to detect anomalies. To alleviate this issue, two contamination-immune bidirectional generative adversarial networks (BiGAN) are developed, which can learn the probability distribution of normal samples from a contaminated dataset under some mild conditions. Rigorous proofs are provided to guarantee the theoretical correctness of the proposed models. Thanks to the removing of negative influences from the contamination samples, the proposed contamination-immune models can thus be applied to detect anomalies accurately for the scenarios with contaminated datasets. Extensive experimental results show that the proposed method outperforms the current state-of-the-art (SOTA) ones significantly under the scenarios with contaminated training datasets.
Qinliang Su, Hai Wan, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.4
2024 Longer Pick-Up for Less Pay: Towards Discount-Based Mobility Services
abstract
With the rapid development of mobile Internet technology, on-demand car-hailing services have become essential for people's daily commuting. Order dispatch is a critical problem in on-demand car-hailing services. However, in most existing works, the service provider is asked to set a unified pick-up distance to prevent long waiting time for requesters. Indeed, different requesters have different tolerance for pick-up distance, and some requesters may accept longer pick-ups if offered discounts for payment. Regarding this fact, we formulate discount-based order dispatch as a coupling of two subproblems, discount determination and order dispatch, aiming to dispatch more orders and thereby more platform profits. We propose customized methods to solve the problems for shared and non-shared mobility services, respectively. We also conduct extensive experiments to evaluate the effectiveness and efficiency of our proposed methods on a real dataset, which shows that our methods can achieve 170% improvements in non-shard services and 43% improvements in ridesharing services on average in terms of attained profit compared to the widely adopted baselines.
Wanyi Xie, Zhijia Chen, Chen Zhang 0013, Libin Zheng 0001, Peng Cheng 0003, Jian Yin 0001, Xuemin Lin 0001
IEEE Trans. Knowl. Data Eng.6
2023 Generating Enlightened Suggestions Based on Mental State Evolution for Emotional Support Conversation
Mengjiao Gan, Jianxing Yu, Wei Liu 0061, Jian Yin 0001
ADMA (1)6
2023 Discovery of Emotion Implicit Causes in Products Based on Commonsense Reasoning
Qiutong Guo, Jianxing Yu, Haowei Jiang, Wei Liu 0061, Jian Yin 0001
ADMA (1)6
2023 Spatial Commonsense Reasoning for Machine Reading Comprehension
Miaopei Lin, Mengxiang Wang, Jianxing Yu, Shiqi Wang 0016, Hanjiang Lai, Wei Liu 0061, Jian Yin 0001
ADMA (2)7
2023 Multi-modal Multi-emotion Emotional Support Conversation
Guangya Liu, Mengxiang Wang, Jianxing Yu, Mengjiao Gan, Wei Liu 0061, Jian Yin 0001
ADMA (1)7
2023 A Knowledge-Enhanced Inferential Network for Cross-Modality Multi-hop VQA
Shiqi Wang 0016, Jianxing Yu, Miaopei Lin, Xiaofeng Luo, Jian Yin 0001
ADMA (2)6
2023 Community Detection in Temporal Biological Metabolic Networks Based on Semi-NMF Method with Node Similarity Fusion
Xuanming Zhang, Jianxing Yu, Miaopei Lin, Shiqi Wang 0016, Wei Liu 0061, Jian Yin 0001
ADMA (4)6
2023 Revisiting Positive and Negative Samples in Variational Autoencoders for Top-N Recommendation
Wei Liu 0061, Leong Hou U, Shangsong Liang, Huaijie Zhu, Jianxing Yu, Jian Yin 0001
DASFAA (2)7
2023 Copula Guided Parallel Gibbs Sampling for Nonparametric and Coherent Topic Discovery (Extended Abstract)
abstract
In terms of the generative process, the Gamma-Gamma-Poisson Process (G2PP) is equivalent to the nonparametric topic model of Hierarchical Dirichlet Process (HDP). Considering the high computational cost of estimating parameters in HDP, a parallel G2PP was developed to generate topics efficiently via multi-threading. Unfortunately, the above model needs to predefine the number of topics. To address this issue, we first propose a Topic Self-Adaptive Model (TSAM) for nonparametric and parallel topic discovery. In TSAM, a monitor-executor mechanism is developed to manage the global topic information using a hierarchical structure of threads. Based on the apparatus of copulas, we further extend our TSAM to TSAMcop for coherent topic modeling by exploiting a copula guided parallel Gibbs sampling algorithm. Extensive experiments validate the effectiveness of both TSAM and TSAMcop.
Lihui Lin, Yanghui Rao, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001, Fu Lee Wang, Qing Li 0001
ICDE5
2023 Opponent-aware Order Pricing towards Hub-oriented Mobility Services
abstract
Hub-oriented mobility services have gained great developments in recent years, enabling riders to simultaneously call vehicles from multiple mobility-supply companies (agents) on a single APP (which we call "hub"). Competing with others on such a hub, to obtain an order, an agent company first needs to get admitted by the requester, which is in turn affected by its quotation. The quotation needs to be attractively low compared to those of the opposing agents. Thus, an opponent-aware pricing strategy is needed for an agent to play well in the hub scenario, which is rarely discussed in existing works. To address the aforementioned issue, in this work, we first propose a quotation prediction model, which employs a neural network with a customized loss function to predict the opponents’ quotations. Based on the predictions, we then propose multi-arm bandit based methods to decide a proper quotation for the agent, in order to obtain orders while retaining profits. We finally conduct extensive experiments on real data, where the quotation-determining method integrated with the prediction model has achieved a remarkable profit improvement up to 85.5% compared to baseline methods, demonstrating their effectiveness.
Zuohan Wu, Libin Zheng 0001, Chen Zhang 0013, Huaijie Zhu, Jian Yin 0001, Di Jiang 0004
ICDE5
2023 Keyword-based Socially Tenuous Group Queries
abstract
Socially tenuous groups (or simply tenuous groups) in a social network/graph refer to subgraphs with few social interactions and weak relationships among members. However, existing studies on tenuous group queries do not consider the user profiles (keywords) of the members whereas in many social network applications, e.g., finding reviewers for paper selection and recommending seed users in social advertising, keywords also need to be considered. Thus, in this paper, we investigate the problem of keywords-based socially tenous group (KTG) queries. A KTG query is to find top N tenuous groups in which the members of each group jointly cover the most number of query keywords. To address the KTG problem, we first propose two exact algorithms, namely KTG-VKC and KTG-VKC-DEG, which give priority to the valid keyword coverage and the combination of valid keyword coverage and degree, respectively, to select members to form a feasible group by adopting a branch and bound (BB) strategy. Moreover, we propose keyword pruning and k-line filtering to accelerate the algorithms. To yield diversified KTG results, we also study the problem of diversified keywords-based socially tenous group (DKTG) queries. To deal with the DKTG problem, we propose a DKTG-Greedy algorithm by exploiting a greedy heuristic in combination with KTG-VKC-DEG. Furthermore, we design two alternative indexes, namely NL and NLRNL, to efficiently check whether the social distance of any two members is greater than the social constraint k in the above algorithms. We conduct extensive experiments using real datasets to validate our ideas and evaluate the proposed algorithms. Experimental results show that the NLRNL index achieves a better performance than the NL index.
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Ningning Cui, Jianliang Xu, Xin Huang 0001, Wang-Chien Lee
ICDE3
2023 Efficient Hash Coding for Image Retrieval Based on Improved Center Generation and Contrastive Pre-training Knowledge Model
Yan Pan 0002, Jian Yin 0001
KSEM (4)3
2023 Semi-Supervised Sentiment Classification and Emotion Distribution Learning Across Domains
abstract
In this study, sentiment classification and emotion distribution learning across domains are both formulated as a semi-supervised domain adaptation problem, which utilizes a small amount of labeled documents in the target domain for model training. By introducing a shared matrix that captures the stable association between document clusters and word clusters, non-negative matrix tri-factorization (NMTF) is robust to the labeled target domain data and has shown remarkable performance in cross-domain text classification. However, the existing NMTF-based models ignore the incompatible relationship of sentiment polarities and the relatedness among emotions. Besides, their applications on large-scale datasets are limited by the high computation complexity. To address these issues, we propose a semi-supervised NMTF framework for sentiment classification and emotion distribution learning across domains. Based on a many-to-many mapping between document clusters and sentiment polarities (or emotions), we first incorporate the prior information of label dependency to improve the model performance. Then, we develop a parallel algorithm based on message passing interface (MPI) to further enhance the model scalability. Extensive experiments on real-world datasets validate the effectiveness of our method.
Yufu Chen, Yanghui Rao, Shurui Chen, Zhiqi Lei, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001
ACM Trans. Knowl. Discov. Data7
2023 Parallel Non-Negative Matrix Tri-Factorization for Text Data Co-Clustering
abstract
As a novel paradigm for data mining and dimensionality reduction, Non-negative Matrix Tri-Factorization (NMTF) has attracted much attention due to its notable performance and elegant mathematical derivation, and it has been applied to a plethora of real-world applications, such as text data co-clustering. However, the existing NMTF-based methods usually involve intensive matrix multiplications, which exhibits a major limitation of high computational complexity. With the explosion at both the size and the feature dimension of texts, there is a growing need to develop a parallel and scalable NMTF-based algorithm for text data co-clustering. To this end, we first show in this paper how to theoretically derive the original optimization problem of NMTF by introducing the Lagrangian multipliers. Then, we propose to solve the Lagrange dual objective function in parallel through an efficient distributed implementation. Extensive experiments on five benchmark corpora validate the effectiveness, efficiency, and scalability of our distributed parallel update algorithm for an NMTF-based text data co-clustering method.
Yufu Chen, Zhiqi Lei, Yanghui Rao, Haoran Xie 0001, Fu Lee Wang, Jian Yin 0001, Qing Li 0001
IEEE Trans. Knowl. Data Eng.6
2023 Multi-Hop Reasoning Question Generation and Its Application
abstract
This article focuses on the topic of multi-hop question generation (QG), which aims to generate the questions requiring multi-hop reasoning skills from the given text. These questions are not only syntactically valid but also logically correlated with the answers. Concretely, we first design a basic QG model and customize several techniques to ensure results' syntactic validity. In order to promote the logical correlations, we use a reasoning chain extracted from the text to regularize the results. Considering that different samples have their own characteristics on the aspects of text contextual structure, the type of question, and logical correlation, we propose a new adaptive meta-learner to optimize the basic QG model. Each case and its similar samples are viewed as a pseudo-QG task. The similar structural contexts contained in the same task are used as guidance to fine-tune the model. To measure the similarity of samples' structured inputs, we propose a data-driven multi-level recognizer. The experimental results on two typical data sets in various domains show the effectiveness of the proposed approach. Moreover, we apply the generated results to the task of machine reading comprehension and achieve significant performance improvements. That demonstrates the capacity of multi-hop QG in facilitating real-world applications.
Jianxing Yu, Qinliang Su, Xiaojun Quan, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.4
2023 Continuous Geo-Social Group Monitoring in Dynamic LBSNs
abstract
Geo-social groupqueries, which return a social cohesive user group with a spatial constraint, have receive significant research interests due to their promising applications for group-based activity planning and scheduling in location-based social networks (LBSNs). However, existing studies on geo-social group queries mostly assume the users are stationary whereas in realistic LBSN application scenarios all users may continuously move over time. Thus, in this paper, we investigate the problem ofcontinuousgeo-socialgroupsmonitoring(CGSGM) over moving users. A challenge in answering CGSGM queries over moving users is how to efficiently update geo-social groups when users are continuously moving. To address the CGSGM problem, we first propose a baseline algorithm, namelyBaseline-BB, which recomputes the new geo-social groups from scratch at each time instance by utilizing a branch and bound (BB) strategy. To improve the inefficiency of BB, we explore a new strategy, called common neighbor or neighbor expanding (CNNE), which expands the common neighbors of edges or the neighbors of users in intermediate groups to quickly produce the valid group combinations. Accordingly, another baseline algorithm, namelyBaseline-CNNE, is proposed. As these baseline algorithms do not maintain intermediate results to facilitate further query processing, we develop an incremental algorithm, calledincremental monitoring algorithm (IMA), which maintains the support, common neighbors and the neighbors of current users when exploring possible user groups for further updates and query processing. Since IMA requires many times of truss decomposition when processing mutiple-users updates, we propose an improved incremental algorithm, calledimproved incremental monitoring algorithm (IIMA), which performs truss decompostion only once. Moreover, we design algorithms for handling the social changes that result in insertion/deletion of some edges in the social network. Owing to the challenge in setting, an appropriate monitoring distance, we further study the top$N$CGSGM problem, which finds top$N$result groups at each time instance. Finally, we conduct extensive experiments using four real datasets to validate our ideas and evaluate the proposed algorithms.
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Libin Zheng 0001, Xin Huang 0001, Jianliang Xu, Wang-Chien Lee
IEEE Trans. Knowl. Data Eng.3
2022 Crowdsourced Fact Validation for Knowledge Bases
abstract
In spite of its wide usage in various applications, existing construction methods for Knowledge Base (KB) are still on their way to obtaining 100% correct facts. Thus, employing crowd workers to validate a KB has been proposed to improve its reliability. Most of the existing works focus on devising games with proper incentives to engage workers in validating more facts, but rarely consider matching facts with proper workers. Facts have diverse domains (topics), which naturally require workers of different expertise. In addition, they also generally have different utilities, i.e., some are more heavily used than others. Thus, distinguishing the facts in terms of utility to give them different validation priorities is meaningful, especially when the budget is limited. To this end, we study the crowdsourced fact validation problem which considers worker domains and fact utilities, and find that with some reductions, it can be solved by the existing minimum cost network flow method. However, directly employing that method requires a huge time cost. We thereby propose an optimized network flow method which reduces the network complexity to save the time cost by properly grouping the facts. Furthermore, we propose an incremental validation method, which utilizes the previous results for validating an evolving KB. We finally conduct extensive experiments to demonstrate the effectiveness of the proposed methods.
Libin Zheng 0001, Peng Cheng 0003, Lei Chen 0002, Jianxing Yu, Xuemin Lin 0001, Jian Yin 0001
ICDE6
2022 Continuous Geo-Social Group Monitoring over Moving Users
abstract
Recently a lot of research works have focused on geo-social group queries for group-based activity planning and scheduling in location-based social networks (LBSNs), which return a social cohesive user group with a spatial constraint. However, existing studies on geo-social group queries assume the users are stationary whereas in real LBSN applications all users may continuously move over time. Thus, in this paper we in-vestigate the problem of continuous geo-social groups monitoring (CGSGM) over moving users. A challenge in answering CGSGM queries over moving users is how to efficiently update geo-social groups when users are continuously moving. To address the CGSGM problem, we first propose a baseline algorithm, namely Baseline-BB, which recomputes the new geo-social groups from scratch at each time instance by utilizing a branch and bound (BB) strategy. To improve the inefficiency of BB, we propose a new strategy, called common neighbor or neighbor expanding (CNNE), which expands the common neighbors of edges or the neighbors of users in intermediate groups to quickly produce the valid group combinations. Based on CNNE, we propose another baseline algorithm, namely Baseline-CNNE. As these baseline algorithms do not maintain any intermediate results to facilitate further query processing, we develop an incremental algorithm, called incremental monitoring algorithm (IMA), which maintains the support, common neighbors and the neighbors of current users when exploring possible user groups for further updates and query processing. Finally, we conduct extensive experiments using three real datasets to validate our ideas and evaluate the proposed algorithms,
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Mengxiang Wang, Jianliang Xu, Xin Huang 0001, Wang-Chien Lee
ICDE3
2022 MRVAE: Variational Autoencoder with Multiple Relationships for Collaborative Filtering
Zhou Pan, Wei Liu 0061, Jian Yin 0001
ICWE3
2022 Availability Attacks Create Shortcuts
abstract
Availability attacks, which poison the training data with imperceptible perturbations, can make the data not exploitable by machine learning algorithms so as to prevent unauthorized use of data. In this work, we investigate why these perturbations work in principle. We are the first to unveil an important population property of the perturbations of these attacks: they are almost linearly separable when assigned with the target labels of the corresponding samples, which hence can work as shortcuts for the learning objective. We further verify that linear separability is indeed the workhorse for availability attacks. We synthesize linearly-separable perturbations as attacks and show that they are as powerful as the deliberately crafted attacks. Moreover, such synthetic perturbations are much easier to generate. For example, previous attacks need dozens of hours to generate perturbations for ImageNet while our algorithm only needs several seconds. Our finding also suggests that the shortcut learning is more widely present than previously believed as deep models would rely on shortcuts even if they are of an imperceptible scale and mixed together with the normal features. Our source code is published at https://github.com/dayu11/Availability-Attacks-Create-Shortcuts.
Huishuai Zhang, Wei Chen 0034, Jian Yin 0001, Tie-Yan Liu
KDD4
2022 Top k Optimal Sequenced Route Query with POI Preferences
abstract
Abstract The optimal sequenced route (OSR) query, as a popular problem in route planning for smart cities, searches for a minimum-distance route passing through several POIs in a specific order from a starting position. In reality, POIs are usually rated, which helps users in making decisions. Existing OSR queries neglect the fact that the POIs in the same category could have different scores, which may affect users’ route choices. In this paper, we study a novel variant of OSR query, namely Rating Constrained Optimal Sequenced Route query (RCOSR), in which the rating score of each POI in the optimal sequenced route should exceed the query threshold. To efficiently process RCOSR queries, we first extend the existing TD-OSR algorithm to propose a baseline method, called MTDOSR. To tackle the shortcomings of MTDOSR, we try to design a new RCOSR algorithm, namely Optimal Subroute Expansion (OSE) Algorithm. To enhance the OSE algorithm, we propose a Reference Node Inverted Index (RNII) to accelerate the distance computation of POI pairs in OSE and quickly retrieve the POIs of each category. To make full use of the OSE and RNII, we further propose a new efficient RCOSR algorithm, called Recurrent Optimal Subroute Expansion (ROSE), which recurrently utilizes OSE to compute the current optimal route as the guiding path and update the distance of POI pairs to guide the expansion. Then, we extend our techniques to handle a variation of RCOSR query, namely RCkOSR query. The experimental results demonstrate that the proposed algorithm significantly outperforms the existing approaches.
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Jianliang Xu
Data Sci. Eng.4
2022 Copula Guided Parallel Gibbs Sampling for Nonparametric and Coherent Topic Discovery
abstract
Hierarchical Dirichlet Process (HDP) has attracted much attention in the research community of natural language processing. Given a corpus, HDP is able to determine the number of topics automatically, possessing an important feature dubbed nonparametric that overcomes the challenging issue of manually specifying a suitable topic number in parametric topic models, such as Latent Dirichlet Allocation (LDA). Nevertheless, HDP requires a much higher computational cost than LDA for parameter estimation. By taking the advantage of multi-threading, a parallel Gibbs sampling algorithm is proposed to estimate parameters for HDP based on the equivalence between HDP and Gamma-Gamma Poisson Process (G2PP) in terms of the generative process. Unfortunately, the above parallel Gibbs sampling algorithm requires to apply the finite approximation on the number of topics manually (i.e., predefine the topic number), thus can not retain the nonparametric feature of HDP. Another drawback of the above models is the lack of capturing the semantic dependencies between words, because the topic assignment of words is independent with each other. Although some works have been done in phrase-based topic modelling, these existing methods are still limited by either enforcing the entire phrase to share a common topic or requiring much complex and time-consuming phrase mining methods. In this paper, we aim to develop a copula guided parallel Gibbs sampling algorithm for HDP which can adjust the number of topics dynamically and capture the latent semantic dependencies between words that compose a coherent segment. Extensive experiments on real-world datasets indicate that our method achieves low perplexities and high topic coherence scores with a small time cost. In addition, we validate the effectiveness of our method on the modelling of word semantic dependencies by comparing the extracted topical phrases with those learned by state-of-the-art phrase-based baselines.
Lihui Lin, Yanghui Rao, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001, Fu Lee Wang, Qing Li 0001
IEEE Trans. Knowl. Data Eng.5
2021 Optimal Sequenced Route Query with POI Preferences
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Jianliang Xu
DASFAA (1)4
2021 Multi-Task Learning with Personalized Transformer for Review Recommendation
Wei Liu 0061, Jian Yin 0001
WISE (2)3
2021 REST: Relational Event-driven Stock Trend Forecasting
abstract
Stock trend forecasting, aiming at predicting the stock future trends, is crucial for investors to seek maximized profits from the stock market. Many event-driven methods utilized the events extracted from news, social media, and discussion board to forecast the stock trend in recent years. However, existing event-driven methods have two main shortcomings: 1) overlooking the influence of event information differentiated by the stock-dependent properties; 2) neglecting the effect of event information from other related stocks. In this paper, we propose a relational event-driven stock trend forecasting (REST) framework, which can address the shortcoming of existing methods. To remedy the first shortcoming, we propose to model the stock context and learn the effect of event information on the stocks under different contexts. To address the second shortcoming, we construct a stock graph and design a new propagation layer to propagate the effect of event information from related stocks. The experimental studies on the real-world data demonstrate the efficiency of our REST framework. The results of investment simulation show that our framework can achieve a higher return of investment than baselines.
Weiqing Liu, Chang Xu 0008, Jiang Bian 0002, Jian Yin 0001, Tie-Yan Liu
WWW5
2021 Querying Optimal Routes for Group Meetup
abstract
Abstract Motivated by location-based social networks which allow people to access location-based services as a group, we study a novel variant of optimal sequenced route (OSR) queries, optimal sequenced route for group meetup (OSR-G) queries. OSR-G query aims to find the optimal meeting POI (point of interest) such that the maximum users’ route distance to the meeting POI is minimized after each user visits a number of POIs of specific categories (e.g., gas stations, restaurants, and shopping malls) in a particular order. To process OSR-G queries, we first propose an OSR-Based (OSRB) algorithm as our baseline, which examines every POI in the meeting category and utilizes existing OSR (called E-OSR) algorithm to compute the optimal route for each user to the meeting POI. To address the shortcomings (i.e., requiring to examine every POI in the meeting category) of OSRB, we propose an upper bound based filtering algorithm, called circle filtering (CF) algorithm, which exploits the circle property to filter the unpromising meeting POIs. In addition, we propose a lower bound based pruning (LBP) algorithm, namely LBP-SP which exploits a shortest path lower bound to prune the unqualified meeting POIs to reduce the search space. Furthermore, we develop an approximate algorithm, namely APS, to accelerate OSR-G queries with a good approximation ratio. Finally the experimental results show that both CF and LBP-SP outperform the OSRB algorithm and have high pruning rates. Moreover, the proposed approximate algorithm runs faster than the exact OSR-G algorithms and has a good approximation ratio.
Huaijie Zhu, Wei Liu 0061, Jian Yin 0001, Wang-Chien Lee, Jianliang Xu
Data Sci. Eng.4
2020 Generating Multi-hop Reasoning Questions to Improve Machine Reading Comprehension
abstract
This paper focuses on the topic of multi-hop question generation, which aims to generate questions needed reasoning over multiple sentences and relations to derive answers. In particular, we first build an entity graph to integrate various entities scattered over text based on their contextual relations. We then heuristically extract the sub-graph by the evidential relations and type, so as to obtain the reasoning chain and textual related contents for each question. Guided by the chain, we propose a holistic generator-evaluator network to form the questions, where such guidance helps to ensure the rationality of generated questions which need multi-hop deduction to correspond to the answers. The generator is a sequence-to-sequence model, designed with several techniques to make the questions syntactically and semantically valid. The evaluator optimizes the generator network by employing a hybrid mechanism combined of supervised and reinforced learning. Experimental results on HotpotQA data set demonstrate the effectiveness of our approach, where the generated samples can be used as pseudo training data to alleviate the data shortage problem for neural network and assist to learn the state-of-the-arts for multi-hop machine comprehension.
Jianxing Yu, Xiaojun Quan, Qinliang Su, Jian Yin 0001
WWW4
2019 Deep Policy Hashing Network with Listwise Supervision
abstract
Deep-networks-based hashing has become a leading approach for large-scale image retrieval, which learns a similarity-preserving network to map similar images to nearby hash codes. The pairwise and triplet losses are two widely used similarity preserving manners for deep hashing. These manners ignore the fact that hashing is a prediction task on the list of binary codes. However, learning deep hashing with listwise supervision is challenging in 1) how to obtain the rank list of whole training set when the batch size of the deep network is always small and 2) how to utilize the listwise supervision. In this paper, we present a novel deep policy hashing architecture with two systems are learned in parallel: aquery network and a shared and slowly changingdatabase network. The following three steps are repeated until convergence: 1) the database network encodes all training samples into binary codes to obtain whole rank list, 2) the query network is trained based on policy learning to maximize a reward that indicates the performance of the whole ranking list of binary codes, e.g., mean average precision (MAP), and 3) the database network is updated as the query network. Extensive evaluations on several benchmark datasets show that the proposed method brings substantial improvements over state-of-the-art hashing methods.
Shaoying Wang, Hanjiang Lai, Jian Yin 0001
ICMR4
2019 Feature Pyramid Hashing
abstract
In recent years, deep-networks-based hashing has become a leading approach for large-scale image retrieval. Most deep hashing approaches use the high layer to extract the powerful semantic representations. However, these methods have limited ability for fine-grained image retrieval because the semantic features extracted from the high layer are difficult in capturing the subtle differences. To this end, we propose a novel two-pyramid hashing architecture to learn both the semantic information and the subtle appearance details for fine-grained image search. Inspired by the feature pyramids of convolutional neural network, avertical pyramid is proposed to capture the high-layer features and ahorizontal pyramid combines multiple low-layer features with structural information to capture the subtle differences. To fuse the low-level features, a novel combination strategy, called consensus fusion, is proposed to capture all subtle information from several low-layers for finer retrieval. Extensive evaluation on two fine-grained datasets CUB-200-2011 and Stanford Dogs demonstrate that the proposed method achieves significant performance compared with the state-of-art baselines.
Libing Geng, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001
ICMR5
2019 Sentiment Classification Using Negative and Intensive Sentiment Supplement Information
abstract
Traditional methods of annotating the sentiment of an unlabeled document are based on sentiment lexicons or machine learning algorithms, which have shown low computational cost or competitive performance. However, these methods ignore the semantic composition problem displaying in several ways such as negative reversing and intensification. In this paper, we propose a new method for sentiment classification using negative and intensive sentiment supplementary information, so as to exploit the linguistic feature of negative and intensive words in conjunction with the context information. Particularly, our method can solve the domain-specific problem without relying on the external sentiment lexicons. Experimental results on two real-world datasets demonstrate the effectiveness of our proposed method.
Xingming Chen, Yanghui Rao, Haoran Xie 0001, Fu Lee Wang, Yingchao Zhao 0001, Jian Yin 0001
Data Sci. Eng.6
2019 Reachable region query and its applications
Mengdie Nie, Zhi-Jie Wang 0009, Jian Yin 0001, Bin Yao 0002
Inf. Sci.3
2018 Learning Dual Preferences with Non-negative Matrix Tri-Factorization for Top-N Recommender System
Xiangsheng Li, Yanghui Rao, Haoran Xie 0001, Yufu Chen, Raymond Y. K. Lau, Fu Lee Wang, Jian Yin 0001
DASFAA (1)7
2018 Geographical Relevance Model for Long Tail Point-of-Interest Recommendation
Wei Liu 0061, Zhi-Jie Wang 0009, Bin Yao 0002, Mengdie Nie, Jing Wang 0030, Rui Mao 0001, Jian Yin 0001
DASFAA (1)7
2018 FastPM: An approach to pattern matching via distributed stream processing
Dingyu Yang, Jianmei Guo, Zhi-Jie Wang 0009, Yuan Wang 0003, Jingsong Zhang, Liang Hu 0004, Jian Yin 0001, Jian Cao 0001
Inf. Sci.7
2017 Cluster-level Emotion Pattern Matching for Cross-Domain Social Emotion Classification
abstract
This paper addresses the task of cross-domain social emotion classification of online documents. The cross-domain task is formulated as using abundant labeled documents from a source domain and a small amount of labeled documents from a target domain, to predict the emotion of unlabeled documents in the target domain. Although several cross-domain emotion classification algorithms have been proposed, they require that feature distributions of different domains share a sufficient overlapping, which is hard to meet in practical applications. This paper proposes a novel framework, which uses the emotion distribution of training documents at the cluster level, to alleviate the aforementioned issue. Experimental results on two datasets show the effectiveness of our proposed model on cross-domain social emotion classification.
Endong Zhu, Yanghui Rao, Haoran Xie 0001, Jian Yin 0001, Fu Lee Wang
CIKM5
2015 Online Personalized Recommendation Based on Streaming Implicit User Feedback
Zhisheng Wang 0001, Wei Liu 0061, Jian Yin 0001
APWeb5
2015 Novel Word Features for Keyword Extraction
Jian Yin 0001, Weiheng Zhu, Shiding Qiu
WAIM2
2013 Recommendations for two-way selections using skyline view queries
Jian Chen 0011, Jin Huang 0007, Bin Jiang 0009, Jian Pei 0001, Jian Yin 0001
Knowl. Inf. Syst.5
2012 Customer Segmentation for Power Enterprise Based on Enhanced-FCM Algorithm
Lihe Song, Weixu Zhan, Shuchai Qian, Jian Yin 0001
ADMA4
2012 Transfer Learning with Local Smoothness Regularizer
Jiaming Hong, Bingchao Chen, Jian Yin 0001
APWeb3
2009 Collaborative Filtering Recommendation Algorithm Using Dynamic Similar Neighbor Probability
Chuangguang Huang, Jian Yin 0001, Jing Wang 0030, Li-Rong Zheng 0004
ADMA2
2008 Face Recognition Using Clustering Based Optimal Linear Discriminant Analysis
Wenxin Yang, Shuqin Rao, Jina Wang, Jian Yin 0001, Jian Chen 0011
ADMA4
2007 Clustering Massive Text Data Streams by Semantic Smoothing Model
Jiarong Cai, Jian Yin 0001, Ada Wai-Chee Fu
ADMA3
2006 Granularity Adaptive Density Estimation and on Demand Clustering of Concept-Drifting Data Streams
Weiheng Zhu, Jian Pei 0001, Jian Yin 0001, Yihuang Xie
DaWaK3
2005 Mining Correlated Rules for Associative Classification
Jian Chen 0011, Jian Yin 0001, Jin Huang 0007
ADMA2
2001 Towards Efficient Data Re-mining (DRM)
Jiming Liu 0001, Jian Yin 0001
PAKDD2