VLDB 2026 Research / reviewers in the wild / expert
Zhoujun Li 0001
dblp:76/2866-1 · also Zhou-Jun Li 0001
· DBLP profile ↗
85ranked-venue papers in the field
0as first author
24since 2021 · last 2026
0000-0002-9603-9713ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 33Database Systems & Data Management · 23Data Mining & Knowledge Discovery · 20Knowledge Engineering, Semantic Web & Information Systems · 6Other / Interdisciplinary · 2Business Process & Enterprise Data · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building and Benchmarking Large Language Models for Machine Translation in Social Network Services
Hongcheng Guo, Fei Zhao 0012, Shaosheng Cao, Xinze Lyu, Zijie Meng, Yao Hu 0002, Zhoujun Li 0001, Zuozhu Liu |
ICDE | 8 |
| 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QAabstractTables serve as a core format for representing structured data on the web, as their two-dimensional layouts effectively encode complex inter-entity relationships. However, real-world web tables often feature heterogeneous structures and rich semantics. Accurately interpreting such tables requires not only spatial layout perception but also multi-step reasoning across rows and columns, posing substantial challenges to web intelligence systems. Multimodal large language models (MLLMs) show promise in table question answering (TableQA) by leveraging visual layouts. However, their performance on complex web tables remains uneven, as existing benchmarks often blur the impact of individual difficulty factors, hindering precise capability analysis. To advance TableQA beyond superficial task difficulty and toward interpretable capability modeling, we introduce MMTableBench, a multi-level benchmark that systematically evaluates MLLMs along two fine-grained dimensions: layout complexity and reasoning complexity. By organizing table-question pairs along these axes, MMTableBench facilitates a detailed evaluation of model performance under varying structural and reasoning challenges, while revealing the respective strengths and limitations of multimodal inputs. Our comprehensive analysis shows that state-of-the-art MLLMs continue to exhibit notable limitations when confronted with complex layouts and deep reasoning tasks, underscoring persistent gaps despite the structural advantages offered by visual inputs. MMTableBench thus provides not only a rigorous evaluation framework but also a diagnostic tool for analyzing and interpreting model behaviors, enabling more transparent and explainable progress in multimodal TableQA development. Xianjie Wu, Xiaohang Xu 0002, Tingyu Jiang, Jian Yang 0030, Di Liang, Xianfu Cheng, Zhenhe Wu, Linzheng Chai, Wei Zhang 0384, Ge Zhang 0009, Bob Simons, Tongliang Li, Zhoujun Li 0001 |
WWW | 14 |
| 2025 | DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI InteractionabstractEmotional development of children fundamentally relies on secure attachment relationships, yet current AI companions lack the theoretical foundation to provide developmentally appropriate emotional support. We introduce DinoCompanion, the first attachment-theory-grounded multimodal robot for emotionally responsive child-AI interaction. We address three critical challenges in child-AI systems: the absence of developmentally-informed AI architectures, the need to balance engagement with safety, and the lack of standardized evaluation frameworks for attachment-based capabilities. Our contributions include: (i) a multimodal dataset of 128 caregiver-child dyads containing 125,382 annotated clips with paired preference-risk labels, (ii) CARPO (Child-Aware Risk-calibrated Preference Optimization), a novel training objective that maximizes engagement while applying epistemic-uncertainty-weighted risk penalties, and (iii) AttachSecure-Bench, a comprehensive evaluation benchmark covering ten attachment-centric competencies with strong expert consensus. AttachSecure-Bench achieves state-of-the-art performance (57.15%), outperforming GPT-4o and Gemini-2.5-Pro, with exceptional secure base behaviors and superior attachment risk detection. Ablations validate the critical importance of multimodal fusion, uncertainty-aware risk modeling, and hierarchical memory for coherent, emotionally attuned interactions. Boyang Wang 0006, Yuhao Song, Jinyuan Cao, Hongcheng Guo, Zhoujun Li 0001 |
CIKM | 6 |
| 2025 | Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network ServicesabstractAs interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce PET-BENCH, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, PET-BENCH emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. PET-BENCH serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions. Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang 0006, Weiting Liu 0001, Zheyu Ye, Zhoujun Li 0001, Zuozhu Liu, Wei Lu 0011 |
CIKM | 7 |
| 2025 | ECLIPSE: Efficient Cross-Lingual Log Intelligence Parser with Semantic Entropy-Enhanced LCS Algorithm
Wei Zhang 0384, Xianfu Cheng, Xiang Li 0117, Jian Yang 0030, Xiangyuan Guan, Zhoujun Li 0001 |
CIKM | 7 |
| 2025 | SCM: Enhancing Large Language Model with Self-Controlled Memory Framework
Xinnian Liang, Jian Yang 0003, Hui Huang 0021, Zhenhe Wu, Shuangzhi Wu, Zejun Ma 0001, Zhoujun Li 0001 |
DASFAA (6) | 8 |
| 2025 | MR-SQL: Multi-level Retrieval Enhances Inference for LLM in Text-to-SQL
Zhenhe Wu, Zhongqiu Li, Mengxiang Li, Zhongjiang He, Jian Yang 0003, Yu Zhao 0007, Ruiyu Fang, Zhoujun Li 0001, Shuangyong Song |
DASFAA (2) | 10 |
| 2025 | Breaking Size Barrier: Enhancing Reasoning for Large-Size Table Question Answering
Xianjie Wu, Di Liang, Jian Yang 0037, Xianfu Cheng, Linzheng Chai, Tongliang Li, Liqun Yang, Zhoujun Li 0001 |
DASFAA (2) | 8 |
| 2024 | SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
Xianfu Cheng, Weixiao Zhou, Xiang Li 0117, Jian Yang 0030, Tao Sun 0016, Wei Zhang 0384, Yuying Mai, Tongliang Li, Xiaoming Chen 0007, Zhoujun Li 0001 |
CIKM | 11 |
| 2024 | EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented GenerationabstractThe integration of multimodal Electronic Health Records (EHR) data has significantly advanced clinical predictive capabilities. Existing models, which utilize clinical notes and multivariate time-series EHR data, often fall short of incorporating the necessary medical context for accurate clinical tasks, while previous approaches with knowledge graphs (KGs) primarily focus on structured knowledge extraction. In response, we propose EMERGE, a Retrieval-Augmented Generation (RAG) driven framework to enhance multimodal EHR predictive modeling. We extract entities from both time-series data and clinical notes by prompting Large Language Models (LLMs) and align them with professional PrimeKG, ensuring consistency. In addition to triplet relationships, we incorporate entities' definitions and descriptions for richer semantics. The extracted knowledge is then used to generate task-relevant summaries of patients' health statuses. Finally, we fuse the summary with other modalities using an adaptive multimodal fusion network with cross-attention. Extensive experiments on the MIMIC-III and MIMIC-IV datasets' in-hospital mortality and 30-day readmission tasks demonstrate the superior performance of the EMERGE framework over baseline models. Comprehensive ablation studies and analysis highlight the efficacy of each designed module and robustness to data sparsity. EMERGE contributes to refining the utilization of multimodal EHR data in healthcare, bridging the gap with nuanced medical contexts essential for informed clinical predictions. We have publicly released the code at https://github.com/yhzhu99/EMERGE. Yinghao Zhu, Changyu Ren, Shiyun Xie, Junlan Feng, Zhoujun Li 0001, Liantao Ma, Chengwei Pan |
CIKM | 8 |
| 2024 | RoNID: New Intent Discovery with Generated-Reliable Labels and Cluster-friendly Representations
Chaoran Yan, Jian Yang 0030, Changyu Ren, Jiaqi Bai 0001, Tongliang Li, Zhoujun Li 0001 |
DASFAA (5) | 7 |
| 2024 | TiNID: A Transfer and Interpretable LLM-Enhanced Framework for New Intent Discovery
Chaoran Yan, Jian Yang 0030, Wei Zhang 0384, Changyu Ren, Tongliang Li, Jiaqi Bai 0001, Zhoujun Li 0001 |
ECML/PKDD (5) | 8 |
| 2024 | Multi-task Hierarchical Heterogeneous Fusion Framework for multimodal summarization
Litian Zhang, Linfeng Han, Zelong Yu, Zhoujun Li 0001 |
Inf. Process. Manag. | 6 |
| 2024 | PPMGS: An efficient and effective solution for distributed privacy-preserving semi-supervised learning
Zhi Li 0045, Chaozhuo Li, Zhoujun Li 0001, Jian Weng 0001, Feiran Huang |
Inf. Sci. | 3 |
| 2023 | Read Then Respond: Multi-granularity Grounding Prediction for Knowledge-Grounded Dialogue Generation
Yiyang Du, Shi-Wei Zhang, Xianjie Wu, Yunbo Cao, Zhoujun Li 0001 |
ADMA (2) | 6 |
| 2023 | GripRank: Bridging the Gap between Retrieval and Generation via the Generative Knowledge Improved Passage RankingabstractRetrieval-enhanced text generation has shown remarkable progress on knowledge-intensive language tasks, such as open-domain question answering and knowledge-enhanced dialogue generation, by leveraging passages retrieved from a large passage corpus for delivering a proper answer given the input query. However, the retrieved passages are not ideal for guiding answer generation because of the discrepancy between retrieval and generation, i.e., the candidate passages are all treated equally during the retrieval procedure without considering their potential to generate a proper answer. This discrepancy makes a passage retriever deliver a sub-optimal collection of candidate passages to generate the answer. In this paper, we propose the GeneRative Knowledge Improved Passage Ranking (GripRank) approach, addressing the above challenge by distilling knowledge from a generative passage estimator (GPE) to a passage ranker, where the GPE is a generative language model used to measure how likely the candidate passages can generate the proper answer. We realize the distillation procedure by teaching the passage ranker learning to rank the passages ordered by the GPE. Furthermore, we improve the distillation quality by devising a curriculum knowledge distillation mechanism, which allows the knowledge provided by the GPE can be progressively distilled to the ranker through an easy-to-hard curriculum, enabling the passage ranker to correctly recognize the provenance of the answer from many plausible candidates. We conduct extensive experiments on four datasets across three knowledge-intensive language tasks. Experimental results show advantages over the state-of-the-art methods for both passage ranking and answer generation on the KILT benchmark. Jiaqi Bai 0001, Hongcheng Guo, Jian Yang 0030, Xinnian Liang, Zhoujun Li 0001 |
CIKM | 7 |
| 2023 | LogLG: Weakly Supervised Log Anomaly Detection via Log-Event Graph Construction
Hongcheng Guo, Yuhui Guo, Jian Yang 0030, Zhoujun Li 0001, Tieqiao Zheng, Liangfan Zheng, Weichao Hou, Bo Zhang 0096 |
DASFAA (4) | 5 |
| 2023 | HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei |
DASFAA (3) | 9 |
| 2023 | Modeling Intra-class and Inter-class Constraints for Out-of-Domain Detection
Jiaqi Bai 0001, Tongliang Li, Zhoujun Li 0001 |
DASFAA (4) | 5 |
| 2023 | KnowPrefix-Tuning: A Two-Stage Prefix-Tuning Framework for Knowledge-Grounded Dialogue Generation
Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Xinnian Liang, Hongcheng Guo, Zhoujun Li 0001 |
ECML/PKDD (2) | 7 |
| 2022 | DialCSP: A Two-Stage Attention-Based Model for Customer Satisfaction Prediction in E-commerce Customer Service
Zhenhe Wu, Liangqing Wu, Shuangyong Song, Jiahao Ji, Zhoujun Li 0001, Xiaodong He 0001 |
ECML/PKDD (3) | 6 |
| 2022 | Generative Adversarial Framework for Cold-Start Item RecommendationabstractThe cold-start problem has been a long-standing issue in recommendation. Embedding-based recommendation models provide recommendations by learning embeddings for each user and item from historical interactions. Therefore, such embedding-based models perform badly for cold items which haven't emerged in the training set. The most common solutions are to generate the cold embedding for the cold item from its content features. However, the cold embeddings generated from contents have different distribution as the warm embeddings are learned from historical interactions. In this case, current cold-start methods are facing an interesting seesaw phenomenon, which improves the recommendation of either the cold items or the warm items but hurts the opposite ones. To this end, we propose a general framework named Generative Adversarial Recommendation (GAR). By training the generator and the recommender adversarially, the generated cold item embeddings can have similar distribution as the warm embeddings that can even fool the recommender. Simultaneously, the recommender is fine-tuned to correctly rank the "fake'' warm embeddings and the real warm embeddings. Consequently, the recommendation of the warms and the colds will not influence each other, thus avoiding the seesaw phenomenon. Additionally, GAR could be applied to any off-the-shelf recommendation model. Experiments on two datasets present that GAR has strong overall recommendation performance in cold-starting both the CF-based model (improved by over 30.18%) and the GNN-based model (improved by over 17.78%). Hao Chen 0062, Zefan Wang, Feiran Huang, Xiao Huang 0001, Yishi Lin, Zhoujun Li 0001 |
SIGIR | 8 |
| 2022 | An improved stochastic gradient descent algorithm based on Rényi differential privacyabstractDeep learning techniques based on the neural network have made significant achievements in various fields of artificial intelligence. However, model training requires large-scale data sets, these data sets are crowd-sourced and model parameters will contain the encoding of private information, resulting in the risk of privacy leakage. With the trend toward sharing pretrained models, the risk of stealing training data sets through member inference attacks and model inversion attacks is further heightened. To tackle the privacy-preserving problems in deep learning tasks, we propose an improved Differential Privacy Stochastic Gradient Descent algorithm, using Simulated Annealing algorithm and Laplace Smooth denoising mechanism to optimize the allocation method of privacy loss, replacing the constant clipping method with adaptive gradient clipping method to improve model accuracy. we also analyze privacy cost under random shuffle data batch processing method in detail within the framework of Subsampled Rényi Differential Privacy. Compared with the existing privacy protection training methods with fixed parameters and dynamic privacy parameters in classification tasks, our implementation and experiments show that we can use less privacy budget train deep neural networks with the nonconvex objective function, obtain a higher model evaluation, and have almost zero additional cost in terms of model complexity, training efficiency, and model quality. Xianfu Cheng, Ao Liu 0006, Zhoujun Li 0001 |
Int. J. Intell. Syst. | 5 |
| 2021 | AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical QueriesabstractModern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users" or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data. Tongliang Li, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001, Dongmei Zhang 0001 |
WSDM | 4 |
| 2020 | Graph Unfolding NetworksabstractThe technique of recursive neighborhood aggregation has dominated the implementation of existing successful Graph Neural Networks (GNNs). However, the recursive information propagation across layers inevitably brings in extra calculations, potentially large variance, and difficulty of parallel computation. In this paper, we propose Graph Unfolding Networks (GUNets) as an alternative mechanism of recursive neighborhood aggregation for graph representation learning. Comparing to generic GNNs, our proposed GUNets are efficient, robust and practically effective. At their core, GUNets unfold the local structure of every node, i.e. the rooted tree, to a set of trajectories, and then adopt set function to capture the topology of the rooted subtree, which is more convenient for parallel computation than the recursive neighborhood aggregation process. More importantly, through a specific design of the set function, our architecture enables efficient and robust learning on large-scale graphs without resorting to any pruning of the rooted subtree that is usually necessary in generic GNNs. Extensive experiments on five large datasets (the number of nodes ranges from 104 to 106) show that our GUNets achieve comparable or even better results than current successful GNNs while gaining significantly more efficiency and lower accuracy variance. Codes can be found at github.com/GUNets/GUNets. Hao Chen 0062, Wenbing Huang 0001, Fuchun Sun 0001, Zhoujun Li 0001 |
CIKM | 5 |
| 2020 | Label-Aware Graph Convolutional NetworksabstractRecent advances in Graph Convolutional Networks (GCNs) have led to state-of-the-art performance on various graph-related tasks. However, most existing GCN models do not explicitly identify whether all the aggregated neighbors are valuable to the learning tasks, which may harm the learning performance. In this paper, we consider the problem of node classification and propose the Label-Aware Graph Convolutional Network (LAGCN) framework which can directly identify valuable neighbors to enhance the performance of existing GCN models. Our contribution is three-fold. First, we propose a label-aware edge classifier that can filter distracting neighbors and add valuable neighbors for each node to refine the original graph into a label-aware (LA) graph. Existing GCN models can directly learn from the LA graph to improve the performance without changing their model architectures. Second, we introduce the concept of positive ratio to evaluate the density of valuable neighbors in the LA graph. Theoretical analysis reveals that using the edge classifier to increase the positive ratio can improve the learning performance of existing GCN models. Third, we conduct extensive node classification experiments on benchmark datasets. The results verify that LAGCN can improve the performance of existing GCN models considerably, in terms of node classification. Hao Chen 0062, Feiran Huang, Zengde Deng, Wenbing Huang 0001, Senzhang Wang, Zhoujun Li 0001 |
CIKM | 8 |
| 2020 | BiGCNN: Bidirectional Gated Convolutional Neural Network for Chinese Named Entity Recognition
Tianyang Zhao 0003, Haoyan Liu 0001, Qianhui Wu, Changzhi Sun, Dongdong Zhang 0001, Zhoujun Li 0001 |
DASFAA (1) | 6 |
| 2020 | A unified framework of identity-based sequential aggregate signatures from 2-level HIBE schemes
Zhoujun Li 0001, Hua Guo 0001 |
Inf. Sci. | 2 |
| 2019 | Partially Shared Adversarial Learning For Semi-supervised Multi-platform User Identity LinkageabstractWith the increasing popularity and diversity of social media, users tend to join multiple social platforms to enjoy different types of services. User identity linkage, which aims to link identical identities across different social platforms, has attracted increasing research attentions recently. Existing methods usually focus on pairwise identity linkage between two platforms, which cannot piece up the information from multi-sources to depict the intrinsic figures of social users. In this paper, we propose a novel adversarial learning based framework MSUIL with partially shared generators to perform Semi-supervised User Identity Linkage across Multiple social networks. The isomorphism across multiple platforms is captured as the complementary to link identities. The insight is that we aim to learn the desirable projection functions (generators) to not only minimize the distance between the distributions of user identities in arbitrary pairs of platforms, but also incorporate the available annotations as the learning guidance. The projection functions of different platform pairs share partial parameters, which ensures MSUIL can capture the interdependencies among multiple platforms and improves the model efficiency. Empirically, we evaluate our proposal over multiple datasets. The experimental results demonstrate the superiority of the proposed MSUIL model. Chaozhuo Li, Senzhang Wang, Hao Wang 0068, Yanbo Liang, Philip S. Yu, Zhoujun Li 0001, Wei Wang 0011 |
CIKM | 6 |
| 2019 | Multi-Hot Compact Network EmbeddingabstractNetwork embedding, as a promising way of the network representation learning, is capable of supporting various subsequent network mining and analysis tasks, and has attracted growing research interests recently. Traditional approaches assign each node with an independent continuous vector, which will cause memory overhead for large networks. In this paper we propose a novel multi-hot compact network embedding framework to effectively reduce memory cost by learning partially shared embeddings. The insight is that a node embedding vector is composed of several basis vectors according to a multi-hot index vector. The basis vectors are shared by different nodes, which can significantly reduce the number of continuous vectors while maintain similar data representation ability. Specifically, we propose a MCNE$_p $ model to learn compact embeddings from pre-learned node features. A novel component named compressor is integrated into MCNE$_p $ to tackle the challenge that popular back-propagation optimization cannot propagate loss through discrete samples. We further propose an end-to-end model MCNE$_t $ to learn compact embeddings from the input network directly. Empirically, we evaluate the proposed models over four real network datasets, and the results demonstrate that our proposals can save about 90% of memory cost of network embeddings without significantly performance decline. Chaozhuo Li, Lei Zheng 0001, Senzhang Wang, Feiran Huang, Philip S. Yu, Zhoujun Li 0001 |
CIKM | 6 |
| 2019 | FGST: Fine-Grained Spatial-Temporal Based Regression for Stationless Bike Traffic Prediction
Hao Chen 0062, Senzhang Wang, Zengde Deng, Xiaoming Zhang 0001, Zhoujun Li 0001 |
PAKDD (1) | 5 |
| 2018 | Event Extraction with Deep Contextualized Word Representation and Multi-attention Layer
Ruixue Ding, Zhoujun Li 0001 |
ADMA | 2 |
| 2018 | Distribution Distance Minimization for Unsupervised User Identity LinkageabstractNowadays, it is common for one natural person to join multiple social networks to enjoy different services. Linking identical users across different social networks, also known as the User Identity Linkage (UIL), is an important problem of great research challenges and practical value. Most existing UIL models are supervised or semi-supervised and a considerable number of manually matched user identity pairs are required, which is costly in terms of labor and time. In addition, existing methods generally rely heavily on some discriminative common user attributes, and thus are hard to be generalized. Motivated by the isomorphism across social networks, in this paper we consider all the users in a social network as a whole and perform UIL from the user space distribution level. The insight is that we convert the unsupervised UIL problem to the learning of a projection function to minimize the distance between the distributions of user identities in two social networks. We propose to use the earth mover's distance (EMD) as the measure of distribution closeness, and propose two models UUIL$_gan $ and UUIL$_omt $ to efficiently learn the distribution projection function. Empirically, we evaluate the proposed models over multiple social network datasets, and the results demonstrate that our proposal significantly outperforms state-of-the-art methods. Chaozhuo Li, Senzhang Wang, Philip S. Yu, Lei Zheng 0001, Xiaoming Zhang 0001, Zhoujun Li 0001, Yanbo Liang |
CIKM | 6 |
| 2018 | Adversarial Learning of Answer-Related Representation for Visual Question AnsweringabstractVisual Question Answering (VQA) aims to learn a joint embedding of the question sentence and the corresponding image to infer the answer. Existing approaches learn the joint embedding don't consider the answer-related information, which results in that the learned representation is not effective to reflect the answer of the question. To address this problem, this paper proposes a novel method, i.e., Adversarial Learning of Answer-Related Representation (ALARR) for visual question answering, which seeks an effective answer-related representation for the question-image pair based on adversarial learning between two processes. The embedding learning process aims to generate modality-invariant joint representations for the question-image and question-answer pairs, respectively. Meanwhile, it tries to confuse the other process, embedding discriminator, which tries to discriminate the two representations from different modalities of pairs. Specifically, the joint embedding of the question-image pair is learned by a three-level attention model, and the joint representation of the question-answer pair is learned by a semantic integration model. Through the adversarial leaning, the answer-related representation are better preserved. Then an answer predictor is proposed to infer the answer from the answer-related representation. Experiments conducted on two widely used VQA benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches. Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhoujun Li 0001 |
CIKM | 4 |
| 2018 | SSDMV: Semi-Supervised Deep Social Spammer Detection by Multi-view Data FusionabstractThe explosive use of social media makes it a popular platform for malicious users, known as social spammers, to overwhelm legitimate users with unwanted content. Most existing social spammer detection approaches are supervised and need a large number of manually labeled data for training, which is infeasible in practice. To address this issue, some semi-supervised models are proposed by incorporating side information such as user profiles and posted tweets. However, these shallow models are not effective to deeply learn the desirable user representations for spammer detection, and the multi-view data are usually loosely coupled without considering their correlations. In this paper, we propose a Semi-Supervised Deep social spammer detection model by Multi-View data fusion (SSDMV). The insight is that we aim to extensively learn the task-relevant discriminative representations for users to address the challenge of annotation scarcity. Under a unified semi-supervised learning framework, we first design a deep multi-view feature learning module which fuses information from different views, and then propose a label inference module to predict labels for users. The mutual refinement between the two modules ensures SSDMV to be able to both generate high quality features and make accurate predictions.Empirically, we evaluate SSDMV over two real social network datasets on three tasks, and the results demonstrate that SSDMV significantly outperforms the state-of-the-art methods. Chaozhuo Li, Senzhang Wang, Lifang He 0001, Philip S. Yu, Yanbo Liang, Zhoujun Li 0001 |
ICDM | 6 |
| 2018 | Multimodal Network Embedding via Attention based Multi-view Variational AutoencoderabstractLearning the embedding for social media data has attracted extensive research interests as well as boomed a lot of applications, such as classification and link prediction. In this paper, we examine the scenario of a multimodal network with nodes containing multimodal contents and connected by heterogeneous relationships, such as social images containing multimodal contents (e.g., visual content and text description), and linked with various forms (e.g., in the same album or with the same tag). However, given the multimodal network, simply learning the embedding from the network structure or a subset of content results in sub-optimal representation. In this paper, we propose a novel deep embedding method, i.e., Attention-based Multi-view Variational Auto-Encoder (AMVAE), to incorporate both the link information and the multimodal contents for more effective and efficient embedding. Specifically, we adopt LSTM with attention model to learn the correlation between different data modalities, such as the correlation between visual regions and the specific words, to obtain the semantic embedding of the multimodal contents. Then, the link information and the semantic embedding are considered as two correlated views. A multi-view correlation learning based Variational Auto-Encoder (VAE) is proposed to learn the representation of each node, in which the embedding of link information and multimodal contents are integrated and mutually reinforced. Experiments on three real-world datasets demonstrate the superiority of the proposed model in two applications, i.e., multi-label classification and link prediction. Feiran Huang, Xiaoming Zhang 0001, Chaozhuo Li, Zhoujun Li 0001, Yueying He, Zhonghua Zhao |
ICMR | 4 |
| 2018 | Marefa: Turning Publishers Catalogs' Data Into Linked DataabstractThis article describes how recently, many new technologies have been introduced to the web; linked data is probably the most important. Individuals and organizations started emerging and publishing their data on the web adhering to a set of best practices. This data is published mostly in English; hence, only English agents can consume it. Meanwhile, although the number of Arabic users on the web is immense, few Arabic datasets are published. Publication catalogs are one of the primary sources of Arabic data that is not being exploited. Arabic catalogs provide a significant amount of meaningful data and metadata that are commonly stored in excel sheets. In this article, an effort has been made to help publishers easily and efficiently share their catalogs' data as linked data. Marefa is the first tool implemented that automatically extracts RDF triples from Arabic catalogs, aligns them to the BIBO ontology and links them with the Arabic chapter of DBpedia. An evaluation of the framework was conducted, and some statistical measures were generated during the different phases of the extraction process. Ahmed Ktob, Zhoujun Li 0001 |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2017 | From Properties to Links: Deep Network Embedding on Incomplete GraphsabstractAs an effective way of learning node representations in networks, network embedding has attracted increasing research interests recently. Most existing approaches use shallow models and only work on static networks by extracting local or global topology information of each node as the algorithm input. It is challenging for such approaches to learn a desirable node representation on incomplete graphs with a large number of missing links or on dynamic graphs with new nodes joining in. It is even challenging for them to deeply fuse other types of data such as node properties into the learning process to help better represent the nodes with insufficient links. In this paper, we for the first time study the problem of network embedding on incomplete networks. We propose a Multi-View Correlation-learning based Deep Network Embedding method named MVC-DNE to incorporate both the network structure and the node properties for more effectively and efficiently perform network embedding on incomplete networks. Specifically, we consider the topology structure of the network and the node properties as two correlated views. The insight is that the learned representation vector of a node should reflect its characteristics in both views. Under a multi-view correlation learning based deep autoencoder framework, the structure view and property view embeddings are integrated and mutually reinforced through both self-view and cross-view learning. As MVC-DNE can learn a representation mapping function, it can directly generate the representation vectors for the new nodes without retraining the model. Thus it is especially more efficient than previous methods. Empirically, we evaluate MVC-DNE over three real network datasets on two data mining applications, and the results demonstrate that MVC-DNE significantly outperforms state-of-the-art methods. Dejian Yang, Senzhang Wang, Chaozhuo Li, Xiaoming Zhang 0001, Zhoujun Li 0001 |
CIKM | 5 |
| 2017 | Semi-Supervised Network Embedding
Chaozhuo Li, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xiaoming Zhang 0001, Jianshe Zhou |
DASFAA (1) | 2 |
| 2017 | PPNE: Property Preserving Network Embedding
Chaozhuo Li, Senzhang Wang, Dejian Yang, Zhoujun Li 0001, Yang Yang 0002, Xiaoming Zhang 0001, Jianshe Zhou |
DASFAA (1) | 4 |
| 2017 | DTRP: A Flexible Deep Framework for Travel Route Planning
Jie Xu 0015, Chaozhuo Li, Senzhang Wang, Feiran Huang, Zhoujun Li 0001, Yueying He, Zhonghua Zhao |
WISE (1) | 5 |
| 2017 | Computing Urban Traffic Congestions by Incorporating Sparse GPS Probe Data and Social Media DataabstractEstimating urban traffic conditions of an arterial network with GPS probe data is a practically important while substantially challenging problem, and has attracted increasing research interests recently. Although GPS probe data is becoming a ubiquitous data source for various traffic related applications currently, they are usually insufficient for fully estimating traffic conditions of a large arterial network due to the low sampling frequency. To explore other data sources for more effectively computing urban traffic conditions, we propose to collect various traffic events such as traffic accident and jam from social media as complementary information. In addition, to further explore other factors that might affect traffic conditions, we also extract rich auxiliary information including social events, road features, Point of Interest (POI), and weather. With the enriched traffic data and auxiliary information collected from different sources, we first study the traffic co-congestion pattern mining problem with the aim of discovering which road segments geographically close to each other are likely to co-occur traffic congestion. A search tree based approach is proposed to efficiently discover the co-congestion patterns. These patterns are then used to help estimate traffic congestions and detect anomalies in a transportation network. To fuse the multisourced data, we finally propose a coupled matrix and tensor factorization model named TCE_R to more accurately complete the sparse traffic congestion matrix by collaboratively factorizing it with other matrices and tensors formed by other data. We evaluate the proposed model on the arterial network of downtown Chicago with 1,257 road segments whose total length is nearly 700 miles. The results demonstrate the superior performance of TCE_R by comprehensive comparison with existing approaches. Senzhang Wang, Xiaoming Zhang 0001, Jianping Cao, Lifang He 0001, Leon Stenneth, Philip S. Yu, Zhoujun Li 0001 |
ACM Trans. Inf. Syst. | 7 |
| 2016 | Query Classification by Leveraging Explicit Concept Information
Fang Wang 0019, Ze Yang 0001, Zhoujun Li 0001, Jianshe Zhou |
ADMA | 3 |
| 2016 | Estimating Urban Traffic Congestions with Multi-sourced DataabstractThis paper studies the novel problem of more accurately estimating urban traffic congestions by integrating sparse probe data and traffic related information collected from social media. Limited by the lack of reliability and low sampling frequency of GPS probes, probe data are usually not sufficient for fully estimating traffic conditions of a large arterial network. To address the data sparsity challenge, we extensively collect and model traffic related data from multiple data sources. Besides the GPS probe data, we also extensively collect traffic related tweets that report various traffic events such as congestion, accident, and road construction from both traffic authority accounts and general user accounts from Twitter. To further explore other factors that might affect traffic conditions, we also extract auxiliary information including road congestion correlations, social events, road features, as well as point of interest (POI) for help. To integrate the different types of data coming from different sources, we finally propose a coupled matrix and tensor factorization model to more accurately complete the very sparse traffic congestion matrix by collaboratively factorizing it with other matrices and tensors formed by other data. We evaluate the proposed model on the arterial network of downtown Chicago with 1257 road segments. The results demonstrate the effectiveness and efficiency of the proposed model by comparison with previous approaches. Senzhang Wang, Lifang He 0001, Leon Stenneth, Philip S. Yu, Zhoujun Li 0001 |
MDM | 5 |
| 2016 | Learning Distributed Representations of Data in Community Question Answering for Question RetrievalabstractWe study the problem of question retrieval in community question answering (CQA). The biggest challenge within this task is lexical gaps between questions since similar questions are usually expressed with different but semantically related words. To bridge the gaps, state-of-the-art methods incorporate extra information such as word-to-word translation and categories of questions into the traditional language models. We find that the existing language model based methods can be interpreted using a new framework, that is they represent words and question categories in a vector space and calculate question-question similarities with a linear combination of dot products of the vectors. The problem is that these methods are either heuristic on data representation or difficult to scale up. We propose a principled and efficient approach to learning representations of data in CQA. In our method, we simultaneously learn vectors of words and vectors of question categories by optimizing an objective function naturally derived from the framework. In question retrieval, we incorporate learnt representations into traditional language models in an effective and efficient way. We conduct experiments on large scale data from Yahoo! Answers and Baidu Knows, and compared our method with state-of-the-art methods on two public data sets. Experimental results show that our method can significantly improve on baseline methods for retrieval relevance. On 1 million training data, our method takes less than 50 minutes to learn a model on a single multicore machine, while the translation based language model needs more than 2 days to learn a translation table on the same machine. Kai Zhang 0038, Wei Wu 0014, Fang Wang 0019, Ming Zhou 0001, Zhoujun Li 0001 |
WSDM | 5 |
| 2016 | CPB: a classification-based approach for burst time prediction in cascades
Senzhang Wang, Xia Ben Hu, Philip S. Yu, Zhoujun Li 0001 |
Knowl. Inf. Syst. | 5 |
| 2016 | Coranking the Future Influence of Multiobjects in Bibliographic Network Through Mutual ReinforcementabstractScientific literature ranking is essential to help researchers find valuable publications from a large literature collection. Recently, with the prevalence of webpage ranking algorithms such as PageRank and HITS, graph-based algorithms have been widely used to iteratively rank papers and researchers through the networks formed by citation and coauthor relationships. However, existing graph-based ranking algorithms mostly focus on ranking the current importance of literature. For researchers who enter an emerging research area, they might be more interested in new papers and young researchers that are likely to become influential in the future, since such papers and researchers are more helpful in letting them quickly catch up on the most recent advances and find valuable research directions. Meanwhile, although some works have been proposed to rank the prestige of a certain type of objects with the help of multiple networks formed of multiobjects, there still lacks a unified framework to rank multiple types of objects in the bibliographic network simultaneously. In this article, we propose a unified ranking framework MRCoRank to corank the future popularity of four types of objects: papers, authors, terms, and venues through mutual reinforcement. Specifically, because the citation data of new publications are sparse and not efficient to characterize their innovativeness, we make the first attempt to extract the text features to help characterize innovative papers and authors. With the observation that the current trend is more indicative of the future trend of citation and coauthor relationships, we then construct time-aware weighted graphs to quantify the importance of links established at different times on both citation and coauthor graphs. By leveraging both the constructed text features and time-aware graphs, we finally fuse the rich information in a mutual reinforcement ranking framework to rank the future importance of multiobjects simultaneously. We evaluate the proposed model through extensive experiments on the ArnetMiner dataset containing more than 1,500,000 papers. Experimental results verify the effectiveness of MRCoRank in coranking the future influence of multiobjects in a bibliographic network. Senzhang Wang, Sihong Xie, Xiaoming Zhang 0001, Zhoujun Li 0001, Philip S. Yu, Yueying He |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2016 | Unsupervised Head-Modifier Detection in Search QueriesabstractInterpreting the user intent in search queries is a key task in query understanding. Query intent classification has been widely studied. In this article, we go one step further to understand the query from the view of head--modifier analysis. For example, given the query “popular iphone 5 smart cover,” instead of using coarse-grained semantic classes (e.g.,find electronic product), we interpret that “smart cover” is the head or the intent of the query and “iphone 5” is its modifier. Query head--modifier detection can help search engines to obtain particularly relevant content, which is also important for applications such as ads matching and query recommendation. We introduce an unsupervised semantic approach for query head--modifier detection. First, we mine a large number of instance level head--modifier pairs from search log. Then, we develop a conceptualization mechanism to generalize the instance level pairs to concept level. Finally, we derive weighted concept patterns that are concise, accurate, and have strong generalization power in head--modifier detection. The developed mechanism has been used in production for search relevance and ads matching. We use extensive experiment results to demonstrate the effectiveness of our approach. Zhongyuan Wang 0006, Fang Wang 0019, Haixun Wang, Zhirui Hu, Jun Yan 0001, Fangtao Li, Ji-Rong Wen, Zhoujun Li 0001 |
ACM Trans. Knowl. Discov. Data | 8 |
| 2015 | Inferring Diffusion Networks with Sparse Cascades by Structure Transfer
Senzhang Wang, Honghui Zhang, Jiawei Zhang 0001, Xiaoming Zhang 0001, Philip S. Yu, Zhoujun Li 0001 |
DASFAA (1) | 6 |
| 2015 | Citywide traffic congestion estimation with social mediaabstractConventional traffic congestion estimation approaches require the deployment of traffic sensors or large-scale probe vehicles. The high cost of deploying and maintaining these equipments largely limits their spatial-temporal coverage. This paper proposes an alternative solution with lower cost and wider spatial coverage by exploring traffic related information from Twitter. By regarding each Twitter user as a traffic monitoring sensor, various real-time traffic information can be collected freely from each corner of the city. However, there are two major challenges for this problem. Firstly, the congestion related information extracted directly from real-time tweets are very sparse due both to the low resolution of geographic location mentioned in the tweets and the inherent sparsity nature of Twitter data. Secondly, the traffic event information coming from Twitter can be multi-typed including congestion, accident, road construction, etc. It is non-trivial to model the potential impacts of diverse traffic events on traffic congestion. We propose to enrich the sparse real-time tweets from two directions: 1) mining the spatial and temporal correlations of the road segments in congestion from historical data, and 2) applying auxiliary information including social events and road features for help. We finally propose a coupled matrix and tensor factorization model to effectively integrate rich information for Citywide Traffic Congestion Eestimation (CTCE). Extensive evaluations on Twitter data and 500 million public passenger buses GPS data on nearly 700 mile roads of Chicago demonstrate the efficiency and effectiveness of the proposed approach. Senzhang Wang, Lifang He 0001, Leon Stenneth, Philip S. Yu, Zhoujun Li 0001 |
SIGSPATIAL/GIS | 5 |
| 2015 | Location Prediction of Social Images via Generative ModelabstractThe vast amount of geo-tagged social images has attracted great attention in research of predicting location using the plentiful content of images, such as visual content and textual description. Most of the existing researches use the text-based or vision-based method to predict location. There still exists a problem: how to effectively exploit the correlation between different types of content as well as their geographical distributions for location prediction. In this paper, we propose to predict image location by learning the latent relation between geographical location and multiple types of image content. In particularly, we propose a geographical topic model GTMSI (geographical topic model of social image) to integrate multiple types of image content as well as the geographical distributions. In GTMI, image topic is modeled on both text vocabulary and visual feature. Each region has its own distribution over topics and hence has its own language model and vision pattern. The location of a new image is estimated based on the joint probability of image content and similarity measure on topic distribution between images. Experiment results demonstrate the performance of location prediction based on GTMSI. Xiaoming Zhang 0001, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xueqiang Lv |
ICMR | 2 |
| 2015 | Friendship Link Recommendation Based on Content Structure Information
Xiaoming Zhang 0001, Zhoujun Li 0001 |
WAIM | 3 |
| 2015 | Exploring Social Network Information for Solving Cold Start in Product Recommendation
Chaozhuo Li, Fang Wang 0019, Yang Yang 0002, Zhoujun Li 0001, Xiaoming Zhang 0001 |
WISE (2) | 4 |
| 2014 | Concept-based Short Text Classification and RankingabstractMost existing approaches for text classification represent texts as vectors of words, namely ``Bag-of-Words.'' This text representation results in a very high dimensionality of feature space and frequently suffers from surface mismatching. Short texts make these issues even more serious, due to their shortness and sparsity. In this paper, we propose using ``Bag-of-Concepts'' in short text representation, aiming to avoid the surface mismatching and handle the synonym and polysemy problem. Based on ``Bag-of-Concepts,'' a novel framework is proposed for lightweight short text classification applications. By leveraging a large taxonomy knowledgebase, it learns a concept model for each category, and conceptualizes a short text to a set of relevant concepts. A concept-based similarity mechanism is presented to classify the given short text to the most similar category. One advantage of this mechanism is that it facilitates short text ranking after classification, which is needed in many applications, such as query or ad recommendation. We demonstrate the usage of our proposed framework through a real online application: Channel-based Query Recommendation. Experiments show that our framework can map queries to channels with a high degree of precision (avg. precision=90.3%), which is critical for recommendation applications. Fang Wang 0019, Zhongyuan Wang 0006, Zhoujun Li 0001, Ji-Rong Wen |
CIKM | 3 |
| 2014 | Exploit Latent Dirichlet Allocation for One-Class Collaborative FilteringabstractPrevious work studied one-class collaborative filtering (OCCF) problems including pointwise methods, pairwise methods, and content-based methods. The fundamental assumptions made on these approaches are roughly the same. They regard all missing values as negative. However, this is unreasonable since the missing values actually are the mixture of negative and positive examples. A user does not give a positive feedback on an item probably only because she/he is unaware of the item, but in fact, she/he is fond of it. Furthermore, content-based methods, e.g. collaborative topic regression (CTR), usually require textual content information of items. This cannot be satisfied in some cases. In this paper, we exploit latent Dirichlet allocation (LDA) model on OCCF problem. It assumes missing values unknown and only models the observed data, and it also does not need content information of items. In our model items are regarded as words and users are considered as documents and the user-item feedback matrix denotes the corpus. Experimental results show that our proposed method outperforms the previous methods on various ranking-oriented evaluation metrics. Haijun Zhang 0007, Zhoujun Li 0001, Yan Chen 0019, Xiaoming Zhang 0001, Senzhang Wang |
CIKM | 2 |
| 2014 | Question Retrieval with High Quality Answers in Community Question AnsweringabstractThis paper studies the problem of question retrieval in community question answering (CQA). To bridge lexical gaps in questions, which is regarded as the biggest challenge in retrieval, state-of-the-art methods learn translation models using answers under an assumption that they are parallel texts. In practice, however, questions and answers are far from "parallel". Indeed, they are heterogeneous for both the literal level and user behaviors. There are a particularly large number of low quality answers, to which the performance of translation models is vulnerable. To address these problems, we propose a supervised question-answer topic modeling approach. The approach assumes that questions and answers share some common latent topics and are generated in a "question language" and "answer language" respectively following the topics. The topics also determine an answer quality signal. Compared with translation models, our approach not only comprehensively models user behaviors on CQA portals, but also highlights the instinctive heterogeneity of questions and answers. More importantly, it takes answer quality into account and performs robustly against noise in answers. With the topic modeling approach, we propose a topic-based language model, which matches questions not only on a term level but also on a topic level. We conducted experiments on large scale data from Yahoo! Answers and Baidu Knows. Experimental results show that the proposed model can significantly outperform state-of-the-art retrieval models in CQA. Kai Zhang 0038, Wei Wu 0014, Haocheng Wu, Zhoujun Li 0001, Ming Zhou 0001 |
CIKM | 4 |
| 2014 | MMRate: inferring multi-aspect diffusion networks with multi-pattern cascadesabstractInferring diffusion networks from traces of cascades has been extensively studied to better understand information diffusion in many domains. A widely used assumption in previous work is that the diffusion network is homogenous and diffusion processes of cascades follow the same pattern. However, in social media, users may have various interests and the connections among them are usually multi-faceted. In addition, different cascades normally diffuse at different speeds and spread to diverse scales, and hence show various diffusion patterns. It is challenging for traditional models to capture the heterogeneous user interactions and diverse patterns of cascades in social media. In this paper, we investigate a novel problem of inferring multi-aspect diffusion networks with multi-pattern cascades. In particular, we study the effects of various diffusion patterns on the information diffusion process by analyzing users' retweeting behavior on a microblogging dataset. By incorporating aspect-level user interactions and various diffusion patterns, a new model for inferring Multi-aspect transmission Rates between users using Multi-pattern cascades (MMRate) is proposed. We also provide an Expectation Maximization algorithm to effectively estimate the parameters. Experimental results on both synthetic and microblogging datasets demonstrate the superior performance of our approach over the state-of-the-art methods in inferring multi-aspect diffusion networks. Senzhang Wang, Xia Ben Hu, Philip S. Yu, Zhoujun Li 0001 |
KDD | 4 |
| 2014 | Future Influence Ranking of Scientific LiteratureabstractResearchers or students entering a emerging research area are particularly interested in what newly published papers will be most cited and which young researchers will become influential in the future, so that they can catch the most recent advances and find valuable research directions. However, predicting the future importance of scientific articles and authors is extremely hard due to the dynamic nature of literature networks and evolving research topics. Different from most previous studies aiming to rank the current importance of literature and authors, we focus on ranking the future popularity of new publications and young researchers by proposing a unified ranking model to combine various available information. Specifically, we first propose to use two kinds of text features, words and words co-occurrence to characterize innovative papers and authors. Then, instead of using static and un-weighted graphs, we construct time-aware weighted graphs to distinguish the various importance of links established at different time. Finally, by leveraging both the constructed text features and graphs, we propose a mutual reinforcement ranking framework called MRFRank to rank the future importance of papers and authors simultaneously. Experimental results on the ArnetMiner dataset show that the proposed approach significantly outperforms the baselines on the metric recommendation intensity. Senzhang Wang, Sihong Xie, Xiaoming Zhang 0001, Zhoujun Li 0001, Philip S. Yu, Xinyu Shu |
SDM | 4 |
| 2014 | Popularity Prediction of Burst Event in Microblogging
Xiaoming Zhang 0001, Zhoujun Li 0001, Wen-Han Chao, Jiali Xia |
WAIM | 2 |
| 2013 | Emerging topic detection for organizations from microblogsabstractMicroblog services have emerged as an essential way to strengthen the communications among individuals and organizations. These services promote timely and active discussions and comments towards products, markets as well as public events, and have attracted a lot of attentions from organizations. In particular, emerging topics are of immediate concerns to organizations since they signal current concerns of, and feedback by their users. Two challenges must be tackled for effective emerging topic detection. One is the problem of real-time relevant data collection and the other is the ability to model the emerging characteristics of detected topics and identify them before they become hot topics. To tackle these challenges, we first design a novel scheme to crawl the relevant messages related to the designated organization by monitoring multi-aspects of microblog content, including users, the evolving keywords and their temporal sequence. We then develop an incremental clustering framework to detect new topics, and employ a range of content and temporal features to help in promptly detecting hot emerging topics. Extensive evaluations on a representative real-world dataset based on Twitter data demonstrate that our scheme is able to characterize emerging topics well and detect them before they become hot topics. Yan Chen 0019, Hadi Amiri, Zhoujun Li 0001, Tat-Seng Chua |
SIGIR | 3 |
| 2013 | Collaborative Filtering Based on Rating Psychology
Haijun Zhang 0007, Zhoujun Li 0001, Xiaoming Zhang 0001 |
WAIM | 3 |
| 2013 | Collaborative Filtering Using Multidimensional Psychometrics Model
Haijun Zhang 0007, Xiaoming Zhang 0001, Zhoujun Li 0001 |
WAIM | 3 |
| 2013 | Efficient and dynamic key management for multiple identities in identity-based systems
Hua Guo 0001, Chang Xu 0004, Zhoujun Li 0001, Yi Mu 0001 |
Inf. Sci. | 3 |
| 2013 | Comparable Entity Mining from Comparative QuestionsabstractComparing one thing with another is a typical part of human decision making process. However, it is not always easy to know what to compare and what are the alternatives. In this paper, we present a novel way to automatically mine comparable entities from comparative questions that users posted online to address this difficulty. To ensure high precision and high recall, we develop a weakly supervised bootstrapping approach for comparative question identification and comparable entity extraction by leveraging a large collection of online question archive. The experimental results show our method achieves F1-measure of 82.5 percent in comparative question identification and 83.3 percent in comparable entity extraction. Both significantly outperform an existing state-of-the-art method. Additionally, our ranking results show highly relevance to user's comparison intents in web. Shasha Li 0001, Chin-Yew Lin, Young-In Song, Zhoujun Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | A Survival Modeling Approach to Biomedical Search Result Diversification Using WikipediaabstractIn this paper, we propose a survival modeling approach to promoting ranking diversity for biomedical information retrieval. The proposed approach concerns with finding relevant documents that can deliver more different aspects of a query. First, two probabilistic models derived from the survival analysis theory are proposed for measuring aspect novelty. Second, a new method using Wikipedia to detect aspects covered by retrieved documents is presented. Third, an aspect filter based on a two-stage model is introduced. It ranks the detected aspects in decreasing order of the probability that an aspect is generated by the query. Finally, the relevance and the novelty of retrieved documents are combined at the aspect level for reranking. Experiments conducted on the TREC 2006 and 2007 Genomics collections demonstrate the effectiveness of the proposed approach in promoting ranking diversity for biomedical information retrieval. Moreover, we further evaluate our approach in the Web retrieval environment. The evaluation results on the ClueWeb09-T09B collection show that our approach can achieve promising performance improvements. Xiaoshi Yin, Jimmy Huang 0001, Zhoujun Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | When Sparsity Meets Noise in Collaborative Filtering
Biyun Hu, Zhoujun Li 0001, Wen-Han Chao |
APWeb | 2 |
| 2012 | Data Sparsity: A Key Disadvantage of User-Based Collaborative Filtering?
Biyun Hu, Zhoujun Li 0001, Wen-Han Chao |
APWeb | 2 |
| 2012 | Tagging image by merging multiple features in a integrated manner
Xiaoming Zhang 0001, Zhoujun Li 0001, Wen-Han Chao |
J. Intell. Inf. Syst. | 2 |
| 2011 | Tagging Image with Informative and Correlative Tags
Xiaoming Zhang 0001, Heng Tao Shen, Zi Huang, Zhoujun Li 0001 |
APWeb | 4 |
| 2011 | Learning to recommend questions based on public interestabstractThis paper is concerned with the problem of question recommendation in the setting of Community Question Answering (CQA). Given a question as query, our goal is to rank all of the retrieved questions according to their likelihood of being good recommendations for the query. In this paper, we propose a notion of public interest, and show how public interest can boost the performance of question recommendation. In particular, to model public interest in question recommendation, we build a language model to combine relevance score to the query and popularity score regarding question popularity. Experimental results on Yahoo!Answers dataset demonstrate the performance of question recommendation can be greatly improved with considering the public interest. Jun Wang 0019, Xia Ben Hu, Zhoujun Li 0001, Wen-Han Chao, Biyun Hu |
CIKM | 3 |
| 2011 | Probabilistic Image Tagging with Tags Expanded By Text-Based Search
Xiaoming Zhang 0001, Zi Huang, Heng Tao Shen, Zhoujun Li 0001 |
DASFAA (1) | 4 |
| 2011 | Tagging Image by Exploring Weighted Correlation between Visual Features and Tags
Xiaoming Zhang 0001, Zhoujun Li 0001 |
WAIM | 2 |
| 2011 | Mining and modeling linkage information from citation context for improving biomedical literature retrieval
Xiaoshi Yin, Jimmy Huang 0001, Zhoujun Li 0001 |
Inf. Process. Manag. | 3 |
| 2011 | Provably secure identity-based authenticated key agreement protocols with malicious private key generators
Hua Guo 0001, Zhoujun Li 0001, Yi Mu 0001, Xiyong Zhang |
Inf. Sci. | 2 |
| 2010 | Hierarchical Classification with Dynamic-Threshold SVM Ensemble for Gene Function Prediction
Yiming Chen 0002, Zhoujun Li 0001, Xiaohua Hu 0001, Junwan Liu |
ADMA (2) | 2 |
| 2010 | User's Latent Interest-Based Collaborative Filtering
Biyun Hu, Zhoujun Li 0001, Jun Wang 0019 |
ECIR | 2 |
| 2010 | Promoting Ranking Diversity for Biomedical Information Retrieval Using Wikipedia
Xiaoshi Yin, Jimmy Huang 0001, Zhoujun Li 0001 |
ECIR | 3 |
| 2010 | The topic-perspective model for social tagging systemsabstractIn this paper, we propose a new probabilistic generative model, called Topic-Perspective Model, for simulating the generation process of social annotations. Different from other generative models, in our model, the tag generation process is separated from the content term generation process. While content terms are only generated from resource topics, social tags are generated by resource topics and user perspectives together. The proposed probabilistic model can produce more useful information than any other models proposed before. The parameters learned from this model include: (1) the topical distribution of each document, (2) the perspective distribution of each user, (3) the word distribution of each topic, (4) the tag distribution of each topic, (5) the tag distribution of each user perspective, (6) and the probabilistic of each tag being generated from resource topics or user perspectives. Experimental results show that the proposed model has better generalization performance or tag prediction ability than other two models proposed in previous research. Caimei Lu, Xiaohua Hu 0001, Xin Chen 0041, Jung-ran Park, Tingting He 0003, Zhoujun Li 0001 |
KDD | 6 |
| 2010 | A survival modeling approach to biomedical search result diversification using wikipediaabstractIn this paper, we propose a probabilistic survival model derived from the survival analysis theory for measuring aspect novelty. The retrieved documents' query-relevance and novelty are combined at the aspect level for re-ranking. Experiments conducted on the TREC 2006 and 2007 Genomics collections demonstrate the effectiveness of the proposed approach in promoting ranking diversity for biomedical information retrieval. Xiaoshi Yin, Jimmy Huang 0001, Zhoujun Li 0001 |
SIGIR | 4 |
| 2010 | A Novel Composite Kernel for Finding Similar Questions in CQA Services
Jun Wang 0019, Zhoujun Li 0001, Xia Ben Hu, Biyun Hu |
WAIM | 2 |
| 2010 | Improving Diversity of Focused Summaries through the Negative Endorsements of Redundant FactsabstractWe present NegativeRank, a novel graph-based sentence ranking model to improve the diversity of focused summary by performing random walks over sentence graph with negative edge weights. Unlike the typical eigenvector centrality ranking, our method models the redundancy among sentence nodes as the negative edges. The negative edges can be thought of as the propagation of disapproval votes which can be used to penalize redundant sentences. As the iterative process continues, the initial ranking score of a given node will be adjusted according to a long-term negative endorsement from other sentence nodes. The evaluation results confirm that our proposed method is very effective in improving the diversity of the focused summary, compared to several well-known text summarization methods. Palakorn Achananuparp, Xiaohua Hu 0001, Lifan Guo, Tingting He 0003, Zhoujun Li 0001 |
Web Intelligence | 6 |
| 2009 | Online New Event Detection Based on IPLSA
Xiaoming Zhang 0001, Zhoujun Li 0001 |
ADMA | 2 |
| 2009 | Boosting Biomedical Information Retrieval Performance through Citation Graph: An Empirical Study
Xiaoshi Yin, Jimmy Huang 0001, Qinmin Hu, Zhoujun Li 0001 |
PAKDD | 4 |
| 2007 | An Effective Gene Selection Method Based on Relevance Analysis and Discernibility Matrix
Zhoujun Li 0001, Huowang Chen |
PAKDD | 2 |
| 2004 | Fast Mining Maximal Frequent ItemSets Based on FP-Tree
Yuejin Yan, Zhoujun Li 0001, Huowang Chen |
ER | 2 |