VLDB 2026 Research / reviewers in the wild / expert
Dong Zhou 0001
dblp:15/2101-1
· DBLP profile ↗
70ranked-venue papers
18as first author
37since 2021 · last 2025
0000-0002-3310-8347ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 11 first-author · 18 since 2021Databases, data management, data science and information retrieval · 19 · 8 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Driven Effective Knowledge Tracing by Integrating Dual-Channel Difficulty
Jiahui Cen, Jianghao Lin, Dong Zhou 0001, Weixuan Zhong, Aimin Yang 0002, Yongmei Zhou |
IEEE Big Data | 3 |
| 2025 | Central-Guided Convolutional Dual Attention for Document-Level Event Argument Extraction
Chengdong Lin, Jianghao Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002 |
IEEE Big Data | 3 |
| 2025 | Rethinking Vocabulary Augmentation: Addressing the Challenges of Low-Resource Languages in Multilingual ModelsabstractThe performance of multilingual language models (MLLMs) is notably inferior for low-resource languages (LRL) compared to high-resource ones, primarily due to the limited available corpus during the pre-training phase. This inadequacy stems from the under-representation of low-resource language words in the subword vocabularies of MLLMs, leading to their misidentification as unknown or incorrectly concatenated subwords. Previous approaches are based on frequency sorting to select words for augmenting vocabularies. However, these methods overlook the fundamental disparities between model representation distributions and frequency distributions. To address this gap, we introduce a novel Entropy-Consistency Word Selection (ECWS) method, which integrates semantic and frequency metrics for vocabulary augmentation. Our results indicate an improvement in performance, supporting our approach as a viable means to enrich vocabularies inadequately represented in current MLLMs. Nankai Lin, Peijian Zeng, Weixiong Zheng, Shengyi Jiang, Dong Zhou 0001, Aimin Yang 0002 |
COLING | 5 |
| 2025 | Global Graph Attention for Contrastive Sequential RecommendationabstractContrastive learning is a primary approach to mitigating data sparsity in sequential recommendation, but its improvements on the item side are limited due to data augmentation impairing user representations. To address these challenges, this paper introduces the Global Graph Attention for Contrastive Sequential Recommendation (GGACSR). GGACSR integrates graph neural networks and self-attention mechanisms, with an Attention Convolution Layer replacing nonlinear transformations in Graph Convolutional Networks (GCNs) with Q, K, and V vector operations, facilitating better handling of sequence dependencies. Leveraging a global user connection graph and projecting embeddings into a lower dimension effectively improves item and user representations. Overall, GGACSR outperforms existing baselines on three public datasets by more accurately capturing complex relationships and adapting user preferences. Zhijin Chen, Dong Zhou 0001, Jianghao Lin, Xingran Zhou, Aimin Yang 0002 |
CSCWD | 2 |
| 2025 | WATER: A Two-Stage in-Context Learning Debiasing Framework for Multilingual Text ClassificationabstractRecently, Large Language Models (LLMs) have shown remarkable success across a variety of tasks, with rapid advancements in supporting multilingual capabilities. However, these models exhibit varying degrees of demographic biases in text classification tasks. Most existing research focuses on debiasing pre-trained models or addressing biases in monolingual text classification, resulting in limited exploration in multilingual contexts. To solve the above problems, this paper introduces a tWo-stAge in-conText learning dEbiasing fRamework (WATER). Our approach does not require updating the model's parameters and is adaptable to any language. It includes three key modules: sample selection, sample filtering, and template filling and prediction. In the first stage, we leverage a sample selection module to identify text that closely matches the model embeddings. In the second stage, we introduce an innovative Contextual Disparity Measure (CDM) in the sample filtering module to filter out samples that effectively address the bias associated with specific attributes. Finally, the template filling and prediction module is used to fill the selected samples into the template and input them into the model to complete the multilingual text classification task. Our experimental results verify the effectiveness of our method in mitigating biases related to four sensitive attributes of gender, age, race, and country, demonstrating its potential to improve the fairness and accuracy of LLMs in multilingual classification tasks. Zeyong Long, Dong Zhou 0001, Zhijin Chen, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 3 |
| 2025 | FairTriplet: Balancing Fairness and Accuracy in Contextual Pre-Trained Models Through Prefix TuningabstractNatural language processing models learn powerful language representation abilities from vast amounts of data, but they also inherit societal biases embedded in that data. Current research on debiasing often struggles to balance the removal of model bias with the preservation of model performance. Most existing approaches depend on fine-tuning model parameters, which can introduce uncertainties in model performance due to the modifications made to these parameters. In this paper, we propose a novel debiasing framework called FairTriplet. First, this framework employs prefix tuning to freeze the parameters of the original pre-trained model. Then, it optimizes the prefix parameters through two debiasing terms. These two debiasing terms function by reducing the semantic distance between social groups (e.g., male and female) and increasing the semantic distance between social groups and neutral attributes (e.g., family and occupation) in the semantic space. This approach not only removes bias from the model but also preserves its performance. Experimental results demonstrate that FairTriplet achieves state-of-the-art (SOTA) levels in debiasing while maintaining model performance on GLUE downstream tasks. Zeyong Long, Weixiong Zheng, Dong Zhou 0001, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 3 |
| 2025 | DomainDiff: Unified Two-Stage Optimization for Text-Video RetrievalabstractThe primary challenge in text-video retrieval lies in achieving cross-modal semantic alignment, particularly the discrepancy between the conciseness of textual descriptions, which often fail to fully encapsulate the breadth of video content, and the redundancy in video data, which introduces noise and masks important semantic features. Current methods align text and video by mapping them into a shared feature space. Despite notable advancements, the inherent differences in modality-specific representations create a bottleneck for fixed-point embedding techniques, making models highly sensitive to dataset distribution and hindering their generalization ability. In this paper, we present DomainDiff, a framework that enhances the embedding space through a two-stage process. In the first stage, stochastic domain modeling, we semantically expand text embeddings to explore potential regions aligned with video content. Simultaneously, we filter video segments to reduce redundancy and highlight key frames. In the second stage, the dynamic agent attention diffusion network, we leverage the generative properties of diffusion models to optimize the embedding space by viewing it from a joint probability distribution perspective. An agent attention mechanism dynamically integrates text and video features, ensuring accurate cross-modal alignment. Experimental results demonstrate that DomainDiff significantly improves retrieval performance across five benchmark datasets, with R@1 improvements ranging from 3% to 7.4%. Moreover, DomainDiff outperforms existing methods in handling long videos and complex textual descriptions, showcasing superior semantic robustness and generalization across varying distributions. Chenxu Wang 0019, Dong Zhou 0001, Jianghao Lin, Yongmei Zhou, Aimin Yang 0002 |
ICMR | 2 |
| 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule RetrievalabstractMolecular retrieval is critical in drug discovery and molecular design. Traditional discriminative methods often model the conditional probability distribution of retrieving candidates, treating the query text as a deterministic input. However, these approaches have notable limitations: (1) They often overlook the statistical properties of the original data distributions of queries and candidates, preventing the recognition of out-of-distribution data. (2) They struggle to balance retrieval accuracy and diversity when processing open-ended semantic queries. To address these challenges, we introduce DiffTMR, a novel framework that reformulates text-molecule retrieval as a reverse denoising process, progressively generating the joint distribution of candidates and queries from noises. DiffTMR uniquely integrates hierarchical diffusion alignment with dynamic perturbation embedding mechanisms. By employing text-anchored perturbations, it enhances the diversity of molecular representations, and through global-local progressive denoising, it achieves cross-modal hierarchical alignment. This leads to significant improvements in retrieval accuracy and out-of-domain generalization. Evaluations on benchmark datasets ChEBI-20 and PCdes demonstrate that DiffTMR surpasses current leading baselines by 4.2%-5.4% in Hits@1 metrics and exhibits superior performance in out-of-domain retrieval tasks. Chenxu Wang 0019, Dong Zhou 0001, Jianghao Lin, Yongmei Zhou, Aimin Yang 0002 |
ACM Multimedia | 2 |
| 2025 | Filter-enhanced Contrast Variational AutoEncoders for sequential recommendationabstractAbstract Data augmentation-based contrastive learning has been successfully employed in Variational AutoEncoders sequence recommendation systems to tackle the issue of data sparsity. Nevertheless, this strategy is generally less advantageous for tail users. The prospective transmission of information from head-to-tail users to alleviate long-tail impact is encouraging. However, data augmentation distorts the original sequence and embeds stochastic noise into latent variables, impeding the decoder’s capacity to accurately identify the user’s true preferences. In addition, contrastive learning seeks to achieve consistency in the latent variables of both the original and augmented data. However, the presence of noise in the augmented data might hamper the encoding of latent variables from the original data, especially impacting head users. In order to address these challenges, this work introduces a new sequence recommendation model called the Filter-enhanced Contrastive Variational Autoencoder (FeCVAE). It employs Fourier filters and adversarial attack training to minimize the impact of stochastic noise, thereby improving the quality of latent variables and facilitating more accurate decoder outputs. Moreover, a user enhancer is introduced to leverage knowledge from head users to empower tail users, thereby alleviating the long-tail effect. The efficacy of FeCVAE is demonstrated through comprehensive experiments across four benchmark datasets. Zhijin Chen, Nankai Lin, Aimin Yang 0002, Dong Zhou 0001 |
Comput. J. | 4 |
| 2025 | A novel curriculum learning framework for multi-label emotion classificationabstractAbstract Curriculum learning (CL) is a training strategy that imitates how humans learn, by gradually introducing more complex samples and information to the model. However, in multi-label emotion classification (MEC) tasks, using a traditional CL approach can result in overfitting on easy samples and lead to biased training. Additionally, the sample difficulty varies as the model trains. To address these challenges, we propose a novel CL framework for MEC tasks called CLF-MEC. Unlike traditional approaches that assess difficulty at the sample level, we utilize category-level assessment to determine the difficulty level of samples. As the model identifies a category well, the score for that category’s samples is reduced, ensuring dynamic changes in the sample difficulty are accounted for. Our CL framework employs two training modes, namely “learning” and “tackling.” These two processes are trained alternatively to imitate the “learning-tackling” process in human learning. This ensures that samples from hard-to-learn categories receive more attention. During the “tackling” process, our method transforms the task of dealing with hard samples into an “easy” learning task by utilizing contrastive learning to enhance the semantic representation of those hard samples. Experimental results demonstrate that our CLF-MEC framework has achieved significant improvements in MEC. Nankai Lin, Peijian Zeng, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Comput. J. | 5 |
| 2025 | GS2F: Multimodal Fake News Detection Utilizing Graph Structure and Guided Semantic FusionabstractThe prevalence of fake news online has become a significant societal concern. To combat this, multimodal detection techniques based on images and text have shown promise. Yet, these methods struggle to analyze complex relationships within and between modalities due to the diverse discriminative elements in the news content. In addition, research on multimodal and multi-class fake news detection remains insufficient. To address the above challenges, in this article, we propose a novel detection model, GS 2 F, leveraging g raph s tructure and g uided s emantic f usion. Specifically, we construct a multimodal graph structure to align two modalities and employ graph contrastive learning for refined fusion representations. Furthermore, a guided semantic fusion module is introduced to maximize the utilization of single-modal information and a dynamic contribution assignment layer is designed to weigh the importance of image, text, and multimodal features. Experimental results on Fakeddit demonstrate that our model outperforms existing methods, marking a step forward in the multimodal and multi-class fake news detection. Dong Zhou 0001, Qiang Ouyang, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2025 | Ensuring accuracy and fairness: a de-biasing framework for sequential recommendation
Qifeng Bai, Nankai Lin, Meiyu Zeng, Guanqiu Qin, Dong Zhou 0001, Aimin Yang 0002 |
User Model. User Adapt. Interact. | 5 |
| 2025 | A contrastive news recommendation framework based on curriculum learning
Xingran Zhou, Nankai Lin, Weixiong Zheng, Dong Zhou 0001, Aimin Yang 0002 |
User Model. User Adapt. Interact. | 4 |
| 2024 | A Retrieval-Augmented Contrastive Framework for Legal Case Retrieval Based on Event Information
Changyong Fan, Nankai Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002 |
ACML | 3 |
| 2024 | MLCL: A Framework for Reducing Language Imbalance in Sino-Tibetan Languages through Adapter Structures
Jiajun Fang, Aimin Yang 0002, Dong Zhou 0001, Nankai Lin |
ACML | 4 |
| 2024 | Enhancing Aspect Sentiment Quad Prediction through Dual-Sequence Data Augmentation and Contrastive Learning
Nankai Lin, Pinmo Wu, Dong Zhou 0001, Aimin Yang 0002 |
ACML | 4 |
| 2024 | HiRAG: A Historical Information-Driven Retrieval-Augmented Generation Framework for Background Summarization
Dong Zhou 0001, Binli Zeng, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACML | 1 |
| 2024 | ADSE: Adversarial Debiasing Framework Based on Sinusoidal Embedding for Sequential RecommendationabstractSequential recommendation plays a key role in recommender systems, where the goal is to predict a user’s future points of interest by analyzing his or her historical interactions. This process not only requires the system to be able to accurately identify and recommend items that are likely to be of interest to the user but also ensures that all items receive equal exposure to prevent over-concentration or marginalization of items due to algorithmic bias. To address these challenges, in this paper, we propose a novel Adversarial Debiasing framework based on Sinusoidal Embedding for sequential recommendation, ADSE. This framework employs sinusoidal position embeddings to extract positional information between sequences more precisely and utilizes a dropout strategy to optimize the handling of cold-start sequences, aiming to resolve the cold-start issue while maintaining the semantics of the original sequences. Additionally, adversarial training was incorporated to reduce implicit bias due to assuming interactions in the calculation of exposure. Qifeng Bai, Nankai Lin, Junheng He, Zhijin Chen, Dong Zhou 0001, Aimin Yang 0002 |
ICWS | 5 |
| 2024 | Knowledge distillation representation and DCNMIX quality prediction-based Web service recommendationabstractSummary Web service recommendation as an emerging topic attracts increasing attention due to its important practical significance. As the number of available Web services continues to grow, users face the challenge of searching the most suitable services that meet their specific needs. Quality of service (QoS)‐based service recommendation becomes a popular approach to address this issue. However, existing QoS‐based service recommendation methods are inability to effectively capture valuable content and structural information from services. These methods often rely solely on low‐order explicit feature intersections in QoS information, do not fully utilize the high‐order implicit feature intersections, and ignore the rich semantic information existing in service descriptions and user preferences. To address this problem, this paper proposes a Web service recommendation method via combining knowledge distillation representation and DCNMIX quality prediction. This method combines content‐based and structure‐based service classification and service prediction based on multi‐dimensional service quality information. First, it builds a service relationship network using semantic features extracted from service descriptions. Second, it designs a graph neural network knowledge distillation framework. The teacher model extracts the knowledge of the graph neural network model, and the student model learns the structure‐based and feature‐based prior knowledge of the service relationship network. Then the student model is used to learn the knowledge of the teacher model, classify Web services, and obtain service representations. Finally, based on service representations and multi‐dimensional QoS information, it exploits the DCNMIX model to learn the explicit and implicit features intersections of Web services and obtain the prediction score and ranking of Web services. The experimental results on the ProgrammableWeb dataset show that the proposed method outperforms the state‐of‐the‐art baselines in terms of Recall, F1, Logloss, and AUC_ROC. Buqing Cao, Shanpeng Liu, Yiping Wen, Dong Zhou 0001, Mingdong Tang |
Concurr. Comput. Pract. Exp. | 6 |
| 2024 | Towards fair decision: A novel representation method for debiasing pre-trained models
Junheng He, Nankai Lin, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Decis. Support Syst. | 5 |
| 2024 | Addressing class-imbalance challenges in cross-lingual aspect-based sentiment analysis: Dynamic weighted loss and anti-decoupling
Nankai Lin, Meiyu Zeng, Xingming Liao, Weizhong Liu, Aimin Yang 0002, Dong Zhou 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Global information enhancement and subgraph-level weakly contrastive learning for lightweight weakly supervised document-level event extraction
Guanqiu Qin, Nankai Lin, Menglan Shen, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Expert Syst. Appl. | 5 |
| 2024 | BiLSTM4DPS: An attention-based BiLSTM approach for detecting phishing scams in ethereum
Mingdong Tang, Mingshun Ye, Weili Chen, Dong Zhou 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Efficient low-rank multi-component fusion with component-specific factors in image-recipe retrieval
Dong Zhou 0001, Buqing Cao, Kai Zhang 0074, Jinjun Chen |
Multim. Tools Appl. | 2 |
| 2024 | Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics RetrievalabstractThe task of retrieving audio content relevant to lyric queries and vice versa plays a critical role in music-oriented applications. In this process, robust feature representations have to be learned for two modalities. Furthermore, interactions between different modalities should be properly captured at a fine-grained level. Existing approaches can effectively extract modal representations and perform retrieving between different modalities through alignment. However, these approaches model interactions between audio and lyrics in a coarse-grained manner. Especially the input features and interactions between enhanced representations produced by the alignment module are largely ignored, resulting in low-quality modality representations for final retrieval. This paper presents a novel method named CMRF that accomplishes cross-modal interactions via a reinforcement feedback procedure to learn high-quality multi-modal embeddings. Initially, we implicitly assimilate representations across distinct modalities via directional pairwise cross-modal attention. Subsequently, our approach recurrently identifies pivotal constituents within these elevated-level attributes to engage with the primary input features via reinforcement learning, thus augmenting the quality of multi-modal embeddings. In addition, we introduce a novel audio-lyrics datasetAL-song, which consists of paired audio with corresponding lyrics for the audio-lyrics retrieval task. The empirical findings derived from theAL-songdataset and the benchmark datasetSounddescssubstantiate the efficacy and efficiency of CMRF when juxtaposed with state-of-the-art methodologies. Dong Zhou 0001, Fang Lei, Lin Li 0001, Yongmei Zhou, Aimin Yang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | An Edge Server Co-deployment Method via Joint Optimization of Communication Delay and Load Balancing in Edge Collaborative EnvironmentabstractIn the edge collaboration environment, the rapid development of smart terminal devices and the huge volume of service requests from smart terminal devices often lead to load imbalance and long communication delay of edge server. In this situation, it's becoming more and more important to efficiently deploy the edge servers to reduce service communication delay and balance the load of edge server, thus fully improving resource utilization of edge server and users' service experience. To this end, this paper proposes a co-deployment method for edge servers in edge collaboration environment by jointly optimizing communication delay and load balancing. It first clusters all communication base stations by the K-means algorithm to derive the most suitable area for edge server deployment. Then, on the premise of balancing the workload between edge servers and minimizing the communication delay between communication base stations and edge servers, the optimal deployment location of edge servers is solved iteratively by the particle swarm algorithm. Finally, validation experiments are conducted based on real data sets of Shanghai Telecom communication base stations, and the experimental results show that the overall performance of the proposed method is better compared with the typical methods such as Top-K, K-means, Random, and Genetic Algorithm. Zilong Zeng, Buqing Cao, Hongfan Ye, Dong Zhou 0001, Mingdong Tang |
CSCWD | 4 |
| 2023 | Simplifying Aspect-Sentiment Quadruple Prediction with Cartesian Product Operation
Jigang Wang, Aimin Yang 0002, Dong Zhou 0001, Nankai Lin, Weifeng Huang |
ICIC (4) | 3 |
| 2023 | Exploring latent weight factors and global information for food-oriented cross-modal retrievalabstractFood-oriented cross-modal retrieval aims to retrieve relevant recipes given food images or vice versa.The modality semantic gap between recipes and food images (text and image modalities) is the main challenge.Though several studies are introduced to bridge this gap, they still suffer from two major limitations: 1) The simple embedding concatenation only can capture the simple interactions rather than complex interactions between different recipe components.2) The image feature extraction based on convolutional neural networks only considers the local features and ignores the global features of an image, as well as the interactions between different extracted features.This paper proposes a novel method based on Latent Component Weight Factors and Global Information (LCWF-GI) to learn the robust recipe and image representations for food-oriented cross-modal retrieval.This proposed method integrates the textual embeddings of different recipe components into a compact embedding to represent the recipes with the latent component-specific weight factors.A transformer encoder is utilised to capture the intra-modality interactions and the importance of different extracted image features for enhanced image representations.Finally, the bi-directional triplet loss is further used to perform retrieval learning.Experimental results on the Recipe 1M dataset show that our LCWF-GI method achieves competent improvements. Dong Zhou 0001, Buqing Cao, Wei Liang 0005, Nitin Sukhija |
Connect. Sci. | 2 |
| 2023 | CL-XABSA: Contrastive Learning for Cross-Lingual Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA), an extensively researched area in the field of natural language processing (NLP), predicts the sentiment expressed in a text relative to the corresponding aspect. Unfortunately, most languages lack sufficient annotation resources; thus, an increasing number of recent researchers have focused on cross-lingual aspect-based sentiment analysis (XABSA). However, most recent studies focus only on cross-lingual data alignment instead of model alignment. Therefore, we propose a novel framework, CL-XABSA: contrastive learning for cross-lingual aspect-based sentiment analysis. Based on contrastive learning, we close the distance between samples with the same label in different semantic spaces, achieving convergence of semantic spaces of different languages. Specifically, we design two contrastive objectives, token-level contrastive learning of token embeddings (TL-CTE) and sentiment-level contrastive learning of token embeddings (SL-CTE), to unify the semantic space of source and target languages. Since CL-XABSA can receive datasets in multiple languages during training, it can be further extended to multilingual aspect-based sentiment analysis (MABSA). To further improve the model performance, we perform knowledge distillation with target-language unlabeled data. In the distillation XABSA task, we further explore the effectiveness of different data (source dataset, translated dataset, and code-switched dataset). The results demonstrate that the proposed method has a certain improvement in the three XABSA tasks, distillation XABSA and MABSA. The source code of this paper is publicly available athttps://github.com/GKLMIP/CL-XABSA. Nankai Lin, Yingwen Fu, Xiaotian Lin, Dong Zhou 0001, Aimin Yang 0002, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Towards the Quantitative Interpretability Analysis of Citizens Happiness PredictionabstractEvaluating the high-effect factors of citizens' happiness is beneficial to a wide range of policy-making for economics and politics in most countries. Benefiting from the high-efficiency of regression models, previous efforts by sociology scholars have analyzed the effect of happiness factors with high interpretability. However, restricted to their research concerns, they are specifically interested in some subset of factors modeled as linear functions. Recently, deep learning shows promising prediction accuracy while addressing challenges in interpretability. To this end, we introduce Shapley value that is inherent in solid theory for factor contribution interpretability to work with deep learning models by taking into account interactions between multiple factors. The proposed solution computes the Shapley value of a factor, i.e., its average contribution to the prediction in different coalitions based on coalitional game theory. Aiming to evaluate the interpretability quality of our solution, experiments are conducted on a Chinese General Social Survey (CGSS) questionnaire dataset. Through systematic reviews, the experimental results of Shapley value are highly consistent with academic studies in social science, which implies our solution for citizens' happiness prediction has 2-fold implications, theoretically and practically. Lin Li 0001, Miao Kong, Dong Zhou 0001, Xiaohui Tao 0001 |
IJCAI | 4 |
| 2022 | Cross-lingual embeddings with auxiliary topic models
Dong Zhou 0001, Xiaoya Peng, Lin Li 0001, Jun-mei Han |
Expert Syst. Appl. | 1 |
| 2022 | Neural topic-enhanced cross-lingual word embeddings for CLIR
Dong Zhou 0001, Lin Li 0001, Mingdong Tang, Aimin Yang 0002 |
Inf. Sci. | 1 |
| 2021 | Utilizing Local Tangent Information for Word Re-embedding
Dong Zhou 0001, Lin Li 0001, Jinjun Chen |
ECIR (1) | 2 |
| 2021 | Multi-modal and Multi-perspective Machine Translation by Collecting Diverse Alignments
Lin Li 0001, Turghun Tayir, Kaixi Hu, Dong Zhou 0001 |
PRICAI (2) | 4 |
| 2021 | A Topic-Sensitive Method for Mashup Tag Recommendation Utilizing Multi-Relational Service DataabstractTagging systems have been widely used as a major way of managing Web service resources. Many portals such as ProgrammableWeb and BioCatalogue allow users to create manual tags annotating Web services and their compositions (e.g., mashups). This is extremely helpful for managing and retrieving enormous Web service data. In the past few years, many tag recommendation approaches have been proposed for Web services that contain few or no tags. Most of them only exploit the textual content or tag service matrix information. Sometimes those approaches suffer from the data sparsity problem, especially when Web services have only few tags or their auxiliary textual contents are hard to be obtained. In real world, a plenty of relationships are available in recommendation systems, e.g., the composition relationship between services and the annotation relationship between mashups and tags. These multi-relational data can be utilized as additional features to improve the recommendation performance. In this paper, we exploit various types of relationships as features and propose a novel topic-sensitive approach based on the Factorization Machines for mashup tag recommendation. Factorization Machines is utilized to model the pair-wise interactions between all features and predict adequate tags for mashups. In this approach, we first obtain the latent topics of all tags as well as the description documents for mashups and APIs based on a novel probabilistic topic model. Then, a multi-relational network by mining various relationships from the Web service data is constructed. Various auxiliary informations are subsequently extracted from the network to train the Factorization Machines. The proposed model is evaluated on three real-world datasets and the experimental results show that it outperforms several state-of-the-art methods. Min Shi 0001, Jianxun Liu 0001, Dong Zhou 0001, Yufei Tang |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Dual-Level Attention Based on a Heterogeneous Graph Convolution Network for Aspect-Based Sentiment ClassificationabstractWith the development of 5G, the advancement of basic infrastructure has led to considerable development in related research and technology. It also promotes the development of various smart devices and social platforms. More and more people are now using smart devices to post their reviews right after something happens. In order to keep pace with this trend, we propose a method to analyze users’ sentiment by using their text data. When analyzing users’ text data, it is noted that a user’s review may contain many aspects. Traditional text classification methods used by smart devices, however, usually ignore the importance of multiple aspects of a review. Additionally, most algorithms usually ignore the network structure information between the words in a sentence and the sentence itself. To address these issues, we propose a novel dual‐level attention‐based heterogeneous graph convolutional network for aspect‐based sentiment classification which minds more context information through information propagation along with graphs. Particularly, we first propose a flexible HIN (heterogeneous information network) framework to model the user‐generated reviews. This framework can integrate various types of additional information and capture their relationships to alleviate semantic sparsity of some labeled data. This framework can also leverage the full advantage of the hidden network structure information through information propagation along with graphs. Then, we propose a dual‐level attention‐based heterogeneous graph convolutional network (DAHGCN), which includes node‐level and type‐level attentions. The attention mechanisms can analyze the importance of different adjacent nodes and the importance of different types of nodes for the current node. The experimental results on three real‐world datasets demonstrated the effectiveness and reliability of our model. Lei Jiang 0007, Jianxun Liu 0001, Dong Zhou 0001, Yang Gao 0039 |
Wirel. Commun. Mob. Comput. | 4 |
| 2021 | A multi-granularity semantic space learning approach for cross-lingual open domain question answering
Lin Li 0001, Miao Kong, Dong Zhou 0001 |
World Wide Web | 4 |
| 2020 | An Attentive Deep Supervision based Semantic Matching Framework For Tag Recommendation in Software Information SitesabstractTag recommendation in software information sites is a popular way to help developers classify software objects. Existing methods mostly consider tag recommendation as a multi-label classification task, which does not adequately leverage the semantic information of tags themselves. It is observed that the information granularity of tags is from abstract to specific and deep learning models have proven capable of automatically learning the features in different layers of an integrated network with different abstraction degrees. In this paper, we propose TagMatchRec, a deep semantic matching framework for tag recommendation instead of being based on classification. In our framework, multiple layers with different information granularities are directly connected to the output layer aiming at improving the quality of tag recommendation. Moreover, because the abstraction levels of semantic features learned by each layer may be different given different software objects and tags, an attentive deep supervision is introduced so that the dense connections from early layers to the output layer have directly weighted impact on loss function optimization. Comprehensive evaluations are conducted the datasets from four software information sites. The experimental results show that TagMatchRec has achieved better performance compared with the state-of-the-art approaches. Xinhao Zheng, Lin Li 0001, Dong Zhou 0001 |
APSEC | 3 |
| 2020 | Manifold Learning-based Word Representation Refinement Incorporating Global and Local InformationabstractRecent studies show that word embedding models often underestimate similarities between similar words and overestimate similarities between distant words.This results in word similarity results obtained from embedding models inconsistent with human judgment.Manifold learningbased methods are widely utilized to refine word representations by re-embedding word vectors from the original embedding space to a new refined semantic space.These methods mainly focus on preserving local geometry information through performing weighted locally linear combination between words and their neighbors twice.However, these reconstruction weights are easily influenced by different selections of neighboring words and the whole combination process is time-consuming.In this paper, we propose two novel word representation refinement methods leveraging isometry feature mapping and local tangent space respectively.Unlike previous methods, our first method corrects pre-trained word embeddings by preserving global geometry information of all words instead of local geometry information between words and their neighbors.Our second method refines word representations by aligning original and refined embedding spaces based on local tangent space instead of performing weighted locally linear combination twice.Experimental results obtained from standard semantic relatedness and semantic similarity tasks show that our methods outperform various state-of-the-art baselines for word representation refinement. Dong Zhou 0001, Lin Li 0001, Jinjun Chen |
COLING | 2 |
| 2020 | Multitask Learning Based on Constrained Hierarchical Attention Network for Multi-aspect Sentiment Classification
Yang Gao 0039, Jianxun Liu 0001, Dong Zhou 0001 |
ICONIP (4) | 4 |
| 2020 | Sentence-based and Noise-robust Cross-modal Retrieval on Cooking Recipes and Food ImagesabstractIn recent years, people are facing with billions of food images, videos and recipes on social medias. An appropriate technology is highly desired to retrieve accurate contents across food images and cooking recipes, like cross-modal retrieval framework. Based on our observations, the order of sequential sentences in recipes and the noises in food images will affect retrieval results. We take into account the sentence-level sequential orders of instructions and ingredients in recipes, and noise portion in food images to propose a new framework for cross-retrieval. In our framework, we propose three new strategies to improve the retrieval accuracy. (1) We encode recipe titles, ingredients, instructions in sentence level, and adopt three attention networks on multi-layer hidden state features separately to capture more semantic information. (2) We apply attention mechanism to select effective features from food images incorporating with recipe embeddings, and adopt an adversarial learning strategy to enhance modality alignment. (3) We design a new triplet loss scheme with an effective sampling strategy to reduce the noise impact on retrieval results. The experimental results show that our framework clearly outperforms the state-of-art methods in terms of median rank and recall rate at top k on the Recipe 1M dataset. Zichen Zan, Lin Li 0001, Jianquan Liu, Dong Zhou 0001 |
ICMR | 4 |
| 2020 | A Cross-Layer Connection Based Approach for Cross-Lingual Open Question Answering
Lin Li 0001, Miao Kong, Dong Zhou 0001 |
NLPCC (1) | 4 |
| 2020 | Predicting the Evolution of Hot Topics: A Solution Based on the Online Opinion Dynamics Model in Social NetworkabstractPredicting and utilizing the evolution trend of hot topics is critical for contingency management and decision-making purposes of government bodies and enterprises. This paper proposes a model named online opinion dynamics (OODs) where any node in a social network has its unique confidence threshold and influence radius. The nodes in the OOD are mainly affected by their neighbors and are also randomly influenced by unfamiliar nodes. In the traditional opinion model, however, each node is affected by all other nodes, including its friends. Furthermore, many traditional opinion evolution approaches are reviewed to see if all nodes (participants) can eventually reach a consensus. On the contrary, OOD is more focused on such details as concluding the overall trend of events and evaluating the support level of each participant through numerical simulation. Experiments show that OOD is superior to the improvement of the original Hegselmann-Krause (HK) model, HK-13 and HK-17, with respect to qualitative predictions of the evolution trend of an event. The quantitative predictions of the HK model cannot be used to make decisions, whereas the results of the OOD model are proved to be acceptable. Lei Jiang 0007, Jujun Liu, Dong Zhou 0001, Xiansheng Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | An iterative method for personalized results adaptation in cross-language search
Dong Zhou 0001, Séamus Lawless, Jianxun Liu 0001 |
Inf. Sci. | 1 |
| 2017 | WE-LDA: A Word Embeddings Augmented LDA Model for Web Services ClusteringabstractDue to the rapid growth in both the number and diversity of Web services on the web, it becomes increasingly difficult for us to find the desired and appropriate Web services nowadays. Clustering Web services according to their functionalities becomes an efficient way to facilitate the Web services discovery as well as the services management. Existing methods for Web services clustering mostly focus on utilizing directly key features from WSDL documents, e.g., input/output parameters and keywords from description text. Probabilistic topic model Latent Dirichlet Allocation (LDA) is also adopted, which extracts latent topic features of WSDL documents to represent Web services, to improve the accuracy of Web services clustering. However, the power of the basic LDA model for clustering is limited to some extent. Some auxiliary features can be exploited to enhance the ability of LDA. Since the word vectors obtained by Word2vec is with higher quality than those obtained by LDA model, we propose, in this paper, an augmented LDA model (named WE-LDA) which leverages the high-quality word vectors to improve the performance of Web services clustering. In WE-LDA, the word vectors obtained by Word2vec are clustered into word clusters by K-means++ algorithm and these word clusters are incorporated to semi-supervise the LDA training process, which can elicit better distributed representations of Web services. A comprehensive experiment is conducted to validate the performance of the proposed method based on a ground truth dataset crawled from ProgrammableWeb. Compared with the state-of-the-art, our approach has an average improvement of 5.3% of the clustering accuracy with various metrics. Min Shi 0001, Jianxun Liu 0001, Dong Zhou 0001, Mingdong Tang, Buqing Cao |
ICWS | 3 |
| 2017 | User Expertise Inference on Twitter: Learning from Multiple Types of User DataabstractThis paper proposes a learning model that tries to infer a user's topical expertise using multiple types of user-related data from Twitter such as tweets posted by the user and the characteristics of their followers. It considers inference consistency of different types of user data in the process of learning and aims to deliver accurate and effective inference results, even in cases where some types of data are missing for a user, e.g. the user has yet to post any tweets. Experiments conducted on a large-scale Twitter dataset show that our model outperforms several baseline approaches which use only a single type of user data for inference. Dong Zhou 0001, Séamus Lawless |
UMAP | 2 |
| 2017 | Inferring your expertise from Twitter: combining multiple types of user activityabstractUnderstanding the expertise of users in social networking sites like Twitter is a key component for many applications such as user recommendation and talent seeking. A range of interactions between users on Twitter can provide important information that implicitly reflects a user's expertise. This paper proposes a learning model that tries to infer a user's topical expertise from Twitter using information such as tweets posted by the user and the characteristics of their followers. The model takes various types of user-related data from Twitter as input and considers their inference consistency in the process of learning. It aims to deliver accurate and effective inference results, even in cases where some types of data are missing for a user, e.g. the user has yet to post any tweets. The experiments reported in the paper were conducted on a large-scale Twitter dataset. Experimental results show that our model outperforms several baseline approaches and outperforms approaches which use only a single type of user data for inference. Dong Zhou 0001, Séamus Lawless |
WI | 2 |
| 2017 | A Hybrid Approach for Automatic Mashup Tag Recommendation
Min Shi 0001, Jianxun Liu 0001, Dong Zhou 0001 |
J. Web Eng. | 3 |
| 2017 | Query Expansion with Enriched User Profiles for Personalized Search Utilizing Folksonomy DataabstractQuery expansion has been widely adopted in Web search as a way of tackling the ambiguity of queries. Personalized search utilizing folksonomy data has demonstrated an extreme vocabulary mismatch problem that requires even more effective query expansion methods. Co-occurrence statistics, tag-tag relationships, and semantic matching approaches are among those favored by previous research. However, user profiles which only contain a user's past annotation information may not be enough to support the selection of expansion terms, especially for users with limited previous activity with the system. We propose a novel model to construct enriched user profiles with the help of an external corpus for personalized query expansion. Our model integrates the current state-of-the-art text representation learning framework, known as word embeddings, with topic models in two groups of pseudo-aligned documents. Based on user profiles, we build two novel query expansion techniques. These two techniques are based on topical weights-enhanced word embeddings, and the topical relevance between the query and the terms inside a user profile, respectively. The results of an in-depth experimental evaluation, performed on two real-world datasets using different external corpora, show that our approach outperforms traditional techniques, including existing non-personalized and personalized query expansion methods. Dong Zhou 0001, Séamus Lawless, Jianxun Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Multi-relation Based Manifold Ranking Algorithm for API Recommendation
Fenfang Xie, Jianxun Liu 0001, Mingdong Tang, Dong Zhou 0001, Buqing Cao, Min Shi 0001 |
APSCC | 4 |
| 2016 | Exploring External Knowledge Base for Personalized Search in Collaborative Tagging Systems
Dong Zhou 0001, Séamus Lawless, Jianxun Liu 0001 |
CollaborateCom | 1 |
| 2016 | Do Your Social Profiles Reveal What Languages You Speak? Language Inference from Social Media Profiles
M. Rami Ghorab, Dong Zhou 0001, Séamus Lawless |
ECIR | 4 |
| 2016 | Enhanced Personalized Search using Social DataabstractSearch personalization that considers the social dimension of the web has attracted a significant volume of research in recent years. A user profile is usually needed to represent a user?s interests in order to tailor future searches. Previous research has typically constructed a profile solely from a user?s usage information. When the user has only limited activities in the system, the effect of the user profile on search is also constrained. This research addresses the setting where a user has only a limited amount of usage information. We build enhanced user profiles from a set of annotations and resources that users have marked, together with an external knowledge base constructed according to usage histories. We present two probabilistic latent topic models to simultaneously incorporate social annotations, documents and the external knowledge base. Our web search strategy is achieved using personalized social query expansion. We introduce a topical query expansion model to enhance the search by utilizing individual user profiles. The proposed approaches have been intensively evaluated on a large public social annotation dataset. Results show that our models significantly outperformed existing personalized query expansion methods which use user profiles solely built from past usage information in personalized search. Dong Zhou 0001, Séamus Lawless, Jianxun Liu 0001 |
EMNLP | 1 |
| 2016 | A Probabilistic Topic Model for Mashup Tag RecommendationabstractMashups are prevalent Service-Oriented Architecture (SOA) based applications consisting of multiple Web Application Programming Interfaces (APIs) and content. Tags have been extensively used to organize and index mashup services. However, people favor manual tags creation in the past. This approach demands user intervention, which is extremely time-consuming and probes to errors. In this paper we propose a novel Mashup-API-Tag model for automatic mashup tag recommendation. The model simultaneously incorporates the composition relationships between mashups and APIs as well as the annotation relationships between APIs and tags to discover the latent topics. Then the semantic similarity between Web APIs and mashups can be acquired. Subsequently, tags of chosen APIs are recommended to a mashup where the mashup and the APIs are most similar. In addition, we develop a tag filtering algorithm to select the most relevant tags for recommendation. The experimental results on a real world dataset prove that our approach outperforms other methods, including frequency-based methods and the methods that only consider the composition relationships and the annotation relationships separately. Min Shi 0001, Jianxun Liu 0001, Dong Zhou 0001, Mingdong Tang, Fenfang Xie |
ICWS | 3 |
| 2016 | Inferring Your Expertise from Twitter: Integrating Sentiment and Topic RelatednessabstractThe ability to understand the expertise of users in Social Networking Sites (SNSs) is a key component for delivering effective information services such as talent seeking and user recommendation. However, users are often unwilling to make the effort to explicitly provide this information, so existing methods aimed at user expertise discovery in SNSs primarily rely on implicit inference. This work aims to infer a user's expertise based on their posts on the popular micro-blogging site Twitter. The work proposes a sentiment-weighted and topic relation-regularized learning model to address this problem. It first uses the sentiment intensity of a tweet to evaluate its importance in inferring a user's expertise. The intuition is that if a person can forcefully and subjectively express their opinion on a topic, it is more likely that the person has strong knowledge of that topic. Secondly, the relatedness between expertise topics is exploited to model the inference problem. The experiments reported in this paper were conducted on a large-scale dataset with over 10,000 Twitter users and 149 expertise topics. The results demonstrate the success of our proposed approach in user expertise inference and show that the proposed approach outperforms several alternative methods. Dong Zhou 0001, Séamus Lawless |
WI | 2 |
| 2015 | A Hybrid Genetic Algorithm for Privacy and Cost Aware Scheduling of Data Intensive Workflow in Cloud
Congyang Chen, Jianxun Liu 0001, Yiping Wen, Jinjun Chen, Dong Zhou 0001 |
ICA3PP (1) | 5 |
| 2014 | Iterative Refinement Methods for Enhanced Information RetrievalabstractInformation retrieval (IR) systems exploit relevant information when tailoring search results to individual information needs. However, current search experience becomes poor without considering similar queries entered by previous searchers. In the following paper, we discuss a solution to this problem, which combines collaborative filtering algorithms with traditional IR models to enable EIR. We also present various iterative refinement methods for improving the raw performance of this system. We validate our theories in an experiment using queries extracted from the click-through log of a commercial search engine. According to our results, an IR system employing iteratively refined, collaborative retrieval significantly outperforms various baseline retrieval models. Dong Zhou 0001, Mark Truran, Jianxun Liu 0001, Wei Li 0054, Gareth J. F. Jones |
Int. J. Intell. Syst. | 1 |
| 2014 | Using multiple query representations in patent prior-art search
Dong Zhou 0001, Mark Truran, Jianxun Liu 0001, Sanrong Zhang |
Inf. Retr. | 1 |
| 2013 | Query Generation Techniques for Patent Prior-Art Search in Multiple Languages
Dong Zhou 0001, Jianxun Liu 0001, Sanrong Zhang |
NLPCC | 1 |
| 2013 | Multilingual vs. Monolingual User Models for Personalized Multilingual Information Retrieval
M. Rami Ghorab, Séamus Lawless, Alexander O'Connor, Dong Zhou 0001, Vincent P. Wade |
UMAP | 4 |
| 2013 | Collaborative pseudo-relevance feedback
Dong Zhou 0001, Mark Truran, Jianxun Liu 0001, Sanrong Zhang |
Expert Syst. Appl. | 1 |
| 2013 | Personalised Information Retrieval: survey and classification
M. Rami Ghorab, Dong Zhou 0001, Alexander O'Connor, Vincent P. Wade |
User Model. User Adapt. Interact. | 2 |
| 2012 | A section title authoring tool for clinical guidelinesabstractProfessional users of medical information often report difficulties when attempting to locate specific information in lengthy documents. Sometimes these difficulties can be attributed to poorly specified section titles which fail to advertise relevant content. In this paper we describe preliminary work on a software plug-in for a document engineering environment that will assist authors when they formulate section-level headings. We describe two different algorithms which can be used to generate section titles. We compare the performance of these algorithms and correlate our experimental results with an evaluation of title quality performed by domain experts. Mark Truran, Gersende Georg, Marc Cavazza, Dong Zhou 0001 |
ACM Symposium on Document Engineering | 4 |
| 2012 | Web Search Personalization Using Social Data
Dong Zhou 0001, Séamus Lawless, Vincent P. Wade |
TPDL | 1 |
| 2012 | Improving search via personalized query expansion using social media
Dong Zhou 0001, Séamus Lawless, Vincent P. Wade |
Inf. Retr. | 1 |
| 2011 | Multilingual Adaptive Search for Digital Libraries
M. Rami Ghorab, Johannes Leveling, Séamus Lawless, Alexander O'Connor, Dong Zhou 0001, Gareth J. F. Jones, Vincent P. Wade |
TPDL | 5 |
| 2010 | A late fusion approach to cross-lingual document re-rankingabstractThe field of information retrieval still strives to develop models which allow semantic information to be integrated in the ranking process to improve performance in comparison to standard bag-of-words based models. Cross-lingual information retrieval is an example of where such a model is required, as content or concepts often need to be matched across languages. To overcome this problem, a conceptual model has been adopted in ranking an entire corpus which normally exploits latent/implicit features of the text. One of the drawbacks of this model is that the computational cost is significant and often intractable in modern test collections. Therefore, approaches utilizing concept-based models for re-ranking initial retrieval results have attracted a considerable amount of study, in particular the latent concept model. However, fitting such a model to a smaller collection is less meaningful than fitting it into the whole corpus. This paper proposes a late fusion method which incorporates scores generated by using external knowledge to enhance the space produced by the latent concept method. This method is further demonstrated to be suitable for multilingual re-ranking purposes. To illustrate the effectiveness of the proposed method, experiments were conducted over test collections across three languages. The results demonstrate that the method can comfortably achieve improvements in retrieval performance over several re-ranking methods. Dong Zhou 0001, Séamus Lawless, Jinming Min, Vincent P. Wade |
CIKM | 1 |
| 2010 | Assessing the readability of clinical documents in a document engineering environmentabstractPrevious work has established that specific linguistic markers present in specialised medical documents (clinical guidelines) can be used to support their automatic structuring within a document engineering environment. This technique is commonly used by the French Health Authority (la Haute Autorite de Sante) during elaboration of clinical guidelines to improve the quality of the final document. In this paper, we explore the readability of clinical guidelines. We discuss a structural measure of document readability that exploits the ratio between these linguistic markers (deontic structures) and the remainder of the text. We describe an experiment in which a corpus of 10 French clinical guidelines is scored for structural readability. We correlate these scores with measures of textual cohesion (computed using latent semantic analysis) and the results of a readability survey performed by a panel of domain experts. Our results suggest an association between the density of deontic structures in a clinical guideline and its overall readability. This implies that certain generic readability measures can henceforth be utilised in our document engineering environment. Mark Truran, Gersende Georg, Marc Cavazza, Dong Zhou 0001 |
ACM Symposium on Document Engineering | 4 |
| 2009 | Latent Document Re-Ranking
Dong Zhou 0001, Vincent P. Wade |
EMNLP | 1 |
| 2008 | A Hybrid Technique for English-Chinese Cross Language Information RetrievalabstractIn this article we describe a hybrid technique for dictionary-based query translation suitable for English-Chinese cross language information retrieval. This technique marries a graph-based model for the resolution of candidate term ambiguity with a pattern-based method for the translation of out-of-vocabulary (OOV) terms. We evaluate the performance of this hybrid technique in an experiment using several NTCIR test collections. Experimental results indicate a substantial increase in retrieval effectiveness over various baseline systems incorporating machine- and dictionary-based translation. Dong Zhou 0001, Mark Truran, Tim J. Brailsford, Helen Ashman |
ACM Trans. Asian Lang. Inf. Process. | 1 |