VLDB 2026 Research / reviewers in the wild / expert
Fuji Ren
dblp:17/4005
· DBLP profile ↗
25ranked-venue papers in the field
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 12Knowledge Engineering, Semantic Web & Information Systems · 8 (2 first)Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReNoRD: Learning from Relations under Noisy Pseudo Labels via Relational Distillation for Multimodal Sentiment
Tiantai Zhai, Yan Zhuang 0002, Fuji Ren, Jiawen Deng 0006 |
ICMR | 3 |
| 2026 | Retrieval-enhanced, Adaptively Collaborative, and Temporal-aware user behavior comprehension for LLM-based sequential recommendation
Zheng Hu 0001, Yongsen Pan, Zetao Li 0002, Satoshi Nakagawa, Jiawen Deng 0006, Shimin Cai, Fuji Ren |
Inf. Process. Manag. | 8 |
| 2025 | MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsabstractLarge language models (LLMs) are expected to offer structured Markdown responses for the sake of readability in web chatbots (e.g., ChatGPT). Although there are a myriad of metrics to evaluate LLMs, they fail to evaluate the readability from the view of output content structure. To this end, we focus on an overlooked yet important metric --- Markdown Awareness, which directly impacts the readability and structure of the content generated by these language models. In this paper, we introduce MDEval, a comprehensive benchmark to assess Markdown Awareness for LLMs, by constructing a dataset with 20K instances covering 10 subjects in English and Chinese. Unlike traditional model-based evaluations, MDEval provides excellent interpretability by combining model-based generation tasks and statistical methods. Our results demonstrate that MDEval achieves a Spearman correlation of 0.791 and an accuracy of 84.1% with human, outperforming existing methods by a large margin. Extensive experimental results also show that through fine-tuning over our proposed dataset, less performant open-source models are able to achieve comparable performance to GPT-4o in terms of Markdown Awareness. To ensure reproducibility and transparency, MDEval is open sourced at https://github.com/SWUFE-DB-Group/MDEval-Benchmark. Zhongpu Chen, Yinfeng Liu, Long Shi 0002, Zhi-Jie Wang 0009, Xingyan Chen, Yu Zhao 0019, Fuji Ren |
WWW | 7 |
| 2025 | ETS-MM: A Multi-Modal Social Bot Detection Model Based on Enhanced Textual Semantic RepresentationabstractSocial bots are becoming increasingly common in social networks, and their activities affect the security and authenticity of social media platforms. Current state-of-the-art social bot detection methods leverage multimodal approaches that analyze various modalities, such as user metadata, text, and social network relationships. However, these methods may not always extract additional dimensions of semantic feature information that could offer a deeper understanding of users' social patterns. To address this issue, we propose ETS-MM, a multimodal detection framework designed to augment multidimensional information from text and extract the semantic feature representation of user text information. We first analyze the user's tweeting behavior based on topic preference and emotion tendency, integrating them into the textual data. Then, we try to extract enhanced semantic representations that reveal the latent relationship between tweeting behavior and tweet content while identifying potential contextual associations and emotional changes. Additionally, to capture the complex interaction between users, we integrate the user's multimodal information, including metadata, textual features, enhanced semantic features, and social network relationships to propagate and aggregate information across various modalities. Experimental results demonstrate that ETS-MM significantly outperforms existing methods across two widely used social bot detection benchmark datasets, validating its effectiveness and superiority. Wei Li 0308, Jiawen Deng 0006, Jiali You 0002, Yan Zhuang 0002, Fuji Ren |
WWW | 6 |
| 2025 | Hierarchical Denoising for Robust Social RecommendationabstractSocial recommendations leverage social networks to augment the performance of recommender systems. However, the critical task of denoising social information has not been thoroughly investigated in prior research. In this study, we introduce a hierarchical denoising robust social recommendation model to tackle noise at two levels: 1) intra-domain noise, resulting from user multi-faceted social trust relationships, and 2) inter-domain noise, stemming from the entanglement of the latent factors over heterogeneous relations (e.g., user-item interactions, user-user trust relationships). Specifically, our model advances a preference and social psychology-aware methodology for the fine-grained and multi-perspective estimation of tie strength within social networks. This serves as a precursor to an edge weight-guided edge pruning strategy that refines the model's diversity and robustness by dynamically filtering social ties. Additionally, we propose a user interest-aware cross-domain denoising gate, which not only filters noise during the knowledge transfer process but also captures the high-dimensional, nonlinear information prevalent in social domains. We conduct extensive experiments on three real-world datasets to validate the effectiveness of our proposed model against state-of-the-art baselines. We perform empirical studies on synthetic datasets to validate the strong robustness of our proposed model. Zheng Hu 0001, Satoshi Nakagawa, Yan Zhuang 0002, Jiawen Deng 0006, Shimin Cai, Tao Zhou 0001, Fuji Ren |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2024 | Enhancing cross-market recommendations by addressing negative transfer and leveraging item co-occurrences
Zheng Hu 0001, Satoshi Nakagawa, Shimin Cai, Fuji Ren, Jiawen Deng 0006 |
Inf. Syst. | 4 |
| 2024 | Prompted and integrated textual information enhancing aspect-based sentiment analysis
Xuefeng Shi, Min Hu 0010, Fuji Ren, Piao Shi, Jiawen Deng 0006, Yiming Tang 0001 |
J. Intell. Inf. Syst. | 3 |
| 2023 | Celebrity-aware Graph Contrastive Learning Framework for Social RecommendationabstractSocial networks exhibit a distinct "celebrity effect" whereby influential individuals have a more significant impact on others compared to ordinary individuals, unlike other network structures such as citation networks and knowledge graphs. Despite its common occurrence in social networks, the celebrity effect is frequently overlooked by existing social recommendation methods when modeling social relationships, thereby hindering the full exploitation of social networks to mine similarities between users. In this paper, we fill this gap and propose a Celebrity-aware Graph Contrastive Learning Framework for Social Recommendation (CGCL), which explicitly models the celebrity effect in the social domain. Technically, we measure the different influences of celebrity and ordinary nodes by mining social network structure features, such as closeness centrality. To model the celebrity effect in social networks, we design a novel user-user impact-aware aggregation method, which incorporates the celebrity-aware influence information into the message propagation process. Additionally, we design a graph neural network-based framework which incorporates social semantics into the user-item interaction modeling with contrastive learning-enhanced data augmentation. The experimental results on three real-world datasets show the effectiveness of the proposed framework. We conduct ablation experiments to prove that the key components of our model benefit the recommendation performance improvement. Zheng Hu 0001, Satoshi Nakagawa, Yu Gu 0003, Fuji Ren |
CIKM | 5 |
| 2023 | Efficient random subspace decision forests with a simple probability dimensionality setting scheme
Fei Wang 0008, Zhongheng Li, Peilin Jiang, Fuji Ren, Feiping Nie 0001 |
Inf. Sci. | 5 |
| 2023 | An Effective Clustering Optimization Method for Unsupervised Linear Discriminant AnalysisabstractThe recent work Unsupervised Linear Discriminant Analysis (Un-LDA) completes its clustering process during the alternating optimization by converting equivalently the objective and finally using the K-means algorithm. However, the K-means algorithm has its inherent drawbacks. It is hard for the K-means algorithm to deal well with some complex clustering cases where there are too many real clusters or non-convex clusters. In this paper, a novel clustering optimization method is presented to accomplish the clustering process in Un-LDA and the resulting method can be named Un-LDA(CD). Specifically, instead of the K-means algorithm, an elaborately designed coordinate descent algorithm is adopted to obtain the clusters after the objective function goes through a series of simple but deft equivalent conversions. Extensive experiments have demonstrated that the coordinate descent clustering solution for Un-LDA can outperform the original K-means based solution on the tested data sets especially those complex data sets with a pretty large number of real clusters. Fei Wang 0008, Fuji Ren, Zhongheng Li, Feiping Nie 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Toward action comprehension for searching: Mining actionable intents in query entitiesabstractUnderstanding search engine users' intents has been a popular study in information retrieval, which directly affects the quality of retrieved information. One of the fundamental problems in this field is to find a connection between the entity in a query and the potential intents of the users, the latter of which would further reveal important information for facilitating the users' future actions. In this article, we present a novel research method for mining the actionable intents for search users, by generating a ranked list of the potentially most informative actions based on a massive pool of action samples. We compare different search strategies and their combinations for retrieving the action pool and develop three criteria for measuring the informativeness of the selected action samples, that is, the significance of an action sample within the pool, the representativeness of an action sample for the other candidate samples, and the diverseness of an action sample with respect to the selected actions. Our experiment, based on the Action Mining (AM) query entity data set from the Actionable Knowledge Graph (AKG) task at NTCIR‐13, suggests that the proposed approach is effective in generating an informative and early‐satisfying ranking of potential actions for search users. Yunong Wu, Fuji Ren |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Weighted high-order hidden Markov models for compound emotions recognition in text
Changqin Quan, Fuji Ren |
Inf. Sci. | 2 |
| 2016 | Role-explicit query extraction and utilization for quantifying user intents
Fuji Ren, Hai-Tao Yu 0003 |
Inf. Sci. | 1 |
| 2015 | Refinement by Filtering Translation Candidates and Similarity Based Approach to Expand Emotion Tagged Corpus
Kazuyuki Matsumoto, Fuji Ren, Minoru Yoshida, Kenji Kita |
IC3K | 2 |
| 2015 | An Approach to Refine Translation Candidates for Emotion Estimation in Japanese-English LanguageabstractResearches on emotion estimation from text mostly use machine learning method. Because machine learning
requires a large amount of example corpora, how to acquire high quality training data has been discussed as
one of its major problems. The existing language resources include emotion corpora; however, they are not
available if the language is different. Constructing bilingual corpus manually is also financially difficult. We
propose a method to convert a training data into different language using an existing Japanese-English parallel
emotion corpus. With a bilingual dictionary, the translation candidates are extracted against every word of
each sentence included in the corpus. Then the extracted translation candidates are narrowed down into a set
of words that highly contribute to emotion estimation and we used the set of words as training data. As the
result of the evaluation experiment using the training data created by our proposed method, the accuracy of
emotion estimation increased up to 66.7% in Naive Bayes.
1 INTRODUCTION
Recently, there have been many researches on emotion
estimation from text in the field of sentiment
analysis or opinion mining (Ren, 2009), (Ren and
Quan, 2015), (Ren and Wu, 2013), (Quan and Ren,
2010), (Quan and Ren, 2014), (Ren and Matsumoto,
2015) and many of them adopted machine learning
methods that used words as a feature. When the type
of the target sentence for emotion estimation and the
type of the sentence prepared as training data are different,
as in the case of terminology in the problem
of domain adaptation for document classification, the
appearance tendency of the emotion words differs.
This causes a problem in fluctuation of accuracy. On
the other hand, when a word is used as a feature for
emotion estimation, the sentence structure does not
have to be considered. As a result, it is easy to apply
the method to other languages. Only if we prepare a
large number of corpora with annotation of emotion
tags on each sentence, emotion would be easily estimated
by using the machine learning method. In the
machine learning method, because manual definition
of a rule is not necessary, we can reduce costs to apply
the method to other languages.
However, just like the problem in the domain, depending
on the Kazuyuki Matsumoto, Minoru Yoshida, Kenji Kita, Fuji Ren |
KEOD | 4 |
| 2014 | Search Result Diversification via Filling Up Multiple KnapsacksabstractResult diversification is a topic of great value for enhancing user experience in many fields, such as web search and recommender systems. Many existing methods generate a diversified result in a sequential manner, but they work well only if the preceding choices are optimal or close to the optimal solution. Moreover, a manually tuned parameter (say,λ) is often required to trade off relevance and diversity. This makes it difficult to know whether the failures are caused by the optimization criterion or the setting of λ. In context of web search, we formulate the result diversification task as a 0-1 multiple subtopic knapsack problem (MSKP), where a subset of documents are optimally chosen like filling up multiple subtopic knapsacks. This formulation yields no trade-off parameters to be specified beforehand. Solving the 0-1 MSKP is NP-hard, we treat the optimization of 0-1 MSKP using a graphical model over latent binary variables as a maximum posterior inference problem, and tackle it with the max-sum belief propagation algorithm. To validate the effectiveness and efficiency of the proposed 0-1 MSKP model, we conduct a series of experiments on two TREC diversity collections. The experimental results show that the proposed model outperforms several state-of-the-art methods significantly, not only in terms of standard diversity metrics (α-nDCG, nERRIA and subtopic recall), but also in terms of efficiency. Hai-Tao Yu 0003, Fuji Ren |
CIKM | 2 |
| 2014 | Subtopic Mining via Modifier Graph Clustering
Hai-Tao Yu 0003, Fuji Ren |
PAKDD (1) | 2 |
| 2014 | Unsupervised product feature extraction for feature-oriented opinion determination
Changqin Quan, Fuji Ren |
Inf. Sci. | 2 |
| 2013 | Class-indexing-based term weighting for automatic text classification
Fuji Ren, Mohammad Golam Sohrab |
Inf. Sci. | 1 |
| 2012 | Role-explicit query identification and intent role annotationabstractUnderstanding the information need or intent encoded within a query has long been regarded as an essential factor of effective information retrieval. For better query representation and understanding, two intent roles (kernel-object and modifier) are introduced to structurally parse a class of role-explicit queries, which constitute a majority of common user queries. Furthermore, we focus on two research problems: RP-1: Given a role-explicit query, how to identify the kernel-object and modifier, namely intent role annotation; RP-2: How to determine whether an arbitrary query is role-explicit or not. To solve RP-1, we propose a simplified word n-gram role model (SWNR), which quantifies the generating probability of a role-explicit query and performs intent role annotation effectively. Using a set of discriminative features, we build classifiers to address RP-2 in a supervised manner. The experimental results show that: (1) SWNR can achieve a satisfactory performance, more than 73% in terms of different metrics; (2) The classifiers can achieve more than 90% precision in identifying role-explicit queries; (3) Compared with traditional techniques for query representation and understanding, e.g., name entity recognition in query and class-level query intent inference, intent role annotation provides a more flexible framework and a number of applications can benefit from annotating role-explicit queries, such as intent mining and diversified document ranking. Hai-Tao Yu 0003, Fuji Ren |
CIKM | 2 |
| 2009 | A Practical System of Domain Ontology Learning Using the Web for ChineseabstractThis paper proposes an ontology learning system model based on the Web search engine and Protege-OWL API, which emphasizes iterative learning approach by the extracted instances. We discuss taxonomic and non-taxonomic relationship learning separately in ontology learning system, and investigate the importance of verb plus noun phrase learning for extraction of activity concepts in Chinese. We also propose an algorithm of relevance measurement for extracting relation instances by binary keywords based on co-occurrence statistics. Finally, we build a practical system of ontology learning through learning relation instances of the Chinese festival ontology, and test the effectiveness of our method. Peilin Jiang, Fuji Ren |
ICIW | 3 |
| 2008 | English-Arabic proper-noun transliteration-pairs creationabstractAbstract Proper nouns may be considered the most important query words in information retrieval. If the two languages use the same alphabet, the same proper nouns can be found in either language. However, if the two languages use different alphabets, the names must be transliterated. Short vowels are not usually marked on Arabic words in almost all Arabic documents (except very important documents like the Muslim and Christian holy books). Moreover, most Arabic words have a syllable consisting of a consonant‐vowel combination (CV), which means that most Arabic words contain a short or long vowel between two successive consonant letters. That makes it difficult to create English‐Arabic transliteration pairs, since some English letters may not be matched with any romanized Arabic letter. In the present study, we present different approaches for extraction of transliteration proper‐noun pairs from parallel corpora based on different similarity measures between the English and romanized Arabic proper nouns under consideration. The strength of our new system is that it works well for low‐frequency proper noun pairs. We evaluate the new approaches presented using two different English‐Arabic parallel corpora. Most of our results outperform previously published results in terms of precision, recall, and F‐Measure. Mohamed Abdel Fattah, Fuji Ren |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2006 | Stemming to improve translation lexicon creation form bitexts
Mohamed Abdel Fattah, Fuji Ren, Shingo Kuroiwa |
Inf. Process. Manag. | 2 |
| 2002 | An information retrieval model based on vector space method by supervised learning
Xiaoying Tai, Fuji Ren, Kenji Kita |
Inf. Process. Manag. | 2 |
| 1999 | Chinese information retrieval: using characters or words?
Jian-Yun Nie, Fuji Ren |
Inf. Process. Manag. | 2 |