VLDB 2026 Research / reviewers in the wild / expert
Jianshu Weng
dblp:98/392
· DBLP profile ↗
15ranked-venue papers
7as first author
1since 2021 · last 2022
0000-0003-2540-3829ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 86% Multi-agent systems · 14% | |
| Databases, data mining, and information retrieval
4 papers |
Web and social media mining · 50% Information retrieval · 29% Data mining · 21% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text segmentation
tweet segmentation |
0.4 | 2 | 2015 | Tweet Segmentation and Its Application to Named Entity Recognition · IEEE Trans. Knowl. Data Eng. 2015 Exploiting hybrid contexts for Tweet segmentation · SIGIR 2013 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.3 | 2 | 2015 | Tweet Segmentation and Its Application to Named Entity Recognition · IEEE Trans. Knowl. Data Eng. 2015 Exploiting hybrid contexts for Tweet segmentation · SIGIR 2013 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.2 | 1 | 2015 | Tweet Segmentation and Its Application to Named Entity Recognition · IEEE Trans. Knowl. Data Eng. 2015 |
Data mining › text mining › information extraction
named entity recognition |
0.1 | 1 | 2012 | TwiNER: named entity recognition in targeted twitter stream · SIGIR 2012 |
Information retrieval
text analysis |
0.1 | 1 | 2012 | TwiNER: named entity recognition in targeted twitter stream · SIGIR 2012 |
Information retrieval
ranking |
0.1 | 2 | 2010 | TwitterRank: finding topic-sensitive influential twitterers · WSDM 2010 What Do People Want in Microblogs? Measuring Interestingness of Hashtags in Twitter · ICDM 2010 |
Knowledge, reasoning and agents › Multi-agent systems
trust modeling |
0.1 | 1 | 2010 | Credibility: How Agents Can Handle Unfair Third-Party Testimonies in Computational Trust Models · IEEE Trans. Knowl. Data Eng. 2010 |
Web and social media mining › social influence analysis
influential user identification |
0.1 | 1 | 2010 | TwitterRank: finding topic-sensitive influential twitterers · WSDM 2010 |
Web and social media mining › social media analysis
microblog analysis |
0.1 | 1 | 2010 | What Do People Want in Microblogs? Measuring Interestingness of Hashtags in Twitter · ICDM 2010 |
Web and social media mining
social network analysis |
0.1 | 1 | 2010 | TwitterRank: finding topic-sensitive influential twitterers · WSDM 2010 |
Web and social media mining
social media analysis |
0.1 | 2 | 2015 | Tweet Segmentation and Its Application to Named Entity Recognition · IEEE Trans. Knowl. Data Eng. 2015 TwiNER: named entity recognition in targeted twitter stream · SIGIR 2012 |
Web and social media mining › social media analysis
twitter analysis |
0.1 | 1 | 2015 | Tweet Segmentation and Its Application to Named Entity Recognition · IEEE Trans. Knowl. Data Eng. 2015 |
Knowledge, reasoning and agents › Multi-agent systems
agent interaction |
0.0 | 1 | 2010 | Credibility: How Agents Can Handle Unfair Third-Party Testimonies in Computational Trust Models · IEEE Trans. Knowl. Data Eng. 2010 |
Data mining › structured data mining › graph mining
community detection |
0.0 | 1 | 2010 | What Do People Want in Microblogs? Measuring Interestingness of Hashtags in Twitter · ICDM 2010 |
Data mining
pattern mining |
0.0 | 1 | 2010 | What Do People Want in Microblogs? Measuring Interestingness of Hashtags in Twitter · ICDM 2010 |
Methods — techniques the papers use, named apart from their topics
stickiness score · 0.4POS tagging · 0.4pseudo feedback · 0.4pseudo-feedback · 0.2weak supervision · 0.2collective inference · 0.2random walk · 0.1dynamic programming · 0.1topic modeling · 0.1supervised ranking · 0.1pagerank · 0.1dempster-shafer theory · 0.1community-based dispersion and divergence · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Federated Learning for Electronic Health RecordsabstractIn data-driven medical research, multi-center studies have long been preferred over single-center ones due to a single institute sometimes not having enough data to obtain sufficient statistical power for certain hypothesis testings as well as predictive and subgroup studies. The wide adoption of electronic health records (EHRs) has made multi-institutional collaboration much more feasible. However, concerns over infrastructures, regulations, privacy, and data standardization present a challenge to data sharing across healthcare institutions. Federated Learning (FL), which allows multiple sites to collaboratively train a global model without directly sharing data, has become a promising paradigm to break the data isolation. In this study, we surveyed existing works on FL applications in EHRs and evaluated the performance of current state-of-the-art FL algorithms on two EHR machine learning tasks of significant clinical importance on a real world multi-center EHR dataset. Trung Kien Dang, Xiang Lan 0004, Jianshu Weng, Mengling Feng |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | Individual Judgments Versus Consensus: Estimating Query-URL RelevanceabstractQuery-URL relevance, measuring the relevance of each retrieved URL with respect to a given query, is one of the fundamental criteria to evaluate the performance of commercial search engines. The traditional way to collect reliable and accurate query-URL relevance requires multiple annotators to provide their individual judgments based on their subjective expertise (e.g., understanding of user intents). In this case, the annotators’ subjectivity reflected in each annotator individual judgment (AIJ) inevitably affects the quality of the ground truth relevance (GTR). But to the best of our knowledge, the potential impact of AIJs on estimating GTRs has not been studied and exploited quantitatively by existing work. This article first studies how multiple AIJs and GTRs are correlated. Our empirical studies find that the multiple AIJs possibly provide more cues to improve the accuracy of estimating GTRs. Inspired by this finding, we then propose a novel approach to integrating the multiple AIJs with the features characterizing query-URL pairs for estimating GTRs more accurately. Furthermore, we conduct experiments in a commercial search engine—Baidu.com—and report significant gains in terms of the normalized discounted cumulative gains. Hengjie Song, Huaqing Min, Qingyao Wu, Wei Wei 0002, Jianshu Weng, Xiaogang Han, Qiang Yang 0001, Jialiang Shi, Jiaqian Gu, Chunyan Miao, Toyoaki Nishida |
ACM Trans. Web | 6 |
| 2015 | Tweet Segmentation and Its Application to Named Entity RecognitionabstractTwitter has attracted millions of users to share and disseminate most up-to-date information, resulting in large volumes of data produced everyday. However, many applications in Information Retrieval (IR) and Natural Language Processing (NLP) suffer severely from the noisy and short nature of tweets. In this paper, we propose a novel framework for tweet segmentation in a batch mode, called HybridSeg. By splitting tweets into meaningful segments, the semantic or context information is well preserved and easily extracted by the downstream applications. HybridSeg finds the optimal segmentation of a tweet by maximizing the sum of the stickiness scores of its candidate segments. The stickiness score considers the probability of a segment being a phrase in English (i.e., global context) and the probability of a segment being a phrase within the batch of tweets (i.e., local context). For the latter, we propose and evaluate two models to derive local context by considering the linguistic features and term-dependency in a batch of tweets, respectively. HybridSeg is also designed to iteratively learn from confident segments as pseudo feedback. Experiments on two tweet data sets show that tweet segmentation quality is significantly improved by learning both global and local contexts compared with using global context alone. Through analysis and comparison, we show that local linguistic features are more reliable for learning local context compared with term-dependency. As an application, we show that high accuracy is achieved in named entity recognition by applying segment-based part-of-speech (POS) tagging. Chenliang Li 0005, Aixin Sun, Jianshu Weng, Qi He 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Exploiting hybrid contexts for Tweet segmentationabstractTwitter has attracted hundred millions of users to share and disseminate most up-to-date information. However, the noisy and short nature of tweets makes many applications in information retrieval (IR) and natural language processing (NLP) challenging. Recently, segment-based tweet representation has demonstrated effectiveness in named entity recognition (NER) and event detection from tweet streams. To split tweets into meaningful phrases or segments, the previous work is purely based on external knowledge bases, which ignores the rich local context information embedded in the tweets. In this paper, we propose a novel framework for tweet segmentation in a batch mode, called HybridSeg. HybridSeg incorporates local context knowledge with global knowledge bases for better tweet segmentation. HybridSeg consists of two steps: learning from off-the-shelf weak NERs and learning from pseudo feedback. In the first step, the existing NER tools are applied to a batch of tweets. The named entities recognized by these NERs are then employed to guide the tweet segmentation process. In the second step, HybridSeg adjusts the tweet segmentation results iteratively by exploiting all segments in the batch of tweets in a collective manner. Experiments on two tweet datasets show that HybridSeg significantly improves tweet segmentation quality compared with the state-of-the-art algorithm. We also conduct a case study by using tweet segments for the task of named entity recognition from tweets. The experimental results demonstrate that HybridSeg significantly benefits the downstream applications. Chenliang Li 0005, Aixin Sun, Jianshu Weng, Qi He 0002 |
SIGIR | 3 |
| 2012 | TwiNER: named entity recognition in targeted twitter streamabstractMany private and/or public organizations have been reported to create and monitor targeted Twitter streams to collect and understand users' opinions about the organizations. Targeted Twitter stream is usually constructed by filtering tweets with user-defined selection criteria e.g. tweets published by users from a selected region, or tweets that match one or more predefined keywords. Targeted Twitter stream is then monitored to collect and understand users' opinions about the organizations. There is an emerging need for early crisis detection and response with such target stream. Such applications require a good named entity recognition (NER) system for Twitter, which is able to automatically discover emerging named entities that is potentially linked to the crisis. In this paper, we present a novel 2-step unsupervised NER system for targeted Twitter stream, called TwiNER. In the first step, it leverages on the global context obtained from Wikipedia and Web N-Gram corpus to partition tweets into valid segments (phrases) using a dynamic programming algorithm. Each such tweet segment is a candidate named entity. It is observed that the named entities in the targeted stream usually exhibit a gregarious property, due to the way the targeted stream is constructed. In the second step, TwiNER constructs a random walk model to exploit the gregarious property in the local context derived from the Twitter stream. The highly-ranked segments have a higher chance of being true named entities. We evaluated TwiNER on two sets of real-life tweets simulating two targeted streams. Evaluated using labeled ground truth, TwiNER achieves comparable performance as with conventional approaches in both streams. Various settings of TwiNER have also been examined to verify our global context + local context combo idea. Chenliang Li 0005, Jianshu Weng, Qi He 0002, Yuxia Yao, Anwitaman Datta, Aixin Sun, Bu-Sung Lee |
SIGIR | 2 |
| 2011 | Comparing Twitter and Traditional Media Using Topic Models
Wayne Xin Zhao, Jing Jiang 0001, Jianshu Weng, Jing He 0010, Ee-Peng Lim, Hongfei Yan, Xiaoming Li 0001 |
ECIR | 3 |
| 2011 | Event Detection in Twitter
Jianshu Weng, Bu-Sung Lee |
ICWSM | 1 |
| 2011 | Trust-based web service selection in virtual communitiesabstractIn the present web service architecture, selecting which web service to use is still a task which requires a high level of human intervention. In this paper, we propose a trust-based web service selection process which aspires to automate part of the Zhiqi Shen 0001, Han Yu 0001, Chunyan Miao, Jianshu Weng |
Web Intell. Agent Syst. | 4 |
| 2010 | Mining interesting link formation rules in social networksabstractLink structures are important patterns one looks out for when modeling and analyzing social networks. In this paper, we propose the task of mining interesting Link Formation rules (LF-rules) containing link structures known as Link Formation patterns (LF-patterns). LF-patterns capture various dyadic and/or triadic structures among groups of nodes, while LF-rules capture the formation of a new link from a focal node to another node as a postcondition of existing connections between the two nodes. We devise a novel LF-rule mining algorithm, known as LFR-Miner, based on frequent subgraph mining for our task. In addition to using a support-confidence framework for measuring the frequency and significance of LF-rules, we introduce the notion of expected support to account for the extent to which LF-rules exist in a social network by chance. Specifically, only LF-rules with higher-than-expected support are considered interesting. We conduct empirical studies on two real-world social networks, namely Epinions and myGamma. We report interesting LF-rules mined from the two networks, and compare our findings with earlier findings in social network analysis. Cane Wing-ki Leung, Ee-Peng Lim, David Lo 0001, Jianshu Weng |
CIKM | 4 |
| 2010 | What Do People Want in Microblogs? Measuring Interestingness of Hashtags in TwitterabstractWhen micro logging becomes a very popular social media, finding interesting posts from high volume stream of user posts is a challenging research problem. To organize large number of posts, users can assign tags to posts so that these posts can be navigated and searched by tag. In this paper, we focus on modeling the interestingness of hash tags in Twitter, the largest and most active micro logging site. We propose to first construct communities based on both follow links and tagged interactions. We then measure the dispersion and divergence of users and tweets using hash tags among the constructed communities. The interestingness of hash tags are then derived from these community-based dispersion and divergence features. We further introduce a supervised approach to rank hash tags by interestingness. Our experiments on a Twitter dataset show that the proposed approach achieves a fairly good performance. Jianshu Weng, Ee-Peng Lim, Qi He 0002, Cane Wing-ki Leung |
ICDM | 1 |
| 2010 | TwitterRank: finding topic-sensitive influential twitterersabstractThis paper focuses on the problem of identifying influential users of micro-blogging services. Twitter, one of the most notable micro-blogging services, employs a social-networking model called "following", in which each user can choose who she wants to "follow" to receive tweets from without requiring the latter to give permission first. In a dataset prepared for this study, it is observed that (1) 72.4% of the users in Twitter follow more than 80% of their followers, and (2) 80.5% of the users have 80% of users they are following follow them back. Our study reveals that the presence of "reciprocity" can be explained by phenomenon of homophily. Based on this finding, TwitterRank, an extension of PageRank algorithm, is proposed to measure the influence of users in Twitter. TwitterRank measures the influence taking both the topical similarity between users and the link structure into account. Experimental results show that TwitterRank outperforms the one Twitter currently uses and other related algorithms, including the original PageRank and Topic-sensitive PageRank. Jianshu Weng, Ee-Peng Lim, Jing Jiang 0001, Qi He 0002 |
WSDM | 1 |
| 2010 | Credibility: How Agents Can Handle Unfair Third-Party Testimonies in Computational Trust ModelsabstractUsually, agents within multiagent systems represent different stakeholders that have their own distinct and sometimes conflicting interests and objectives. They would behave in such a way so as to achieve their own objectives, even at the cost of others. Therefore, there are risks in interacting with other agents. A number of computational trust models have been proposed to manage such risk. However, the performance of most computational trust models that rely on third-party recommendations as part of the mechanism to derive trust is easily deteriorated by the presence of unfair testimonies. There have been several attempts to combat the influence of unfair testimonies. Nevertheless, they are either not readily applicable since they require additional information which is not available in realistic settings, or ad hoc as they are tightly coupled with specific trust models. Against this background, a general credibility model is proposed in this paper. Empirical studies have shown that the proposed credibility model is more effective than related work in mitigating the adverse influence of unfair testimonies. Jianshu Weng, Zhiqi Shen 0001, Chunyan Miao, Angela Goh, Cyril Leung |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2006 | A Robust Reputation System for the GridabstractIn the grid, the reliability of the providers who manage the resources is usually unknown to the consumers. Unreliable providers usually fail to complete consumers' task. Therefore, consumers take the risk that they lose their payment for the resource access. Establishing reputation systems is an alternative to manage such risk. However, existing attempts to establish reputation systems in the grid mainly focus on enabling reputation evaluation to measure the resource providers' reliability, but do not effectively address the inaccurate testimonies and malicious referrers, which generally deteriorate the performance of reputation systems. Against this background, this paper proposes a reputation system with the ability to mitigate the influence of the inaccurate testimonies and malicious referrers. Experimental studies show that the proposed system is efficient in mitigating the influence of these inaccurate testimonies and malicious referrers. Jianshu Weng, Chunyan Miao, Angela Goh |
CCGRID | 1 |
| 2005 | Trust-based collaborative filteringabstractNo abstract available. Jianshu Weng, Chunyan Miao, Angela Goh, Dongtao Li |
CIKM | 1 |
| 2005 | Protecting Online Rating Systems from Unfair Ratings
Jianshu Weng, Chunyan Miao, Angela Goh |
TrustBus | 1 |