Zhong Su

dblp:87/1363 · DBLP profile ↗
← Back
70ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-2303-9787ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 34 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Information extraction and text analysis · 58% Transfer learning and domain adaptation · 18% Graph learning · 6%
Databases, data mining, and information retrieval
24 papers
Information retrieval · 36% Data mining · 24% Web and social media mining · 23%
Computer graphics and multimedia
6 papers
Multimedia analysis and retrieval · 96% Visualization and visual analytics · 4%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 30 heaviest of 79, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
sentiment analysis
1.252019
CAN: Constrained Attention Networks for Multi-Aspect Sentiment Analysis · EMNLP/IJCNLP (1) 2019
Domain-Invariant Feature Distillation for Cross-Domain Sentiment Classification · EMNLP/IJCNLP (1) 2019
Sentence Compression for Aspect-Based Sentiment Analysis · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Natural language and speech › Information extraction and text analysis › document understanding
mind-map generation
0.922021
Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph Refinement · EMNLP (1) 2021
Revealing Semantic Structures of Texts: Multi-grained Framework for Automatic Mind-map Generation · IJCAI 2019
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect category detection
0.512021
Multi-Label Few-Shot Learning for Aspect Category Detection · ACL/IJCNLP (1) 2021
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.512021
Multi-Label Few-Shot Learning for Aspect Category Detection · ACL/IJCNLP (1) 2021
Machine learning › Graph learning › graph structure learning
graph refinement
0.512021
Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph Refinement · EMNLP (1) 2021
Machine learning › Transfer learning and domain adaptation › few-shot learning
multi-label few-shot learning
0.512021
Multi-Label Few-Shot Learning for Aspect Category Detection · ACL/IJCNLP (1) 2021
Data mining › text mining
topic modeling
0.432014
Probabilistic text modeling with orthogonalized topics · SIGIR 2014
Mining topics on participations for community discovery · SIGIR 2011
Joint Emotion-Topic Modeling for Social Affective Text Mining · ICDM 2009
Machine learning › Deep learning architectures and training
attention mechanism
0.412019
CAN: Constrained Attention Networks for Multi-Aspect Sentiment Analysis · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-domain sentiment classification
0.412019
Domain-Invariant Feature Distillation for Cross-Domain Sentiment Classification · EMNLP/IJCNLP (1) 2019
Machine learning › Transfer learning and domain adaptation
domain-invariant representation learning
0.412019
Domain-Invariant Feature Distillation for Cross-Domain Sentiment Classification · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis
relation extraction
0.412019
Revealing Semantic Structures of Texts: Multi-grained Framework for Automatic Mind-map Generation · IJCAI 2019
Medical and health informatics › computational pathology
pathological image classification
0.412019
P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification · CVPR 2019
Privacy and data protection
differential privacy
0.412019
P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification · CVPR 2019
Privacy and data protection
privacy-preserving machine learning
0.412019
P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification · CVPR 2019
Multimedia analysis and retrieval › image retrieval › instance-level image retrieval
landmark search
0.322012
Searching for diversified landmarks by photo · ACM Multimedia 2012
DLMSearch: diversified landmark search by photo · ACM Multimedia 2012
Recommender systems › e-commerce recommendation
product recommendation
0.212016
Boosting Recommendation in Unexplored Categories by User Price Preference · ACM Trans. Inf. Syst. 2016
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.212015
Sentence Compression for Aspect-Based Sentiment Analysis · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Web and social media mining
social network analysis
0.212015
Sampling Representative Users from Large Social Networks · AAAI 2015
Graph algorithms and graph theory
graph sampling
0.212015
Sampling Representative Users from Large Social Networks · AAAI 2015
Recommender systems
cold-start recommendation
0.212014
Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014
Recommender systems
cross-domain recommendation
0.212014
Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014
Information retrieval
text analysis
0.212014
Probabilistic text modeling with orthogonalized topics · SIGIR 2014
Information retrieval
diversified retrieval
0.222012
DLMSearch: diversified landmark search by photo · ACM Multimedia 2012
Searching for diversified landmarks by photo · ACM Multimedia 2012
Multimedia analysis and retrieval › object recognition
landmark recognition
0.212013
Tell me what happened here in history · ACM Multimedia 2013
Machine learning › Reinforcement learning
reinforcement learning for structured prediction
0.112021
Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph Refinement · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
topic model
0.112012
Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012
Natural language and speech › Question answering and dialogue systems
community question answering
0.112011
Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011
Natural language and speech › Information extraction and text analysis
text classification
0.112011
Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011
Web and social media mining
co-authorship networks
0.112011
Mining topics on participations for community discovery · SIGIR 2011
Data mining › structured data mining › graph mining
community detection
0.112011
Mining topics on participations for community discovery · SIGIR 2011

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 1.1noise injection · 1.1differential privacy · 1.1sequence-to-graph · 0.5reinforcement learning · 0.5graph refinement · 0.5few-shot learning · 0.5tree search · 0.4objective function optimization · 0.4approximation algorithm · 0.4NP-hardness proof · 0.4graph neural network · 0.4attention mechanism · 0.4feature distillation · 0.4visualization · 0.2utility function · 0.2regularization · 0.2joint factorization · 0.2
YearPublicationVenuePosition
2024 Analyzing source code vulnerabilities in the D2A dataset with ML ensembles and C-BERT
abstract
Abstract Static analysis tools are widely used for vulnerability detection as they can analyze programs with complex behavior and millions of lines of code. Despite their popularity, static analysis tools are known to generate an excess of false positives. The recent ability of Machine Learning models to learn from programming language data opens new possibilities of reducing false positives when applied to static analysis. However, existing datasets to train models for vulnerability identification suffer from multiple limitations such as limited bug context, limited size, and synthetic and unrealistic source code. We propose Differential Dataset Analysis or D2A, a differential analysis based approach to label issues reported by static analysis tools. The dataset built with this approach is called the D2A dataset. The D2A dataset is built by analyzing version pairs from multiple open source projects. From each project, we select bug fixing commits and we run static analysis on the versions before and after such commits. If some issues detected in a before-commit version disappear in the corresponding after-commit version, they are very likely to be real bugs that got fixed by the commit. We use D2A to generate a large labeled dataset. We then train both classic machine learning models and deep learning models for vulnerability identification using the D2A dataset. We show that the dataset can be used to build a classifier to identify possible false alarms among the issues reported by static analysis, hence helping developers prioritize and investigate potential true positives first. To facilitate future research and contribute to the community, we make the dataset generation pipeline and the dataset publicly available. We have also created a leaderboard based on the D2A dataset, which has already attracted attention and participation from the community.
Saurabh Pujar, Yunhui Zheng, Luca Buratti, Burn L. Lewis, Yunchung Chen, Jim Laredo, Alessandro Morari, Edward A. Epstein, Tsungnan Lin, Bo Yang 0013, Zhong Su
Empir. Softw. Eng.11
2023 Fine-Grained Domain Adaptation for Aspect Category Level Sentiment Analysis
abstract
Aspect category level sentiment analysis aims to identify the sentiment polarities towards the aspect categories discussed in a sentence. It usually suffers from a lack of labeled data. A popular solution is to transfer knowledge from a labeled source domain to an unlabeled target domain by unsupervised domain adaptation. However, most domain adaptation methods in sentiment analysis are coarse-grained, considering the source or target domain as a whole during the adaptation. We argue that these single-source single-target methods are inefficient since they ignore the difference between different aspect categories. In this article, we propose a fine-grained domain adaptation method to address the aspect category level sentiment analysis task by considering the adaptation between subdomains. Specifically, the source/target domain is divided into multiple subdomains according to the hierarchical structure of the aspect categories. We then design a multi-source multi-target transfer network to achieve fine-grained transfer. Extensive experimental results demonstrate the effectiveness of our fine-grained domain adaptation method on aspect category level sentiment analysis.
Mengting Hu 0002, Hang Gao 0003, Yike Wu 0002, Zhong Su, Shiwan Zhao
IEEE Trans. Affect. Comput.4
2023 Hybrid Regularizations for Multi-Aspect Category Sentiment Analysis
abstract
Aspect level sentiment classification aims to identify the sentiment polarity towards a particular aspect in a sentence. Previous attention-based methods generate an aspect-specific representation for each aspect and employ it to classify the sentiment polarity. However, normalized attention scores scatter over every word in the sentence, resulting in two issues. First, the attention may inherently introduce noise and downgrade the performance. Second, the opinion words may be “diluted” by other words, while the opinion feature should dominate for sentiment analysis. The issues become more severe in multi-aspect sentences. In this paper, we address the above two issues via hybrid regularizations, i.e.,aspect-levelandtask-level regularizations. Concretely, the aspect-level regularizations constrain the attention weights to alleviate noise. Among them, orthogonal regularization is designed for multi-aspect sentences and sparse regularization is for single-aspect sentences. To extract sentiment-dominant features, task-level regularization is proposed by introducing an orthogonal auxiliary task, i.e., aspect category detection. This regularization can allocate task-oriented context information for specific downstream tasks. Extensive experimental results on three public datasets demonstrate the effectiveness of the proposed approach in both single-task and multi-task scenarios.
Mengting Hu 0002, Shiwan Zhao, Zhong Su
IEEE Trans. Affect. Comput.4
2022 An incremental algorithm for concept lattice based on structural similarity index
Yu Hu 0009, Yan Zhu Hu, Zhong Su, Xiao Li Li, Wenjia Tian, Yan Ying Yang, Jia Feng Chai
Soft Comput.3
2022 When Pairs Meet Triplets: Improving Low-Resource Captioning via Multi-Objective Optimization
abstract
Image captioning for low-resource languages has attracted much attention recently. Researchers propose to augment the low-resource caption dataset into (image, rich-resource language, and low-resource language) triplets and develop the dual attention mechanism to exploit the existence of triplets in training to improve the performance. However, datasets in triplet form are usually small due to their high collecting cost. On the other hand, there are already many large-scale datasets, which contain one pair from the triplet, such as caption datasets in the rich-resource language and translation datasets from the rich-resource language to the low-resource language. In this article, we revisit the caption-translation pipeline of the translation-based approach to utilize not only the triplet dataset but also large-scale paired datasets in training. The caption-translation pipeline is composed of two models, one caption model of the rich-resource language and one translation model from the rich-resource language to the low-resource language. Unfortunately, it is not trivial to fully benefit from incorporating both the triplet dataset and paired datasets into the pipeline, due to the gap between the training and testing phases and the instability in the training process. We propose to jointly optimize the two models of the pipeline in an end-to-end manner to bridge the training and testing gap, and introduce two auxiliary training objectives to stabilize the training process. Experimental results show that the proposed method improves significantly over the state-of-the-art methods.
Yike Wu 0002, Shiwan Zhao, Ying Zhang 0015, Xiaojie Yuan, Zhong Su
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Multi-Label Few-Shot Learning for Aspect Category Detection
abstract
Mengting Hu, Shiwan Zhao, Honglei Guo, Chao Xue, Hang Gao, Tiegang Gao, Renhong Cheng, Zhong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mengting Hu 0002, Shiwan Zhao, Chao Xue 0003, Hang Gao 0003, Tiegang Gao, Renhong Cheng, Zhong Su
ACL/IJCNLP (1)8
2021 Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph Refinement
abstract
A mind-map is a diagram that represents the central concept and key ideas in a hierarchical way.Converting plain text into a mindmap will reveal its key semantic structure and be easier to understand.Given a document, the existing automatic mind-map generation method extracts the relationships of every sentence pair to generate the directed semantic graph for this document.The computation complexity increases exponentially with the length of the document.Moreover, it is difficult to capture the overall semantics.To deal with the above challenges, we propose an efficient mind-map generation network that converts a document into a graph via sequenceto-graph.To guarantee a meaningful mindmap, we design a graph refinement module to adjust the relation graph in a reinforcement learning manner.Extensive experimental results demonstrate that the proposed approach is more effective and efficient than the existing methods.The inference time is reduced by thousands of times compared with the existing methods.The case studies verify that the generated mind-maps better reveal the underlying semantic structures of the document.
Mengting Hu 0002, Shiwan Zhao, Hang Gao 0003, Zhong Su
EMNLP (1)5
2021 Design of sparse Bayesian echo state network for time series prediction
Lei Wang 0173, Zhong Su, Junfei Qiao 0001, Cuili Yang
Neural Comput. Appl.2
2020 Deep Semantic Compliance Advisor for Unstructured Document Compliance Checking
abstract
Unstructured document compliance checking is always a big challenge for banks since huge amounts of contracts and regulations written in natural language require professionals' interpretation and judgment. Traditional rule-based or keyword-based methods cannot precisely characterize the deep semantic distribution in the unstructured document semantic compliance checking due to the semantic complexity of contracts and regulations. Deep Semantic Compliance Advisor (DSCA) is an unstructured document compliance checking platform which provides multi-level semantic comparison by deep learning algorithms. In the statement-level semantic comparison, a Graph Neural Network (GNN) based syntactic sentence encoder is proposed to capture the complicate syntactic and semantic clues of the statement sentences. This GNN-based encoder outperforms existing syntactic sentence encoders in deep semantic comparison and is more beneficial for long sentences. In the clause-level semantic comparison, an attention-based semantic relatedness detection model is applied to find the most relevant legal clauses. DSCA significantly enhances the productivity of legal professionals in the unstructured document compliance checking for banks.
Zhili Guo, Zhong Su
IJCAI4
2019 Learning to Detect Opinion Snippet for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is to predict the sentiment polarity towards a particular aspect in a sentence.Recently, this task has been widely addressed by the neural attention mechanism, which computes attention weights to softly select words for generating aspect-specific sentence representations.The attention is expected to concentrate on opinion words for accurate sentiment prediction.However, attention is prone to be distracted by noisy or misleading words, or opinion words from other aspects.In this paper, we propose an alternative hard-selection approach, which determines the start and end positions of the opinion snippet, and selects the words between these two positions for sentiment prediction.Specifically, we learn deep associations between the sentence and aspect, and the long-term dependencies within the sentence by leveraging the pre-trained BERT model.We further detect the opinion snippet by selfcritical reinforcement learning.Especially, experimental results demonstrate the effectiveness of our method and prove that our hardselection approach outperforms soft-selection approaches when handling multi-aspect sentences.
Mengting Hu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su
CoNLL5
2019 P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification
abstract
Recently, deep convolutional neural networks (CNNs) have achieved great success in pathological image classification. However, due to the limited number of labeled pathological images, there are still two challenges to be addressed: (1) overfitting: the performance of a CNN model is undermined by the overfitting due to its huge amounts of parameters and the insufficiency of labeled training data. (2) privacy leakage: the model trained using a conventional method may involuntarily reveal the private information of the patients in the training dataset. The smaller the dataset, the worse the privacy leakage. To tackle the above two challenges, we introduce a novel stochastic gradient descent (SGD) scheme, named patient privacy preserving SGD (P3SGD), which performs the model update of the SGD in the patient level via a large-step update built upon each patient's data. Specifically, to protect privacy and regularize the CNN model, we propose to inject the well-designed noise into the updates. Moreover, we equip our P3SGD with an elaborated strategy to adaptively control the scale of the injected noise. To validate the effectiveness of P3SGD, we perform extensive experiments on a real-world clinical dataset and quantitatively demonstrate the superior ability of P3SGD in reducing the risk of overfitting. We also provide a rigorous analysis of the privacy cost under differential privacy. Additionally, we find that the models trained with P3SGD are resistant to the model-inversion attack compared with those trained using non-private SGD.
Bingzhe Wu, Shiwan Zhao, Guangyu Sun 0003, Zhong Su, Caihong Zeng
CVPR5
2019 Domain-Invariant Feature Distillation for Cross-Domain Sentiment Classification
abstract
Mengting Hu, Yike Wu, Shiwan Zhao, Honglei Guo, Renhong Cheng, Zhong Su. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mengting Hu 0002, Yike Wu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su
EMNLP/IJCNLP (1)6
2019 CAN: Constrained Attention Networks for Multi-Aspect Sentiment Analysis
abstract
Mengting Hu, Shiwan Zhao, Li Zhang, Keke Cai, Zhong Su, Renhong Cheng, Xiaowei Shen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mengting Hu 0002, Shiwan Zhao, Li Zhang 0007, Keke Cai, Zhong Su, Renhong Cheng
EMNLP/IJCNLP (1)5
2019 Improving Captioning for Low-Resource Languages by Cycle Consistency
abstract
Improving the captioning performance on low-resource languages by leveraging English caption datasets has received increasing research interest in recent years. Existing works mainly fall into two categories: translation-based and alignment-based approaches. In this paper, we propose to combine the merits of both approaches in one unified architecture. Specifically, we use a pre-trained English caption model to generate high-quality English captions, and then take both the image and generated English captions to generate low-resource language captions. We improve the captioning performance by adding the cycle consistency constraint on the cycle of image regions, English words, and low-resource language words. Moreover, our architecture has a flexible design which enables it to benefit from large monolingual English caption datasets. Experimental results demonstrate that our approach outperforms the state-of-the-art methods on common evaluation metrics. The attention visualization also shows that the proposed approach really improves the fine-grained alignment between words and image regions.
Yike Wu 0002, Shiwan Zhao, Jia Chen 0001, Ying Zhang 0015, Xiaojie Yuan, Zhong Su
ICME6
2019 Revealing Semantic Structures of Texts: Multi-grained Framework for Automatic Mind-map Generation
abstract
A mind-map is a diagram used to represent ideas linked to and arranged around a central concept. It’s easier to visually access the knowledge and ideas by converting a text to a mind-map. However, highlighting the semantic skeleton of an article remains a challenge. The key issue is to detect the relations amongst concepts beyond intra-sentence. In this paper, we propose a multi-grained framework for automatic mind-map generation. That is, a novel neural network is taken to detect the relations at first, which employs multi-hop self-attention and gated recurrence network to reveal the directed semantic relations via sentences. A recursive algorithm is then designed to select the most salient sentences to constitute the hierarchy. The human-like mind-map is automatically constructed with the key phrases in the salient sentences. Promising results have been achieved on the comparison with manual mind-maps. The case studies demonstrate that the generated mind-maps reveal the underlying semantic structures of the articles.
Jinmao Wei 0001, Zhong Su
IJCAI4
2016 BigNet 2016: First Workshop on Big Network Analytics
abstract
The first ACM international workshop on big network analytics is held in Indianapolis, Indiana, USA on October 24, 2016 and co-located with the ACM 25th Conference on Information and Knowledge Management (CIKM). The main objective of the workshop is to provide a forum for presenting the most recent advances in mining big networks to unearth rich knowledge. It is related to information retrieval, Web mining, social network analysis, and computational advertising. The anticipated outcome includes a fruitful discussion about the emerging challenges in this field, the development of novel theories for mining big networks, and motivating the interesting applications. The broader anticipated outcome includes: fostering future research directions, publishing high quality papers, attracting new researchers to this field, and concrete solutions to the existing problems.
Jie Tang 0001, Keke Cai, Zhong Su, Hanghang Tong, Michalis Vazirgiannis, Yang Yang 0009
CIKM3
2016 Boosting Recommendation in Unexplored Categories by User Price Preference
abstract
State-of-the-art methods for product recommendation encounter a significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this article, we investigate the challenging problem of product recommendation in unexplored categories and discover that the price, a factor comparable across categories, can improve the recommendation performance significantly. We introduce the price utility concept to characterize users’ sense of price and propose three different utility functions. We show that user price preference in a category is a distribution and we mine typical user price preference patterns based on three different types of distance between distributions. We fuse user price preference through regularization and joint factorization to boost recommendation performance in both browsing and buying shopping orientations. Experimental results show that fusing user price preference improves performance in a series of recommendation tasks: unexplored category recommendation, product recommendation under a given unexplored category, and product recommendation under generic unexplored categories.
Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001
ACM Trans. Inf. Syst.6
2015 Sampling Representative Users from Large Social Networks
abstract
Finding a subset of users to statistically represent the original social network is a fundamental issue in Social Network Analysis (SNA). The problem has not been extensively studied in existing literature. In this paper, we present a formal definition of the problem of \textbf{sampling representative users} from social network. We propose two sampling models and theoretically prove their NP-hardness. To efficiently solve the two models, we present an efficient algorithm with provable approximation guarantees. Experimental results on two datasets show that the proposed models for sampling representative users significantly outperform (+6\%-23\% in terms of Precision@100) several alternative methods using authority or structure information only. The proposed algorithms are also effective in terms of time complexity. Only a few seconds are needed to sampling ~300 representative users from a network of 100,000 users.All data and codes are publicly available.
Jie Tang 0001, Keke Cai, Li Zhang 0007, Zhong Su
AAAI5
2015 Lead curve detection in drawings with complex cross-points
Jia Chen 0001, Min Li 0022, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001
Neurocomputing5
2015 Sentence Compression for Aspect-Based Sentiment Analysis
abstract
Sentiment analysis, which addresses the computational treatment of opinion, sentiment, and subjectivity in text, has received considerable attention in recent years. In contrast to the traditional coarse-grained sentiment analysis tasks, such as document-level sentiment classification, we are interested in the fine-grained aspect-based sentiment analysis that aims to identify aspects that users comment on and these aspects' polarities. Aspect-based sentiment analysis relies heavily on syntactic features. However, the reviews that this task focuses on are natural and spontaneous, thus posing a challenge to syntactic parsers. In this paper, we address this problem by proposing a framework of adding a sentiment sentence compression (Sent_Comp) step before performing the aspect-based sentiment analysis. Different from the previous sentence compression model for common news sentences, Sent_Comp seeks to remove the sentiment-unnecessary information for sentiment analysis, thereby compressing a complicated sentiment sentence into one that is shorter and easier to parse. We apply a discriminative conditional random field model, with certain special features, to automatically compress sentiment sentences. Using the Chinese corpora of four product domains, Sent_Comp significantly improves the performance of the aspect-based sentiment analysis. The features proposed for Sent_Comp, especially the potential semantic features, are useful for sentiment sentence compression.
Wanxiang Che, Zhong Su, Ting Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Exploitation and Exploration Balanced Hierarchical Summary for Landmark Images
abstract
While we have made significant progress over image understanding and search, how to meet the ultimate goal of satisfying both exploration and exploitation in one single system is still an open challenge. In the context of landmark images, it means that a system should not only be able to help users to quickly locate the photo they are interested in (exploitation), but also to discover different parts of the landmark which have never been seen before (exploration), which is a common request as evidenced by many recent multimedia studies. To the best of our knowledge, existing systems mainly focus on either exploration (e.g., photo browsing) or exploitation (e.g., representative photo identification), while users' need of exploration and exploitation is dynamically mixed. In this paper, we tackle the challenge by organizing landmark images into a hierarchical summary which gives user the flexibility of conducting both exploration and exploitation. In the hierarchical summary construction, we introduce two principles: the coherence principle and the diversity principle. Behind these two principles, the intrinsic concept is “detail-level,” which measures how much detail that an image reflects for a certain landmark. A new objective function is derived from the definition of both exploration and exploitation experience on detail-level. The problem of finding an optimal hierarchical summary is formulated as searching over a space of trees for the one that achieves the best objective score. Extensive quantitative experimental results and comprehensive user studies show that the optimized hierarchical summary is able to satisfy both experiences simultaneously.
Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Shimin Chen, Yong Yu 0001
IEEE Trans. Multim.4
2014 Sentence Compression for Target-Polarity Word Collocation Extraction
Wanxiang Che, Bing Qin 0001, Zhong Su, Ting Liu 0001
COLING5
2014 Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help
abstract
State-of-the-art methods for product recommendation encounter significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this paper, we investigate the challenge problem of product recommendation in unexplored categories and discover that the price, a factor transferrable across categories, can improve the recommendation performance significantly. Through our investigation, we address four research questions progressively: 1) what is the impact of unexplored category on recommendation performance? 2) How to represent the price factor from the recommendation point of view? 3) What does price factor across categories mean to recommendation? 4) How to utilize price factor across categories for recommendation in unexplored categories? Based on a series of experiments and analysis conducted on a dataset collected from a leading E-commerce website, we discover valuable findings for the above four questions: first, unexplored categories cause performance drop by 40% relatively for current recommendation systems; second, the price factor can be represented as either a quantity for a product or a distribution for a user to improve performance; third, consumer behavior with respect to price factor across categories is complicated and needs to be carefully modeled; finally and most importantly, we propose a new method which encodes the two perspectives of the price factor. The proposed method significantly improves the recommendation performance in unexplored categories over the state-of-the-art baseline systems and shortens the performance gap by 43% relatively.
Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001
SIGIR6
2014 Probabilistic text modeling with orthogonalized topics
abstract
Topic models have been widely used for text analysis. Previous topic models have enjoyed great success in mining the latent topic structure of text documents. With many efforts made on endowing the resulting document-topic distributions with different motivations, however, none of these models have paid any attention on the resulting topic-word distributions.Since topic-word distribution also plays an important role in the modeling performance,topic models which emphasize only the resulting document-topic representations but pay less attention to the topic-term distributions are limited. In this paper, we propose the Orthogonalized Topic Model(OTM) which imposes an orthogonality constraint on the topic-term distributions. We also propose a novel model fitting algorithm based on the generalized Expectation-Maximization algorithm and the Newthon-Raphson method. Quantitative evaluation of text classification demonstrates that OTM outperforms other baseline models and indicates the important role played by topic orthogonalizing.
Enpeng Yao, Guoqing Zheng, Ou Jin, Shenghua Bao, Kailong Chen, Zhong Su, Yong Yu 0001
SIGIR6
2014 How Do People Communicate through Different Social Connections?
Keke Cai, Jie Tang 0001, Li Zhang 0007, Zhong Su
WAIM5
2013 Head-shoulder based gender recognition
abstract
This paper proposes a novel gender recognition method based on the head-shoulder part of human body. The head-shoulder area contains much information that could be cues to infer the gender of a person, such as hair-style, face, neckline style and so on. A rich high-dimensional feature descriptor is designed to extract gradient, texture and orientation information from the head-shoulder area, then Partial Least Squares (PLS) is employed to learn a very low dimensional discriminative subspace. Features are projected into the low dimensional subspace and linear SVM is employed to learn an efficient classification model between the male and female categories. Experimental results on a large real-world dataset demonstrate the effectiveness of the proposed method.
Min Li 0022, Shenghua Bao, Weishan Dong, Yu Wang 0021, Zhong Su
ICIP5
2013 Tell me what happened here in history
abstract
This demo shows our system that takes a landmark image as input, recognizes the landmark from the image and returns historical events of the landmark with related photos. Different from existing landmark related researches, we focus on the temporal dimension of a landmark. Our system automatically recognizes the landmark, shows historical events chronologically and provides detailed photos for the events. To build these functions, we fuse information from multiple online resources.
Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001
ACM Multimedia5
2012 DLMSearch: diversified landmark search by photo
abstract
This paper focuses on the problem of searching for diversified landmarks with photos. More particularly, we propose a system called DLMSearch which handles image query, searches for diversified landmarks and provides representative visual summaries. DLMSearch allows a user to upload a query photo and searches for landmarks with high relevance and diversity in real time. Then DLMSearch presents a delicate photo summary for each returned landmark, considering both visual representativeness and diversity. Quantative evaluations on a web-scale landmark photo collection demonstrate the effectiveness of the DLMSearch system. Experimental results verify the merits of the proposed system.
Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001
ACM Multimedia6
2012 Searching for diversified landmarks by photo
abstract
This demo focuses on the problem of searching for diversified landmarks with photos as input. More particularly, we propose a system called DLMSearch that allows a user to upload a photo as a query and searches for a diverse set of relevant landmarks in real time. It also presents a photo summary for each retrieved landmark, considering both visual representativeness and diversity. Our online demo is available at http://lm.apexlab.org/landmark/demo.
Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001
ACM Multimedia6
2012 Introduction to the Special Section on Computational Models of Collective Intelligence in the Social Web
abstract
No abstract available.
Evgeniy Gabrilovich, Zhong Su, Jie Tang 0001
ACM Trans. Intell. Syst. Technol.2
2012 Mining Social Emotions from Affective Text
abstract
This paper is concerned with the problem of mining social emotions from text. Recently, with the fast development of web 2.0, more and more documents are assigned by social users with emotion labels such as happiness, sadness, and surprise. Such emotions can provide a new aspect for document categorization, and therefore help online users to select related documents based on their emotional preferences. Useful as it is, the ratio with manual emotion labels is still very tiny comparing to the huge amount of web/enterprise documents. In this paper, we aim to discover the connections between social emotions and affective terms and based on which predict the social emotion from text content automatically. More specifically, we propose a joint emotion-topic model by augmenting Latent Dirichlet Allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model.
Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001
IEEE Trans. Knowl. Data Eng.5
2011 Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services
abstract
This paper focuses on analyzing and predicting not-answered questions in Community based Question Answering (CQA) services, such as Yahoo! Answers. In CQA services, users express their information needs by submitting natural language questions and await answers from other human users. Comparing to receiving results from web search engines using keyword queries, CQA users are likely to get more specific answers, because human answerers may catch the main point of the question. However, one of the key problems of this pattern is that sometimes no one helps to give answers, while web search engines hardly fail to response. In this paper, we analyze the not-answered questions and give a first try of predicting whether questions will receive answers. More specifically, we first analyze the questions of Yahoo Answers based on the features selected from different perspectives. Then, we formalize the prediction problem as supervised learning – binary classification problem and leverage the proposed features to make predictions. Extensive experiments are made on 76,251 questions collected from Yahoo! Answers. We analyze the specific characteristics of not-answered questions and try to suggest possible reasons for why a question is not likely to be answered. As for prediction, the experimental results show that classification based on the proposed features outperforms the simple word-based approach significantly.
Lichun Yang, Shenghua Bao, Qingliang Lin, Xian Wu 0001, Dingyi Han, Zhong Su, Yong Yu 0001
AAAI6
2011 Domain customization for aspect-oriented opinion analysis with multi-level latent sentiment clues
abstract
Aspect-oriented opinion mining detects the reviewers' sentiment orientation (e.g. positive, negative or neutral) towards different product-features. Domain customization is a big challenge for opinion mining due to the accuracy loss across domains. In this paper, we show our experiences and lessons learned in the domain customization for the aspect-oriented opinion analysis system OpinionIt. We present a customization method for sentiment classification with multi-level latent sentiment clues. We first construct Latent Semantic Association model to capture latent association among product-features from the unlabeled corpus. Meanwhile, we present an unsupervised method to effectively extract various domain-specific sentiment clues from the unlabeled corpus. In the customization, we tune the sentiment classifier on the labeled source domain data by incorporating the multi-level latent sentiment clues (e.g. latent association among product-features, domain-specific and generic sentiment clues). Experimental results show that the proposed method significantly reduces the accuracy loss of sentiment classification without any labeled target domain data.
Huijia Zhu, Zhili Guo, Zhong Su
CIKM4
2011 Visualizing anomalies in sensor networks
abstract
Diagnosing a large-scale sensor network is a crucial but challenging task due to the spatiotemporally dynamic network behaviors of sensor nodes. In this demo, we present Sensor Anomaly Visualization Engine (SAVE), an integrated system that tackles the sensor network diagnosis problem using both visualization and anomaly detection analytics to guide the user quickly and accurately diagnose sensor network failures. Temporal expansion model, correlation graphs and dynamic projection views are proposed to effectively interpret the topological, correlational and dimensional sensor data dynamics and their anomalies. Through a real-world large-scale wireless sensor network deployment (GreenOrbs), we demonstrate that SAVE is able to help better locate the problem and further identify the root cause of major sensor network failures.
Qi Liao 0002, Lei Shi 0002, Yuan He 0004, Rui Li 0047, Zhong Su, Aaron Striegel, Yunhao Liu 0001
SIGCOMM5
2011 Social context summarization
abstract
We study a novel problem of social context summarization for Web documents. Traditional summarization research has focused on extracting informative sentences from standard documents. With the rapid growth of online social networks, abundant user generated content (e.g., comments) associated with the standard documents is available. Which parts in a document are social users really caring about? How can we generate summaries for standard documents by considering both the informativeness of sentences and interests of social users? This paper explores such an approach by modeling Web documents and social contexts into a unified framework. We propose a dual wing factor graph (DWFG) model, which utilizes the mutual reinforcement between Web documents and their associated social contexts to generate summaries. An efficient algorithm is designed to learn the proposed factor graph model.Experimental results on a Twitter data set validate the effectiveness of the proposed model. By leveraging the social context information, our approach obtains significant improvement (averagely +5.0%-17.3%) over several alternative methods (CRF, SVM, LR, PR, and DocLead) on the performance of summarization.
Keke Cai, Jie Tang 0001, Li Zhang 0007, Zhong Su, Juan-Zi Li
SIGIR5
2011 Mining topics on participations for community discovery
abstract
Community discovery on large-scale linked document corpora has been a hot research topic for decades. There are two types of links. The first one, which we call d2d-link, indicates connectiveness among different documents, such as blog references and research paper citations. The other one, which we call u2u-link, represents co-occurrences or simultaneous participations of different users in one document and typically each document from u2u-link corpus has more than one user/author. Examples of u2u-link data covers email archives and research paper co-authorship networks. Community discovery in d2d-link data has achieved much success, while methods for that in u2u-link data either make no use of the textual content of the documents or make oversimplified assumptions about the users and the textual content. In this paper we propose a general approach of community discovery for u2u-link data, i.e., multiple user data, by placing topical variables on multiple authors' participations in documents. Experiments on a research proceeding co-authorship corpus and a New York Times news corpus show the effectiveness of our model.
Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001
SIGIR6
2011 A Classification Framework for Disambiguating Web People Search Result Using Feedback
Ou Jin, Shenghua Bao, Zhong Su, Yong Yu 0001
WAIM3
2011 Finding Appropriate Experts for Collaboration
Zhenjiang Zhan, Lichun Yang, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001
WAIM5
2011 OOLAM: an opinion oriented link analysis model for influence persona discovery
abstract
Social influence is a complex and subtle force that governs the dynamics of social networks. In the past years, a lot of research work has been conducted to understand the spread patterns of social influence. However, most of approaches assume that influence exists between users with active social interactions, but ignore the question of what kind of influence happens between them. As such one interesting and also fundamental question is raised here: "in a social network, could the social connection reflect users'influence from both positive and negative aspects?". To this end, an Opinion Oriented Link Analysis Model (OOLAM) is proposed in this paper to characterize users' influence personae in order to exhibit their distinguishing influence ability in the social network. In particular, three types of influence personae are generalized and the problem of influence persona discovery is formally defined. Within the OOLAM model, two factors, i.e., opinion consistency and opinion creditability, are defined to capture the persona information from public opinion perspective. Extensive experimental studies have been performed to demonstrate the effectiveness of the proposed approach on influence persona analysis using real web data sets.
Keke Cai, Shenghua Bao, Jie Tang 0001, Li Zhang 0007, Zhong Su
WSDM7
2011 Topic level expertise search over heterogeneous networks
Jie Tang 0001, Jing Zhang 0001, Ruoming Jin, Keke Cai, Li Zhang 0007, Zhong Su
Mach. Learn.7
2010 OpinionIt: a text mining system for cross-lingual opinion analysis
abstract
Opinion mining focuses on extracting customers' opinions from the reviews and predicting their sentiment orientation. Reviewers usually praise a product in some aspects and bemoan it in other aspects. With the business globalization, it is very important for enterprises to extract the opinions toward different aspects and find out cross-lingual/cross-culture difference in opinions. Cross-lingual opinion mining is a very challenging task as amounts of opinions are written in different languages, and not well structured. Since people usually use different words to describe the same aspect in the reviews, product-feature (PF) categorization becomes very critical in cross-lingual opinion mining. Manual cross-lingual PF categorization is time consuming, and practically infeasible for the massive amount of data written in different languages. In order to effectively find out cross-lingual difference in opinions, we present an aspect-oriented opinion mining method with Cross-lingual Latent Semantic Association (CLaSA). We first construct CLaSA model to learn the cross-lingual latent semantic association among all the PFs from multi-dimension semantic clues in the review corpus. Then we employ CLaSA model to categorize all the multilingual PFs into semantic aspects, and summarize cross-lingual difference in opinions towards different aspects. Experimental results show that our method achieves better performance compared with the existing approaches. With CLaSA model, our text mining system OpinionIt can effectively discover cross-lingual difference in opinions.
Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su
CIKM5
2010 Understanding retweeting behaviors in social networks
abstract
Retweeting is an important action (behavior) on Twitter, indicating the behavior that users re-post microblogs of their friends. While much work has been conducted for mining textual content that users generate or analyzing the social network structure, few publications systematically study the underlying mechanism of the retweeting behaviors. In this paper, we perform an interesting analysis for the problem on Twitter. We have found that almost 25.5% of the tweets posted by users are actually retweeted from friends' blog spaces. Our investigation unveils that for the retweet behaviors, some statistics still follows the power law distribution, while some others violate the state-of-the-art distribution for Web. Based on these important observations, we propose a factor graph model to predict users' retweeting behaviors. Experimental results on the Twitter data set show that our method can achieve a precision of 28.81% and recall of 37.33% for prediction of the retweet behaviors.
Jingyi Guo, Keke Cai, Jie Tang 0001, Juan-Zi Li, Li Zhang 0007, Zhong Su
CIKM7
2010 CasJoin: a cascade chain for text similarity joins
abstract
We are concerned with the problem of similarity joins of text data, where the task is to find all pairs of documents above an expected similarity. Such a problem often serves as an indispensable step in many web applications. A crucial issue is to preclude unnecessary candidate pairs as many as possible ahead of expensive similarity evaluation. In this paper, we initiate an idea of adopting a cascade structure in text joins for a large speedup, where a latter stage can exclude a considerable number of invalid pairs survived in former stages. The proposed algorithm is shortly referred to as CasJoin. We further adopt a prefix filter to build the stage of CasJoin by introducing a novel vision to the dynamic generation of document vector. Specifically, a vector is partitioned into a chain of multiple prefixes that are appended one by one for cascade joining. We evaluate our CasJoin on a typical web corpus, ODP. Experiments indicate that, comparing to the state-of-the-art prefix algorithms, CasJoin can achieve a drastic reduction of candidates by as much as 98.15% and a dramatic speedup of joining by up to 13.34x.
Xiaoxun Zhang, Zhili Guo, Huijia Zhu, Zhong Su
CIKM5
2010 A topical link model for community discovery in textual interaction graph
abstract
This paper is concerned with community discovery in textual interaction graph, where the links between entities are indicated by textual documents. Specifically, we propose a Topical Link Model(TLM), which leverages Hierarchical Dirichlet Process(HDP) to introduce hidden topical variable of the links. Other than the use of links, TLM can look into the documents on the links in detail to recover sound communities. Moreover, TLM is a nonparametric model, which is able to learn the number of communities from the data. Extensive experiments on two real world corpora show TLM outperforms two state-of-the-art baseline models, which verify the effectiveness of TLM in determining the proper number of communities and generating sound communities.
Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001
CIKM6
2009 Product feature categorization with multilevel latent semantic association
abstract
In recent years, the number of freely available online reviews is increasing at a high speed. Aspect-based opinion mining technique has been employed to find out reviewers' opinions toward different product aspects. Such finer-grained opinion mining is valuable for the potential customers to make their purchase decisions. Product-feature extraction and categorization is very important for better mining aspect-oriented opinions. Since people usually use different words to describe the same aspect in the reviews, product-feature extraction and categorization becomes more challenging. Manually product-feature extraction and categorization is tedious and time consuming, and practically infeasible for the massive amount of products. In this paper, we propose an unsupervised product-feature categorization method with multilevel latent semantic association. After extracting product-features from the semi-structured reviews, we construct the first latent semantic association (LaSA) model to group words into a set of concepts according to their virtual context documents. It generates the latent semantic structure for each product-feature. The second LaSA model is constructed to categorize the product-features according to their latent semantic structures and context snippets in the reviews. Experimental results demonstrate that our method achieves better performance compared with the existing approaches. Moreover, the proposed method is language- and domain-independent.
Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su
CIKM5
2009 A study of information retrieval on accumulative social descriptions using the generation features
abstract
This paper is concerned with the study of information retrieval (IR) on Accumulative Social Descriptions (ASDs). ASDs refer to Web texts that accumulated by many Web users describing certain Web resources, such as anchor texts, search logs and social annotations. There have been some studies working on leveraging ASDs for improving search performance in both internet and intranet. However, to the best of our knowledge, no prior study has concerned the specific generation features of ASDs, which are the focus point of this paper. Specifically, we consider the generation features from two perspectives, the generation processes and the generated distributions. Further, three probabilistic IR models are derived based on them. The three models are first demonstrated with one toy dataset and then empirically evaluated with two real datasets: an internet dataset consisting of 90,295 Web pages, with 25,845,818 social annotations crawled from Del.icio.us and 31,320,005 pieces of anchor texts crawled through Yahoo! API, and an intranet dataset consisting of 179,835 Web pages with 1,245,522 annotations dumped from the intranet tagging system in IBM, named as Dogear. Extensive experimental results show that the proposed methods, which fully leverage the generation features of ASDs, improve the performance of both internet and intranet search significantly.
Lichun Yang, Shengliang Xu, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001
CIKM5
2009 sDoc: exploring social wisdom for document enhancement in web mining
abstract
Web document could be seen to be composed of textual content as well as social metadata of various forms (e.g., anchor text, search query and social annotation), both of which are valuable to indicate the semantic content of the document. However, due to the free nature of the web, the two streams of web data suffer from the serious problems of noise and sparseness, which have actually become the major challenges to the success of many web mining applications. Previous work has shown that it could enhance the content of web document by integrating anchor text and search query. In this paper, we study the problem of exploring emergent social annotation for document enhancement and propose a novel reinforcement framework to generate "social representation" of document. Distinguishing from prior work, textual content and social annotation are enhanced simultaneously in our framework, which is achieved by exploiting a kind of mutual reinforcement relationship behind them. Two convergent models, social content model and social annotation model, are symmetrically derived from the framework to represent enhanced textual content and enhanced social annotation respectively. The enhanced document is referred to as Social Document or sDoc in that it could embed complementary viewpoints from many web authors and many web visitors. In this sense, the document semantics is enhanced exactly by exploring social wisdom. We build the framework on a large Del.icio.us data and evaluate it through three typical web mining applications: annotation, classification and retrieval. Experimental results demonstrate that social representation of web document could boost the performance of these applications significantly.
Xiaoxun Zhang, Lichun Yang, Xian Wu 0001, Zhili Guo, Shenghua Bao, Yong Yu 0001, Zhong Su
CIKM8
2009 Joint Emotion-Topic Modeling for Social Affective Text Mining
abstract
This paper is concerned with the problem of social affective text mining, which aims to discover the connections between social emotions and affective terms based on user-generated emotion labels. We propose a joint emotion-topic model by augmenting latent Dirichlet allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model.
Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001
ICDM5
2009 Topic Distributions over Links on Web
abstract
It is well known that Web users create links with different intentions. However, a key question, which is not well studied, is how to categorize the links and how to quantify the strength of the influence of a Web page on another if there is a link between the two linked Web pages. In this paper, we focus on the problem of link semantics analysis, and propose a novel supervised learning approach to build a model, based on a training link-labeled and link-weighted graph where a link-label represents the category of a link and a link-weight represents the influence of one web page on the other in a link. Based on the model built, we categorize links and quantify the influence of Web pages on the others in a large graph in the same application domain. We discuss our proposed approach, namely pairwise restricted Boltzmann machines (PRBMs), and conduct extensive experimental studies to demonstrate the effectiveness of our approach using large real datasets.
Jie Tang 0001, Jing Zhang 0001, Jeffrey Xu Yu, Keke Cai, Li Zhang 0007, Zhong Su
ICDM8
2009 Address standardization with latent semantic association
abstract
Address standardization is a very challenging task in data cleansing. To provide better customer relationship management and business intelligence for customer-oriented cooperates, millions of free-text addresses need to be converted to a standard format for data integration, de-duplication and householding. Existing commercial tools usually employ lots of hand-craft, domain-specific rules and reference data dictionary of cities, states etc. These rules work better for the region they are designed. However, rule-based methods usually require more human efforts to rewrite these rules for each new domain since address data are very irregular and varied with countries and regions. Supervised learning methods usually are more adaptable than rule-based approaches. However, supervised methods need large-scale labeled training data. It is a labor-intensive and time-consuming task to build a large-scale annotated corpus for each target domain. For minimizing human efforts and the size of labeled training data set, we present a free-text address standardization method with latent semantic association (LaSA). LaSA model is constructed to capture latent semantic association among words from the unlabeled corpus. The original term space of the target domain is projected to a concept space using LaSA model at first, then the address standardization model is active learned from LaSA features and informative samples. The proposed method effectively captures the data distribution of the domain. Experimental results on large-scale English and Chinese corpus show that the proposed method significantly enhances the performance of standardization with less efforts and training data.
Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su
KDD5
2009 Domain Adaptation with Latent Semantic Association for Named Entity Recognition
Huijia Zhu, Zhili Guo, Xiaoxun Zhang, Zhong Su
HLT-NAACL6
2009 Social Propagation: Boosting Social Annotations for Web Mining
Shenghua Bao, Bohai Yang, Ben Fei, Shengliang Xu, Zhong Su, Yong Yu 0001
World Wide Web5
2008 Boosting social annotations using propagation
abstract
This paper is concerned with the problem of boosting social annotations using propagation, which is also called social propagation. In particular, we focus on propagating social annotations of web pages (e.g., annotations in Del.icio.us). Although social annotations are developing fast, they cover only a small proportion of Web pages on the World Wide Web. To alleviate the low coverage problem, a general propagation model based on Random Surfer is proposed. Specifically, four steps are included: basic propagation, multiple-annotation propagation, multiple-link-type propagation, and constraint-guided propagation. Experimental results show that the proposed model is very effective in increasing coverage of annotations as well as preserving property of social annotations.
Shenghua Bao, Bohai Yang, Ben Fei, Shengliang Xu, Zhong Su, Yong Yu 0001
CIKM5
2008 ArnetMiner: extraction and mining of academic social networks
abstract
This paper addresses several key issues in the ArnetMiner system, which aims at extracting and mining academic social networks. Specifically, the system focuses on: 1) Extracting researcher profiles automatically from the Web; 2) Integrating the publication data into the network from existing digital libraries; 3) Modeling the entire academic network; and 4) Providing search services for the academic network. So far, 448,470 researcher profiles have been extracted using a unified tagging approach. We integrate publications from online Web databases and propose a probabilistic framework to deal with the name ambiguity problem. Furthermore, we propose a unified modeling approach to simultaneously model topical aspects of papers, authors, and publication venues. Search services such as expertise search and people association search have been provided based on the modeling results. In this paper, we describe the architecture and main features of the system. We also present the empirical evaluation of the proposed methods.
Jie Tang 0001, Jing Zhang 0001, Limin Yao, Juan-Zi Li, Li Zhang 0007, Zhong Su
KDD6
2008 Exploring folksonomy for personalized search
abstract
As a social service in Web 2.0, folksonomy provides the users the ability to save and organize their bookmarks online with "social annotations" or "tags". Social annotations are high quality descriptors of the web pages' topics as well as good indicators of web users' interests. We propose a personalized search framework to utilize folksonomy for personalized search. Specifically, three properties of folksonomy, namely the categorization, keyword, and structure property, are explored. In the framework, the rank of a web page is decided not only by the term matching between the query and the web page's content but also by the topic matching between the user's interests and the web page's topics. In the evaluation, we propose an automatic evaluation framework based on folksonomy data, which is able to help lighten the common high cost in personalized search evaluations. A series of experiments are conducted using two heterogeneous data sets, one crawled from Del.icio.us and the other from Dogear. Extensive experimental results show that our personalized search approach can significantly improve the search quality.
Shengliang Xu, Shenghua Bao, Ben Fei, Zhong Su, Yong Yu 0001
SIGIR4
2008 Hidden sentiment association in chinese web opinion mining
abstract
The boom of product review websites, blogs and forums on the web has attracted many research efforts on opinion mining. Recently, there was a growing interest in the finer-grained opinion mining, which detects opinions on different review features as opposed to the whole review level. The researches on feature-level opinion mining mainly rely on identifying the explicit relatedness between product feature words and opinion words in reviews. However, the sentiment relatedness between the two objects is usually complicated. For many cases, product feature words are implied by the opinion words in reviews. The detection of such hidden sentiment association is still a big challenge in opinion mining. Especially, it is an even harder task of feature-level opinion mining on Chinese reviews due to the nature of Chinese language. In this paper, we propose a novel mutual reinforcement approach to deal with the feature-level opinion mining problem. More specially, 1) the approach clusters product features and opinion words simultaneously and iteratively by fusing both their content information and sentiment link information. 2) under the same framework, based on the product feature categories and opinion word groups, we construct the sentiment association set between the two groups of data objects by identifying their strongest n sentiment links. Moreover, knowledge from multi-source is incorporated to enhance clustering in the procedure. Based on the pre-constructed association set, our approach can largely predict opinions relating to different product features, even for the case without the explicit appearance of product feature words in reviews. Thus it provides a more accurate opinion evaluation. The experimental results demonstrate that our method outperforms the state-of-art algorithms.
Qi Su 0001, Xinying Xu, Zhili Guo, Xian Wu 0001, Xiaoxun Zhang, Bing Swen, Zhong Su
WWW8
2008 Floatcascade learning for fast imbalanced web mining
abstract
This paper is concerned with the problem of Imbalanced Classification (IC) in web mining, which often arises on the web due to the "Matthew Effect". As web IC applications usually need to provide online service for user and deal with large volume of data, classification speed emerges as an important issue to be addressed. In face detection, Asymmetric Cascade is used to speed up imbalanced classification by building a cascade structure of simple classifiers, but it often causes a loss of classification accuracy due to the iterative feature addition in its learning procedure. In this paper, we adopt the idea of cascade classifier in imbalanced web mining for fast classification and propose a novel asymmetric cascade learning method called FloatCascade to improve the accuracy. To the end, FloatCascade selects fewer yet more effective features at each stage of the cascade classifier. In addition, a decision-tree scheme is adopted to enhance feature diversity and discrimination capability for FloatCascade learning. We evaluate FloatCascade through two typical IC applications in web mining: web page categorization and citation matching. Experimental results demonstrate the effectiveness and efficiency of FloatCascade comparing to the state-of-the-art IC methods like Asymmetric Cascade, Asymmetric AdaBoost and Weighted SVM.
Xiaoxun Zhang, Zhili Guo, Xian Wu 0001, Zhong Su
WWW6
2007 Optimizing web search using social annotations
abstract
This paper explores the use of social annotations to improve web search. Nowadays, many services, e.g. del.icio.us, have been developed for web users to organize and share their favorite web pages on line by using social annotations. We observe that the social annotations can benefit web search in two aspects: 1) the annotations are usually good summaries of corresponding web pages; 2) the count of annotations indicates the popularity of web pages. Two novel algorithms are proposed to incorporate the above information into page ranking: 1) SocialSimRank (SSR) calculates the similarity between social annotations and web queries; 2) SocialPageRank (SPR) captures the popularity of web pages. Preliminary experimental results show that SSR can find the latent semantic association between queries and annotations, while SPR successfully measures the quality (popularity) of a web page from the web users ’ perspective. We further evaluate the proposed methods empirically with 50 manually constructed queries and 3000 auto-generated queries on a dataset crawled from del.icio.us. Experiments show that both SSR and SPR benefit web search significantly.
Shenghua Bao, Gui-Rong Xue, Xiaoyuan Wu, Yong Yu 0001, Ben Fei, Zhong Su
WWW6
2007 Towards effective browsing of large scale social annotations
abstract
This paper is concerned with the problem of browsing social annotations. Today, a lot of services (e.g., Del.icio.us, Filckr) have been provided for helping users to manage and share their favorite URLs and photos based on social annotations. Due to the exponential increasing of the social annotations, more and more users, however, are facing the problem how to effectively find desired resources from large annotation data. Existing methods such as tag cloud and annotation matching work well only on small annotation sets. Thus, an effective approach for browsing large scale annotation sets and the associated resources is in great demand by both ordinary users and service providers. In this paper, we propose a novel algorithm, namely Effective Large Scale Annotation Browser (ELSABer), to browse large-scale social annotation data. ELSABer helps the users browse huge number of annotations in a semantic, hierarchical and efficient way. More specifically, ELSABer has the following features: 1) the semantic relations between annotations are explored for browsing of similar resources; 2) the hierarchical relations between annotations are constructed for browsing in a top-down fashion; 3) the distribution of social annotations is studied for efficient browsing. By incorporating the personal and time information, ELSABer can be further extended for personalized and time-related browsing. A prototype system is implemented and shows promising results.
Rui Li 0049, Shenghua Bao, Yong Yu 0001, Ben Fei, Zhong Su
WWW5
2006 Empirical Study on the Performance Stability of Named Entity Recognition Model across Domains
Li Zhang 0007, Zhong Su
EMNLP3
2004 RStar: an RDF storage and query system for enterprise resource management
abstract
Modern corporations operate in an extremely complex environment and strongly depend on all kinds of information resources across the enterprise. Unfortunately, with the growth of an enterprise, its information resources are not only heterogeneous but also distributed in physically different systems and databases. How to effectively exploit information across the enterprise is becoming a critical but hard problem. In recent years, metadata which is the detailed description of the data is used to efficiently exploit information resources in the web. The World Wide Web Consortium (W3C) recommends the resource description framework (RDF) as a standard for the definition and use of metadata descriptions of resources in the web. In this paper, we present an RDF storage and query system called RStar for enterprise resource management. RStar uses a relational database as the persistent data store and defines RStar Query Language (RSQL) for resource retrieval. Currently, most of existing RDF storage and query systems are evaluated on small data sets and no detailed performance analysis is given for such systems. Therefore, we conduct extensive experiments on a large scale data set to investigate the performance problem in RDF storage. Such analysis will be helpful for designing RDF storage and query systems as well as for understanding not well-solved issues in RDF based enterprise resource management. In addition, experiences and lessons learned in our implementation are presented for further research and development.
Li Ma 0002, Zhong Su, Li Zhang 0007
CIKM2
2004 ORIENT: Integrate Ontology Engineering into Industry Tooling Environment
Lei Zhang 0007, Yong Yu 0001, Kewei Tu, MingChuan Guo, Guo Tong Xie, Zhong Su
ISWC9
2003 Document Clustering Based on Vector Quantization and Growing-Cell Structure
Zhong Su, Li Zhang 0007
IEA/AIE1
2003 Relevance feedback in content-based image retrieval: Bayesian framework, feature subspaces, and progressive learning
abstract
Research has been devoted in the past few years to relevance feedback as an effective solution to improve performance of content-based image retrieval (CBIR). In this paper, we propose a new feedback approach with progressive learning capability combined with a novel method for the feature subspace extraction. The proposed approach is based on a Bayesian classifier and treats positive and negative feedback examples with different strategies. Positive examples are used to estimate a Gaussian distribution that represents the desired images for a given query; while the negative examples are used to modify the ranking of the retrieved candidates. In addition, feature subspace is extracted and updated during the feedback process using a principal component analysis (PCA) technique and based on user's feedback. That is, in addition to reducing the dimensionality of feature spaces, a proper subspace for each type of features is obtained in the feedback process to further improve the retrieval accuracy. Experiments demonstrate that the proposed method increases the retrieval speed, reduces the required memory and improves the retrieval accuracy significantly.
Zhong Su, HongJiang Zhang, Stan Z. Li, Shaoping Ma
IEEE Trans. Image Process.1
2003 Relevance Feedback and Learning in Content-Based Image Search
HongJiang Zhang, Zheng Chen 0001, Mingjing Li, Zhong Su
World Wide Web4
2002 Correlation-Based Web Document Clustering for Adaptive Web Interface Design
Zhong Su, Qiang Yang 0001, HongJiang Zhang, Xiaowei Xu 0001, Yu Hen Hu, Shaoping Ma
Knowl. Inf. Syst.1
2001 Extraction of feature subspaces for content-based retrieval using relevance feedback
abstract
In the past few years, relevance feedback (RF) has been used as an effective solution for content-based image retrieval (CBIR). Although effective, the RF-CBIR framework does not address the issue of feature extraction for dimension reduction and noise reduction. In this paper, we propose a novel method for extracting features for the class of images represented by the positive images provided by subjective RF. Principal Component Analysis (PCA) is used to reduce both noise contained in the original image features and dimensionality of feature spaces. The method increases the retrieval speed and reduces the memory significantly without sacrificing the retrieval accuracy.
Zhong Su, Stan Z. Li, HongJiang Zhang
ACM Multimedia1
2000 Using Bayesian classifier in relevant feedback of image retrieval
abstract
Relevance feedback is a powerful technique in content-based image retrieval (CBIR) and has been an active research area for the past few years. In this paper, we propose a new relevance feedback approach based on a Bayesian classifier, and it treats positive and negative feedback examples with different strategies. For positive examples, a Bayesian classifier is used to determine the distribution of the query space. A 'dibbling' process is applied to penalize images that are near the negative examples in the query and retrieval refinement process. The proposed algorithm also has a progressive learning capability that utilizes past feedback information to help the current query. Experimental results show that our algorithm is effective.
Zhong Su, HongJiang Zhang, Shaoping Ma
ICTAI1
2000 A prediction system for multimedia pre-fetching in Internet
abstract
The rapid development of Internet has resulted in more and more multimedia in Web content. However, due to the limitation in the bandwidth and huge size of the multimedia data, users always suffer from long time waiting. On the other hand, if we can predict the web object or page that the user most likely will view next while the user is viewing the current page, and pre-fetch the content, then the perceived network latency can be significantly reduced. In this paper, we present an n-gram based model to utilize path profiles of users from very large web log to predict the users' future requests. Our model is based on a simple extension of existing point-based models for such predictions, but our results show that by sacrificing the applicability somewhat one can gain a great deal in prediction precision. Also we present an efficient method to compress the prediction model size so that it can be fitted into the main memory. Our result can potentially be applied to a wide range of applications on the web, including pro-fetching, enhancement of recommendation systems as well as web caching policies. The experiments based on three realistic web logs have proved the effectiveness of the proposed scheme.
Zhong Su, Qiang Yang 0001, HongJiang Zhang
ACM Multimedia1
2000 WhatNext: A Prediction System for Web Requests Using N-gram Sequence Models
abstract
As an increasing number of users access information on the Web, there is a great opportunity to learn from the server logs to learn about the users' probable actions in the future. We present an n-gram based model to utilize path profiles of users from very large data sets to predict the users' future requests. Since this is a prediction system, we cannot measure the recall in a traditional sense. We, therefore, present the notion of applicability to give a measure of the ability to predict the next document. Our model is based on a simple extension of existing point-based models for such predictions, but our results show for n-gram based prediction when n is greater than three, we can increase precision by 20% or more for two realistic Web logs. Also we present an efficient method that can compress our model to 30% of its original size so that the model can be loaded in main memory. Our result can potentially be applied to a wide range of applications on the Web, including pre-sending, pre-fetching, enhancement of recommendation systems as well as Web caching policies. Our tests are based on three realistic Web logs. Our algorithm is implemented in a prediction system called WhatNext, which shows a marked improvement in precision and applicability over previous approaches.
Zhong Su, Qiang Yang 0001, HongJiang Zhang
WISE1