VLDB 2026 Research / reviewers in the wild / expert
Bin Wu 0001
dblp:98/4432-1
· DBLP profile ↗
104ranked-venue papers in the field
4as first author
25since 2021 · last 2026
0000-0002-7112-126XORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 50 (2 first)Information Retrieval & Web Search · 21Database Systems & Data Management · 10Knowledge Engineering, Semantic Web & Information Systems · 9Big Data, Cloud & Distributed Data Systems · 8Other / Interdisciplinary · 6 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DenseSpeech: Dense Multi-Segment Temporal Grounding in Public Speaking VideosabstractTemporal video grounding aims to localize temporal segments in untrimmed videos based on natural language queries. While dense grounding methods accept paragraph queries, they retain a restrictive one-to-one mapping between sentences and segments. This assumption fails in real-world behavioral analysis like public speaking, where a single query often corresponds to multiple, scattered, and overlapping segments due to event repetition, co-occurrence, and cross-modal complexity. To address this gap, we extend dense video grounding to a multi-segment setting and present DenseSpeech, the first densely annotated video grounding dataset in the speech domain. DenseSpeech comprises 1,800 authentic classroom recordings with 13,880 annotated segments spanning over 52 hours, featuring high annotation density, frequent boundary overlap, and pervasive event repetition with form variability. We further propose DenseMSG, an end-to-end dense multi-scale grounding model. Its Cross-modal Consistent Interaction module employs text-guided cross-attention to highlight query-relevant temporal regions and performs hierarchical bidirectional audiovisual fusion with adaptive gating to build a multi-scale feature pyramid. A Multi-Scale Integration module enables bidirectional feature propagation across scales for complementary global-local modeling, while similarity-based classification and query-aware regression heads jointly predict boundaries at all scales, naturally supporting one-to-many retrieval. Experiments show DenseMSG achieves state-of-the-art performance on DenseSpeech and public benchmarks (TACoS, ActivityNet Captions), validating the dataset and model. Jiachen Tan, Guangyao Su, Jianwei Fang, Bin Wu 0001, Chunping Zheng |
ICMR | 6 |
| 2026 | Collective-level propagation analysis via LLM-enhanced hypergraph transformer for fake news detection
Bin Wu 0001, Xuanning Liu, Jiachen Tan, Di Liu 0032, Guangyao Su |
Inf. Process. Manag. | 2 |
| 2025 | Sketch-Based Poetry Retrieval with Unsupervised Vision-and-Language Pre-training
Yangfu Zhu, Bin Wu 0001 |
DASFAA (3) | 4 |
| 2025 | Demonstration Meets Typed Events: Type Specific Video Semantic Role Labeling via Multimodal Prompting and RetrievalabstractVideo Semantic Role Labeling (VidSRL) aims to detect salient events and label their semantic roles from videos. Existing methods often struggle with missing, redundant, misordered, or incorrectly understood argument roles due to ignoring the semantics of the argument roles and semantic structure within events. Inspired by the labeling logic of human annotators, we address this by presenting a novel method, Type Specific Video Semantic Role Labeling with Demonstration-Enhanced Vision-Language Model (TypesDeV). Our method is based on a three-stage framework: (1) Selection involves predicting potential event types using a video classifier, (2) Generation uses a vision-language model (VLM) to generate argument role descriptions (semantic role labels) for each candidate event type, and (3) Retrieval selects the best matching descriptions through video-to-text retrieval. We enhance the VLM with event-type-specific prompts and demonstrations to explicitly capture the semantics of and relationships between the verb (event type) and argument roles. Our approach achieves new state-of-the-art on multiple metrics, and for the first time achieves human-level performance on most metrics when using ground-truth verbs. Hanxiao Wei, Bin Wu 0001, Chunjia Wang, Guangyao Su |
ICMR | 2 |
| 2025 | A diachronic language model for long-time span classical Chinese
Yangfu Zhu, Yuanxing Xu, Bin Wu 0001 |
Inf. Process. Manag. | 6 |
| 2025 | DCCMA-Net: Disentanglement-based cross-modal clues mining and aggregation network for explainable multimodal fake news detection
Xuanning Liu, Bin Wu 0001 |
Inf. Process. Manag. | 5 |
| 2024 | Enhancing Temporal and Geographical Named Entity Recognition in Chinese Ancient Texts with External Time-series Knowledge BasesabstractIn the field of ancient Chinese text, extracting and analysing temporal and geographic information are crucial for understanding the personal experiences of historical figures, the development of historical events, and the overall historical background. Currently, named entity recognition(NER) strategies such as BERT+CRF are used to extract temporal and geographic information from ancient Chinese text. However, ancient Chinese text covers a vast time span, and the temporal and geographic entities constantly evolve and change, making it difficult to extract these entities from text. This paper proposes a temporal and geographic extraction model for ancient Chinese text, enhanced by time-series external knowledge base. The extraction of proprietary nouns and general structures are divided into two independent networks. An external database is applied to enhance extraction of proprietary nouns and reduce noise for general structure inference. We constructed address trees and chronological tables containing commonly used places and time-related keywords from different periods and collected 12,000 texts spanning 3,000 years for extensive training. Overall, our research highlights the importance of external knowledge base for ancient Chinese NER, and provides new ideas for research in related fields. Xuanning Liu, Shuai Zhong, Xinming Chen, Bin Wu 0001 |
CIKM | 5 |
| 2024 | A Dynamic pre-trained Model for Chinese Classical Poetry
Xuanning Liu, Haorui Wang, Bin Wu 0001 |
DASFAA (2) | 4 |
| 2024 | Intra and Inter-modality Incongruity Modeling and Adversarial Contrastive Learning for Multimodal Fake News DetectionabstractMultimodal fake news detection (FND) is significant in safeguarding network security and societal safety. Most existing studies only focus on common semantic features between different modalities and utilize simple cross-entropy loss for model training. However, these studies overlook the incongruent semantic features in multimodal news data, which can arise within or between modalities. Moreover, the utilization of simple cross-entropy loss may not provide the model with robustness against well-designed forged fake news. To address the above issues, we propose a novel approach named Signed Attention-based Graph Transformer with Adversarial Contrastive Learning (SAGT-ACL) for the detection of multimodal fake news. SAGT-ACL models fine-grained semantic associations in multimodal news articles by constructing a fully connected multimodal graph and reframes the fake news classification task as a graph classification problem. Additionally, SAGT-ACL incorporates a signed attention-based graph transformer module to identify both common and incongruent semantics within and across modalities. Finally, SAGT-ACL proposes an adversarial data augmentation mechanism to simulate malicious forgeries by fake news creators and designs an auxiliary adversarial contrastive learning task to help the model learn more discriminative news representations from the adversarial samples for robust and effective detection. Extensive experiments demonstrate that SAGT-ACL outperforms existing methods, with detection accuracy improvements of 4.95%, 6.01%, and 5.68% on Weibo, Twitter, and Gossipcop datasets, respectively. Bin Wu 0001 |
ICMR | 2 |
| 2024 | Enhancing multimodal depression detection with intra- and inter-sample contrastive learning
Yangfu Zhu, Bin Wu 0001 |
Inf. Sci. | 5 |
| 2024 | Graph Alignment Neural Network Model With Graph to Sequence LearningabstractNetwork alignment aims at detecting the corresponding entities across multiple networks, which is an essential basis for the fusion and analysis of multiple network information. Moreover, embedding-based network alignment has gradually become one of the promising methods. However, existing methods ignore the confusing selection problem caused by the similarity-orientated principle of network embedding and over-dependence on the hypothesis of structural consistency. In this paper, we propose an end-to-end Graph Alignment Neural Network (GANN) model with graph-to-sequence learning. GANN mainly consists of two modules: Graph encoder and Sequence decoder. In graph encoder module, we present a restricted network embedding method, which can not only capture the local structure and attribute information of nodes but also realize the constraint of node embedding and space reconciliation. In sequence decoder module, we propose a graph-to-sequence learning model to address large graphs' structural consistency hypothesis problem. In this model, an attention-based LSTM mechanism is introduced to infer a node in the source network corresponding to the candidate node sequence in target networks. In this candidate sequence, the correct aligned node is placed at the top. We demonstrate that GANN outperforms the state-of-the-art methods in network alignment tasks on various real-world datasets. Nianwen Ning, Bin Wu 0001, Haoqing Ren, Qiuyue Li |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | PCENet: Psychological Clues Exploration Network for Multimodal Personality AssessmentabstractMultimodal personality assessment aims to identify and express human personality traits in videos. Existing methods primarily focus on multimodal fusion while ignoring the inherent psychological clues essential for this interdisciplinary task. Modality clues: personality traits are stable over time due to their genetic and environmental origins, resulting in stable personality traits in the multimodal data. Trait clues: multiple traits often co-occur with non-negligible correlations, which can collectively aid trait identification. To simultaneously capture the above psychological clues, we propose a novel Psychological Clues Exploration Network (PCENet) for multimodal personality assessment, which is a human-like judgment paradigm with more generalization capability. Specifically, we first devise a multimodal hierarchical disentanglement, which clearly aligns stable representations among different modalities and separates the mutability of each modality. Subsequently, a Transformer-backbone decoder equipped with modality-to-trait attention is exploited to adaptively generate a tailored representation for each trait with the guidance of trait semantics. The trait semantics are obtained by exploiting trait correlations through self-attention. Extensive experiments on the First Impression V2 dataset demonstrate that our PCENet outperforms the state-of-the-art methods for multimodal personality assessment. Yangfu Zhu, Bin Wu 0001 |
CIKM | 6 |
| 2023 | ReGR: Relation-aware graph reasoning framework for video question answeringabstractAs one of the challenging cross-modal tasks, video question answering (VideoQA) aims to fully understand video content and answer relevant questions. The mainstream approach in current work involves extracting appearance and motion features to characterize videos separately, ignoring the interactions between them and with the question. Furthermore, some crucial semantic interaction details between visual objects are overlooked. In this paper, we propose a novel Relation-aware Graph Reasoning (ReGR) framework for video question answering, which first combines appearance–motion and location–semantic multiple interaction relations between visual objects. For the interaction between appearance and motion, we design the Appearance–Motion Block, which is question-guided to capture the interdependence between appearance and motion. For the interaction between location and semantics, we design the Location–Semantic Block, which utilizes the constructed Multi-Relation Graph Attention Network to capture the geometric position and semantic interaction between objects. Finally, the question-driven Multi-Visual Fusion captures more accurate multimodal representations. Extensive experiments on three benchmark datasets, TGIF-QA, MSVD-QA, and MSRVTT-QA, demonstrate the superiority of our proposed ReGR compared to the state-of-the-art methods. Fangtao Li, Kaoru Ota, Mianxiong Dong, Bin Wu 0001 |
Inf. Process. Manag. | 5 |
| 2022 | Learning Social Influence from Network Structure for Recommender Systems
Ting Bai 0004, Bin Wu 0001 |
DASFAA (2) | 3 |
| 2022 | Semi-supervised Graph Learning with Few Labeled Nodes
Ting Bai 0004, Bin Wu 0001 |
DASFAA (2) | 3 |
| 2022 | Learning Advisor-Advisee Relationship from Multiplex Network Structure
Xiangchong Cui, Ting Bai 0004, Bin Wu 0001, Xinkai Meng |
KSEM (3) | 3 |
| 2022 | Open Relation Extraction via Query-Based Span Prediction
Huifan Yang, Zekun Li 0003, Donglin Yang, Jinsheng Qi, Bin Wu 0001 |
KSEM (2) | 6 |
| 2022 | A Contrastive Sharing Model for Multi-Task RecommendationabstractMulti-Task Learning (MTL) has attracted increasing attention in recommender systems. A crucial challenge in MTL is to learn suitable shared parameters among tasks and to avoid negative transfer of information. The most recent sparse sharing models use independent parameter masks, which only activate useful parameters for a task, to choose the useful subnet for each task. However, as all the subnets are optimized in parallel for each task independently, it is faced with the problem of conflict between parameter gradient updates (i.e, parameter conflict problem). To address this challenge, we propose a novel Contrastive Sharing Recommendation model in MTL learning (CSRec). Each task in CSRec learns from the subnet by the independent parameter mask as in sparse sharing models, but a contrastive mask is carefully designed to evaluate the contribution of the parameter to a specific task. The conflict parameter will be optimized relying more on the task which is more impacted by the parameter. Besides, we adopt an alternating training strategy in CSRec, making it possible to self-adaptively update the conflict parameters by fair competitions. We conduct extensive experiments on three real-world large scale datasets, i.e., Tencent Kandian, Ali-CCP and Census-income, showing better effectiveness of our model over state-of-the-art methods for both offline and online MTL recommendation scenarios. Ting Bai 0004, Yudong Xiao, Bin Wu 0001, Guojun Yang, Hongyong Yu, Jian-Yun Nie |
WWW | 3 |
| 2022 | What happens next? Combining enhanced multilevel script learning and dual fusion strategies for script event predictionabstractScript event prediction (SEP), aiming at predicting next event from context event sequences (i.e., scripts), has played an important role in many real-world applications such as government decision-making. While most of the existing research only depend on the top-level event prediction, they ignore the influence of other bottom levels or other relationship modeling manners. In this paper, we focus on the problem of SEP via multilevel script learning where the goal of is to explore a multistage, multiprediction and multilevel information fusion model for SEP. This is challenging in (1) simultaneously modeling of the multilevel event relationship semantic information and (2) effectively designing multilevel information fusion strategies. In this paper, we propose a new script event prediction model based on Enhanced Multilevel script learning and Dual Fusion strategies, named EMDF-Net. Specifically, EMDF-Net designs the multilevel (event/chain/segment level) script learning to model both temporal and casual information as well as the rich structural relevance via neural stacking of self-attention mechanism and graph neural networks. Then it proposes dual fusion strategies to fully integrate different-level information by nonlinear feature composition and weighted score fusion. Finally, a deep supervision strategy is utilized to end-to-end train the whole model and provide a good initialization for information fusion. Experimental results on the popular NYT corpus demonstrate the effectiveness and superiority of EMDF-Net. Pengpeng Zhou, Bin Wu 0001, Caiyong Wang, Hao Peng 0001, Juwei Yue, Song Xiao 0004 |
Int. J. Intell. Syst. | 2 |
| 2021 | Dependency Parsing Representation Learning for Open Information Extraction
Zekun Li 0003, Nianwen Ning, Chengcheng Peng, Bin Wu 0001 |
KSEM | 4 |
| 2021 | Embedding-Based Network Alignment Using Neural Tensor Networks
Qiuyue Li, Nianwen Ning, Bin Wu 0001, Wenying Guo |
KSEM | 3 |
| 2021 | Relation-aware Hierarchical Attention Framework for Video Question AnsweringabstractVideo Question Answering (VideoQA) is a challenging video understanding task since it requires a deep understanding of both question and video. Previous studies mainly focus on extracting sophisticated visual and language embeddings, fusing them by delicate hand-crafted networks. However, the relevance of different frames, objects, and modalities to the question are varied along with the time, which is ignored in most of existing methods. Lacking understanding of the the dynamic relationships and interactions among objects brings a great challenge to VideoQA task. To address this problem, we propose a novel Relation-aware Hierarchical Attention (RHA) framework to learn both the static and dynamic relations of the objects in videos. In particular, videos and questions are embedded by pre-trained models firstly to obtain the visual and textual features. Then a graph-based relation encoder is utilized to extract the static relationship between visual objects. To capture the dynamic changes of multimodal objects in different video frames, we consider the temporal, spatial, and semantic relations, and fuse the multimodal features by hierarchical attention mechanism to predict the answer. We conduct extensive experiments on a large scale VideoQA dataset, and the experimental results demonstrate that our RHA outperforms the state-of-the-art methods. Fangtao Li, Ting Bai 0004, Chenyu Cao, Chenghao Yan, Bin Wu 0001 |
ICMR | 6 |
| 2021 | Social Relation Analysis from Videos via Multi-entity ReasoningabstractVideos contain rich semantic information. Analyzing social relations in video semantics can help machines interpret the behavior of human beings. However, most of the work related to social relationship recognition is based on still images, while video-based social relationship analysis tasks are less concerned. Here we propose a Multi-entity Relation Reasoning (MRR) framework that can be used for recognizing or predicting social relations in videos. To capture temporal features and contextual cues in videos, and use richer information to represent the person in the video, we track each person's appearance timeline and design a multi-entity representation method to build a social relationship knowledge graph. Then we use graph attention networks to gather information from the entity's neighborhood. Besides, situation information is helpful to identify relationships, we design a situation information extraction module to generate situation embedding from the video clip. Finally, a decoder is adopted to predict relationships between character entities. We evaluate the model on the MovieGraphs dataset and verify the effectiveness of the proposed framework. Chenghao Yan, Fangtao Li, Chenyu Cao, Bin Wu 0001 |
ICMR | 6 |
| 2021 | A multimodal fake news detection model based on crossmodal attention residual and multichannel convolutional neural networks
Chenguang Song, Nianwen Ning, Bin Wu 0001 |
Inf. Process. Manag. | 4 |
| 2021 | Temporally evolving graph neural network for fake news detection
Chenguang Song, Kai Shu, Bin Wu 0001 |
Inf. Process. Manag. | 3 |
| 2020 | Temporal Graph Neural Networks for Social RecommendationabstractIn social recommendation, the purchase decision of users is influenced by their basic preference of items, as well as the social influence of peers. Such social connections had been proved to be effective in modeling users' preference of items. However, most models in social recommender literature only considered two types of relations, i.e., user-item relation in interaction network and user-user relation in social network. The temporal sequential information of items, i.e., item-item relation, can also be utilized to infer the preference of users, but had been ignored in almost all of the graph based recommendation models. Two issues of such temporal information had not been well studied in social recommender systems: the temporal strength information, i.e., the real purchase time of an item, and its influence on social relations. To address the above issues, we propose a novel Temporal Enhanced Graph Model for Social Recommendation (TGRec). In TGRec, the purchase time information between items is characterized as a special temporal relation, and the purchase decision of users depends on three factors: (1) a user's basic preference of items, (2) the collaborative influence of peers, (3) the temporal impact of previous items bought by the user. Experimental results on three real-world commerce datasets demonstrate the effectiveness of our model for social recommendation, showing the usefulness of modeling the temporal information in heterogeneous graph. Ting Bai 0004, Youjie Zhang, Bin Wu 0001, Jian-Yun Nie |
IEEE BigData | 3 |
| 2020 | Detecting Rumor on Microblogging Platforms via a Hybrid Stance Attention Mechanism
Lingyu Zeng, Bin Wu 0001, Bai Wang 0001 |
ICWE | 2 |
| 2020 | A Time Interval Aware Approach for Session-Based Social Recommendation
Youjie Zhang, Ting Bai 0004, Bin Wu 0001, Bai Wang 0001 |
KSEM (2) | 3 |
| 2020 | Deep Adversarial Completion for Sparse Heterogeneous Information Network EmbeddingabstractHeterogeneous information network (HIN) contains multiple types of entities and relations. Most of existing HIN embedding methods learn the semantic information based on the heterogeneous structures between different entities, which are implicitly assumed to be complete. However, in real world, it is common that some relations are partially observed due to privacy or other reasons, resulting in a sparse network, in which the structure may be incomplete, and the ”unseen” links may also be positive due to the missing relations in data collection. To address this problem, we propose a novel and principled approach: a Multi-View Adversarial Completion Model (MV-ACM). Each relation space is characterized in a single viewpoint, enabling us to use the topological structural information in each view. Based on the multi-view architecture, an adversarial learning process is utilized to learn the reciprocity (i.e., complementary information) between different relations: In the generator, MV-ACM generates the complementary views by computing the similarity of the semantic representation of the same node in different views; while in the discriminator, MV-ACM discriminates whether the view is complementary by the topological structural similarity. Then we update the node’s semantic representation by aggregating neighborhoods information from the syncretic views. We conduct systematical experiments1 on six real-world networks from varied domains: AMiner, PPI, YouTube, Twitter, Amazon and Alibaba. Empirical results show that MV-ACM significantly outperforms the state-of-the-art approaches for both link prediction and node classification tasks. Kai Zhao 0009, Ting Bai 0004, Bin Wu 0001, Bai Wang 0001, Youjie Zhang, Yuanyu Yang, Jian-Yun Nie |
WWW | 3 |
| 2019 | Spatial-Temporal Recurrent Neural Network for Anomalous Trajectories Detection
Yunyao Cheng 0001, Bin Wu 0001, Chuan Shi 0001 |
ADMA | 2 |
| 2019 | MLCA: A Multi-label Competency Analysis Method Based on Deep Neural Network
Guohao Qiao, Bin Wu 0001, Bai Wang 0001, Baoli Zhang |
ADMA | 2 |
| 2019 | Strengthening social networks analysis by networks fusionabstractThe relationship extraction and fusion of networks are the hotspots of current research in social network mining. Most previous work is based on single-source data. However, the relationships portrayed by single-source data are not sufficient to characterize the relationships of the real world. To solve this problem, a Semi-supervised Fusion framework for Multiple Network (SFMN), using gradient boosting decision tree algorithm (GBDT) to fuse the information of multi-source networks into a single network, is proposed in this paper. Our framework aims to take advantage of multi-source networks fusion to enhance the accuracy of the network construction. The experiment shows that our method optimizes the structural and community accuracy of social networks which makes our framework outperforms several state-of-the-art methods. Feiyu Long, Nianwen Ning, Chenguang Song, Bin Wu 0001 |
ASONAM | 4 |
| 2019 | A hierarchical insurance recommendation framework using GraphOLAM approachabstractGraph has been widely used for modeling complex relationship datasets in different application fields. Social networks based recommendation system have obtained satisfactory results in Business Intelligence(BI). However, current personalized recommendation methods based on graph structure generally lack interactivity and seldom consider efficient data management. To address these problems, Graph OnLine Analytical Mining (GraphOLAM) is a promising method, which combines OLAP technology with social networks. We first propose an efficient recommendation framework based on GraphOLAM data cube technology for the recommendation in the insurance service. Based on this framework, a new algorithm framework named RU-GOLAM for insurance is proposed, which combines GraphOLAM dimensional aggregation operation and specific recommendation methods. A series of graphs can be generated by GraphOLAM dimensional aggregation operations, which reflect the relationships of nodes under the constraints of different hierarchical dimensions. Node similarities are calculated to generate the Top-N sequential recommendation based on all of these graphs, which can achieve the balance between the topology of the original graph and high-dimensional information of the nodes. Experiments show that our approach outperforms other baseline algorithms on an insurance service dataset. Sirui Sun, Bin Wu 0001, Zixing Zhang 0002, Nianwen Ning, Bai Wang 0001 |
ASONAM | 2 |
| 2019 | DRAM: A Deep Reinforced Intra-attentive Model for Event Prediction
Shuqi Yu, Linmei Hu, Bin Wu 0001 |
KSEM (1) | 3 |
| 2019 | Jointly Modeling Community and Topic in Social Network
Nianwen Ning, Jinna Lv, Chenguang Song, Bin Wu 0001 |
KSEM (1) | 5 |
| 2019 | Complaint Classification Using Hybrid-Attention GRU Neural Network
Bin Wu 0001, Bai Wang 0001, Xuesong Tong |
PAKDD (1) | 2 |
| 2018 | Power Equipment Fault Diagnosis Model Based on Deep Transfer Learning with Balanced Distribution Adaptation
Bin Wu 0001 |
ADMA | 2 |
| 2018 | A Parallel Community Detection Algorithm Based on Incremental Clustering in Dynamic NetworkabstractDynamic community detection is a key method for the research of network evolution. However, most existing dynamic community detection algorithms are time-consuming in dealing with large-scale networks. Moreover, most current parallel community detection algorithms are static and they ignore the changes of network structure over time. In this paper, we propose a novel parallel algorithm based on incremental vertices, which is able to process large-scale dynamic networks, called PICD. In PICD algorithm, the revised Parallel Weighted Community Clustering (PWCC) metric is conductive to a convenient calculation, which is more sensitive to community structure compared to other metrics. The PICD approach consists of two main steps. Firstly, it identifies the incremental vertices in the dynamic network. Secondly, it maximizes the PWCC of the entire network by merely adjusting the community membership of incremental vertices to capture community structure in high quality. The results of experiments on both the synthetic and real world networks demonstrate that the PICD algorithm achieves a higher accuracy and efficiency. Moreover, it performs more stable than most of the baseline methods. The experiments also show that PICD algorithm takes an almost linear time with the growth of the network scale. Cuiyun Zhang, Bin Wu 0001 |
ASONAM | 3 |
| 2018 | Road Damage Detection and Classification with Faster R-CNNabstractThis technical paper presents the method that we use in the Road Damage Detection and Classification Challenge, which is designed to detect damages contained in road images photographed by a vehicle-mounted smartphone. In this task, we apply Faster R-CNN to detect and classify damaged roads. Through analyses of aspect ratios and sizes of the damaged areas in the training dataset, we adjust relevant parameters of the model. In order to solve the problem of unbalanced data distribution of different classes, we introduce some data augmentation techniques (contrast transformation, brightness adjustment, and Gaussian blur) before training. Experimental results demonstrate that our method can achieve a Mean F1-Score of 0.6255 in the competition. The source code and model are publicly available at https://github.com/zhezheey/tf-faster-rcnn-rddc. Wenzhe Wang, Bin Wu 0001, Sixiong Yang |
IEEE BigData | 2 |
| 2018 | A Heterogeneous Information Network Method for Entity Set Expansion in Knowledge Graph
Xiaohuan Cao, Chuan Shi 0001, Yuyan Zheng, Xiaoli Li 0001, Bin Wu 0001 |
PAKDD (2) | 6 |
| 2017 | CES: A System for Community EvaluationabstractWith the development of Internet, we are gradually entering the era of big data. The size of network is increasing and the structure of network is becoming more complex. Community is a unique network structure with great research value. The task of community analysis goes through two separate phases: first, detection of meaningful community structure from a network, and second, evaluation of the appropriateness of the detected community structure. With the popularity of network research, many community detection algorithms emerged which can be grouped in categories, based on different criteria. In order to applicate community detection algorithms in real-world network analysis, we need to measure the performance of the algorithm. The performance depends on two points, that is, whether the algorithm can give the result of community division in an acceptable time, and whether the algorithm can reveal the community structure of the network with high quality. In recent years, systems used to analyze network and detect community are mushrooming. However, existing systems rarely have evaluation function, either providing social network analysis or providing data analysis service. We need a tool to evaluate community detection algorithms. In response to the challenge, the Community Evaluation System (CES) is proposed to meet the demands of community detection algorithms' evaluation. CES can evaluate community detection algorithms with multiple metrics. It uses B/S mode, integrates Spark, Yarn and HDFS technology to support the operation of large-scale data, and experiments prove that it is effective. Bin Wu 0001, Xuesong Tong |
ASONAM | 1 |
| 2017 | A Parallel Network Community Detection Algorithm Based on Distance DynamicsabstractIn recent years, community detection has drawn more and more researchers' attention. With the development of Internet, the scale of network data is growing fast. It is necessary to find an effective parallel community detection algorithm for large-scale network. In this paper, we propose a novel and parallel community detection algorithm, PCDU algorithm, based on distance dynamics. We send distances information to nodes and update distances of edges constantly, based on previous values and the unified model, which is introduced to quantify different influences from nodes and edges. It ends until the distances are stable. Then we remove some special edges from the original graph and get all subgraphs, which are the community partitions. It still inherits the advantage of uncovering small communities and outliers. Experiments based on synthetic networks and real world networks, show that our algorithm execute more efficient than stand-alone version. Since it is based on the Spark platform and designed in parallelization, the algorithm is very suitable for large datasets. We also provide a novel method taking use of double summation to calculate the NMI value of community partition result and the embedded community structure. Compared with the traditional way, it is not only as accurate as the traditional way and more efficient, but also has less space complexity. Experiments show that it is suitable for evaluating community division results in large-scale network. Bin Wu 0001, Cuiyun Zhang |
ASONAM | 1 |
| 2017 | A Parallel Framework for Large-scale Multidimensional Heterogeneous Network AnalysisabstractHow to manage and analyze interconnected, multidimensional and heterogeneous information network data has become the focus of current research. Graph On-Line Analytical Processing (GraphOLAP) can process a quick online analysis and query operation of graph data. With the existing achievement of GraphOLAP we propose a new graph cube framework according to the multidimensional heterogeneous informational network. We introduce the concept of relation path which is the guidance of the relation path aggregate network. We propose the concept of derived dimension to support more analysis with clustering algorithm. We also propose some traditional operations and new operations based on our model. Then we discuss the materialization strategies and implement the framework in Spark. The result of experiments has proved the efficiency and effectiveness of our framework. Zixing Zhang 0002, Bin Wu 0001, Zeao Wang |
ASONAM | 2 |
| 2017 | An entity disambiguation method based on LeaderRankabstractEntity Disambiguation is commonly faced in semantic search and knowledge base population. However, it is a challenging task because of the diversity of mentions. Previous methods can be classified into two main groups. One focuses on disambiguating mentions in a document independently and mainly relies on the local context similarity. The other collectively disambiguates mentions only taking into account link information. These are not appropriate when the context and the link information are poor or misleading. In this paper, we propose a new method to collectively disambiguate mentions in documents. Our proposed framework considers three features, including text similarity, entity popularity, and entity relationship. First we adopt LeaderRank algorithm on the graph model to rank entities according to the link information among entities. Then we combine with global text similarity between entity and document to disambiguate mentions. Our detailed experimental evaluation on two benchmark datasets demonstrates our methods is effective. Bingjing Jia, Bin Wu 0001, Jinna Lv, Pengpeng Zhou, Yao Bu |
IEEE BigData | 2 |
| 2017 | DSBPR: Dual Similarity Bayesian Personalized Ranking
Longfei Shi, Bin Wu 0001, Chuan Shi 0001, Mengxin Li |
PAKDD (1) | 2 |
| 2017 | Entity Set Expansion with Meta Path in Knowledge Graph
Yuyan Zheng, Chuan Shi 0001, Xiaohuan Cao, Xiaoli Li 0001, Bin Wu 0001 |
PAKDD (1) | 5 |
| 2016 | Event Evolution Model Based on Random Walk Model with Hot Topic Extraction
Chunzi Wu, Bin Wu 0001, Bai Wang 0001 |
ADMA | 2 |
| 2016 | Dynamic community detection based on distance dynamicsabstractDynamic community detection has been of great significance on analyzing network structure and community evolution. Among state-of-the-art methods, incremental algorithms based on modularity have been used widely, for the fully utilization of both current and historical information. Unfortunately, they are difficult to uncover small community due to problem called “resolution limit” and also sensitive to the sequence of network increments' arrival. In this paper, we propose a novel dynamic community detection algorithm based on distance dynamics, which detects community in near-linear time. Meanwhile, the proposed algorithm overcomes the shortcomings of traditional methods by replacing modularity with local interaction model. In a sense, it is assumed that increments can be treated as disturbance of network. It can be limited to a certain “local area” by disturbing factor. Experiments show that the proposed algorithm has achieved well balance between efficiency and effectiveness both in synthetic and real world networks. Lei Zhang 0049, Bin Wu 0001, Xuelin Zeng |
ASONAM | 3 |
| 2016 | Personalized recommendation for new questions in community question answeringabstractCommunity question answering(CQA) websites such as Yahoo! Answers and Stack Overflow provide a new way of asking and answering questions which are not well served by general web search engines. Due to the huge volume and ever-increasing number of questions, not all new questions can get fully answered in required time. Therefore, it is of great significance to design some effective strategies of recommending experts for new questions. In this paper, we propose a novel personalized recommendation method for routing new questions to a group of experts. Different from prior work which only considers the topic modeling or the link structure, we aim at recommending new questions to more appropriate experts by considering both of these two factors. Moreover, we design a new strategy of network construction with the personalization fully considered. The comparison experiments are conducted with Stack Overflow data and the experimental results demonstrate that the proposed method improves the recommendation performance over other methods in expert recommendation. Bin Wu 0001, Juan Yang 0002, Shuang Peng 0003 |
ASONAM | 2 |
| 2016 | Forming a research team of experts in expert-skill co-occurrence network of research newsabstractThe team formation problem is required to find a group of individuals that can match the skills required by a collaborative task. Large-scale and comprehensive scientific research tasks need skilled experts from various fields to form a research team and work for it. This paper constructs a dataset and proposes team formation algorithms to find out research teams, which provides decision support for the research projects. The size of existing datasets is relatively small and fields of experts in it are less diversified. This paper extracts information of experts and skills from research news to construct a co-occurrence network with heterogeneous network structure. Based on the dataset, this work designs approximate algorithms regarding skill as the priority to find near optimum teams with provable guarantees. On heterogeneous structure, the proposed algorithms directly search requested skills to form the subgraph of team, which achieve significant improvement in time efficiency. Experimental results suggest that our methods can form the high-quality research team, and have better efficiently compared to naive strategies and scale well with the size of the data. Juan Yang 0002, Mengxin Li, Bin Wu 0001 |
ASONAM | 3 |
| 2016 | Efficient large scale near-duplicate video detection base on sparkabstractWith the huge amount of web video data and its exponential growth in recent years, there are new challenges in Near-Duplicate Video Detection (NDVD) which have attracted much attention owing to its wide applications. One of the problems is how to extract discriminative features to achieve higher precision, and the other problem is how to improve the efficiency of large scale video analysis. Existing methods have predominantly focused on improving the precision and achieve some success on NDVD. However, it is an extremely challenging task to guarantee the higher precision and efficiency simultaneously. To balance the efficiency and precision, we propose a Multi-Feature based Parallel System (MFPS) for large scale NDVD to overcome these challenges. For feature extraction, we combine local and global features to accurately represent the critical information of video. In our work, not only Scale-Invariant Feature Transform (SIFT) but also Local Maximal Occurrence (LOMO) is introduced as local features and Color Name (CN) is adopted as global feature. Meanwhile, for SIFT and CN features, Bag-of-Visual-Words (BoVW) method is used to quantize the features due to its accuracy and efficiency in NDVD. Furthermore, we design a similarity measure for video pairs which is inspired by hough transform and sliding window idea. Finally, we implement the system on parallel cloud computing platform based on spark. A comprehensive evaluation is conducted on CC_WEB_VIDEO dataset which includes 12790 videos and 27% near-duplicates. Experimental results show that the proposed method is able to achieve some improvements in term of mAP against other methods. Moreover, our approach is efficient and achieves 5.8X speedup. Jinna Lv, Bin Wu 0001, Bingjing Jia, Peigang Qiu |
IEEE BigData | 2 |
| 2016 | Link Prediction in Schema-Rich Heterogeneous Information Network
Xiaohuan Cao, Yuyan Zheng, Chuan Shi 0001, Jingzhi Li 0001, Bin Wu 0001 |
PAKDD (1) | 5 |
| 2016 | Dual Similarity Regularization for Recommendation
Jian Liu 0001, Chuan Shi 0001, Fuzhen Zhuang, Jingzhi Li 0001, Bin Wu 0001 |
PAKDD (2) | 6 |
| 2016 | Aspect Mining with Rating Bias
Chuan Shi 0001, Fuzhen Zhuang, Bin Wu 0001 |
ECML/PKDD (2) | 5 |
| 2016 | Constrained-meta-path-based ranking in heterogeneous information network
Chuan Shi 0001, Philip S. Yu, Bin Wu 0001 |
Knowl. Inf. Syst. | 4 |
| 2016 | Integrating heterogeneous information via flexible regularization framework for recommendation
Chuan Shi 0001, Jian Liu 0001, Fuzhen Zhuang, Philip S. Yu, Bin Wu 0001 |
Knowl. Inf. Syst. | 5 |
| 2015 | Characterizing super spreading in microblog: An epidemic-based modelabstractMicroblogs play an important role in online social communications. Different from ordinary pieces of information, some hot topics and emerging news will become much more popular in a very short time with the help of this information spreading platform of microblogs. In these "super spreading events", messages are transmitted to a vast range of individuals through a small portion of users engaged in the information diffusion process, a.k.a. super spreaders. Gaining an awareness of super spreading phenomena and an understanding of patterns of vast-ranged information diffusion process is worthy for several tasks such as hot topic detection, predictions of information propagation, harmful information monitoring and intervention. In this paper, inspired by the analogous patterns of super spreading in both information diffusion and spread of a contagious disease, we build a parameterized model based on well-known epidemic models to characterize super spreading phenomenon of tweet message diffusion accompanied with super spreaders. Through a fitting process, parameter settings under different scenarios are obtained and the corresponding basic reproduction number is also analyzed, which indicates the degree super spreader will affect the spreading. With the help of the SAIR model, some feasible applications can be exploited. Bin Wu 0001, Bai Wang 0001 |
IEEE BigData | 2 |
| 2015 | A community detection method based on K-shellabstractWe propose a community detection method based on K-shell. Our method determines some core nodes of the graph according to the K-shell value of these nodes. These core nodes constitute a subgraph on which we use the community detection algorithm to divide the core nodes into communities. Compared to classical methods, by this way, our proposed method removes the non-core nodes which may impact the quality of the solutions, and the running time would be reduced by 45% at most because the graph scale is reduced. Then we use the idea of LPA to infer community labels for the non-core nodes. Our experiments demonstrate that our method can reach the quality of solutions of CNM algorithms on the network constructed by Planted l-partition model which is also better than it on dataset of the Zachary karate club network. Meanwhile, we run our method on large-scale datasets, which has a better performance than the CNM algorithm. Liutong Xu, Bin Wu 0001 |
IEEE BigData | 3 |
| 2015 | Finding community structure via rough K-means in social networkabstractMuch of the data of scientific interest, particularly when independence of data is not assumed, can be represented in the form of networks where data nodes are joined together to form edges corresponding to some kind of associations or relationships. Such information networks abound, like protein interactions in biology, web page hyperlink connections in information retrieval on the Web, cellphone call graphs in telecommunication, co-authorships in bibliometrics, crime event connections in criminology, etc. All these networks, also known as social networks, share a common property, the formation of connected groups of information nodes, called community structures. These groups are densely connected nodes with sparse connections outside the group. Finding these communities is an important task for the discovery of underlying structures in social networks, and has attracted much attention in data mining research. In this paper, we present rough k-means method (RKM), a new community mining approach that, simply put, regards a community as a set of nodes, these communities have their lower and upper approximation sets. Our algorithm starts by selecting k nodes as the center nodes of communities in a given network then iteratively assembles node to their closest center node to form communities, and subsequently calculates new center node in each group around which to gather nodes again until convergence. Our intuitions are based on proven observations in social networks and the results are promising. Experimental results on benchmark networks verify the feasibility and effectiveness of our new community mining approach. Bin Wu 0001 |
IEEE BigData | 2 |
| 2015 | Semantic Path based Personalized Recommendation on Weighted Heterogeneous Information NetworksabstractRecently heterogeneous information network (HIN) analysis has attracted a lot of attention, and many data mining tasks have been exploited on HIN. As an important data mining task, recommender system includes a lot of object types (e.g., users, movies, actors, and interest groups in movie recommendation) and the rich relations among object types, which naturally constitute a HIN. The comprehensive information integration and rich semantic information of HIN make it promising to generate better recommendations. However, conventional HINs do not consider the attribute values on links, and the widely used meta path in HIN may fail to accurately capture semantic relations among objects, due to the existence of rating scores (usually ranging from 1 to 5) between users and items in recommender system. In this paper, we are the first to propose the weighted HIN and weighted meta path concepts to subtly depict the path semantics through distinguishing different link attribute values. Furthermore, we propose a semantic path based personalized recommendation method SemRec to predict the rating scores of users on items. Through setting meta paths, SemRec not only flexibly integrates heterogeneous information but also obtains prioritized and personalized weights representing user preferences on paths. Experiments on two real datasets illustrate that SemRec achieves better recommendation performance through flexibly integrating information with the help of weighted meta paths. Chuan Shi 0001, Zhiqiang Zhang 0012, Ping Luo 0001, Philip S. Yu, Yading Yue, Bin Wu 0001 |
CIKM | 6 |
| 2015 | TSMH Graph Cube: A novel framework for large scale multi-dimensional network analysisabstractAs a representation of information, Multi-dimension network is more and more popular, such as web data and social network. With the increment of data source, the entities of the network become diverse. How to analyze these multi-dimensional heterogeneous networks effectively and efficiently is a big challenge. In this paper, we propose a Two-Step Multi-dimensional Heterogeneous (TSMH) Graph Cube framework. We use the meta path in heterogeneous network to guide the aggregation of the network and build the Entity Hyper Cube. For the cuboid in Entity Hyper Cube, we do dimension roll-up/drill down to build the Dimension Cube. In Entity Hyper Cube, we design meta path aggregation algorithms and propose materialization strategy. In Dimension Cube, we use hierarchical coding for entities and dimensions and it saves the process of join operations of entities and dimensions which greatly improve the efficiency of dimension operations. In addition, we propose more new Graph OLAP operations which can make network analysis more diverse. At last, we implement the framework in Spark. The results of experiments on real data set and synthetic data set confirm the efficiency and effectiveness of our framework. Pengsen Wang, Bin Wu 0001, Bai Wang 0001 |
DSAA | 2 |
| 2014 | Relevance Measure in Large-Scale Heterogeneous Networks
Xiaofeng Meng 0004, Chuan Shi 0001, Lei Zhang 0049, Bin Wu 0001 |
APWeb | 5 |
| 2014 | Location inference using microblog text and friendshipsabstractIn this paper, we proposed a novel scheme to infer user's location using microblog text and friendships, without known geo information. The major part of our research is identifying local words, words that associated with some particular location. With local words we identified, we use conditional random fields (CRF), to detect location specific microblog. Then we can estimate the most possible location of a user. And we take advantage of users' friendships to improve the result. Another key feature of our approach is that we consider timeliness of local words, as some local words are descriptions of local events and they are only associated with location during a certain period of time. Experimental evidence suggests that our algorithm works well in practice and outperforms the existing algorithms for estimating the location of microblog users. Chuanyang Li, Xiuqin Lin, Bin Wu 0001, Chuan Shi 0001 |
ASONAM | 3 |
| 2014 | SDHM: A hybrid model for spammer detection in WeiboabstractAs the microblogging service (such as Weibo) is becoming popular, spam becomes a serious problem of affecting the credibility and readability of Online Social Networks. Most existing studies took use of a set of features to identify spam, but without the consideration of the overlap and dependency among different features. In this study, we investigate the problem of spam detection by analyzing real spam dataset collections of Weibo and propose a novel hybrid model of spammer detection, called SDHM, which utilizing significant features, i.e. user behavior information, online social network attributes and text content characteristics, in an organic way. Experiments on real Weibo dataset demonstrate the power of the proposed hybrid model and the promising performance. Bin Wu 0001, Bai Wang 0001 |
ASONAM | 2 |
| 2014 | Overlapping community detection in large networks from a data fusion viewabstractCommunity detection is one of the most important problems in social network analysis in the context of the structure of the underlying graphs. Many researchers have proposed their own methods for discovering dense regions in social networks. Such methods are only designed with links of the underlying social network. However, with the development of recent applications, rich edge content can be available to give another view to the community detection process. In this study, we focus on improving community detection with the edge content in social networks. In order to regulate the effect of both linkage structure and edge content, we propose two feature integration strategies. Experiment results illustrate that the presence of edge content provides unprecedented opportunities and flexibility for the community detection process. Bin Wu 0001, Shuai Zhao 0001, Bai Wang 0001 |
ASONAM | 2 |
| 2014 | Ranking-based Clustering on General Heterogeneous Information Networks by Network ProjectionabstractRecently there is an increasing attention in heterogeneous information network analysis, which models networked data as networks including different types of objects and relations. Many data mining tasks have been exploited in heterogeneous networks, among which clustering and ranking are two basic tasks. These two tasks are usually done separately, whereas recent researches show that they can mutually enhance each other. Unfortunately, these works are limited to heterogeneous networks with special structures (e.g. bipartite or star-schema network). However, real data are more complex and irregular, so it is desirable to design a general method to manage objects and relations in heterogeneous networks with arbitrary schema. In this paper, we study the ranking-based clustering problem in a general heterogeneous information network and propose a novel solution HeProjI. HeProjI projects a general heterogeneous network into a sequence of sub-networks and an information transfer mechanism is designed to keep the consistency among sub-networks. For each sub-network, a path-based random walk model is built to estimate the reachable probability of objects which can be used for clustering and ranking analysis. Iteratively analyzing each sub-network leads to effective ranking-based clustering. Extensive experiments on three real datasets illustrate that HeProjI can achieve better clustering and ranking performances compared to other well-established algorithms. Chuan Shi 0001, Philip S. Yu, Bin Wu 0001 |
CIKM | 5 |
| 2014 | A Two-Phase Model for Retweet Number Prediction
Gang Liu 0008, Chuan Shi 0001, Bin Wu 0001, Jiayin Qi |
WAIM | 4 |
| 2014 | Maximizing the spread of influence ranking in social networks
Tian Zhu 0001, Bai Wang 0001, Bin Wu 0001, Chuanxi Zhu |
Inf. Sci. | 3 |
| 2014 | Multi-Label Classification Based on Multi-Objective OptimizationabstractMulti-label classification refers to the task of predicting potentially multiple labels for a given instance. Conventional multi-label classification approaches focus on single objective setting, where the learning algorithm optimizes over a single performance criterion (e.g., Ranking Loss ) or a heuristic function. The basic assumption is that the optimization over one single objective can improve the overall performance of multi-label classification and meet the requirements of various applications. However, in many real applications, an optimal multi-label classifier may need to consider the trade-offs among multiple inconsistent objectives, such as minimizing Hamming Loss while maximizing Micro F1 . In this article, we study the problem of multi-objective multi-label classification and propose a novel solution (called M oml ) to optimize over multiple objectives simultaneously. Note that optimization objectives may be inconsistent, even conflicting, thus one cannot identify a single solution that is optimal on all objectives. Our M oml algorithm finds a set of non-dominated solutions which are optimal according to different trade-offs among multiple objectives. So users can flexibly construct various predictive models from the solution set, which provides more meaningful classification results in different application scenarios. Empirical studies on real-world tasks demonstrate that the M oml can effectively boost the overall performance of multi-label classification by optimizing over multiple objectives simultaneously. Chuan Shi 0001, Xiangnan Kong, Di Fu, Philip S. Yu, Bin Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | HeteSim: A General Framework for Relevance Measure in Heterogeneous NetworksabstractSimilarity search is an important function in many applications, which usually focuses on measuring the similarity between objects with the same type. However, in many scenarios, we need to measure the relatedness between objects with different types. With the surge of study on heterogeneous networks, the relevance measure on objects with different types becomes increasingly important. In this paper, we study the relevance search problem in heterogeneous networks, where the task is to measure the relatedness of heterogeneous objects (including objects with the same type or different types). A novel measure HeteSim is proposed, which has the following attributes: (1) a uniform measure: it can measure the relatedness of objects with the same or different types in a uniform framework; (2) a path-constrained measure: the relatedness of object pairs are defined based on the search path that connects two objects through following a sequence of node types; (3) a semi-metric measure: HeteSim has some good properties (e.g., self-maximum and symmetric), which are crucial to many data mining tasks. Moreover, we analyze the computation characteristics of HeteSim and propose the corresponding quick computation strategies. Empirical studies show that HeteSim can effectively and efficiently evaluate the relatedness of heterogeneous objects. Chuan Shi 0001, Xiangnan Kong, Yue Huang 0001, Philip S. Yu, Bin Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2013 | Role discovery based on sociology attributes clustering in Sina MicroblogabstractUnderstanding and mastering users role plays an important part in online public opinions tracking and electronic commerce marketing, etc. Different groups of users have different sociology attributes. Thus, it is very important and interesting to discover user role based on their sociology attributes. The present user role discovery methods are generally based on the structural features or static coarse-grained behavior features. In this paper, by analyzing a large number of real social network data, we propose a novel method for social role discovery based on sociology attributes features: we first mining and define several properties on behalf of sociology attributes; then, to deal with the sociology attributes features clustering, we use Bayesian information criterion as our stopping criterion; at last, the experimental results show that using this method can better understand user role in Sina Microblog. Besides, the methodology in this paper for user role discovery also can be applied to other social network in general. Xinglong Yu, Bin Wu 0001 |
ASONAM | 2 |
| 2013 | Topic model-based link community detection with adjustable range of overlappingabstractComplex networks have attracted much research attentions. Community detection is an important problem in complex network which is useful in a variety of applications such as information propagation, link prediction, recommendations and marketing. In this paper, we focus on discovering overlapping community structure using link partition. We proposed a LDA-based link partition (LBLP) method which can find communities with adjustable range of overlapping. This method employs topic model to detect link partition, which can calculate the community belonging factor for each link. Based on the belonging factor, link partitions with bridge links can be found efficiently. We validate the effectiveness of our solution on both real-world and synthesized networks. The experiment results demonstrate that the approach can find meaningful and relevant link community structure. Bin Wu 0001, Bai Wang 0001 |
ASONAM | 2 |
| 2013 | A unified framework for predicting attributes and links in social networksabstractNode attributes and their associated links are two crucial components of user profiling in social networks. Recently, several research works show that both attributes and links are partially predictable by various classification methods. However, most of these works suffer from an isolation of attributes and links. In this paper, we propose a novel unified framework to predict attributes and links simultaneously, by a two-layer artificial neural network that encodes them in the same network. We obtain a better predictive accuracy on both attributes and links than the previous state-of-the-art methods, tested on different real-world datasets. Our results show that we can discover much more useful information in the whole social networks than in node attributes or their associated links alone. Xusen Yin, Bin Wu 0001, Xiuqin Lin |
IEEE BigData | 2 |
| 2013 | Integrating Clustering and Ranking on Hybrid Heterogeneous Information Network
Chuan Shi 0001, Philip S. Yu, Bin Wu 0001 |
PAKDD (1) | 4 |
| 2013 | How Long Will She Call Me? Distribution, Social Theory and Duration Prediction
Yuxiao Dong, Jie Tang 0001, Tiancheng Lou, Bin Wu 0001, Nitesh V. Chawla |
ECML/PKDD (2) | 4 |
| 2013 | A link clustering based overlapping community detection algorithm
Chuan Shi 0001, Yanan Cai, Di Fu, Yuxiao Dong, Bin Wu 0001 |
Data Knowl. Eng. | 5 |
| 2012 | Detecting Probabilistic Community with Topic Modeling on Sampling SubGraphsabstractDetecting communities plays a great important role in sociology, biology and computer science, disciplines where systems are often modeled as graphs. Such inherent community structures make us deeply understand about the networks and therefore have drawn significant interests among researchers. This paper describes a probabilistic community detection algorithm by modeling topic on sampling sub graphs. In this algorithm, the communities are modeled as latent topic variables of an LDA topic model and the vertices of sampling sub graphs are drawn from these topics with different probabilities. This paper also proposes a sub graph sampling algorithm and explores its impact on community detection performance. Our algorithm is evaluated by extensive experiments using many computer-generated artificial graphs and real-world networks. The results show that our algorithm is effective in detecting probabilistic community. Zengfeng Zeng, Bin Wu 0001 |
ASONAM | 2 |
| 2012 | A Method for Local Community Detection by Finding Core NodesabstractCurrently, the detection of global community structure in networks has gathered a lot of attention. Most of the methods need global knowledge of the graphs which would be unrealistic to get when the graphs are too large or evolve too quickly. Moreover, sometimes we are only interested in the community structures of some given nodes, not all nodes. So detecting the community of a given node i.e. local community detection is more appropriate. Most of the proposed solutions for local community detection built upon the source nodes are sensitive to the position of source nodes. In this paper, we propose a method to detect local community of a given node by finding the core node of the community firstly. Then expand the core node's cliques to get community of the given node. We validate our method on real-world networks whose community structures are available. The result shows that our method can get high recall and precision score and is quite effective and flexible to identify local communities, irrespective of the source node position. Bin Wu 0001 |
ASONAM | 2 |
| 2012 | HMGraph OLAP: a novel framework for multi-dimensional heterogeneous network analysisabstractAs information continues to grow at an explosive rate, more and more heterogeneous network data sources are coming into being. While OLAP (On-Line Analytical Processing) techniques have been proven effective for analyzing and mining structured data, unfortunately, to our best knowledge, there are no OLAP tools available that are able to analyze multi-dimensional heterogeneous networks from different perspectives and with multiple granularities. Therefore, we have developed a novel HMGraph OLAP (Heterogeneous and Multi-dimensional Graph OLAP) framework for the purpose of providing more dimensions and operations to mine multi-dimensional heterogeneous information network. After information dimensions and topological dimensions, we have been the first to propose entity dimensions, which represent an important dimension for heterogeneous network analysis. On the basis of this notion, we designed HMGraph OLAP operations named (Rotate and Stretch for entity dimensions, which are able to mine relationships between different entities. We then proposed the HMGraph Cube, which is an efficient data warehousing model for HMGraph OLAP. In addition, through comparison with common strategies, we have shown that the optimizations we have proposed deliver better performance. Finally, we have implemented a HMGraph OLAP prototype, LiterMiner, which has proven effective for the analysis of multi-dimensional heterogeneous networks. Mu Yin, Bin Wu 0001, Zengfeng Zeng |
DOLAP | 2 |
| 2012 | Relevance search in heterogeneous networksabstractConventional research on similarity search focuses on measuring the similarity between objects with the same type. However, in many real-world applications, we need to measure the relatedness between objects with different types. For example, in automatic expert profiling, people are interested in finding the most relevant objects to an expert, where the objects can be of various types, such as research areas, conferences and papers, etc. With the surge of study on heterogeneous networks, the relatedness measure on objects with different types becomes increasingly important. In this paper, we study the relevance search problem in heterogeneous networks, where the task is to measure the relatedness of heterogeneous objects (including objects with the same type or different types). We propose a novel measure, called HeteSim, with the following attributes: (1) a path-constrained measure: the relatedness of object pairs are defined based on the search path that connect two objects through following a sequence of node types; (2) a uniform measure: it can measure the relatedness of objects with the same or different types in a uniform framework; (3) a semi-metric measure: HeteSim has some good properties (e.g., self-maximum and symmetric), that are crucial to many tasks. Empirical studies show that HeteSim can effectively evaluate the relatedness of heterogeneous objects. Moreover, in the query and clustering tasks, it can achieve better performances than conventional measures. Chuan Shi 0001, Xiangnan Kong, Philip S. Yu, Sihong Xie, Bin Wu 0001 |
EDBT | 5 |
| 2012 | BC-PDM: data mining, social network analysis and text mining system based on cloud computingabstractTelecom BI(Business Intelligence) system consists of a set of application programs and technologies for gathering, storing, analyzing and providing access to data, which contribute to manage business information and make decision precisely. However, traditional analysis algorithms meet new challenges as the continued exponential growth in both the volume and the complexity of telecom data. With the Cloud Computing development, some parallel data analysis systems have been emerging. However, existing systems have rarely comprehensive function, either providing data analysis service or providing social network analysis. We need a comprehensive tool to store and analysis large scale data efficiently. In response to the challenge, the SaaS (Software-as-a-Service) BI system, BC-PDM (Big Cloud-Parallel Data Mining), are proposed. BC-PDM supports parallel ETL process, statistical analysis, data mining, text mining and social network analysis which are based on Hadoop. This demo introduces three tasks: business recommendation, customer community detection and user preference classification by employing a real telecom data set. Experimental results show BC-PDM is very efficient and effective for intelligence data analysis. Wei Chong Shen, Bin Wu 0001, Bai Wang 0001, Bo Ren Zhang |
KDD | 4 |
| 2011 | A Novel Genetic Algorithm for Overlapping Community Detection
Yanan Cai, Chuan Shi 0001, Yuxiao Dong, Qing Ke, Bin Wu 0001 |
ADMA (1) | 5 |
| 2011 | Link Prediction Based on Local InformationabstractLink prediction in complex networks is an important issue in graph mining. It aims at estimating the likelihood of the existence of links between nodes by the know network structure information. Currently, most link prediction algorithms based on local information consider only the individual characteristics of common neighbors. In this paper, first, we study the link prediction results as the change of the exponent on the degree of common neighbors, and find some regular pattern between different networks and different exponent. After that, we come up with a new algorithm exploiting the interactions between common neighbors, namely Individual Attraction Index. To reduce the time complexity, we design a simple edition, called Simple Individual Attraction Index. We compare nine well-known local information metrics on eight real networks. The result proves well the best overall performance of these two new algorithms. Yuxiao Dong, Qing Ke, Bai Wang 0001, Bin Wu 0001 |
ASONAM | 4 |
| 2011 | Efficient Search in Networks Using ConductanceabstractDecentralized search in networks is an important algorithmic problem in the study of complex networks and social networks analysis. It has a large number of practical applications, from shortest paths search in social network relationship, web pages search in WWW to querying files in peer-to-peer file sharing networks and so on. In this paper, we explore this problem from a perspective of community structure. We first find that through maximizing sample conductance, we can get high coverage sample. Based on this result, then, we propose a new decentralized search strategy named Conductance Search which tries to efficiently find the nodes belonging to different communities. We compare the strategy with other common strategies. And the results show that the conductance search outperforms others in number of steps to find the target and time complexity. Finally, we find some previous conclusions fail in many real-world networks and discuss network search-ability from the perspective of various structural properties. Qing Ke, Yuxiao Dong, Bin Wu 0001 |
ASONAM | 3 |
| 2011 | Detecting Link Communities in Massive NetworksabstractMost of the existing literature which has entirely focused on clustering nodes in large-scale networks. To discover multi-scale overlapping communities quickly, we propose a highly efficient multi-resolution link community detection algorithm to detect the link communities in massive networks based on the idea of edge labeling. First, we will get the node partition of the network based on a new multi-resolution node detection algorithm. After that, we can find the link community in a linear time by the labels of nodes. Its time complexity is near linear and its space complexity is linear. The effectiveness of our algorithm is demonstrated by extensive experiments on lots of computer generated artificial graphs and real-world networks. The results show that our algorithm is very fast and highly reliable. Tests on real and artificial networks also give excellent results comparing with the newly proposed link partition algorithm. Qi Ye 0008, Bin Wu 0001, Zhixiong Zhao, Bai Wang 0001 |
ASONAM | 2 |
| 2011 | On selection of objective functions in multi-objective community detectionabstractThere is a surge of community detection of complex networks in recent years. Different from conventional single-objective community detection, this paper formulates community detection as a multi-objective optimization problem and proposes a general algorithm NSGA-Net based on evolutionary multi-objective optimization. Interested in the effect of optimization objectives on the performance of the multi-objective community detection, we further study the correlations (i.e., positively correlated, independent, or negatively correlated) of 11 objective functions that have been used or can potentially be used for community detection. Our experiments show that NSGA-Net optimizing over a pair of negatively correlated objectives usually performs better than the single-objective algorithm optimizing over either of the original objectives, and even better than other well-established community detection approaches. Chuan Shi 0001, Philip S. Yu, Yanan Cai, Zhenyu Yan 0001, Bin Wu 0001 |
CIKM | 5 |
| 2010 | A Novel Algorithm for Hierarchical Community Structure Detection in Complex Networks
Chuan Shi 0001, Liangliang Shi, Yanan Cai, Bin Wu 0001 |
ADMA (1) | 5 |
| 2010 | Distance Distribution and Average Shortest Path Length Estimation in Real-World Networks
Qi Ye 0008, Bin Wu 0001, Bai Wang 0001 |
ADMA (1) | 2 |
| 2010 | Multiple Level Views on the Adherent Cohesive Subgraphs in Massive Temporal Call Graphs
Qi Ye 0008, Bin Wu 0001, Bai Wang 0001 |
ADMA (1) | 2 |
| 2010 | Detecting Communities in Massive Networks Based on Local Community Attractive Force OptimizationabstractCurrently, community detection has led to a huge interest in data analysis on real-world networks. However, the high computationally demanding of most community detection algorithms limits their applications. In this paper, we propose a heuristic algorithm to extract the community structure in large networks based on local community attractive force optimization whose time complexity is near linear and space complexity is linear. The effectiveness of our algorithm is demonstrated by extensive experiments on lots of computer generated graphs and public available real-world graphs. The result shows our algorithm is extremely fast, and it is easy for us to explore massive networks interactively. Qi Ye 0008, Bin Wu 0001, Bai Wang 0001 |
ASONAM | 2 |
| 2010 | Mining program workflow from interleaved tracesabstractSuccessful software maintenance is becoming increasingly critical due to the increasing dependence of our society and economy on software systems. One key problem of software maintenance is the difficulty in understanding the evolving software systems. Program workflows can help system operators and administrators to understand system behaviors and verify system executions so as to greatly facilitate system maintenance. In this paper, we propose an algorithm to automatically discover program workflows from event traces that record system events during system execution. Different from existing workflow mining algorithms, our approach can construct concurrent workflows from traces of interleaved events. Our workflow mining approach is a three-step coarse-to-fine algorithm. At first, we mine temporal dependencies for each pair of events. Then, based on the mined pair-wise tem-poral dependencies, we construct a basic workflow model by a breadth-first path pruning algorithm. After that, we refine the workflow by verifying it with all training event traces. The re-finement algorithm tries to find out a workflow that can interpret all event traces with minimal state transitions and threads. The results of both simulation data and real program data show that our algorithm is highly effective. Jian-Guang Lou, Qiang Fu 0015, Shengqi Yang, Jiang Li 0008, Bin Wu 0001 |
KDD | 5 |
| 2010 | Group-Level Analysis by Extracting Semantic Relations from Query GraphabstractRecently, a growing number of researches have focused on the issues raised by the knowledge discovery of online information, particularly the problems of tracking topics, ideas, and users' spreading influence across the Web. In this paper, the search-engine query logs on Topic Detection and Tracking (TDT) is analyzed other than study of the quality of the search result or query recommendation. By constructing a novel bi-type heterogeneous query graph, the queries' semantic similarity and query-URL relation are combined together. Utilizing social network analysis (SNA) method to analyze the query graph with optimization of the community discovery algorithm LPA by grouping the nodes who are linked with the same URL initially, we can find the topics in the query logs. To evaluate the topic evolution pattern, we group the similar communities over each adjacent time stamps into clusters. Extensive experiments demonstrate the effectiveness and efficiency of the methods. Bin Wu 0001, Tian Zhu 0001, Weiduo Wang, Qi Ye 0008, Bai Wang 0001 |
Web Intelligence | 1 |
| 2010 | Empirical Analysis and Multiple Level Views in Massive Social NetworksabstractWith the emergence of massive social media, massive social networks have led to a huge interest in data analysis. In this paper, we propose an empirical study on several massive social networks including 4 mobile call graphs, a fixed-line call graph, two co-authorship networks and two Email networks. We find that call graphs tend to be more locality than the co-authorship networks and Email networks. To our surprise, we even find that there is no significant relations between community sizes and their quality scores for most extracted communities. We also find that some very huge community with high mean quality values, and we can not find the universal "V" shape in their mean quality values. Qi Ye 0008, Bin Wu 0001, Bai Wang 0001 |
Web Intelligence | 2 |
| 2009 | Structure Correlation in Mobile Call Networks
Deyong Hu, Bin Wu 0001, Qi Ye 0008, Bai Wang 0001 |
ADMA | 2 |
| 2009 | VisNetMiner: An Integration Tool for Visualization and Analysis of Networks
Chuan Shi 0001, Bin Wu 0001, Jian Liu 0001 |
ADMA | 3 |
| 2009 | Social Influence and Role Analysis Based on Community Structure in Social Network
Tian Zhu 0001, Bin Wu 0001, Bai Wang 0001 |
ADMA | 2 |
| 2009 | TeleComVis: Exploring Temporal Communities in Telecom Networks
Qi Ye 0008, Bin Wu 0001, Lijun Suo, Tian Zhu 0001, Bai Wang 0001 |
ECML/PKDD (2) | 2 |
| 2009 | CosDic: Towards a Comprehensive System for Knowledge Discovery in Large-Scale Data: Architecture, Implementation and Case StudiesabstractThe continued exponential growth in both the volume and the complexity of information is giving birth to a new challenge to the specific requirements of analysts, researchers and intelligence providers. In this paper, to move the scientific activity forward to practice, we elaborate a prototype of our on-going constructed system, CosDic, for knowledge discovery from extremely large-scale datasets. The major infrastructure of CosDic is deployed on a distributed cluster environment using MapReduce platform. To undertake the mining tasks from gigabytes to petabytes, we carefully devised our system, from architecture to particular algorithms, from under layer construction to upper layer public service interface, from effectiveness to efficiency. Moreover, to illustrate its functionality, we employ CosDic to a real-world huge dataset and demonstrate an integrated analysis procedure from initial raw data preprocessing to finally knowledge discovering. We show that CosDic has a good performance in such cloud-scale data computing. Bin Wu 0001, Shengqi Yang, Haizhou Zhao, Lijun Suo |
Web Intelligence | 1 |
| 2008 | CommTracker: A Core-Based Algorithm of Tracking Community Evolution
Yi Wang 0010, Bin Wu 0001, Xin Pei |
ADMA | 2 |
| 2008 | JSNVA: A Java Straight-Line Drawing Framework for Network Visual Analysis
Qi Ye 0008, Bin Wu 0001, Bai Wang 0001 |
ADMA | 2 |
| 2008 | Overlapping community structure detection in networksabstractMany systems in nature and human society take the form of networks with community structures. In this paper, we describe a simple algorithm COCD(Clique-based Overlapping Community Detection) to efficiently mine the overlapping communities in large-scale networks, which is useful for us to have a better understanding of the nested sub-structures embedded in the whole network. Bai Wang 0001, Bin Wu 0001 |
CIKM | 3 |
| 2008 | Overlapping Community Detection in Bipartite NetworksabstractResearches have discovered that rich interactions among entities in nature and human society bring about complex networks with community structures. In this paper, we propose a novel algorithm BiTector (bi-community detector) to mine the overlapping communities in large-scale sparse bipartite networks. We apply the algorithm to various real-world datasets, showing that BiTector can identify the overlapping community structures in the bipartite networks efficiently and effectively. Bai Wang 0001, Bin Wu 0001, Yi Wang 0010 |
Web Intelligence | 3 |
| 2007 | Backbone Discovery in Social NetworksabstractRecent years have seen a thriving development of the World Wide Web as the most visible social media which enables people to share opinions, experiences and expertise with each other across the world. People now get involved in many different social networks simultaneously, which are often large intricate web of connections among the massive entities they are made of. As a result, the challenge of collecting and analyzing large-scale data among social members has left most basic questions about the global composition and function of such networks largely unresolved: What is the essential organization of a social network? who are the influential individuals whose voice is echoed by others? To address these questions, this paper presents an algorithm called sketcher to discover and describe the overall backbone of a specific network. Experimental results on the American College Football, Scientific Collaboration, and Telecommunications Call networks show that sketcher can extract the essential composition of a social network both efficiently and intuitively. Bin Wu 0001, Bai Wang 0001 |
Web Intelligence | 2 |
| 2006 | A New Algorithm for Enumerating All Maximal Cliques in Complex Network
Bin Wu 0001, Qi Ye 0008 |
ADMA | 2 |