Xiaoming Zhang 0001

dblp:86/2120-1 · also Xiao-Ming Zhang 0001 · DBLP profile ↗
← Back
81ranked-venue papers
20as first author
28since 2021 · last 2026
0000-0002-6662-4102ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 30 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 7 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Counterfactual Question Generation Uncovering Learner Contradictions
abstract
Conventional feedback, even when accompanied by brief explanations, rarely uncovers the hidden contradictions that trigger a learner's mistake. We bridge this gap with counterfactual question generation (CFQG): given a learner's answer, generate a follow-up question that deliberately contradicts it, compelling the learner to confront the underlying conflict. CFQG thus transforms assessment from passive scoring into an interactive and contradiction-centered dialogue that supports knowledge repair. To automate CFQG, we propose GapProbe, which probes the knowledge gap between a learner’s belief and curated facts through a knowledge graph (KG), then designs counterfactual questions (CFQs) that negate the belief. Identifying contradiction-aware triples, and more importantly, selecting those most likely to confuse the learner, are highly challenging in large-scale KGs. GapProbe tackles these challenges with an iterative ProConB cycle coupled with a schema-aware KGMap. By caching one- and multi-hop schema patterns of the KG, KGMap provides ``roadmap'' to guide LLMs jump to deep and contradiction-aware triples, beyond traditional step-wise graph traversal. We present the CFQG benchmark and corresponding metrics for evaluating how generated CFQs trigger, focus, and deepen learner reflection through explicit contradictions. Experiments on multiple datasets and LLMs show that GapProbe boosts LLM reasoning over KGs and generates follow-up questions that consistently promote deeper and more focused learner reflection.
Bo Zhang 0096, Yvhang Yang, Dezhuang Miao, Fengyi Song, Yanhui Gu, Xiaoming Zhang 0001, Junsheng Zhou
AAAI8
2026 A multi-scale representation and multi-level decision learning network for multimodal sentiment analysis
Xiang Li 0117, Zhiqiang Dong, Xianfu Cheng, Dezhuang Miao, Haijun Zhang 0007, Tianbo Wang 0001, Xiaoming Zhang 0001, Zhoujun Li 0001
Expert Syst. Appl.7
2025 What Is a Good Question? Assessing Question Quality via Meta-Fact Checking
abstract
Knowledge-based questions are typically employed to evaluate LLM's knowledge boundaries; meanwhile, numerous studies focus on question generation as a means to enhance the capabilities of both models and individuals. However, there is a lack of in-depth exploration about what constitutes a good question from the perspective of knowledge cognition. This paper proposes aligning the complete knowledge underlying questions with educational criteria effectively employed in physics courses, thereby developing novel knowledge-intensive metrics of question quality. To this end, we propose Meta-Fact Checking (MFC), which transforms questions into knowledge graph (KG) triples utilizing LLMs through few-shot prompting, thereby quantifying question quality based on the patterns observed within these triples. MFC introduces a novel interaction mechanism for KGs that communicates meta-facts, illustrating the types of knowledge that KGs can offer to the LLM for reasoning questions, rather than relying solely on the original triples. This strategy ensures that MFC remains unaffected by unexplored triples that LLM has not yet encountered within KGs compared to the retrieve-while-reasoning routine. Experiments across multiple datasets and LLMs demonstrate that MFC significantly improves the accuracy and efficiency of both question answering and assessing. This research marks a pioneering effort to automate the evaluation of question quality based on cognitive capabilities.
Bo Zhang 0096, Jianghua Zhu, Chaozhuo Li, Dezhuang Miao, Xiaoming Zhang 0001, Junsheng Zhou
AAAI8
2025 Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection
abstract
The proliferation of fake news on social media platforms has exerted a substantial influence on society, leading to discernible impacts and deleterious consequences. Conventional deep learning methodologies employing small language models (SLMs) suffer from the necessity for extensive supervised training and the challenge of adapting to rapidly evolving circumstances. Large language models (LLMs), despite their robust zero-shot capabilities, have fallen short in effectively identifying fake news due to a lack of pertinent demonstrations and the dynamic nature of knowledge. In this paper, a novel framework Multi-Round Collaboration Detection (MRCD) is proposed to address these aforementioned limitations. The MRCD framework is capable of enjoying the merits from both LLMs and SLMs by integrating their generalization abilities and specialized functionalities, respectively. Our approach features a two-stage retrieval module that selects relevant and up-to-date demonstrations and knowledge, enhancing in-context learning for better detection of emerging news events. We further design a multi-round learning framework to ensure more reliable detection results. Our framework MRCD achieves SOTA results on two real-world datasets Pheme and Twitter16, with accuracy improvements of 7.4% and 12.8% compared to using only SLMs, which effectively addresses the limitations of current models and improves the detection of emergent fake news detection.
Ziyi Zhou 0003, Xiaoming Zhang 0001, Shenghan Tan, Litian Zhang, Chaozhuo Li
AAAI2
2025 Learning fine-grained representation with token-level alignment for multimodal sentiment analysis
Xiang Li 0117, Haijun Zhang 0007, Zhiqiang Dong, Xianfu Cheng, Yun Liu 0017, Xiaoming Zhang 0001
Expert Syst. Appl.6
2025 Moment matching of joint distributions for unsupervised domain adaptation
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Yancong Li, Feiran Huang
Inf. Process. Manag.2
2025 Image captioning with residual swin transformer and Actor-Critic
Zhoujun Li 0001, Xiaoming Zhang 0001, Feiran Huang
Neural Comput. Appl.4
2025 Early Detection of Multimodal Fake News via Reinforced Propagation Path Generation
abstract
Amidst the rapid propagation of multimodal fake news across social media platforms, the detection of fake news has emerged as a prime research pursuit. To detect heightened level of meticulous fabrications, propagation paths are introduced to provide nuanced social context that enhances the basic semantic analysis of the news content. However, existing propagation-enhanced models encounter a dilemma between detection efficacy and social hazard. In this paper, we explore the innovative problem of early fake news detection through the generation of propagation paths, capable of benefiting from the extensive social context within propagation paths while mitigating potential social hazards. To address these challenges, we propose a novel Reinforced Propagation Path Generation Fake News Detection model,RPPG-Fake. Departing from conventional discriminative approaches,RPPG-Fakecaptures the propagation topology pattern from a heterogeneous social graph and generates the propagation paths to detect fake news effectively under a reinforcement learning paradigm. Our proposal is extensively evaluated over three popular datasets, and experimental results demonstrate the superiority of our proposal.
Litian Zhang, Xiaoming Zhang 0001, Ziyi Zhou 0003, Xi Zhang 0008, Senzhang Wang, Philip S. Yu, Chaozhuo Li
IEEE Trans. Knowl. Data Eng.2
2024 Reinforced Adaptive Knowledge Learning for Multimodal Fake News Detection
abstract
Nowadays, detecting multimodal fake news has emerged as a foremost concern since the widespread dissemination of fake news may incur adverse societal impact. Conventional methods generally focus on capturing the linguistic and visual semantics within the multimodal content, which fall short in effectively distinguishing the heightened level of meticulous fabrications. Recently, external knowledge is introduced to provide valuable background facts as complementary to facilitate news detection. Nevertheless, existing knowledge-enhanced endeavors directly incorporate all knowledge contexts through static entity embeddings, resulting in the potential noisy and content-irrelevant knowledge. Moreover, the integration of knowledge entities makes it intractable to model the sophisticated correlations between multimodal semantics and knowledge entities. In light of these limitations, we propose a novel Adaptive Knowledge-Aware Fake News Detection model, dubbed AKA-Fake. For each news, AKA-Fake learns a compact knowledge subgraph under a reinforcement learning paradigm, which consists of a subset of entities and contextual neighbors in the knowledge graph, restoring the most informative knowledge facts. A novel heterogeneous graph learning module is further proposed to capture the reliable cross-modality correlations via topology refinement and modality-attentive pooling. Our proposal is extensively evaluated over three popular datasets, and experimental results demonstrate the superiority of AKA-Fake.
Litian Zhang, Xiaoming Zhang 0001, Ziyi Zhou 0003, Feiran Huang, Chaozhuo Li
AAAI2
2024 Hyperbolic Graph Neural Network for Temporal Knowledge Graph Completion
abstract
Temporal Knowledge Graphs (TKGs) represent a crucial source of structured temporal information and exhibit significant utility in various real-world applications. However, TKGs are susceptible to incompleteness, necessitating Temporal Knowledge Graph Completion (TKGC) to predict missing facts. Existing models have encountered limitations in effectively capturing the intricate temporal dynamics and hierarchical relations within TKGs. To address these challenges, HyGNet is proposed, leveraging hyperbolic geometry to effectively model temporal knowledge graphs. The model comprises two components: the Hyperbolic Gated Graph Neural Network (HGGNN) and the Hyperbolic Convolutional Neural Network (HCNN). HGGNN aggregates neighborhood information in hyperbolic space, effectively capturing the contextual information and dependencies between entities. HCNN interacts with embeddings in hyperbolic space, effectively modeling the complex interactions between entities, relations, and timestamps. Additionally, a consistency loss is introduced to ensure smooth transitions in temporal embeddings. The extensive experimental results conducted on four benchmark datasets for TKGC highlight the effectiveness of HyGNet. It achieves state-of-the-art performance in comparison to previous models, showcasing its potential for real-world applications that involve temporal reasoning and knowledge prediction.
Yancong Li, Xiaoming Zhang 0001, Shuai Ma 0001
LREC/COLING2
2024 Mitigating Social Hazards: Early Detection of Fake News via Diffusion-Guided Propagation Path Generation
abstract
The detection of fake news has emerged as a pressing issue in the era of online social media. To detect meticulously fabricated fake news, propagation paths are introduced to provide nuanced social context to complement the pure semantics within news content. However, existing propagation-enhanced models face a dilemma between detection efficacy and social hazard. In this paper, we investigate the novel problem of early fake news detection via propagation path generation, capable of enjoying the merits of rich social context within propagation paths while alleviating potential social hazards. In contrast to previous discriminative detection models, we further propose a novel generative model, DGA-Fake, by simulating realistic propagation paths based on news content before actual spreading. A guided diffusion module is integrated into DGA-Fake to generate simulated user interaction sequences, guided by historical interactions and news content. Evaluation across three datasets demonstrates the superiority of our proposal.
Litian Zhang, Xiaoming Zhang 0001, Chaozhuo Li, Ziyi Zhou 0003, Feiran Huang, Xi Zhang 0008
ACM Multimedia2
2024 Blockchain-Based Lightweight and Privacy-Preserving Quality Assurance Framework in Crowdsensing Systems
abstract
The novel sensing paradigm known as crowdsensing leverages ubiquitous smart devices to collect data in Internet of Things (IoT) applications. Traditional crowdsensing schemes assume a central framework to execute truth discovery algorithm to assure data quality, which may introduce reliability and privacy issues. Blockchain is a promising technology that provides a decentralized, transparent, and immutable platform. However, designing a blockchain-based quality assurance scheme in crowdsensing is not a trivial problem. First, truth discovery is a time-consuming iterative algorithm, which is not practical to execute on blockchain. Second, privacy-preserving schemes always require that the participants join in multiround communications, which is not acceptable in open blockchain because of users’ highly unpredictable behaviors. Finally, on-chain data are publicly accessible, and achieving a good balance between data utility and privacy is an important issue. In this article, we propose a lightweight quality assurance framework atop blockchain to build a reliable, privacy preserving, and fair crowdsensing system. Specifically, we carefully design two kinds of smart contracts to cooperatively maintain a long-term reliable platform to execute crowdsensing tasks. In the contracts, we devise a reputation-based aggregator selection algorithm to reach the consensus on truthful results while avoiding expensive on-chain iterative processes. The participant selection scheme and reward policy are further utilized to filter appropriate participants to complete the task. Our scheme also protects data privacy and does not require communications between participants. Finally, we implement and deploy the contracts on Ethereum and conduct extensive experiments to demonstrate that the contracts can practically execute crowdsensing tasks.
Chang Xu 0004, Liehuang Zhu, Rongxing Lu, Yunguo Guan, Xiaoming Zhang 0001
IEEE Internet Things J.6
2024 Cross-domain knowledge collaboration for blending-target domain adaptation
Bo Zhang 0096, Xiaoming Zhang 0001, Feiran Huang, Dezhuang Miao
Inf. Process. Manag.2
2024 SANe: Space adaptation network for temporal knowledge graph completion
Yancong Li, Xiaoming Zhang 0001, Bo Zhang 0096, Feiran Huang, Shuai Ma 0001
Inf. Sci.2
2024 Deep hyperbolic convolutional model for knowledge graph embedding
Yancong Li, Jiangxiao Zhang, Haiying Ren, Xiaoming Zhang 0001
Knowl. Based Syst.5
2023 Illegal Accounts Detection on Ethereum Using Heterogeneous Graph Transformer Networks
Chang Xu 0004, Liehuang Zhu, Xiaoming Zhang 0001
ICICS5
2023 CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization
abstract
Multimodal summarization (MS) aims to generate a summary from multimodal input. Previous works mainly focus on textual semantic coverage metrics such as ROUGE, which considers the visual content as supplemental data. Therefore, the summary is ineffective to cover the semantics of different modalities. This paper proposes a multi-task cross-modality learning framework (CISum) to improve multimodal semantic coverage by learning the cross-modality interaction in the multimodal article. To obtain the visual semantics, we translate images into visual descriptions based on the correlation with text content. Then, the visual description and text content are fused to generate the textual summary to capture the semantics of the multimodal content, and the most relevant image is selected as the visual summary. Furthermore, we design an automatic multimodal semantics coverage metric to evaluate the performance. Experimental results show that CISum outperforms baselines in multimodal semantics coverage metrics while maintaining the excellent performance of ROUGE and BLEU.
Litian Zhang, Xiaoming Zhang 0001, Ziming Guo
SDM2
2023 How to Find a Bitcoin Mixer: A Dual Ensemble Model for Bitcoin Mixing Service Detection
abstract
Bitcoin is the first decentralized peer-to-peer cryptocurrency that has gained popularity by providing users with transaction anonymity. With the development of Bitcoin and the higher privacy requirements of users, mixing services have emerged to enhance Bitcoin anonymity by obfuscating the flow of funds. However, they are also widely used for illegal activities due to its strong anonymity, especially for money laundering. Therefore, detecting mixing services has great significance for Bitcoin anti-money laundering. In this article, we propose a novel detection scheme to identify the addresses belonging to Bitcoin mixing services. Specifically, we first construct the Bitcoin mixing data set, which summarizes a total of 26 features to describe the transaction behavior of addresses. Next, we design a new classification model, called the Dual Ensemble Classification Model. The model combines the advantages of multiple models based on different algorithms and obtains better classification performance. In order to detect more complex mixing patterns, we also extract transaction subgraphs from the established Bitcoin address-transaction network. The subgraphs are then classified using a kernel-based graph classification method, which is embedded in the model. Comprehensive experiments on three data sets demonstrate the effectiveness of our scheme, and the proposed model has a detection accuracy of 99.84% for the Bitcoin mixing service.
Chang Xu 0004, Ruting Xiong, Liehuang Zhu, Xiaoming Zhang 0001
IEEE Internet Things J.5
2023 Deep Kernel Network Embedding
abstract
This paper concerns the problem of network embedding (NE), whose aim is to learn a low-dimensional representation for each node in networks. We provide a new train to solve the sparsity problem where most of nodes including the new arrival nodes have little knowledge with respect to the network. A novel paradigm is proposed to integrate the multiple information from the subgraph covering the target node instead of only the target node. Particularly, to epxress the distinctive feature over the vertex domain, a probabiltiy distribution over subgraph space is constructed for each node. The distribution is more effective to express the distinctive characteristic and feature in a higher dimension compared to the latent representation vectors. One of the primary goals of this paradigm is to define the convolution operation over the distributions, which are efficient to evaluate and learn. Experiments on four real-world network datasets demonstrate that our approach significantly outperforms state-of-the-art methods, especially on the representation learning for the nodes newly joining in the network.
Bo Zhang 0096, Xiaoming Zhang 0001, Feiran Huang, Shuai Ma 0001
IEEE Trans. Knowl. Data Eng.2
2022 Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization
abstract
Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observed that different modalities of data in the news report correlate hierarchically. Traditional MSMO methods indistinguishably handle different modalities of data by learning a representation for the whole data, which is not directly adaptable to the heterogeneous contents and hierarchical correlation. In this paper, we propose a hierarchical cross-modality semantic correlation learning model (HCSCL) to learn the intra- and inter-modal correlation existing in the multimodal data. HCSCL adopts a graph network to encode the intra-modal correlation. Then, a hierarchical fusion framework is proposed to learn the hierarchical correlation between text and images. Furthermore, we construct a new dataset with relevant image annotation and image object label information to provide the supervision information for the learning procedure. Extensive experiments on the dataset show that HCSCL significantly outperforms the baseline methods in automatic summarization metrics and fine-grained diversity tests.
Litian Zhang, Xiaoming Zhang 0001, Junshu Pan
AAAI2
2022 Dynamic self-attention with vision synchronization networks for video question answering
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Shixun Shen, Zhoujun Li 0001
Pattern Recognit.2
2022 ALSA: Adversarial Learning of Supervised Attentions for Visual Question Answering
abstract
Visual question answering (VQA) has gained increasing attention in both natural language processing and computer vision. The attention mechanism plays a crucial role in relating the question to meaningful image regions for answer inference. However, most existing VQA methods: 1) learn the attention distribution either from free-form regions or detection boxes in the image, which is intractable in answering questions about the foreground object and background form, respectively and 2) neglect the prior knowledge of human attention and learn the attention distribution with an unguided strategy. To fully exploit the advantages of attention, the learned attention distribution should focus more on the question-related image regions, such as human attention for both the questions, about the foreground object and background form. To achieve this, this article proposes a novel VQA model, called adversarial learning of supervised attentions (ALSAs). Specifically, two supervised attention modules: 1) free form-based and 2) detection-based, are designed to exploit the prior knowledge for attention distribution learning. To effectively learn the correlations between the question and image from different views, that is, free-form regions and detection boxes, an adversarial learning mechanism is implemented as an interplay between two supervised attention modules. The adversarial learning reinforces the two attention modules mutually to make the learned multiview features more effective for answer inference. The experiments performed on three commonly used VQA datasets confirm the favorable performance of ALSA.
Yun Liu 0017, Xiaoming Zhang 0001, Zhiyun Zhao, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Cybern.2
2022 Cross-Attentional Spatio-Temporal Semantic Graph Networks for Video Question Answering
abstract
Due to the rich spatio-temporal visual content and complex multimodal relations, Video Question Answering (VideoQA) has become a challenging task and attracted increasing attention. Current methods usually leverage visual attention, linguistic attention, or self-attention to uncover latent correlations between video content and question semantics. Although these methods exploit interactive information between different modalities to improve comprehension ability, inter- and intra-modality correlations cannot be effectively integrated in a uniform model. To address this problem, we propose a novel VideoQA model called Cross-Attentional Spatio-Temporal Semantic Graph Networks (CASSG). Specifically, a multi-head multi-hop attention module with diversity and progressivity is first proposed to explore fine-grained interactions between different modalities in a crossing manner. Then, heterogeneous graphs are constructed from the cross-attended video frames, clips, and question words, in which the multi-stream spatio-temporal semantic graphs are designed to synchronously reasoning inter- and intra-modality correlations. Last, the global and local information fusion method is proposed to coalesce the local reasoning vector learned from multi-stream spatio-temporal semantic graphs and the global vector learned from another branch to infer the answer. Experimental results on three public VideoQA datasets confirm the effectiveness and superiority of our model compared with state-of-the-art methods.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Image Process.2
2021 Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation
abstract
Bo Zhang, Xiaoming Zhang, Yun Liu, Lei Cheng, Zhoujun Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Zhoujun Li 0001
ACL/IJCNLP (1)2
2021 Discriminative Feature Adaptation via Conditional Mean Discrepancy for Cross-Domain Text Classification
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017
DASFAA (2)2
2021 Dual self-attention with co-attention networks for visual question answering
Yun Liu 0017, Xiaoming Zhang 0001, Qianyun Zhang 0001, Chaozhuo Li, Feiran Huang, Xianghong Tang, Zhoujun Li 0001
Pattern Recognit.2
2021 Multimodal Learning of Social Image Representation by Exploiting Social Relations
abstract
Learning the representation for social images has recently made remarkable achievements for many tasks, such as cross-modal retrieval and multilabel classification. However, since social images contain both multimodal contents (e.g., visual images and textual descriptions) and social relations among images, simply modeling the content information may lead to suboptimal embedding. In this paper, we propose a novel multimodal representation learning model for social images, that is, correlational multimodal variational autoencoder (CMVAE) via triplet network. Specifically, in order to mine the highly nonlinear correlation between the visual content and the textual content, a CMVAE is proposed to learn a unified representation for the multiple modalities of social images. Both common information in all modalities and private information in each modality are encoded for the representation learning. To incorporate the social relations among images, we employ the triplet network to embed multiple types of social links in the representation learning. Then, a joint embedding model is proposed to combine the social relations for representation learning of the multimodal contents. Comprehensive experiment results on four datasets confirm the effectiveness of our method in two tasks, namely, multilabel classification and cross-modal retrieval. Our method outperforms the state-of-the-art multimodal representation learning methods with significant improvement of performance.
Feiran Huang, Xiaoming Zhang 0001, Jie Xu 0015, Zhonghua Zhao, Zhoujun Li 0001
IEEE Trans. Cybern.2
2021 Adversarial Learning With Multi-Modal Attention for Visual Question Answering
abstract
Visual question answering (VQA) has been proposed as a challenging task and attracted extensive research attention. It aims to learn a joint representation of the question-image pair for answer inference. Most of the existing methods focus on exploring the multi-modal correlation between the question and image to learn the joint representation. However, the answer-related information is not fully captured by these methods, which results that the learned representation is ineffective to reflect the answer of the question. To tackle this problem, we propose a novel model, i.e., adversarial learning with multi-modal attention (ALMA), for VQA. An adversarial learning-based framework is proposed to learn the joint representation to effectively reflect the answer-related information. Specifically, multi-modal attention with the Siamese similarity learning method is designed to build two embedding generators, i.e., question-image embedding and question-answer embedding. Then, adversarial learning is conducted as an interplay between the two embedding generators and an embedding discriminator. The generators have the purpose of generating two modality-invariant representations for the question-image and question-answer pairs, whereas the embedding discriminator aims to discriminate the two representations. Both the multi-modal attention module and the adversarial networks are integrated into an end-to-end unified framework to infer the answer. Experiments performed on three benchmark data sets confirm the favorable performance of ALMA compared with state-of-the-art approaches.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhoujun Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 A11 Your PLCs Belong to Me: ICS Ransomware Is Realistic
abstract
Ransomware is a new business model for cybercrime which mainly targets individual users and machines. Many events have shown how profitable the technique can be. Industrial control systems (ICS) are becoming the next domain. More and more researchers and attackers have become the focus on this field and presented some ICS ransomware. But existing ICS ransomware is theoretically feasible and has a limited effect on real ICS. In this work, we present ICS-BROCK, a full-fledged ICS ransomware that can compromise a real-world. To demonstrate the capability of ICS-BROCK, we use SIEMENS S7-300 PLC, one of the most widely used devices in ICSs, to build a real water treatment environment. The results empirically demonstrate the feasibility of launching ICS ransomware attacks in a practical setting. In the end, we give some suggestions on ICS ransomware to aid in future study and defenses.
Liqun Yang, Zhoujun Li 0001, Qiang Zeng 0001, Yueying He, Xiaoming Zhang 0001
TrustCom7
2020 Detecting bi-level false data injection attack based on time series analysis method in smart grid
Liqun Yang, Xiaoming Zhang 0001, Zhi Li 0045, Zhoujun Li 0001, Yueying He
Comput. Secur.2
2020 Visual Question Answering via Combining Inferential Attention and Semantic Space Mapping
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhonghua Zhao, Zhoujun Li 0001
Knowl. Based Syst.2
2019 FGST: Fine-Grained Spatial-Temporal Based Regression for Stationless Bike Traffic Prediction
Hao Chen 0062, Senzhang Wang, Zengde Deng, Xiaoming Zhang 0001, Zhoujun Li 0001
PAKDD (1)4
2019 Network embedding by fusing multimodal contents and links
Feiran Huang, Xiaoming Zhang 0001, Jie Xu 0015, Chaozhuo Li, Zhoujun Li 0001
Knowl. Based Syst.2
2019 Image-text sentiment analysis via deep multimodal attentive fusion
Feiran Huang, Xiaoming Zhang 0001, Zhonghua Zhao, Jie Xu 0015, Zhoujun Li 0001
Knowl. Based Syst.2
2019 Visual-textual sentiment classification with bi-directional multi-level attention networks
Jie Xu 0015, Feiran Huang, Xiaoming Zhang 0001, Senzhang Wang, Chaozhuo Li, Zhoujun Li 0001, Yueying He
Knowl. Based Syst.3
2019 Bi-Directional Spatial-Semantic Attention Networks for Image-Text Matching
abstract
Image-text matching by deep models has recently made remarkable achievements in many tasks, such as image caption and image search. A major challenge of matching the image and text lies in that they usually have complicated underlying relations between them and simply modeling the relations may lead to suboptimal performance. In this paper, we develop a novel approach Bi-directional Spatial-Semantic Attention Networks (BSSAN), which leverages both the word to regions (W2R) relation and image object to words (O2W) relation in a holistic deep framework for more effectively matching. Specifically, to effectively encode the W2R relation, we adopt LSTM with bilinear attention function to infer the image regions which are more related to the particular words, which is referred as the W2R attention network. On the other side, the O2W attention network is proposed to discover the semanticallyclose words for each visual object in the image, i.e., the visual object to words (O2W) relation. Then a deep model unifying both of the two directional attention networks into a holistic learning framework is proposed to learn the matching scores of image and text pairs. Compared to existing image-text matching methods, our approach achieves state-of-the-art performance on the datasets of Flickr30K and MSCOCO.
Feiran Huang, Xiaoming Zhang 0001, Zhonghua Zhao, Zhoujun Li 0001
IEEE Trans. Image Process.2
2019 Efficient Traffic Estimation With Multi-Sourced Data by Parallel Coupled Hidden Markov Model
abstract
Traffic congestion estimation in arterial networks with sparse GPS probe data is a practically important while substantially challenging research issue. The effectiveness and efficiency of the existing GPS probe data-based traffic estimation models are largely limited due to the following two challenges. First, due to the low sampling frequency of GPS probes, probe data are usually sparse, especially for some road links not located in the central urban areas. Second, due to the very complex temporal and spatial dependencies among the road links, the variable space of the existing traffic estimation models is huge. It is time consuming to get an accurate estimation of a large arterial road network with thousands of road links. To address the above-mentioned issues, this paper proposes to extract traffic event signals from social media and incorporate them with GPS probe data to alleviate the data sparse issue. We first collect traffic-related posts that report various traffic events, including traffic jam, accident, and road construction from Twitter. By considering the GPS probe readings and the traffic event tweets as two types of observations, we next extend the conventional coupled hidden Markov model for integrating the two types of data to obtain a more accurate estimation of traffic conditions. To address the computational challenge, a parallel importance sampling-based electromagnetic algorithm is further introduced. We evaluate our model on the arterial network of downtown Chicago. The experimental results demonstrate the superior performance of the model in both effectiveness and efficiency.
Senzhang Wang, Xiaoming Zhang 0001, Fengxiang Li, Philip S. Yu
IEEE Trans. Intell. Transp. Syst.2
2018 Distribution Distance Minimization for Unsupervised User Identity Linkage
abstract
Nowadays, it is common for one natural person to join multiple social networks to enjoy different services. Linking identical users across different social networks, also known as the User Identity Linkage (UIL), is an important problem of great research challenges and practical value. Most existing UIL models are supervised or semi-supervised and a considerable number of manually matched user identity pairs are required, which is costly in terms of labor and time. In addition, existing methods generally rely heavily on some discriminative common user attributes, and thus are hard to be generalized. Motivated by the isomorphism across social networks, in this paper we consider all the users in a social network as a whole and perform UIL from the user space distribution level. The insight is that we convert the unsupervised UIL problem to the learning of a projection function to minimize the distance between the distributions of user identities in two social networks. We propose to use the earth mover's distance (EMD) as the measure of distribution closeness, and propose two models UUIL$_gan $ and UUIL$_omt $ to efficiently learn the distribution projection function. Empirically, we evaluate the proposed models over multiple social network datasets, and the results demonstrate that our proposal significantly outperforms state-of-the-art methods.
Chaozhuo Li, Senzhang Wang, Philip S. Yu, Lei Zheng 0001, Xiaoming Zhang 0001, Zhoujun Li 0001, Yanbo Liang
CIKM5
2018 Adversarial Learning of Answer-Related Representation for Visual Question Answering
abstract
Visual Question Answering (VQA) aims to learn a joint embedding of the question sentence and the corresponding image to infer the answer. Existing approaches learn the joint embedding don't consider the answer-related information, which results in that the learned representation is not effective to reflect the answer of the question. To address this problem, this paper proposes a novel method, i.e., Adversarial Learning of Answer-Related Representation (ALARR) for visual question answering, which seeks an effective answer-related representation for the question-image pair based on adversarial learning between two processes. The embedding learning process aims to generate modality-invariant joint representations for the question-image and question-answer pairs, respectively. Meanwhile, it tries to confuse the other process, embedding discriminator, which tries to discriminate the two representations from different modalities of pairs. Specifically, the joint embedding of the question-image pair is learned by a three-level attention model, and the joint representation of the question-answer pair is learned by a semantic integration model. Through the adversarial leaning, the answer-related representation are better preserved. Then an answer predictor is proposed to infer the answer from the answer-related representation. Experiments conducted on two widely used VQA benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhoujun Li 0001
CIKM2
2018 Keyphrase Generation with Correlation Constraints
abstract
In this paper, we study automatic keyphrase generation.Although conventional approaches to this task show promising results, they neglect correlation among keyphrases, resulting in duplication and coverage issues.To solve these problems, we propose a new sequence-to-sequence architecture for keyphrase generation named CorrRNN, which captures correlation among multiple keyphrases in two ways.First, we employ a coverage vector to indicate whether the word in the source document has been summarized by previous phrases to improve the coverage for keyphrases.Second, preceding phrases are taken into account to eliminate duplicate phrases and improve result coherence.Experiment results show that our model significantly outperforms the state-of-the-art method on benchmark datasets in terms of both accuracy and diversity.
Xiaoming Zhang 0001, Yu Wu 0012, Zhoujun Li 0001
EMNLP2
2018 Multimodal Network Embedding via Attention based Multi-view Variational Autoencoder
abstract
Learning the embedding for social media data has attracted extensive research interests as well as boomed a lot of applications, such as classification and link prediction. In this paper, we examine the scenario of a multimodal network with nodes containing multimodal contents and connected by heterogeneous relationships, such as social images containing multimodal contents (e.g., visual content and text description), and linked with various forms (e.g., in the same album or with the same tag). However, given the multimodal network, simply learning the embedding from the network structure or a subset of content results in sub-optimal representation. In this paper, we propose a novel deep embedding method, i.e., Attention-based Multi-view Variational Auto-Encoder (AMVAE), to incorporate both the link information and the multimodal contents for more effective and efficient embedding. Specifically, we adopt LSTM with attention model to learn the correlation between different data modalities, such as the correlation between visual regions and the specific words, to obtain the semantic embedding of the multimodal contents. Then, the link information and the semantic embedding are considered as two correlated views. A multi-view correlation learning based Variational Auto-Encoder (VAE) is proposed to learn the representation of each node, in which the embedding of link information and multimodal contents are integrated and mutually reinforced. Experiments on three real-world datasets demonstrate the superiority of the proposed model in two applications, i.e., multi-label classification and link prediction.
Feiran Huang, Xiaoming Zhang 0001, Chaozhuo Li, Zhoujun Li 0001, Yueying He, Zhonghua Zhao
ICMR2
2018 Learning Joint Multimodal Representation with Adversarial Attention Networks
abstract
Recently, learning a joint representation for the multimodal data (e.g., containing both visual content and text description) has attracted extensive research interests. Usually, the features of different modalities are correlational and compositive, and thus a joint representation capturing the correlation is more effective than a subset of the features. Most of existing multimodal representation learning methods suffer from lack of additional constraints to enhance the robustness of the learned representations. In this paper, a novel Adversarial Attention Networks (AAN) is proposed to incorporate both the attention mechanism and the adversarial networks for effective and robust multimodal representation learning. Specifically, a visual-semantic attention model with siamese learning strategy is proposed to encode the fine-grained correlation between visual and textual modalities. Meanwhile, the adversarial learning model is employed to regularize the generated representation by matching the posterior distribution of the representation to the given priors. Then, the two modules are incorporated into a integrated learning framework to learn the joint multimodal representation. Experimental results in two tasks, i.e., multi-label classification and tag recommendation, show that the proposed model outperforms state-of-the-art representation learning methods.
Feiran Huang, Xiaoming Zhang 0001, Zhoujun Li 0001
ACM Multimedia2
2018 From content to links: Social image embedding with deep multimodal model
Feiran Huang, Xiaoming Zhang 0001, Zhoujun Li 0001, Zhonghua Zhao, Yueying He
Knowl. Based Syst.2
2018 Incorporating temporal dynamics into LDA for one-class collaborative filtering
Haijun Zhang 0007, Xiaoming Zhang 0001, Jianye Yu, Feng Li 0058
Knowl. Based Syst.2
2018 Unsupervised geographically discriminative feature learning for landmark tagging
Xiaoming Zhang 0001, Zhonghua Zhao, Haijun Zhang 0007, Senzhang Wang, Zhoujun Li 0001
Knowl. Based Syst.1
2018 Classify social image by integrating multi-modal content
Xiaoming Zhang 0001, Xiong Li 0002, Zhoujun Li 0001, Senzhang Wang
Multim. Tools Appl.1
2018 Landmark Image Retrieval by Jointing Feature Refinement and Multimodal Classifier Learning
abstract
Landmark retrieval is to return a set of images with their landmarks similar to those of the query images. Existing studies on landmark retrieval focus on exploiting the geometries of landmarks for visual similarity matches. However, the visual content of social images is of large diversity in many landmarks, and also some images share common patterns over different landmarks. On the other side, it has been observed that social images usually contain multimodal contents, i.e., visual content and text tags, and each landmark has the unique characteristic of both visual content and text content. Therefore, the approaches based on similarity matching may not be effective in this environment. In this paper, we investigate whether the geographical correlation among the visual content and the text content could be exploited for landmark retrieval. In particular, we propose an effective multimodal landmark classification paradigm to leverage the multimodal contents of social image for landmark retrieval, which integrates feature refinement and landmark classifier with multimodal contents by a joint model. The geo-tagged images are automatically labeled for classifier learning. Visual features are refined based on low rank matrix recovery, and multimodal classification combined with group sparse is learned from the automatically labeled images. Finally, candidate images are ranked by combining classification result and semantic consistence measuring between the visual content and text content. Experiments on real-world datasets demonstrate the superiority of the proposed approach as compared to existing methods.
Xiaoming Zhang 0001, Senzhang Wang, Zhoujun Li 0001, Shuai Ma 0001
IEEE Trans. Cybern.1
2017 From Properties to Links: Deep Network Embedding on Incomplete Graphs
abstract
As an effective way of learning node representations in networks, network embedding has attracted increasing research interests recently. Most existing approaches use shallow models and only work on static networks by extracting local or global topology information of each node as the algorithm input. It is challenging for such approaches to learn a desirable node representation on incomplete graphs with a large number of missing links or on dynamic graphs with new nodes joining in. It is even challenging for them to deeply fuse other types of data such as node properties into the learning process to help better represent the nodes with insufficient links. In this paper, we for the first time study the problem of network embedding on incomplete networks. We propose a Multi-View Correlation-learning based Deep Network Embedding method named MVC-DNE to incorporate both the network structure and the node properties for more effectively and efficiently perform network embedding on incomplete networks. Specifically, we consider the topology structure of the network and the node properties as two correlated views. The insight is that the learned representation vector of a node should reflect its characteristics in both views. Under a multi-view correlation learning based deep autoencoder framework, the structure view and property view embeddings are integrated and mutually reinforced through both self-view and cross-view learning. As MVC-DNE can learn a representation mapping function, it can directly generate the representation vectors for the new nodes without retraining the model. Thus it is especially more efficient than previous methods. Empirically, we evaluate MVC-DNE over three real network datasets on two data mining applications, and the results demonstrate that MVC-DNE significantly outperforms state-of-the-art methods.
Dejian Yang, Senzhang Wang, Chaozhuo Li, Xiaoming Zhang 0001, Zhoujun Li 0001
CIKM4
2017 Semi-Supervised Network Embedding
Chaozhuo Li, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xiaoming Zhang 0001, Jianshe Zhou
DASFAA (1)5
2017 PPNE: Property Preserving Network Embedding
Chaozhuo Li, Senzhang Wang, Dejian Yang, Zhoujun Li 0001, Yang Yang 0002, Xiaoming Zhang 0001, Jianshe Zhou
DASFAA (1)6
2017 Computing Urban Traffic Congestions by Incorporating Sparse GPS Probe Data and Social Media Data
abstract
Estimating urban traffic conditions of an arterial network with GPS probe data is a practically important while substantially challenging problem, and has attracted increasing research interests recently. Although GPS probe data is becoming a ubiquitous data source for various traffic related applications currently, they are usually insufficient for fully estimating traffic conditions of a large arterial network due to the low sampling frequency. To explore other data sources for more effectively computing urban traffic conditions, we propose to collect various traffic events such as traffic accident and jam from social media as complementary information. In addition, to further explore other factors that might affect traffic conditions, we also extract rich auxiliary information including social events, road features, Point of Interest (POI), and weather. With the enriched traffic data and auxiliary information collected from different sources, we first study the traffic co-congestion pattern mining problem with the aim of discovering which road segments geographically close to each other are likely to co-occur traffic congestion. A search tree based approach is proposed to efficiently discover the co-congestion patterns. These patterns are then used to help estimate traffic congestions and detect anomalies in a transportation network. To fuse the multisourced data, we finally propose a coupled matrix and tensor factorization model named TCE_R to more accurately complete the sparse traffic congestion matrix by collaboratively factorizing it with other matrices and tensors formed by other data. We evaluate the proposed model on the arterial network of downtown Chicago with 1,257 road segments whose total length is nearly 700 miles. The results demonstrate the superior performance of TCE_R by comprehensive comparison with existing approaches.
Senzhang Wang, Xiaoming Zhang 0001, Jianping Cao, Lifang He 0001, Leon Stenneth, Philip S. Yu, Zhoujun Li 0001
ACM Trans. Inf. Syst.2
2016 Integrating multiple types of features for event identification in social images
Xiaoming Zhang 0001, Zhoujun Li 0001, Xueqiang Lv, Xiaoming Chen 0007
Multim. Tools Appl.1
2016 Geographical Topics Learning of Geo-Tagged Social Images
abstract
With the availability of cheap location sensors, geotagging of images in online social media is very popular. With a large amount of geo-tagged social images, it is interesting to study how these images are shared across geographical regions and how the geographical language characteristics and vision patterns are distributed across different regions. Unlike textual document, geo-tagged social image contains multiple types of content, i.e., textual description, visual content, and geographical information. Existing approaches usually mine geographical characteristics using a subset of multiple types of image contents or combining those contents linearly, which ignore correlations between different types of contents, and their geographical distributions. Therefore, in this paper, we propose a novel method to discover geographical characteristics of geo-tagged social images using a geographical topic model called geographical topic model of social images (GTMSIs). GTMSI integrates multiple types of social image contents as well as the geographical distributions, in which image topics are modeled based on both vocabulary and visual features. In GTMSI, each region of the image would have its own topic distribution, and hence have its own language model and vision pattern. Experimental results show that our GTMSI could identify interesting topics and vision patterns, as well as provide location prediction and image tagging.
Xiaoming Zhang 0001, Shufan Ji, Senzhang Wang, Zhoujun Li 0001, Xueqiang Lv
IEEE Trans. Cybern.1
2016 Coranking the Future Influence of Multiobjects in Bibliographic Network Through Mutual Reinforcement
abstract
Scientific literature ranking is essential to help researchers find valuable publications from a large literature collection. Recently, with the prevalence of webpage ranking algorithms such as PageRank and HITS, graph-based algorithms have been widely used to iteratively rank papers and researchers through the networks formed by citation and coauthor relationships. However, existing graph-based ranking algorithms mostly focus on ranking the current importance of literature. For researchers who enter an emerging research area, they might be more interested in new papers and young researchers that are likely to become influential in the future, since such papers and researchers are more helpful in letting them quickly catch up on the most recent advances and find valuable research directions. Meanwhile, although some works have been proposed to rank the prestige of a certain type of objects with the help of multiple networks formed of multiobjects, there still lacks a unified framework to rank multiple types of objects in the bibliographic network simultaneously. In this article, we propose a unified ranking framework MRCoRank to corank the future popularity of four types of objects: papers, authors, terms, and venues through mutual reinforcement. Specifically, because the citation data of new publications are sparse and not efficient to characterize their innovativeness, we make the first attempt to extract the text features to help characterize innovative papers and authors. With the observation that the current trend is more indicative of the future trend of citation and coauthor relationships, we then construct time-aware weighted graphs to quantify the importance of links established at different times on both citation and coauthor graphs. By leveraging both the constructed text features and time-aware graphs, we finally fuse the rich information in a mutual reinforcement ranking framework to rank the future importance of multiobjects simultaneously. We evaluate the proposed model through extensive experiments on the ArnetMiner dataset containing more than 1,500,000 papers. Experimental results verify the effectiveness of MRCoRank in coranking the future influence of multiobjects in a bibliographic network.
Senzhang Wang, Sihong Xie, Xiaoming Zhang 0001, Zhoujun Li 0001, Philip S. Yu, Yueying He
ACM Trans. Intell. Syst. Technol.3
2016 Learning Geographical Hierarchy Features via a Compositional Model
abstract
Image location prediction is used to estimate the geolocation where an image is taken, which is important for many image applications, such as image retrieval, image browsing, and organization. Since a social image contains heterogeneous contents, such as visual content and textual content, effectively incorporating these contents to predict location is nontrivial. Moreover, it is observed that image content patterns and the locations where they may appear correlate hierarchically. Traditional image location prediction methods mainly adopt a single-level architecture and assume images are independently distributed in geographical space, which is not directly adaptable to the hierarchical correlation. In this paper, we propose a geographically hierarchical bi-modal deep belief network (GH-BDBN) model, which is a compositional learning architecture that integrates multi-modal deep learning model with a non-parametric hierarchical prior model. GH-BDBN learns a joint representation capturing the correlations among different types of image content using a bi-modal DBN, with a geographically hierarchical prior over the joint representation to model the hierarchical correlation between image content and location. Then, an efficient inference algorithm is proposed to learn the parameters and the geographical hierarchical structure of geographical locations. Experimental results demonstrate the superiority of our model for image location prediction.
Xiaoming Zhang 0001, Xia Ben Hu, Senzhang Wang, Yang Yang 0002, Zhoujun Li 0001, Jianshe Zhou
IEEE Trans. Multim.1
2015 Inferring Diffusion Networks with Sparse Cascades by Structure Transfer
Senzhang Wang, Honghui Zhang, Jiawei Zhang 0001, Xiaoming Zhang 0001, Philip S. Yu, Zhoujun Li 0001
DASFAA (1)4
2015 Learning Geographical Hierarchy Features for Social Image Location Prediction
Xiaoming Zhang 0001, Xia Ben Hu, Zhoujun Li 0001
IJCAI1
2015 Location Prediction of Social Images via Generative Model
abstract
The vast amount of geo-tagged social images has attracted great attention in research of predicting location using the plentiful content of images, such as visual content and textual description. Most of the existing researches use the text-based or vision-based method to predict location. There still exists a problem: how to effectively exploit the correlation between different types of content as well as their geographical distributions for location prediction. In this paper, we propose to predict image location by learning the latent relation between geographical location and multiple types of image content. In particularly, we propose a geographical topic model GTMSI (geographical topic model of social image) to integrate multiple types of image content as well as the geographical distributions. In GTMI, image topic is modeled on both text vocabulary and visual feature. Each region has its own distribution over topics and hence has its own language model and vision pattern. The location of a new image is estimated based on the joint probability of image content and similarity measure on topic distribution between images. Experiment results demonstrate the performance of location prediction based on GTMSI.
Xiaoming Zhang 0001, Zhoujun Li 0001, Senzhang Wang, Yang Yang 0002, Xueqiang Lv
ICMR1
2015 Friendship Link Recommendation Based on Content Structure Information
Xiaoming Zhang 0001, Zhoujun Li 0001
WAIM1
2015 Exploring Social Network Information for Solving Cold Start in Product Recommendation
Chaozhuo Li, Fang Wang 0019, Yang Yang 0002, Zhoujun Li 0001, Xiaoming Zhang 0001
WISE (2)5
2015 Search engine reinforced semi-supervised classification and graph-based summarization of microblogs
Yan Chen 0019, Xiaoming Zhang 0001, Zhoujun Li 0001, Jun-Ping Ng
Neurocomputing2
2015 Event detection and popularity prediction in microblogging
Xiaoming Zhang 0001, Xiaoming Chen 0007, Yan Chen 0019, Senzhang Wang, Zhoujun Li 0001, Jiali Xia
Neurocomputing1
2015 Exploiting social circle broadness for influential spreaders identification in social networks
Senzhang Wang, Fang Wang 0019, Yan Chen 0019, Zhoujun Li 0001, Xiaoming Zhang 0001
World Wide Web6
2014 Exploit Latent Dirichlet Allocation for One-Class Collaborative Filtering
abstract
Previous work studied one-class collaborative filtering (OCCF) problems including pointwise methods, pairwise methods, and content-based methods. The fundamental assumptions made on these approaches are roughly the same. They regard all missing values as negative. However, this is unreasonable since the missing values actually are the mixture of negative and positive examples. A user does not give a positive feedback on an item probably only because she/he is unaware of the item, but in fact, she/he is fond of it. Furthermore, content-based methods, e.g. collaborative topic regression (CTR), usually require textual content information of items. This cannot be satisfied in some cases. In this paper, we exploit latent Dirichlet allocation (LDA) model on OCCF problem. It assumes missing values unknown and only models the observed data, and it also does not need content information of items. In our model items are regarded as words and users are considered as documents and the user-item feedback matrix denotes the corpus. Experimental results show that our proposed method outperforms the previous methods on various ranking-oriented evaluation metrics.
Haijun Zhang 0007, Zhoujun Li 0001, Yan Chen 0019, Xiaoming Zhang 0001, Senzhang Wang
CIKM4
2014 Future Influence Ranking of Scientific Literature
abstract
Researchers or students entering a emerging research area are particularly interested in what newly published papers will be most cited and which young researchers will become influential in the future, so that they can catch the most recent advances and find valuable research directions. However, predicting the future importance of scientific articles and authors is extremely hard due to the dynamic nature of literature networks and evolving research topics. Different from most previous studies aiming to rank the current importance of literature and authors, we focus on ranking the future popularity of new publications and young researchers by proposing a unified ranking model to combine various available information. Specifically, we first propose to use two kinds of text features, words and words co-occurrence to characterize innovative papers and authors. Then, instead of using static and un-weighted graphs, we construct time-aware weighted graphs to distinguish the various importance of links established at different time. Finally, by leveraging both the constructed text features and graphs, we propose a mutual reinforcement ranking framework called MRFRank to rank the future importance of papers and authors simultaneously. Experimental results on the ArnetMiner dataset show that the proposed approach significantly outperforms the baselines on the metric recommendation intensity.
Senzhang Wang, Sihong Xie, Xiaoming Zhang 0001, Zhoujun Li 0001, Philip S. Yu, Xinyu Shu
SDM3
2014 Popularity Prediction of Burst Event in Microblogging
Xiaoming Zhang 0001, Zhoujun Li 0001, Wen-Han Chao, Jiali Xia
WAIM1
2014 Training data reduction to speed up SVM training
Senzhang Wang, Zhoujun Li 0001, Xiaoming Zhang 0001, Haijun Zhang 0007
Appl. Intell.4
2013 From Interest to Function: Location Estimation in Social Media
abstract
Recent years have witnessed the tremendous development of social media, which attracts a vast number of Internet users. The high-dimension content generated by these users provides an unique opportunity to understand their behavior deeply. As one of the most fundamental topics, location estimation attracts more and more research efforts. Different from the previous literature, we find that user's location is strongly related to user interest. Based on this, we first build a detection model to mine user interest from short text. We then establish the mapping between location function and user interest before presenting an efficient framework to predict the user's location with convincing fidelity. Thorough evaluations and comparisons on an authentic data set show that our proposed model significantly outperforms the state-of-the-arts approaches. Moreover, the high efficiency of our model also guarantees its applicability in real-world scenarios.
Yan Chen 0019, Jichang Zhao, Xia Ben Hu, Xiaoming Zhang 0001, Zhoujun Li 0001, Tat-Seng Chua
AAAI4
2013 Collaborative Filtering Based on Rating Psychology
Haijun Zhang 0007, Zhoujun Li 0001, Xiaoming Zhang 0001
WAIM4
2013 Collaborative Filtering Using Multidimensional Psychometrics Model
Haijun Zhang 0007, Xiaoming Zhang 0001, Zhoujun Li 0001
WAIM2
2013 Improving image tags by exploiting web search results
Xiaoming Zhang 0001, Zhoujun Li 0001, Wen-Han Chao
Multim. Tools Appl.1
2013 Social image tagging using graph-based reinforcement on multi-type interrelated objects
Xiaoming Zhang 0001, Xiaojian Zhao, Zhoujun Li 0001, Jiali Xia, Ramesh Jain 0001, Wen-Han Chao
Signal Process.1
2012 A Semi-Supervised Bayesian Network Model for Microblog Topic Classification
Yan Chen 0019, Zhoujun Li 0001, Liqiang Nie, Xia Ben Hu, Tat-Seng Chua, Xiaoming Zhang 0001
COLING7
2012 Bootstrap Sampling Based Data Cleaning and Maximum Entropy SVMs for Large Datasets
abstract
Support Vector Machines (SVMs) is a popular machine learning algorithm based on Statistical Learning Theory (SLT). However, traditional solutions suffer from O(n2) time complexity. In this paper, a novel two-stage informative pattern abstraction algorithm is proposed. The first stage of the algorithm is data cleaning based on bootstrap sampling. A bundle of weak SVM classifiers are trained based on the sampled small datasets. Training data correctly classified by all the weak classifiers are cleaned. In the second stage, to further improve performance of final classifier and reduce training time, two novel informative pattern extraction algorithms based on entropy maximization SVMs are proposed. Empirical studies show our approach is effective in reducing size of training datasets and the computational cost, outperforming the state-of-the-art SVM training algorithms PEGASOS, RSVM and LIBLINEAR SVM with comparable classification accuracy.
Senzhang Wang, Zhoujun Li 0001, Xiaoming Zhang 0001
ICTAI3
2012 Tagging image by merging multiple features in a integrated manner
Xiaoming Zhang 0001, Zhoujun Li 0001, Wen-Han Chao
J. Intell. Inf. Syst.1
2012 Automatic tagging by exploring tag information capability and correlation
Xiaoming Zhang 0001, Zi Huang, Heng Tao Shen, Yang Yang 0002, Zhoujun Li 0001
World Wide Web1
2011 Tagging Image with Informative and Correlative Tags
Xiaoming Zhang 0001, Heng Tao Shen, Zi Huang, Zhoujun Li 0001
APWeb1
2011 Probabilistic Image Tagging with Tags Expanded By Text-Based Search
Xiaoming Zhang 0001, Zi Huang, Heng Tao Shen, Zhoujun Li 0001
DASFAA (1)1
2011 Tagging Image by Exploring Weighted Correlation between Visual Features and Tags
Xiaoming Zhang 0001, Zhoujun Li 0001
WAIM1
2009 Online New Event Detection Based on IPLSA
Xiaoming Zhang 0001, Zhoujun Li 0001
ADMA1
2007 Research of Routing Algorithm in Hierarchy-Adaptive P2P Systems
Xiaoming Zhang 0001, Yijie Wang 0001, Zhoujun Li 0001
ISPA1