EDBT 2026 Demo / reviewers in the wild / expert
Zhixin Li 0001
dblp:82/7406-1
· DBLP profile ↗
51ranked-venue papers in the field
3as first author
40since 2021 · last 2027
0000-0002-5313-6134ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 20 (1 first)Data Mining & Knowledge Discovery · 12 (2 first)Database Systems & Data Management · 10Knowledge Engineering, Semantic Web & Information Systems · 8Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Generative shortcut probing with multi-level margins for robust VQA
Lieyuan Ni, Zhixin Li 0001 |
Inf. Process. Manag. | 2 |
| 2026 | Image-Text Matching via Graph Counterfactual Learning and Trustworthy Alignment
Bo Li 0144, Jingyi Yan, Qinghui Hu, Zhixin Li 0001 |
KSEM (3) | 5 |
| 2026 | Question-guided attention and cross-modal alignment for knowledge-based visual question answering
Wei Li 0233, Fuyun Deng, Zhixin Li 0001 |
Inf. Process. Manag. | 3 |
| 2026 | Adaptive confidence-driven learning and cross-modal hard sample mining for unsupervised visible-infrared person re-identification
Canlong Zhang, Haifei Ma, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
Inf. Process. Manag. | 4 |
| 2025 | QGCMA: A Framework for Knowledge-Based Visual Question Answering
Wei Li 0233, Zhixin Li 0001 |
CIKM | 2 |
| 2025 | Multimodal Sentiment Analysis via Progressive Fusion of Audio-Visual Affective DescriptionsabstractMultimodal Sentiment Analysis (MSA) holds significant research value in the fields of intelligent human-computer interaction and affective computing. Although existing MSA approaches have made considerable progress, challenges remain in identifying subtle emotional distinctions within audio and visual expressions. In particular, conventional fusion methods have not effectively addressed the difficulty of integrating heterogeneous modality information. To tackle these challenges, we propose a progressive fusion framework based on audio-visual affective descriptions for MSA. Specifically, we design an audio-visual emotional description generator that transforms raw audiovisual data into textual emotional descriptions, thereby effectively highlighting affective features. Subsequently, this emotional description is integrated with the original multimodal features to obtain a richer feature representation. Building upon this, we introduce a three-stage progressive fusion architecture. First, we employ cross-modal transformers to facilitate interactions among modalities and to learn inter-modal dependencies. Second, a gated fusion mechanism is incorporated to effectively eliminate redundant information and further promote interaction and compatibility among features. Finally, an attention mechanism is utilized to dynamically adjust the weights of features from different modalities, enabling effective multimodal sentiment information fusion. Experimental results on widely used sentiment analysis benchmark datasets, including MOSI, MOSEI, and CH-SIMS, underscore significant enhancements compared to state-of-the-art models. Lisong Ou, Zhixin Li 0001 |
CIKM | 2 |
| 2025 | Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
Zengli Luo, Canlong Zhang, Zhixin Li 0001, Zhiwen Wang 0001, Chunrong Wei |
KSEM (2) | 3 |
| 2025 | DMR-XNet: Dynamic Multi-Relation Cross-Fusion Network for Aspect-Based Multimodal Sentiment AnalysisabstractMulti-modal aspect-based sentiment analysis (MABSA) has been attracting growing interest, encompassing two main tasks: multimodal aspect term extraction (MATE) and multimodal aspect-level sentiment classification (MASC). Existing methods primarily focus on aligning targets with aspect terms. Still, this strict alignment approach has three major limitations: (i) it overlooks the semantic information of the target's context, which can be crucial for sentiment analysis. (ii) Existing models struggle to capture the higher-order relationships in multimodal data, which may involve interactions beyond simple pairwise relationships. (iii) existing models fail to exploit the rich syntactic information embedded in the text fully. To address these challenges, we propose a Dynamic Multi-Relation Cross-Fusion Network (DMR-XNet). By introducing a dynamic hypergraph-based multi-relation capturing method, we can learn many-to-many multimodal associations, enabling a more comprehensive representation of complex relationships. Additionally, we propose a Dual-stream Cross-excitation Layer, which separately emphasizes syntactic dependencies and aspect relevance, while cross-activating the ignored complementary features from the other stream. Extensive experimental results prove that DMR-XNet significantly outperforms existing state-of-the-art methods in terms of performance. Fengling Zhou, Zhixin Li 0001 |
ICMR | 2 |
| 2024 | Beyond Language Bias: Overcoming Multimodal Shortcut and Distribution Biases for Robust Visual Question AnsweringabstractRecent studies have found that many VQA models are influenced by biases, preventing them from effectively using multimodal information for reasoning. Consequently, these methods, which perform well on standard VQA datasets, exhibit underwhelming performance on the bias-sensitive VQA-CP dataset. Although numerous studies in the past have focused on mitigating biases in VQA models, most have only considered language bias. In this paper, we address the issue of bias in VQA task by targeting the various sources of bias. Specifically, to counteract shortcut biases, we integrate a bias detector capable of capturing both vision and language biases, and we reinforce its ability to capture biases using a generative adversarial network and knowledge distillation. To combat distribution bias, we use a cosine classifier to obtain a cosine feature branch from the base model, training it with an adaptive angular margin loss based on answer frequency and difficulty, along with a supervised contrastive loss to enhance the model's classification ability in the feature space. In the prediction stage, we fuse the cosine features with the prediction of the base model to obtain the final prediction of our model. Finally, extensive experiments demonstrate that our approach SD-VQA achieves state-of-the-art performance on the VQA-CPv2 dataset without using any data balancing, and achieves competitive results on the VQAv2 dataset. Jingliang Gu, Zhixin Li 0001 |
CIKM | 2 |
| 2024 | MV-BART: Multi-view BART for Multi-modal Sarcasm DetectionabstractUnderstanding emotions in dialogue is an essential part of human communication, its an extremely complex cognitive process involving cross-modal interactions and cross-emotional associations, and multimodal sarcasm detection is an emerging but challenging research task in this process aiming at video discourse incorporating appropriate contextual information and external knowledge and identifying sarcasm by understanding both verbal and non-verbal components. However, existing research primarily focuses on constructing multimodal fusion representations and capturing incongruity between modalities as indicative cues for recognizing sarcasm, which relies on a fixed network design architecture that is difficult to cope with complex and diverse satirical scenarios in real life. As humans, we rely on the combination of visual and auditory cues, such as facial expressions and intonations, to understand information. Our brains are implicitly trained to integrate information from multiple senses to form a comprehensive understanding of conveyed messages, a process known as multi-sensory integration. The combination of different modalities not only provides additional information but also amplifies the information conveyed by each modality relative to others. Therefore, dynamic variations in the weights of different modalities play a crucial role in multi-modal understanding. From this perspective, we propose a new framework called Multi-view BART(MV-BART), which is capable of exploiting multi-granularity cues from multiple viewpoints and dynamically adjusting the view weights, applied to different sarcastic scenarios. It is worth mentioning that we analyze the proposed framework by testing it on several benchmark datasets, and the results outperform the existing state-of-the-art. Xingjie Zhuang, Fengling Zhou, Zhixin Li 0001 |
CIKM | 3 |
| 2024 | Path-Aware Co-contrastive Learning for Signed Directed Network Embedding
Yuechen Tang, Huifang Ma, Ke Shu, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (6) | 4 |
| 2024 | Time-aware Session Modeling for Knowledge Tracing
Huifang Ma, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (4) | 4 |
| 2024 | Dual-Channel Dual-Scale Interactive Learning for the Prediction of Compound-Protein Interaction
Zheyu Wu, Huifang Ma, Bin Deng 0014, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (7) | 4 |
| 2024 | Multi-Modal Sarcasm Detection via Dual Synergetic Perception Graph Convolutional NetworksabstractMulti-modal Sarcasm Detection (MSD) combines multiple modalities to identify implicit sarcastic sentiment, but commonsense knowledge's role in emotion recognition is often overlooked. Visual emotions tied to sarcastic cues in text are usually dispersed across the image, complicating detection. we propose a Dual Synergetic Perception Graph Convolutional Networks (DSP-GCN) to address these issues. First, we create a cross-modal knowledge incongruity graph linking key visual sentiments and relevant text tokens. Then, we enhance feature focus using a transformer encoder with the Convolutional Block Attention Module. Finally, the Global Modality Synergistic Fusion (GMSF) block models global relationships in each modality for improved sarcasm detection. Note that, our framework outperforms state-of-the-art methods on benchmark datasets. Xingjie Zhuang, Zhixin Li 0001 |
ICDM | 2 |
| 2024 | Team HUGE: Image-Text Matching via Hierarchical and Unified Graph EnhancingabstractGraph structures can represent rich semantic relationships, but currently, image-text matching methods have not been well applied. How to efficiently achieve graph learning and prevent overfitting of complex graph-based models are the challenges for all graph-based methods. Besides, single perspective similarity representation learning may overlook some potential correlations. For these issues, we develop Hierarchical and Unified Graph Enhancing (HUGE) to effectively extract the modality variations and collect the corresponding semantics of cross-modal features. Specifically, hierarchical graph learning (HGL) promotes cross-modal learning via training different graph sub-modules from multiple perspectives, while unified graph enhancing (UGE) aims to integrate the plausible alignments from different graph sub-module feedback. Besides, we designed a two-stage similarity representation learning strategy that combines the advantages of cosine similarity and vector similarity. Extensive experiments evidence that our HUGE approach can outperform the SoTA image-text matching methods. Sufficient ablation experiments verify the effectiveness of each component of HUGE. Bo Li 0144, Zhixin Li 0001 |
ICMR | 3 |
| 2024 | Modeling Multi-Task Joint Training of Aggregate Networks for Multi-Modal Sarcasm DetectionabstractWith the continuous emergence of various types of social media, which people often use to express their emotions in daily life, the multi-modal sarcasm detection (MSD) task has attracted more and more attention. However, due to the unique nature of sarcasm itself, there are still two main challenges on the way to achieving robust MSD: 1) existing mainstream methods often fail to take into account the problem of multi-modal weak correlation, thus ignoring the important sarcasm information of the uni-modal itself; 2) inefficiency in modeling cross-modal interactions in unaligned multi-modal data. Therefore, this paper proposes a multi-task jointly trained aggregation network (MTAN), which mainly adopts networks adapted to different modalities according to different modality processing tasks. Specifically, we design a multi-task CLIP framework that includes an uni-modal text task, an uni-modal image task, and a cross-modal interaction task, which can utilize sentiment cues from multiple tasks for multi-modal sarcasm detection. In addition, we design a global-local cross-modal interaction learning method that utilizes discourse-level representations from each modality as the global multi-modal context to interact with local uni-modal features, which not only avoids the secondary scaling cost of previous local-local cross-modal interaction methods but also allows the global multi-modal context and local uni-modal features to be mutually reinforcing and progressively improved through multi-layer superposition. After extensive experimental results and in-depth analysis, our model achieves state-of-the-art performance in multi-modal sarcasm detection. Lisong Ou, Zhixin Li 0001 |
ICMR | 2 |
| 2024 | Robust Video Hashing with Non-negative Tensor Factorization for Copy DetectionabstractCopy detection is a key task of video copyright protection. This paper presents a robust video hashing with non-negative tensor factorization (NTF) for copy detection. In the presented video hashing scheme, secondary frames are computed from the preprocessed video by assigning weights to all frames within a video group based on color entropy. Next, the secondary frames are fed into the pre-trained MobileNetV2 and then NTF is exploited to compress the three-order tensor constructed by stacking the output feature maps for hash construction. Experiments conducted on publicly available video datasets indicate that the presented hashing scheme outperforms the evaluated hashing schemes in the performances of classification and copy detection. Mengzhu Yu, Zhenjun Tang, Huijiang Zhuang, Xiaoping Liang, Zhixin Li 0001, Xianquan Zhang |
ICMR | 5 |
| 2024 | Multi-Interest Network with Simple Diffusion for Multi-Behavior Sequential RecommendationabstractMulti-behavior sequential recommendation (MBSR) aims to learn dynamic user preference from historical heterogeneous user interactions for identifying the next item under target behavior (i.e., purchase). Although significant efforts have been devoted to modeling users over observed multi-behavior interaction sequences, user modeling with dynamic behavior-aware multiple interests and elimination of inherent noises within these interactions are still underexplored. This limits user representations' awareness of true preference evolution and further constrains recommendation performance. To address the aforementioned issues, we propose a Multi-Interest Network with Simple Diffusion (MISD) via a combination of multi-interest learning and diffusion generative process for MBSR. Concretely, the dynamic multi-interest network is proposed to generate time-evolving personalized interests from the encoded dual-granularity user sequential patterns, leading to more accurate user preference learning. Additionally, simple diffusion is proposed to model the complex latent preference generation procedures in an iterative denoising manner, thereby alleviating the effect of noisy interactions. Extensive experiments on three real-world datasets demonstrate that MISD consistently outperforms various state-of-the-art recommendation methods under multiple settings (e.g., clean and noisy training). Qingfeng Li 0001, Huifang Ma, Wangyu Jin, Yugang Ji, Zhixin Li 0001 |
SDM | 5 |
| 2024 | Pre-training Question Embeddings for Improving Knowledge Tracing with Self-supervised Bi-graph Co-contrastive LearningabstractLearning high-quality vector representations (aka. embeddings) of educational questions lies at the core of knowledge tracing (KT), which defines a task of estimating students’ knowledge states by predicting the probability that they correctly answer questions. Although existing KT efforts have leveraged question information to achieve remarkable improvements, most of them learn question embeddings by following the supervised learning paradigm. In this article, we propose a novel question embedding pre-training method for improving knowledge tracing with self-supervised Bi -graph Co -contrastive learning ( BiCo ). Technically, on the basis of self-supervised learning paradigm, we first select two similar but distinct views (i.e., representing objective and subjective semantic perspectives) as the semantic source of question embeddings. Then, we design a primary task (structure recovery) together with two auxiliary tasks (question difficulty recovery and contrastive learning) to further enhance the representativeness of questions. Finally, extensive experiments conducted on two real-world datasets show BiCo has a higher expressive power that enables KT methods to effectively predict students’ performances. Huifang Ma, Zhixin Li 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Multiresolution Local Spectral Attributed Community SearchabstractCommunity search has become especially important in graph analysis task, which aims to identify latent members of a particular community from a few given nodes. Most of the existing efforts in community search focus on exploring the community structure with a single scale in which the given nodes are located. Despite promising results, the following two insights are often neglected. First, node attributes provide rich and highly related auxiliary information apart from network interactions for characterizing the node properties. Attributes may indicate the community assignment of a node with very few links, which would be difficult to determine from the network structure alone. Second, the multiresolution community affords latent information to depict the hierarchical relation of the network and ensure that one of them is closest to the real one. It is essential for users to understand the underlying structure of the network and explore the community with strong structure and attribute cohesiveness at disparate scales. These aspects motivate us to develop a new community search framework called Multiresolution Local Spectral Attributed Community Search (MLSACS). Specifically, inspired by the local modularity, graph wavelets, and scaling functions, we propose a new Multiresolution Local modularity (MLQ) based on a reconstructed node attribute graph. Furthermore, to detect local communities with cohesive structures and attributes at different scales, a sparse indicator vector is developed based on MLQ by solving a linear programming problem. Extensive experimental results on both synthetic and real-world attributed graphs have demonstrated the detected communities are meaningful and the scale can be changed reasonably. Huifang Ma, Zhixin Li 0001, Liang Chang 0003 |
ACM Trans. Web | 3 |
| 2023 | Synergistic Disease Similarity Measurement via Unifying Hierarchical Relation Perception and Association CapturingabstractQuantifying similarities among human diseases is crucial to enhance our understanding of disease biology. Deep learning efforts have been devoted to quantifying disease similarity by integrating multi-view data sources from disparate biological data. However, disease data are often sparse, leading to suboptimal representation of disease given biological entity relationships and labeled disease data are not adequately modeled. In this paper, we propose an effective Synergistic disease Similarity measurement model called SynerSim. SynerSim possesses two key components: a hierarchical biological entity relation perception module to capture disease features from various biological entities, and a disease association capturing module based on signed random walk to model precious disease data. Additionally, SynerSim leverages dual granularity contrastive learning to enhance the representation of diverse biological entities, owing to the ability to enable the synergistic supervision of diseases represented by both homogeneous and heterogeneous information. Experimental results demonstrate that SynerSim achieves outstanding performance in the disease similarity measurement. Zihao Gao 0001, Huifang Ma, Yike Wang 0001, Zhixin Li 0001, Liang Chang 0003 |
CIKM | 4 |
| 2023 | Co-guided Random Walk for Polarized Communities SearchabstractPolarized Communities Search (PCS) aims to identify query-dependent communities where positive links predominantly connect nodes within each community, while negative links primarily connect nodes across different communities. Existing solutions primarily focus on modeling network topology, disregarding the crucial factor of node attributes. However, it is non-trivial to incorporate node attributes into PCS. In this paper, we propose a novel method called CO-guided RAndom walk in attributed signed networks (CORA) for PCS. Our approach involves constructing an attribute-based signed network to represent the auxiliary relations between nodes. We introduce a weight assignment mechanism to assess the reliability of edges in the signed network. Then, we design a co-guided random walk scheme that operates on two signed networks to model the connections between network topology and node attributes, thereby enhancing the search outcomes. Finally, we identify polarized communities using the Rayleigh quotient in the signed network. Extensive experiments conducted on three public datasets demonstrate the superior performance of CORA compared to state-of-the-art baselines for polarized communities search. Fanyi Yang, Huifang Ma, Cairui Yan, Zhixin Li 0001, Liang Chang 0003 |
CIKM | 4 |
| 2023 | Intra- and Inter-behavior Contrastive Learning for Multi-behavior Recommendation
Qingfeng Li 0001, Huifang Ma, Ruoyi Zhang, Wangyu Jin, Zhixin Li 0001 |
DASFAA (2) | 5 |
| 2023 | Multi-scale Community Detection in Subspace of Attribute
Cairui Yan, Huifang Ma, Yuechen Tang, Xiaohong Li 0012, Zhixin Li 0001 |
DASFAA (3) | 5 |
| 2023 | Local Spectral for Polarized Communities Search in Attributed Signed Network
Fanyi Yang, Huifang Ma, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (3) | 4 |
| 2023 | Dual-View Self-supervised Co-training for Knowledge Graph Recommendation
Ruoyi Zhang, Huifang Ma, Qingfeng Li 0001, Yike Wang 0001, Zhixin Li 0001 |
DASFAA (2) | 5 |
| 2023 | Balancing and Contrasting Biased Samples for Debiased Visual Question AnsweringabstractThe goal of Visual Question Answering (VQA) is to test the reasoning ability of an intelligent agent by evaluating visual and textual information. However, recent studies suggest that many VQA models may only capture the correlation between questions and answers in the dataset rather than demonstrating true reasoning ability. To address this issue, we propose a new training approach called Balancing and Contrasting Biased Samples for Debiased VQA (BC-VQA) to build a robust VQA model. In our approach, we first generate two types of negative samples to balance the biased data and use self-supervised auxiliary tasks to help the base VQA model overcome language priors. Our method does not require any additional annotations. We then filter out biased training samples, construct positive samples by eliminating spurious correlations in biased samples, and perform auxiliary training through contrastive learning. Our approach is straightforward to implement and compatible with various VQA backbones. The experimental results demonstrate that BC-VQA achieves higher accuracy on VQA-CP v2 compared to the current state-of-the-art approaches. Runlin Cao, Zhixin Li 0001 |
ICDM | 2 |
| 2023 | Multi-label Image Classification with Multi-scale Global-Local Semantic Graph Network
Wenlan Kuang, Qiangxi Zhu, Zhixin Li 0001 |
ECML/PKDD (3) | 3 |
| 2023 | Graph Rebasing and Joint Similarity Reconstruction for Cross-Modal Hash Retrieval
Zhixin Li 0001 |
ECML/PKDD (2) | 2 |
| 2023 | Unifying knowledge iterative dissemination and relational reconstruction network for image-text matching
Xiumin Xie, Zhixin Li 0001, Zhenjun Tang, Huifang Ma |
Inf. Process. Manag. | 2 |
| 2023 | Discriminative feature mining with relation regularization for person re-identification
Jing Yang 0046, Canlong Zhang, Zhixin Li 0001, Yanping Tang, Zhiwen Wang 0001 |
Inf. Process. Manag. | 3 |
| 2023 | Attributed multi-query community search via random walk similarity
Huifang Ma, Ju Li 0004, Zhixin Li 0001, Liang Chang 0003 |
Inf. Sci. | 4 |
| 2022 | Multi-behavior Recommendation with Two-Level Graph Attentional Networks
Yunhe Wei, Huifang Ma, Yike Wang 0001, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (2) | 4 |
| 2022 | Enhancing Session-Based Recommendation with Global Context Information and Knowledge Graph
Xiaohui Zhang 0020, Huifang Ma, Zihao Gao 0001, Zhixin Li 0001, Liang Chang 0003 |
DASFAA (2) | 4 |
| 2022 | An Effective Two-way Metapath Encoder over Heterogeneous Information Network for RecommendationabstractHeterogeneous information networks (HINs) are widely used in recommender system research due to their ability to model complex auxiliary information beyond historical interactions to alleviate data sparsity problem. Existing HIN-based recommendation studies have achieved great success via performing graph convolution operators between pairs of nodes on predefined metapath induced graphs, but they have the following major limitations. First, existing heterogeneous network construction strategies tend to exploit item attributes while failing to effectively model user relations. In addition, previous HIN-based recommendation models mainly convert heterogeneous graph into homogeneous graphs by defining metapaths ignoring the complicated relation dependency involved on the metapath. To tackle these limitations, we propose a novel recommendation model with two-way metapath encoder for top-N recommendation, which models metapath similarity and sequence relation dependency in HIN to learn node representations. Specifically, our model first learns the initial node representation through a pre-training module, and then identifies potential friends and item relations based on their similarity to construct a unified HIN. We then develop the two-way encoder module with similarity encoder and instance encoder to capture the similarity collaborative signals and relational dependency on different metapaths. Finally, the representations on different meta-paths are aggregated through the attention fusion layer to yield rich representations. Extensive experiments on three real datasets demonstrate the effectiveness of our method. Yanbin Jiang, Huifang Ma, Xiaohui Zhang 0020, Zhixin Li 0001, Liang Chang 0003 |
ICMR | 4 |
| 2022 | Fine-Grained Bidirectional Attention-Based Generative Networks for Image-Text Matching
Zhixin Li 0001, Jianwei Zhu, Jiahui Wei, Yufei Zeng |
ECML/PKDD (3) | 1 |
| 2022 | Flexible Image Captioning via Internal Understanding and External ReasoningabstractImage captioning aims to generate a grammatically correct and semantically accurate natural language description of a given image. In order to capture the more complex information contained in the image and expand the relevant external knowledge outside the image to generate better image caption, this paper proposes an end-to-end image captioning framework Flexible Image Captioning via Internal Understanding and External Reasoning (IUER) based on the Transformer model. IUER enhances visual understanding ability and caption reasoning ability to improve image captioning performance. To achieve this goal, we use the semantic features of the core objects detected from the image to guide the visual feature, where the visual feature incorporate the spatial positional relationship information between the objects, then we introduce external knowledge network to obtain information other than the intuitive content from the image. In this way, a high-quality image caption sentence about the given image is generated. Experiments prove that our method is superior to the baseline model and comparable to other state-of-the-art methods. Jiahui Wei, Zhixin Li 0001, Jianwei Zhu, Huifang Ma |
SDM | 2 |
| 2022 | Exploiting cross-session information for knowledge-aware session-based recommendation via graph attention networksabstractSession-based recommendation (SBR) aims to predict the next item based on anonymous behavior session, which has become increasingly essential in various online services. Prior efforts mainly focus on modeling user preference based on the current session. Although some of them have been proven effective, they fail to address two main challenges in SBR. First, SBR suffers more from the problem of data sparsity due to the very limited user–item interactions, and hence it cannot sufficiently capture complicated item dependency relationships. Second, most of the user-item interaction sequences may be with noisy preference signals due to the uncertainty of user's behaviors, and it is difficult to distill high-quality item for recommendation. In this study, we propose a novel SBR model that exploits Cross-session information for Knowledge-aware Session-based Recommendation (CKSR) to address these two issues. Specifically, cross-session graph and knowledge graph are combined to model a cross-session knowledge graph, based on which a knowledge-aware attention mechanism is performed to capture the complicated transition pattern among interacted items. Each session is then represented as the composition of the global preference and the current interest of that session. Moreover, we leverage the similar sessions for the target session to establish a similar session referral circle and apply an influence coupler to judge the significance of different session referrals. An attentive network is designed to distill session preferences from its unique session referral circle. It dynamically extracts high-quality item from noisy session. Experiments on two benchmark data sets demonstrate that CKSR outperforms the state-of-the-art methods consistently. Xiaohui Zhang 0020, Huifang Ma, Zihao Gao 0001, Zhixin Li 0001, Liang Chang 0003 |
Int. J. Intell. Syst. | 4 |
| 2021 | Exploring Implicit Relationships in Social Network for Recommendation Systems
Yunhe Wei, Huifang Ma, Ruoyi Zhang, Zhixin Li 0001, Liang Chang 0003 |
PAKDD (2) | 4 |
| 2021 | A novel joint biomedical event extraction framework via two-level modeling of documents
Weizhong Zhao, Jinyong Zhang, Jincai Yang, Huifang Ma, Zhixin Li 0001 |
Inf. Sci. | 6 |
| 2020 | Image Captioning with Internal and External KnowledgeabstractAutomatically generating a human-like description for a given image is a potential research in artificial intelligence, which has attracted a great of attention recently. Most of the existing attention methods explore the mapping relationships between words in sentence and regions in image, such unpredictable matching manner sometimes causes inharmonious alignments that may reduce the quality of generated captions. In this paper, we make our efforts to reason about more accurate and meaningful captions. We first propose word attention to improve the correctness of visual attention when generating sequential descriptions word-by-word. The special word attention emphasizes on word importance when focusing on different regions of the input image, and makes full use of the internal annotation knowledge to assist the calculation of visual attention. Then, in order to reveal those incomprehensible intentions that cannot be expressed straightforwardly by machines, we inject external knowledge extracted from knowledge graph into the encoder-decoder framework to facilitate meaningful captioning. We validate our model on two freely available captioning benchmarks: Microsoft COCO dataset and Flickr30k dataset. The results demonstrate that our approach achieves state-of-the-art performance and outperforms many of the existing approaches. Feicheng Huang, Zhixin Li 0001, Shengjia Chen, Canlong Zhang, Huifang Ma |
CIKM | 2 |
| 2020 | Improving Object Detection with Relation Mining NetworkabstractDue to the deteriorated quality of feature in the propagation process of the neural network, it may be hard for traditional detector to identify a small object by just utilizing information within one region proposal. To overcome the limitation of the traditional object detector, we proposed a graph based relation mining network, to capture the relation information from labels and images. The semantic relation network is proposed to mine the global semantic relation in labels, and the spatial relation network is proposed to capture the local spatial relation in images. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the relation mining network explores more advanced feature information. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that key relation information significantly improve the performance of object detection with better ability to detect small objects and reasonable bounding box. The results on COCO dataset demonstrate our method can detect objects robustly, increasing the detection performance of small objects from average precision and average recall by 4.7% and 7.6% respectively in performance relative to Faster R-CNN. Shengjia Chen, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma |
ICDM | 2 |
| 2020 | Improving social and behavior recommendations via network embedding
Weizhong Zhao, Huifang Ma, Zhixin Li 0001, Xiang Ao 0001 |
Inf. Sci. | 3 |
| 2019 | SBRNE: An Improved Unified Framework for Social and Behavior Recommendations with Network Embedding
Weizhong Zhao, Huifang Ma, Zhixin Li 0001, Xiang Ao 0001 |
DASFAA (2) | 3 |
| 2019 | Cross-Media Image-Text Retrieval Combined with Global Similarity and Local SimilarityabstractIn this paper, we study the problem of image-text matching in order to make the image and text have better semantic matching. In the previous work, people just simply used the pre-training network to extract image and text features and project directly into a common subspace, or change various loss functions on this basis, or use the attention mechanism to directly match the image region proposals and the text phrases. This is not a good match for the semantics of the image and the text. In this study, we propose a method of cross-media retrieval based on global representation and local representation. We constructed a cross-media two-level network to explore better semantic matching between images and text, which contains subnets that handle both global and local features. Specifically, we not only use the self-attention network to obtain a macro representation of the global image but also use the local fine-grained patch with the attention mechanism. Then, we use a two-level alignment framework to promote each other to learn different representations of cross-media retrieval. The innovation of this study lies in the use of more comprehensive features of image and text to design the two kinds of similarity and add them up in some way. Experimental results show that this method is effective in image-text retrieval. Experimental results on the Flickr30K and MS-COCO datasets show that this model has a better recall rate than many of the current advanced cross-media retrieval models. Zhixin Li 0001, Canlong Zhang |
DSAA | 1 |
| 2019 | Leveraging User Preferences for Community Search via Attribute Subspace
Haijiao Liu, Huifang Ma, Yang Chang, Zhixin Li 0001, Wenjuan Wu |
KSEM (1) | 4 |
| 2019 | Object Detection by Combining Deep Dilated Convolutions Network and Light-Weight Network
Yu Quan, Zhixin Li 0001, Canlong Zhang |
KSEM (1) | 2 |
| 2019 | Effectively Classify Short Texts with Sparse Representation Using Entropy Weighted Constraint
Ting Tuo, Huifang Ma, Zhixin Li 0001, Xianghong Lin |
KSEM (2) | 3 |
| 2019 | Collaborating CNN and SVM for Automatic Image AnnotationabstractTo learn a well-performed image annotation model, a large number of labeled samples are usually required. In this paper, we propose a novel semi-supervised approach based on adaptive weighted fusion for automatic image annotation, which can utilize the labeled data and unlabeled data simultaneously. Firstly, two different classifiers, namely the CNN (convolutional neural network) and the LDA-SVM, are constructed by all the labeled data. These two classifiers are independently represented as different feature views. Then, the most confident data with relevant pseudo-labels are chosen and amalgamated with the whole labeled dataset. After that, the two classifiers are retrained with the new labeled dataset until a stop condition is reached. In each iteration process, the unlabeled samples are labeled by high confidence pseudo-labels that are estimated by an adaptive weighted fusion strategy. Finally, we conduct experiments on two datasets, namely IAPR TC12 and NUS-WIDE, and measure the performance of the model with standard criteria, including precision, recall, F-measure, N+ and mAP. The experimental results show that our approach outperforms many state-of-the-art automatic image annotation approaches. Zhixin Li 0001, Canlong Zhang, Huifang Ma, Weizhong Zhao |
ICMR | 1 |
| 2019 | Short Text Similarity Measurement Based on Coupled Semantic Relation and Strong Classification Features
Huifang Ma, Zhixin Li 0001, Xianghong Lin |
PAKDD (1) | 3 |
| 2010 | Combining the Missing Link: An Incremental Topic Model of Document Content and HyperlinkabstractThe content and structure of linked information such as sets of web pages or research paper archives are dynamic and keep on changing. Even though different methods are proposed to exploit both the link structure and the content information, no existing approach can effectively deal with this evolution. We propose a novel joint model, called Link-IPLSI, to combine texts and links in a topic modeling framework incrementally. The model takes advantage of a novel link updating technique that can cope with dynamic changes of online document streams in a faster and scalable way. Furthermore, an adaptive asymmetric learning method is adopted to freely control the assignment of weights to terms and citations. Experimental results on two different sources of online information demonstrate the time saving strength of our method and indicate that our model leads to systematic improvements in the quality of classification and link prediction. Huifang Ma, Zhixin Li 0001, Zhongzhi Shi |
APWeb | 2 |