EDBT 2026 Demo / reviewers in the wild / expert
Jie Ma 0001
dblp:62/5110-1
· DBLP profile ↗
25ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-7432-3238ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Securely Answering Multi-Hop Questions on the Joint of Private and Public Knowledge GraphsabstractKnowledge Graph Question Answering (KGQA) plays an important role in modern information service systems. However, due to the challenges in constructing Knowledge Graphs (KGs) and their rapid update pace, KGs often remain incomplete. Previous works on multi-hop KGQA on incomplete KGs tend to explore potential relations between topic entities and answers, so they are constrained by the limited set of entities in KGs. To address the issue, we propose a method that executes multi-hop KGQA over private-public KGs to expand the size of the available entity set while being privacy-preserving to the private KGs and the user query. To be specific, by embedding entities and relations in both private and public KGs and using the triple scoring function as a distance metric, we reduce multi-hop KGQA into the iterative information retrieval task and formulate the iterative inference module for this task. Extensive experiments demonstrate that our method achieves the absolute average accuracy increase of 15.4% on WebQuestionsSP and 24.4% on SimpleQuestions over state-of-the-art methods under various private-public settings while achieving no privacy leakage of user queries. Shuaipeng Li, Baoyu An, Xueyang Huo, Jie Ma 0001, Yuansi Zhang, Linxi Cai, Pinghui Wang, Chao-Bo Yan |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingabstractChart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension. While existing Visual Language Models (VLMs) have shown progress in chart understanding, the lack of high-quality training data and comprehensive evaluation benchmarks hinders VLM chart comprehension. In this paper, we introduce EvoChart, a novel self-training method for generating synthetic chart data to enhance VLMs' capabilities in real-world chart comprehension. We also propose EvoChart-QA, a noval benchmark for measuring models' chart comprehension abilities in real-world scenarios. Specifically, EvoChart is a unique self-training data synthesis approach that simultaneously produces high-quality training corpus and a high-performance chart understanding model. EvoChart-QA consists of 650 distinct real-world charts collected from 140 different websites and 1,250 expert-curated questions that focus on chart understanding. Experimental results on various open-source and proprietary VLMs tested on EvoChart-QA demonstrate that even the best proprietary model, GPT-4o, achieves only 49.8% accuracy. Moreover, the EvoChart method significantly boosts the performance of open-source VLMs on real-world chart understanding tasks, achieving 54.2% accuracy on EvoChart-QA. Muye Huang, Han Lai, Xinyu Zhang 0021, Jie Ma 0001, Lingling Zhang 0005, Jun Liu 0002 |
AAAI | 5 |
| 2025 | Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language ModelsabstractLarge Language Models (LLMs) may suffer from hallucinations in real-world applications due to the lack of relevant knowledge. In contrast, knowledge graphs encompass extensive, multi-relational structures that store a vast array of symbolic facts. Consequently, integrating LLMs with knowledge graphs has been extensively explored, with Knowledge Graph Question Answering (KGQA) serving as a critical touchstone for the integration. This task requires LLMs to answer natural language questions by retrieving relevant triples from knowledge graphs. However, existing methods face two significant challenges: *excessively long reasoning paths distracting from the answer generation*, and *false-positive relations hindering the path refinement*. In this paper, we propose an iterative interactive KGQA framework that leverages the interactive learning capabilities of LLMs to perform reasoning and Debating over Graphs (DoG). Specifically, DoG employs a subgraph-focusing mechanism, allowing LLMs to perform answer trying after each reasoning step, thereby mitigating the impact of lengthy reasoning paths. On the other hand, DoG utilizes a multi-role debate team to gradually simplify complex questions, reducing the influence of false-positive relations. This debate mechanism ensures the reliability of the reasoning process. Experimental results on five public datasets demonstrate the effectiveness and superiority of our architecture. Notably, DoG outperforms the state-of-the-art method ToG by 23.7% and 9.1% in accuracy on WebQuestions and GrailQA, respectively. Furthermore, the integration experiments with various LLMs on the mentioned datasets highlight the flexibility of DoG. Jie Ma 0001, Zhitao Gao 0003, Qi Chai, Wangchun Sun, Pinghui Wang, Hongbin Pei, Lingyun Song, Jun Liu 0002 |
AAAI | 1 |
| 2025 | Non-Stationary Predictions May Be More Informative: Exploring Pseudo-Labels with a Two-Phase Pattern of Training DynamicsabstractPseudo-labeling is a widely used strategy in semi-supervised learning. Existing methods typically select predicted labels with high confidence scores and high training stationarity, as pseudo-labels to augment training sets. In contrast, this paper explores the pseudo-labeling potential of predicted labels that **do not** exhibit these characteristics. We discover a new type of predicted labels suitable for pseudo-labeling, termed *two-phase labels*, which exhibit a two-phase pattern during training: *they are initially predicted as one category in early training stages and switch to another category in subsequent epochs.* Case studies show the two-phase labels are informative for decision boundaries. To effectively identify the two-phase labels, we design a 2-*phasic* metric that mathematically characterizes their spatial and temporal patterns. Furthermore, we propose a loss function tailored for two-phase pseudo-labeling learning, allowing models not only to learn correct correlations but also to eliminate false ones. Extensive experiments on eight datasets show that **our proposed 2-*phasic* metric acts as a powerful booster** for existing pseudo-labeling methods by additionally incorporating the two-phase labels, achieving an average classification accuracy gain of 1.73% on image datasets and 1.92% on graph datasets. Hongbin Pei, Jingxin Hai, Huiqi Deng, Denghao Ma, Jie Ma 0001, Pinghui Wang, Xiaohong Guan |
ICML | 6 |
| 2025 | Multi-Scale Temporal Neural Network for Stock Trend Prediction Enhanced by Temporal Hyepredge LearningabstractExisting research in Stock Trend Prediction (STP) focuses on temporal features extracted from a temporal sequence of stock data with a look-back window, which frequently leads to the omission of important periodic patterns, such as weekly and monthly variations in stock prices. Furthermore, these methods examine stocks individually, ignoring the temporal variation patterns among stocks that share higher-order relationships, like those within the same industry. These relationships typically provide contextual insights into market investments influencing stock price fluctuations. To tackle these issues, we propose a Multi-Scale Temporal Neural Network (MSTNN) framework tailored for STP. This architecture explores the periodic fluctuation behaviors of individual stocks through an innovative 3D convolutional neural network, alongside examining temporal variation patterns of stocks linked to specific industries via a temporal hypergraph attention mechanism. Empirical results from two real-world benchmark datasets show that MSTNN significantly outperforms prior state-of-the-art STP methods. The code of our MSTNN is available at https://github.com/sunlitsong/MSTNN. Lingyun Song, Siyu Chen 0024, Xinbiao Gan, Binze Shi, Jie Ma 0001, Yudai Pan, Xuequn Shang 0001 |
IJCAI | 6 |
| 2025 | Metapath and Hypergraph Structure-based Multi-Channel Graph Contrastive Learning for Student Performance PredictionabstractConsiderable attention has been paid to predicting student performance on exercises. The performance of prior studies is determined by the quality of the trait features of students and exercises. Nevertheless, most of the prior study primarily examines simple pairwise interactions in learning trait features, like those between students and exercises or exercises and concepts, while disregarding the complex higher-order interactions that typically exist among these components, which in turn hinders the prediction results. In this paper, we using an innovative Multi-Channel Graph Contrastive Learning (MCGCL) framework that integrates various high-order interactions for predicting student performance. MCGCL characterizes graph structures reflecting various high-order relationships among students, exercises, and concepts through multiple channels, thereby enhancing the trait features of both students and exercises. Moreover, graph contrastive learning is employed to enhance the representation of trait features acquired from high-order graph structures in diverse views. Extensive experiments on real-world datasets show that MCGCL achieves state-of-the-art results on the task of predicting student performance. The code is available at https://github.com/sunlitsong/MCGCL. Lingyun Song, Xiaofan Sun, Xinbiao Gan, Yudai Pan, Xiaolin Han 0002, Jie Ma 0001, Jun Liu 0002, Xuequn Shang 0001 |
IJCAI | 6 |
| 2025 | ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingabstractCharts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension. Muye Huang, Lingling Zhang 0005, Jie Ma 0001, Han Lai, Fangzhi Xu, Yifei Li 0006, Yaqiang Wu, Jun Liu 0002 |
NeurIPS | 3 |
| 2025 | Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge GraphsabstractKnowledge graph-based retrieval-augmented generation seeks to mitigate hallucinations in Large Language Models (LLMs) caused by insufficient or outdated knowledge. However, existing methods often fail to fully exploit the prior knowledge embedded in knowledge graphs (KGs), particularly their structural information and explicit or implicit constraints. The former can enhance the faithfulness of LLMs' reasoning, while the latter can improve the reliability of response generations. Motivated by these, we propose a trustworthy reasoning framework, termed Deliberation over Priors (\texttt{DP}), which sufficiently utilizes the priors contained in KGs. Specifically, \texttt{DP} adopts a progressive knowledge distillation strategy that integrates structural priors into LLMs through a combination of supervised fine-tuning and Kahneman-Tversky Optimization, thereby improving the faithfulness of relation path generation. Furthermore, our framework employs a reasoning-introspection strategy, which guides LLMs to perform refined reasoning verification based on extracted constraint priors, ensuring the reliability of response generation. Extensive experiments on three benchmark datasets demonstrate that \texttt{DP} achieves new state-of-the-art performance, especially a H@1 improvement of 13% on the ComplexWebQuestions dataset, and generates highly trustworthy responses. We also conduct various analyses to verify its flexibility and practicality. Code is available at [https://github.com/mira-ai-lab/Deliberation-on-Priors](https://github.com/mira-ai-lab/Deliberation-on-Priors). Jie Ma 0001, Ning Qu, Zhitao Gao 0003, Jun Liu 0002, Hongbin Pei, Jiang Xie 0002, Lingyun Song, Pinghui Wang |
NeurIPS | 1 |
| 2024 | Multi-Track Message Passing: Tackling Oversmoothing and Oversquashing in Graph Learning via Preventing Heterophily MixingabstractThe advancement toward deeper graph neural networks is currently obscured by two inherent issues in message passing, oversmoothing and oversquashing. We identify the root cause of these issues as information loss due to heterophily mixing in aggregation, where messages of diverse category semantics are mixed. We propose a novel multi-track graph convolutional network to address oversmoothing and oversquashing effectively. Our basic idea is intuitive: if messages are separated and independently propagated according to their category semantics, heterophilic mixing can be prevented. Consequently, we present a novel multi-track message passing scheme capable of preventing heterophilic mixing, enhancing long-distance information flow, and improving separation condition. Empirical validations show that our model achieved state-of-the-art performance on several graph datasets and effectively tackled oversmoothing and oversquashing, setting a new benchmark of $86.4$% accuracy on Cora. Hongbin Pei, Huiqi Deng, Jingxin Hai, Pinghui Wang, Jie Ma 0001, Yuheng Xiong, Xiaohong Guan |
ICML | 6 |
| 2024 | Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robustness. Furthermore, current datasets may not provide a precise diagnostic for these methods. To tackle these challenges, firstly, we propose a novel dataset, *MUSIC-AVQA-R*, crafted in two steps: rephrasing questions within the test split of a public dataset (*MUSIC-AVQA*) and subsequently introducing distribution shifts to split questions. The former leads to a large, diverse test space, while the latter results in a comprehensive robustness evaluation on rare, frequent, and overall questions. Secondly, we propose a robust architecture that utilizes a multifaceted cycle collaborative debiasing strategy to overcome bias learning. Experimental results show that this architecture achieves state-of-the-art performance on MUSIC-AVQA-R, notably obtaining a significant improvement of 9.32\%. Extensive ablation experiments are conducted on the two datasets mentioned to analyze the component effectiveness within the debiasing strategy. Additionally, we highlight the limited robustness of existing multi-modal QA methods through the evaluation on our dataset. We also conduct experiments combining various baselines with our proposed strategy on two datasets to verify its plug-and-play capability. Our dataset and code are available at <https://github.com/reml-group/MUSIC-AVQA-R>. Jie Ma 0001, Pinghui Wang, Wangchun Sun, Lingyun Song, Hongbin Pei, Jun Liu 0002, Youtian Du |
NeurIPS | 1 |
| 2024 | Memory Disagreement: A Pseudo-Labeling Measure from Training Dynamics for Semi-supervised Graph Learning
Hongbin Pei, Yuheng Xiong, Pinghui Wang, Jialun Liu, Huiqi Deng, Jie Ma 0001, Xiaohong Guan |
WWW | 7 |
| 2024 | Diagram Perception Networks for Textbook Question Answering via Joint Optimization
Jie Ma 0001, Jun Liu 0002, Qi Chai, Pinghui Wang |
Int. J. Comput. Vis. | 1 |
| 2024 | A Multi-Group Multi-Stream attribute Attention network for fine-grained zero-shot learning
Lingyun Song, Xuequn Shang 0001, Ruizhi Zhou, Jun Liu 0002, Jie Ma 0001, Zhanhuai Li, Mingxuan Sun 0001 |
Neural Networks | 5 |
| 2024 | Robust Visual Question Answering: Datasets, Methods, and Future ChallengesabstractVisual question answering requires a system to provide an accurate natural language answer given an image and a natural language question. However, it is widely recognized that previous generic VQA methods often tend to memorize biases present in the training data rather than learning proper behaviors, such as grounding images before predicting answers. Therefore, these methods usually achieve high in-distribution but poor out-of-distribution performance. In recent years, various datasets and debiasing methods have been proposed to evaluate and enhance the VQA robustness, respectively. This paper provides the first comprehensive survey focused on this emerging fashion. Specifically, we first provide an overview of the development process of datasets from in-distribution and out-of-distribution perspectives. Then, we examine the evaluation metrics employed by these datasets. Third, we propose a typology that presents the development process, similarities and differences, robustness comparison, and technical features of existing debiasing methods. Furthermore, we analyze and discuss the robustness of representative vision-and-language pre-training models on VQA. Finally, through a thorough review of the available literature and experimental analysis, we discuss the key areas for future research from various viewpoints. Jie Ma 0001, Pinghui Wang, Dechen Kong, Jun Liu 0002, Hongbin Pei, Junzhou Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | XTQA: Span-Level Explanations for Textbook Question AnsweringabstractTextbook question answering (TQA) is the task of correctly answering diagram or nondiagram (ND) questions given large multimodal contexts consisting of abundant essays and diagrams. In real-world scenarios, an explainable TQA system plays a key role in deepening humans' understanding of learned knowledge. However, there is no work to investigate how to provide explanations currently. To address this issue, we devise a novel architecture toward span-level eXplanations for TQA (XTQA). In this article, spans are the combinations of sentences within a paragraph. The key idea is to consider the entire textual context of a lesson as candidate evidence and then use our proposed coarse-to-fine grained explanation extracting (EE) algorithm to narrow down the evidence scope and extract the span-level explanations with varying lengths for answering different questions. The EE algorithm can also be integrated into other TQA methods to make them explainable and improve the TQA performance. Experimental results show that XTQA obtains the best overall explanation result [mean intersection over union (mIoU)] of 52.38% on the first 300 questions of CK12-QA test splits, demonstrating the explainability of our method (ND: 150 and diagram: 150). The results also show that XTQA achieves the best TQA performance of 36.46% and 36.95% on the aforementioned splits. We have released our code in https://github.com/dr-majie/opentqa. Jie Ma 0001, Qi Chai, Jun Liu 0002, Qingyu Yin, Pinghui Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | GBTTE: Graph Attention Network Based Bus Travel Time EstimationabstractReal-time bus travel time is crucial for the smart public transportation system and is beneficial for improving user satisfaction for online map services. However, it faces great challenges due to fine-grained spatial dependencies and dynamic temporal dependencies. To address the above problem, we propose GBTTE, a novel end-to-end graph attention network framework to estimate bus travel time. Specifically, we construct a novel graph structure of bus routes and use a graph attention network to capture the fine-grained spatial features of bus routes. Then, we fully exploit the joint spatial-temporal relations of bus stops through a spatial-temporal graph attention network and also capture the dynamic correlation between the route and the bus transportation network with a cross graph attention network. Finally, we integrate the route representation, the spatial-temporal representation and contextual information to estimate bus travel time. Extensive experiments carried out on two large-scale real-world datasets demonstrate the effectiveness of GBTTE. In addition, GBTTE has been deployed in production at Baidu Maps, handling tens of millions of requests every day. Yuecheng Rong, Juntao Yao, Jun Liu 0002, Yifan Fang, Wei Luo 0013, Hao Liu 0026, Jie Ma 0001, Zepeng Dan, Jinzhu Lin, Yan Zhang 0156, Chuanming Zhang |
CIKM | 7 |
| 2023 | Dynamic dual graph networks for textbook question answering
Yaxian Wang, Jun Liu 0002, Jie Ma 0001, Hongwei Zeng 0001, Lingling Zhang 0005 |
Pattern Recognit. | 3 |
| 2023 | Multitask Learning for Visual Question AnsweringabstractVisual question answering (VQA) is a task that machines should provide an accurate natural language answer given an image and a question about the image. Many studies have found that the current VQA methods are heavily driven by the surface correlation or statistical bias in the training data, and lack sufficient image grounding. To address this issue, we devise a novel end-to-end architecture that uses multitask learning to promote more sufficient image grounding and learn effective multimodality representations. The tasks consist of VQA and our proposed image cloze (IC) task requires machines to fill in the blanks accurately given an image and a textual description of the image. To ensure our model performs sufficient image grounding as much as possible, we propose a novel word-masking algorithm to develop the multimodal IC task based on the part-of-speech of words. Our model predicts the VQA answer and fills in the blanks after the multimodality representation learning that is shared by the two tasks. Experimental results show that our model achieves almost the equivalent, state-of-the-art, second-best performance on the VQA v2.0, VQA-changing priors (CP) v2, and grounded question answering (GQA) datasets, respectively, with fewer parameters and without additional data compared with baselines. Jie Ma 0001, Jun Liu 0002, Qika Lin, Bei Wu 0003, Yaxian Wang, Yang You 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Relation-Aware Fine-Grained Reasoning Network for Textbook Question AnsweringabstractTextbook question answering (TQA) is a task that one should answer non-diagram and diagram questions accurately, given a large context which consists of abundant diagrams and essays. Although lots of studies have made significant progress in the natural image question answering (QA), they are not applicable to comprehending diagrams and reasoning over the long multimodal context. To address the above issues, we propose a relation-aware fine-grained reasoning (RAFR) network that performs fine-grained reasoning over the nodes of relation-based diagram graphs. Our method uses semantic dependencies and relative positions between nodes in the diagram to construct relation graphs and applies graph attention networks to learn diagram representations. To extract and reason over the multimodal knowledge, we first extract the text that is the most relevant to questions, options, and the instructional diagram which is the most relevant to question diagrams at the word-sentence level and the node-diagram level, respectively. Then, we apply instructional-diagram-guided attention and question-guided attention to reason over the node of question diagrams, respectively. The experimental results show that our proposed method achieves the best performance on the TQA dataset compared with baselines. We also conduct extensive ablation studies to comprehensively analyze the proposed method. Jie Ma 0001, Jun Liu 0002, Yaxian Wang, Tongliang Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Weakly Supervised Learning for Textbook Question AnsweringabstractTextbook Question Answering (TQA) is the task of answering diagram and non-diagram questions given large multi-modal contexts consisting of abundant text and diagrams. Deep text understandings and effective learning of diagram semantics are important for this task due to its specificity. In this paper, we propose a Weakly Supervised learning method for TQA (WSTQ), which regards the incompletely accurate results of essential intermediate procedures for this task as supervision to develop Text Matching (TM) and Relation Detection (RD) tasks and then employs the tasks to motivate itself to learn strong text comprehension and excellent diagram semantics respectively. Specifically, we apply the result of text retrieval to build positive as well as negative text pairs. In order to learn deep text understandings, we first pre-train the text understanding module of WSTQ on TM and then fine-tune it on TQA. We build positive as well as negative relation pairs by checking whether there is any overlap between the items/regions detected from diagrams using object detection. The RD task forces our method to learn the relationships between regions, which are crucial to express the diagram semantics. We train WSTQ on RD and TQA simultaneously, i.e., multitask learning, to obtain effective diagram semantics and then improve the TQA performance. Extensive experiments are carried out on CK12-QA and AI2D to verify the effectiveness of WSTQ. Experimental results show that our method achieves significant accuracy improvements of 5.02% and 4.12% on test splits of the above datasets respectively than the current state-of-the-art baseline. We have released our code on https://github.com/dr-majie/WSTQ. Jie Ma 0001, Qi Chai, Jingyue Huang, Jun Liu 0002, Yang You 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Rule-enhanced iterative complementation for knowledge graph reasoning
Qika Lin, Jun Liu 0002, Yudai Pan, Lingling Zhang 0005, Jie Ma 0001 |
Inf. Sci. | 6 |
| 2020 | Stochastic Batch Augmentation with An Effective Distilled Dynamic Soft Label RegularizerabstractData augmentation have been intensively used in training deep neural network to improve the generalization, whether in original space (e.g., image space) or representation space. Although being successful, the connection between the synthesized data and the original data is largely ignored in training, without considering the distribution information that the synthesized samples are surrounding the original sample in training. Hence, the behavior of the network is not optimized for this. However, that behavior is crucially important for generalization, even in the adversarial setting, for the safety of the deep learning system. In this work, we propose a framework called Stochastic Batch Augmentation (SBA) to address these problems. SBA stochastically decides whether to augment at iterations controlled by the batch scheduler and in which a ''distilled'' dynamic soft label regularization is introduced by incorporating the similarity in the vicinity distribution respect to raw samples. The proposed regularization provides direct supervision by the KL-Divergence between the output soft-max distributions of original and virtual data. Our experiments on CIFAR-10, CIFAR-100, and ImageNet show that SBA can improve the generalization of the neural networks and speed up the convergence of network training. Qian Li 0024, Yong Qi 0001, Saiyu Qi, Jie Ma 0001, Jian Zhang 0087 |
IJCAI | 5 |
| 2020 | Jointly Optimized Neural Coreference Resolution with Mutual AttentionabstractCoreference resolution aims at recognizing different forms in a document which refer to the same entity in the real world. Although many models have been proposed and achieved success, there still exist some challenges. Recent models that use recurrent neural networks to obtain mention representations ignore dependencies between spans and their proceeding distant spans, which will lead to predicted clusters that are locally consistent but globally inconsistent. In addition, these models are trained only by maximizing the marginal likelihood of gold antecedent spans from coreference clusters, which will make some gold mentions undetectable and cause unsatisfactory coreference results. To address these challenges, we propose a neural coreference resolution model. It employs mutual attention to take into account the dependencies between spans and their proceeding spans directly (use attention mechanism to capture global information between spans and their proceeding spans). And our model is trained by jointly optimizing mention clustering and imbalanced mention detection, which enables it to detect more gold mentions in a document to make more accurate coreference decisions. Experimental results on the CoNLL-2012 English dataset show that our model can detect the most gold mentions and achieve the state-of-the-art coreference performance compared with baselines. Jie Ma 0001, Jun Liu 0002, Yufei Li 0002, Yudai Pan, Shen Sun, Qika Lin |
WSDM | 1 |
| 2020 | Fine-Grained 3D-Attention Prototypes for Few-Shot LearningabstractIn the real world, a limited number of labeled finely grained images per class can hardly represent the class distribution effectively. Due to the more subtle visual differences in fine-grained images than simple images with obvious objects, that is, there exist smaller interclass and larger intraclass variations. To solve these issues, we propose an end-to-end attention-based model for fine-grained few-shot image classification (AFG) with the recent episode training strategy. It is composed mainly of a feature learning module, an image reconstruction module, and a label distribution module. The feature learning module mainly devises a 3D-Attention mechanism, which considers both the spatial positions and different channel attentions of the image features, in order to learn more discriminative local features to better represent the class distribution. The image reconstruction module calculates the mappings between local features and the original images. It is constrained by a designed loss function as auxiliary supervised information, so that the learning of each local feature does not need extra annotations. The label distribution module is used to predict the label distribution of a given unlabeled sample, and we use the local features to represent the image features for classification. By conducting comprehensive experiments on Mini-ImageNet and three fine-grained data sets, we demonstrate that the proposed model achieves superior performance over the competitors. Jun Liu 0002, Jie Ma 0001, Yudai Pan, Lingling Zhang 0005 |
Neural Comput. | 3 |
| 2018 | Density core-based clustering algorithm with dynamic scanning radius
Jiang Xie 0002, Zhongyang Xiong, Yu-Fang Zhang, Yong Feng 0002, Jie Ma 0001 |
Knowl. Based Syst. | 5 |