EDBT 2026 Demo / reviewers in the wild / expert
Jin Xu 0014
dblp:97/3265-14
· DBLP profile ↗
25ranked-venue papers
0as first author
23since 2021 · last 2027
0009-0001-8735-3532ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | RTBAgent-X: A unified framework for transparent real-time bidding
Leng Cai, Junxuan He, Yuanping Lin, Yawen Zeng, Jin Xu 0014 |
Expert Syst. Appl. | 5 |
| 2026 | WeSEAL: Well-calibrated Search for Eliminating Attention-sink Leakage
Juyuan Wang, Chenxing Wang 0001, Aolin Li, Huiyun Hu, Yuchen Fang 0001, Haijun Wu, Jin Xu 0014, Dongliang Liao |
SIGIR | 7 |
| 2026 | MMRepAgent: Explainable stock earnings forecasting via multimodal report agent framework
Xiangyu Li 0010, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001 |
Knowl. Based Syst. | 4 |
| 2026 | Multi-View Graph Clustering via Dual View-Cluster-Order Interactivity MiningabstractMulti-view Graph Clustering (MGC) is a crucial approach for uncovering complex data structures by leveraging multiple perspectives of data. However, existing MGC methods face two key challenges: (1) limitations in graph structure that neglect long-range dependencies, and (2) overlooking the view-cluster local structure when mining view discrepancies. To address these issues, we propose a Multi-view Graph Clustering approach based on Dual View-Cluster-Order Interactivity (DVCOI-MGC). This approach consists of three modules: (1) Multi-View Multi-Order Graph Construction, where high-order graphs are generated using matrix exponentiation to capture long-range dependencies; (2) Dual View-Cluster-Order Interactivity, which utilizes a discrete graph cut model to separately learn order-specific and view-specific clustering results from the sets of order-specific multi-view graphs and view-specific multi-order graphs, with a separate View-Cluster-Order tensor weight for each learning direction; and (3) Bidirectional Truncation Consistency Learning, which applies a sparse boolean weight vector to locally select and integrate clustering results while preserving both the view-cluster and order-cluster local structures. Additionally, we introduce an efficient iterative optimization method to solve the discrete graph cut problem and provide a theoretical analysis of its convergence and computational complexity. Extensive experiments on 8 real-world datasets demonstrate that our approach significantly improves clustering performance over 11 state-of-the-art methods. Xia Dong, Penglei Wang, Jin Xu 0014, Danyang Wu, Feiping Nie 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | HEAR: A Holistic Extraction and Agentic Reasoning Framework for Document UnderstandingabstractThe automated comprehension of complex, multi-modal documents is fundamentally hampered by a disconnect between information extraction and reasoning. Existing systems suffer from inherent limitations. Monolithic models embed reasoning as a black box process, sacrificing transparency and depth. Meanwhile, current agent-based frameworks follow a passive, non-interactive paradigm; they handle static, global inputs rather than information derived from active exploration, which fundamentally restricts their ability to achieve structural understanding and complex reasoning. To bridge this critical gap, we introduce HEAR, a framework for Holistic Extraction and Agentic Reasoning. This innovative framework establishes a synergistic, closed-loop between a deep Vision-Language Model (VLM) driven holistic parsing engine and a collaborative multi-agent reasoning system. Our HEAR initially transforms unstructured documents into a semantically-rich, structured representation, preserving complex layouts and reconstituting multi-page tables. Subsequently, a multi-agent system performs cross-modal analysis, governed by a crucial verification protocol that forces agents to validate findings across textual and visual modalities. A conflict driven re-evaluation mechanism enables the system to dynamically re-engage the document to resolve ambiguities, thereby unifying the perception-cognition cycle. HEAR achieved first place in the ACM MM 2025 Grand Challenge on Large Vision–Language Model Learning and Applications. Longfeng Chen, Juyuan Wang, Yawen Zeng, Jin Xu 0014 |
ACM Multimedia | 6 |
| 2025 | MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical InstructionsabstractGiven the significant advances in Large Vision Language Models (LVLMs) in reasoning and visual understanding, mobile agents are rapidly emerging to meet users' automation needs. However, existing evaluation benchmarks are disconnected from the real world and fail to adequately address the diverse and complex requirements of users. From our extensive collection of user questionnaire, we identified five tasks: Multi-App, Vague, Interactive, Single-App, and Unethical Instructions. Around these tasks, we present MVISU-Bench, a bilingual benchmark that includes 404 tasks across 137 mobile applications. Furthermore, we propose Aider, a plug-and-play module that acts as a dynamic prompt prompter to mitigate risks and clarify user intent for mobile agents. Our Aider is easy to integrate into several frameworks and has successfully improved overall success rates by 19.55% compared to the current state-of-the-art (SOTA) on MVISU-Bench. Specifically, it achieves success rate improvements of 53.52% and 29.41% for unethical and interactive instructions, respectively. Through extensive experiments and analysis, we highlight the gap between existing mobile agents and real-world user expectations. Juyuan Wang, Longfeng Chen, Boyi Xiao, Leng Cai, Yawen Zeng, Jin Xu 0014 |
ACM Multimedia | 7 |
| 2025 | Unsupervised Cross-view Message Passing Method for Multi-view Graph ClusteringabstractIn recent years, multi-view graph clustering (MVGC) has attracted increasing attention from researchers. However, many existing MVGC methods focus on view-level integration through strategies like assigning weights to different views, for example, ignoring cross-view interactions between nodes. In fact, cross-view interactions at node level are crucial for extraction and fusion of semantic information. Additionally, some methods separate representation learning from clustering, which results in suboptimal clustering performance. To address these problems, we propose a novel unsupervised cross-view message passing method for MVGC. The kernel of our method is the cross-view interaction mechanism, which dynamically constructs node-specific cross-view edges based on node features and structural information. The mechanism enables adaptive interactions of informative nodes from different views, which promotes the extraction and propagation of complementary information. Besides, our method unifies representation learning and hyperspherical clustering in an end-to-end framework, which projects node representations into a hypersphere space, thereby enabling direct acquisition of balanced clustering results without dependence on external clustering methods. We provide comprehensive analyses on our method, and evaluate our method on six multi-view datasets. The results show that our method consistently achieves superior performance than existing state-of-the-art multi-view clustering methods. Ziming Quan, Penglei Wang, Danyang Wu, Jin Xu 0014 |
ACM Multimedia | 4 |
| 2025 | Cluster-Aware Contrastive Multi-View Clustering Based on Masked ViewsabstractIn this paper, we present a novel Self-Supervised Learning (SSL) framework tailored for Multi-View Clustering (MVC), which learns cross-view semantic representations with clear clustering boundaries and derives balanced clustering in an end-to-end manner. Concretely, we propose a generative SSL module that learns high-level semantic representations by recovering randomly masked views from observed views. Then the extracted representations are unified via a sample-level local fusion mechanism and projected into a unit-hypersphere space with evenly distributed cluster prototypes such that the pseudo labels can be directly retrieved using cosine similarity. For each sample, we define highly credible positive pairs of the same cluster and negative pairs of different clusters and design a contrastive SSL module to force the sample to move toward its cluster prototype while farther from the other prototypes in the embedding space. Consequently, the representations exhibit clearer clustering boundaries, and the two SSL modules benefit each other. Finally, we further introduce a clustering regularizer to prevent trivial solutions and derive balanced clustering with theoretical guarantees. Comprehensive evaluations over eight benchmark datasets validate the effectiveness of our proposals against ten state-of-the-art MVC methods. Penglei Wang, Ziming Quan, Danyang Wu, Jin Xu 0014 |
ACM Multimedia | 4 |
| 2025 | Comprehensive Information Extraction With Separable Representation Learning for Multi-View ClusteringabstractDeep Multi-View Clustering (MVC) methods partition multi-view data into disjoint clusters in an unsupervised manner, showing significant promise across various domains. However, current MVC methods primarily focus on capturing the consistency information shared across all views and undervalue the specificity information inherent in each view that reflects its unique characteristics. Furthermore, the underexploration of the separability of learned representations limits the overall clustering performance of existing MVC methods and leads to undesirable clustering results. In this paper, we propose a fully differentiable and end-to-end deep MVC framework, named Comprehensive Information Extraction with Separable Representation Learning (CIRSEL), to address these issues. CIRSEL recasts specificity information extraction as a high-order graph pooling process to capture the view-specific characteristics of individual views. Utilizing the cross-attention mechanism, CIRSEL adaptively fuses the consistent and view-specific representations to achieve comprehensive information extraction. Subsequently, CIRSEL maps representations into a unit hypersphere space with evenly distributed prototypes and maximizes the variational estimation of Mutual Information, which enhances the inter-cluster separability and intra-cluster compactness in the embedding space and further benefits the following clustering learning. Finally, CIRSEL introduces a nuclear norm-based balance regularization, which ensures balanced clustering results can be directly retrieved by the cosine similarity between the representations and prototypes. Extensive experiments on ten benchmark datasets demonstrate the effectiveness of CIRSEL compared to sixteen current MVC methods. Penglei Wang, Danyang Wu, Jin Xu 0014, Feiping Nie 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | RetrievalMMT: Retrieval-Constrained Multi-Modal Prompt Learning for Multi-Modal Machine TranslationabstractAs an extension of machine translation, the primary objective of multi-modal machine translation is to optimize the utilization of visual information. Technically, image information is integrated into multi-modal fusion and alignment as an auxiliary modality through concepts or latent semantics, which are typically based on the Transformer framework. However, current approaches often ignore one modality to design numerous handcrafted features (e.g. visual concept extraction) and require training of all parameters in their framework. Therefore, it is worthwhile to explore multi-modal concepts or features to enhance performance and an efficient approach to incorporate visual information with minimal cost. Meanwhile, with the development of multi-modal large language models (MLLMs), they are faced with the visual hallucination issue of compromising performance, despite their powerful capabilities. Inspired by pioneering techniques in the multi-modal field, such as prompt learning and MLLMs, this paper innovatively explores the possibility of applying multi-modal prompt learning to this multi-modal machine translation task. Yan Wang 0140, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001 |
ICMR | 5 |
| 2024 | Contrastive topic-enhanced network for video captioning
Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Hong Man, Xiangmin Xu 0001 |
Expert Syst. Appl. | 5 |
| 2024 | EBMGC-GNF: Efficient Balanced Multi-View Graph Clustering via Good Neighbor FusionabstractExploiting consistent structure from multiple graphs is vital for multi-view graph clustering. To achieve this goal, we propose an Efficient Balanced Multi-view Graph Clustering via Good Neighbor Fusion (EBMGC-GNF) model which comprehensively extracts credible consistent neighbor information from multiple views by designing a Cross-view Good Neighbors Voting module. Moreover, a novel balanced regularization term based on p-power function is introduced to adjust the balance property of clusters, which helps the model adapt to data with different distributions. To solve the optimization problem of EBMGC-GNF, we transform EBMGC-GNF into an efficient form with graph coarsening method and optimize it based on accelareted coordinate descent algorithm. In experiments, extensive results demonstrate that, in the majority of scenarios, our proposals outperform state-of-the-art methods in terms of both effectiveness and efficiency. Danyang Wu, Jitao Lu, Jin Xu 0014, Xiangmin Xu 0001, Feiping Nie 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Confidence-Aware Sentiment Quantification via Sentiment Perturbation ModelingabstractSentiment Quantification aims to detect the overall sentiment polarity of users from a set of reviews corresponding to a target. Existing methods equally treat and aggregate individual reviews' sentiment to judge the overall sentiment polarity. However, the confidence of each review is not equal in sentiment quantification where sentiment perturbation arising from high- and low-confidence reviews may degrade the accuracy of Sentiment Quantification. Specifically, fake reviews with deceptive sentiments are low confidence, which perturbs the overall sentiment prediction. Whereas, some reviews generated by responsible users are high confidence. They contain authoritative suggestions so they should be emphasized in Sentiment Quantification. In this paper, we design and build COSE, a confidence-aware sentiment quantification framework, which can measure the confidence of individual reviews to eliminate sentiment perturbation and facilitate sentiment quantification. We design a Review Graph that achieves review confidence modeling in an unsupervised manner and obtains review confidence representations. Moreover, we develop a dynamic fusion attention mechanism, which produces sentiment “de-perturbation” vectors to eliminate the sentiment perturbation based on the confidence representations. Extensive experiments on large-scale review datasets validate the significant superiority of COSE over the state-of-the-art. Xiangyun Tang, Dongliang Liao, Meng Shen 0001, Liehuang Zhu, Shen Huang, Gongfu Li, Hong Man, Jin Xu 0014 |
IEEE Trans. Affect. Comput. | 8 |
| 2023 | Keyword-Based Diverse Image Retrieval With Variational Multiple Instance GraphabstractThe task of cross-modal image retrieval has recently attracted considerable research attention. In real-world scenarios, keyword-based queries issued by users are usually short and have broad semantics. Therefore, semantic diversity is as important as retrieval accuracy in such user-oriented services, which improves user experience. However, most typical cross-modal image retrieval methods based on single point query embedding inevitably result in low semantic diversity, while existing diverse retrieval approaches frequently lead to low accuracy due to a lack of cross-modal understanding. To address this challenge, we introduce an end-to-end solution termed variational multiple instance graph (VMIG), in which a continuous semantic space is learned to capture diverse query semantics, and the retrieval task is formulated as a multiple instance learning problems to connect diverse features across modalities. Specifically, a query-guided variational autoencoder is employed to model the continuous semantic space instead of learning a single-point embedding. Afterward, multiple instances of the image and query are obtained by sampling in the continuous semantic space and applying multihead attention, respectively. Thereafter, an instance graph is constructed to remove noisy instances and align cross-modal semantics. Finally, heterogeneous modalities are robustly fused under multiple losses. Extensive experiments on two real-world datasets have well verified the effectiveness of our proposed solution in both retrieval accuracy and semantic diversity. Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Da Cao, Hong Man |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Procedural Text Understanding via Scene-Wise EvolutionabstractProcedural text understanding requires machines to reason about entity states within the dynamical narratives. Current procedural text understanding approaches are commonly entity-wise, which separately track each entity and independently predict different states of each entity. Such an entity-wise paradigm does not consider the interaction between entities and their states. In this paper, we propose a new scene-wise paradigm for procedural text understanding, which jointly tracks states of all entities in a scene-by-scene manner. Based on this paradigm, we propose Scene Graph Reasoner (SGR), which introduces a series of dynamically evolving scene graphs to jointly formulate the evolution of entities, states and their associations throughout the narrative. In this way, the deep interactions between all entities and states can be jointly captured and simultaneously derived from scene graphs. Experiments show that SGR not only achieves the new state-of-the-art performance but also significantly accelerates the speed of reasoning. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
AAAI | 8 |
| 2022 | Hybrid Contrastive Quantization for Efficient Cross-View Video RetrievalabstractWith the recent boom of video-based social platforms (e.g., YouTube and TikTok), video retrieval using sentence queries has become an important demand and attracts increasing research attention. Despite the decent performance, existing text-video retrieval models in vision and language communities are impractical for large-scale Web search because they adopt brute-force search based on high-dimensional embeddings. To improve efficiency, Web search engines widely apply vector compression libraries (e.g., FAISS [26]) to post-process the learned embeddings. Unfortunately, separate compression from feature encoding degrades the robustness of representations and incurs performance decay. To pursue a better balance between performance and efficiency, we propose the first quantized representation learning method for cross-view video retrieval, namely Hybrid Contrastive Quantization (HCQ). Specifically, HCQ learns both coarse-grained and fine-grained quantizations with transformers, which provide complementary understandings for texts and videos and preserve comprehensive semantic information. By performing Asymmetric-Quantized Contrastive Learning (AQ-CL) across views, HCQ aligns texts and videos at coarse-grained and multiple fine-grained levels. This hybrid-grained learning strategy serves as strong supervision on the cross-view video quantization model, where contrastive learning at different levels can be mutually promoted. Extensive experiments on three Web video benchmark datasets demonstrate that HCQ achieves competitive performance with state-of-the-art non-compressed retrieval methods while showing high efficiency in storage and computation. Code and configurations are available at https://github.com/gimpong/WWW22-HCQ. Jinpeng Wang 0002, Bin Chen 0011, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shutao Xia, Jin Xu 0014 |
WWW | 7 |
| 2022 | Two-phase Multi-document Event Summarization on Core Event GraphsabstractSuccinct event description based on multiple documents is critical to news systems as well as search engines. Different from existing summarization or event tasks, Multi-document Event Summarization (MES) aims at the query-level event sequence generation, which has extra constraints on event expression and conciseness. Identifying and summarizing the key event from a set of related articles is a challenging task that has not been sufficiently studied, mainly because online articles exhibit characteristics of redundancy and sparsity, and a perfect event summarization needs high level information fusion among diverse sentences and articles. To address these challenges, we propose a two-phase framework for the MES task, that first performs event semantic graph construction and dominant event detection via graph-sequence matching, then summarizes the extracted key event by an event-aware pointer generator. For experiments in the new task, we construct two large-scale real-world datasets for training and assessment. Extensive evaluations show that the proposed framework significantly outperforms the related baseline methods, with the most dominant event of the articles effectively identified and correctly summarized. Zengjian Chen, Jin Xu 0014, Meng Liao, Tong Xue, Kun He 0001 |
J. Artif. Intell. Res. | 2 |
| 2021 | Fully Exploiting Cascade Graphs for Real-time Forwarding PredictionabstractReal-time forwarding prediction for predicting online contents' popularity is beneficial to various social applications for enhancing interactive social behaviors. Cascade graphs, formed by online contents' propagation, play a vital role in real-time forwarding prediction. Existing cascade graph modeling methods are inadequate to embed cascade graphs that have hub structures and deep cascade paths, or they fail to handle the short-term outbreak of forwarding amount. To this end, we propose a novel real-time forwarding prediction method that includes an effective approach for cascade graph embedding and a short-term variation sensitive method for time-series modeling, making the best of cascade graph features. Using two real world datasets, we demonstrate the significant superiority of the proposed method compared with the state-of-the-art. Our experiments also reveal interesting implications hidden in the performance differences between cascade graph embedding and time-series modeling. Xiangyun Tang, Dongliang Liao, Jin Xu 0014, Liehuang Zhu, Meng Shen 0001 |
AAAI | 4 |
| 2021 | Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event ExtractionabstractYaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, Shaoyi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yaojie Lu 0001, Jin Xu 0014, Xianpei Han, Jialong Tang, Annan Li, Le Sun 0001, Meng Liao, Shaoyi Chen |
ACL/IJCNLP (1) | 3 |
| 2021 | Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge BasesabstractBoxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, Jin Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Boxi Cao, Xianpei Han, Le Sun 0001, Lingyong Yan, Meng Liao, Tong Xue, Jin Xu 0014 |
ACL/IJCNLP (1) | 8 |
| 2021 | From Discourse to Narrative: Knowledge Projection for Event Relation ExtractionabstractJialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu, Xianpei Han, Le Sun, Weijian Xie, Jin Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
ACL/IJCNLP (1) | 8 |
| 2021 | Adaptive Feature Weight Learning For Robust Clustering Problem with Sparse ConstraintabstractClustering task has been greatly developed in recent years like partition-based and graph-based methods. However, in terms of improving robustness, most existing algorithms only focus on noise and outliers between data, while ignoring the noise in feature space. To deal with this situation, we propose a novel weight learning mechanism to adaptively reweight each feature in the data. Combining with the clustering task, we further propose a robust fuzzy K-Means model based on the auto-weighted feature learning, which can effectively reduce the proportion of noisy features. Besides, a regularization term is introduced into our model to make the sample-to-clusters memberships of each sample have suitable sparsity. Specifically, we design an effective strategy to determine the value of the regularization parameter. The experimental results on both synthetic and real-world datasets demonstrate that our model has better performance than other classical algorithms. Feiping Nie 0001, Wei Chang 0002, Xuelong Li 0001, Jin Xu 0014, Gongfu Li |
ICASSP | 4 |
| 2021 | GSPL: A Succinct Kernel Model for Group-Sparse Projections Learning of Multiview DataabstractThis paper explores a succinct kernel model for Group-Sparse Projections Learning (GSPL), to handle multiview feature selection task completely. Compared to previous works, our model has the following useful properties: 1) Strictness: GSPL innovatively learns group-sparse projections strictly on multiview data via ‘2;0-norm constraint, which is different with previous works that encourage group-sparse projections softly. 2) Adaptivity: In GSPL model, when the total number of selected features is given, the numbers of selected features of different views can be determined adaptively, which avoids artificial settings. Besides, GSPL can capture the differences among multiple views adaptively, which handles the inconsistent problem among different views. 3) Succinctness: Except for the intrinsic parameters of projection-based feature selection task, GSPL does not bring extra parameters, which guarantees the applicability in practice. To solve the optimization problem involved in GSPL, a novel iterative algorithm is proposed with rigorously theoretical guarantees. Experimental results demonstrate the superb performance of GSPL on synthetic and real datasets. Danyang Wu, Jin Xu 0014, Xia Dong, Meng Liao, Rong Wang 0001, Feiping Nie 0001, Xuelong Li 0001 |
IJCAI | 2 |
| 2020 | Transfer Value Iteration NetworksabstractValue iteration networks (VINs) have been demonstrated to have a good generalization ability for reinforcement learning tasks across similar domains. However, based on our experiments, a policy learned by VINs still fail to generalize well on the domain whose action space and feature space are not identical to those in the domain where it is trained. In this paper, we propose a transfer learning approach on top of VINs, termed Transfer VINs (TVINs), such that a learned policy from a source domain can be generalized to a target domain with only limited training data, even if the source domain and the target domain have domain-specific actions and features. We empirically verify that our proposed TVINs outperform VINs when the source and the target domains have similar but not identical action and feature spaces. Furthermore, we show that the performance improvement is consistent across different environments, maze sizes, dataset sizes as well as different values of hyperparameters such as number of iteration and kernel size. Hankui Zhuo, Jin Xu 0014, Bin Zhong, Sinno Jialin Pan |
AAAI | 3 |
| 2020 | Active Learning with Query Generation for Cost-Effective Text ClassificationabstractLabeling a text document is usually time consuming because it requires the annotator to read the whole document and check its relevance with each possible class label. It thus becomes rather expensive to train an effective model for text classification when it involves a large dataset of long documents. In this paper, we propose an active learning approach for text classification with lower annotation cost. Instead of scanning all the examples in the unlabeled data pool to select the best one for query, the proposed method automatically generates the most informative examples based on the classification model, and thus can be applied to tasks with large scale or even infinite unlabeled data. Furthermore, we propose to approximate the generated example with a few summary words by sparse reconstruction, which allows the annotators to easily assign the class label by reading a few words rather than the long document. Experiments on different datasets demonstrate that the proposed approach can effectively improve the classification performance while significantly reduce the annotation cost. Sheng-Jun Huang, Shaoyi Chen, Meng Liao, Jin Xu 0014 |
AAAI | 5 |