EDBT 2026 Demo / reviewers in the wild / expert
Jun Wang 0023
dblp:125/8189-23
· DBLP profile ↗
34ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-8932-6661ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-grained multi-label feature selection: Jointly towards semantic-aware and instance-specific features
Hengpeng Xu, Zhenglu Yang, Jun Wang 0023 |
Knowl. Based Syst. | 4 |
| 2026 | Adaptive knowledge selection in dialogue systems: Accommodating diverse knowledge types, requirements, and generation models
Zhongtian Bao, Hongru Liang, Jun Wang 0023, Zhenglu Yang, Zhe Sun 0009, Andrzej Cichocki |
Neural Networks | 5 |
| 2025 | Enhancing Cross-Lingual Dialogue Summarization Through Interpretable Chain-of-Thought
Zhongtian Bao, Jun Wang 0023, Adam Jatowt, Zhenglu Yang |
DASFAA (2) | 3 |
| 2024 | Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question AnsweringabstractMultiple-choice visual question answering (MC VQA) requires an answer picked from a list of distractors, based on a question and an image. This research has attracted wide interest from the fields of visual question answering, visual question generation, and visual distractor generation. However, these fields still stay in their own territories, and how to jointly generate meaningful questions, correct answers, and challenging distractors remains unexplored. In this paper, we introduce a novel task, Visual Question-Answer-Distractors Generation (VQADG), which can bridge this research gap as well as take as a cornerstone to promote existing VQA models. Specific to the VQADG task, we present a novel framework consisting of a vision-and-language model to encode the given image and generate QADs jointly, and contrastive learning to ensure the consistency of the generated question, answer, and distractors. Empirical evaluations on the benchmark dataset validate the performance of our model in the VQADG task. Wenjian Ding, Jun Wang 0023, Adam Jatowt, Zhenglu Yang |
LREC/COLING | 3 |
| 2024 | Exploring Union and Intersection of Visual Regions for Generating Questions, Answers, and DistractorsabstractMultiple-choice visual question answering (VQA) is to automatically choose a correct answer from a set of choices after reading an image.Existing efforts have been devoted to a separate generation of an image-related question, a correct answer, or challenge distractors.By contrast, we turn to a holistic generation and optimization of questions, answers, and distractors (QADs) in this study.This integrated generation strategy eliminates the need for human curation and guarantees information consistency.Furthermore, we first propose to put the spotlight on different image regions to diversify QADs.Accordingly, a novel framework ReBo is formulated in this paper.ReBo cyclically generates each QAD based on a recurrent multimodal encoder, and each generation is focusing on a different area of the image compared to those already concerned by the previously generated QADs.In addition to traditional VQA comparisons with state-of-the-art approaches, we also validate the capability of ReBo in generating augmented data to benefit VQA models. Wenjian Ding, Jun Wang 0023, Adam Jatowt, Zhenglu Yang |
EMNLP | 3 |
| 2024 | Unsupervised feature selection by learning exponential weights
Jun Wang 0023, Zhichen Gu, Jinmao Wei 0001, Jian Liu 0040 |
Pattern Recognit. | 2 |
| 2023 | Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded DialogueabstractAccurate knowledge selection is critical in knowledge-grounded dialogue systems.Towards a closer look at it, we offer a novel perspective to organize existing literature, i.e., knowledge selection coupled with, after, and before generation.We focus on the third underexplored category of study, which can not only select knowledge accurately in advance, but has the advantage to reduce the learning, adjustment, and interpretation burden of subsequent response generation models, especially LLMs.We propose GATE, a generator-agnostic knowledge selection method, to prepare knowledge for subsequent response generation models by selecting context-related knowledge among different knowledge structures and variable knowledge requirements.Experimental results demonstrate the superiority of GATE, and indicate that knowledge selection before generation is a lightweight yet effective way to facilitate LLMs (e.g., ChatGPT) to generate more informative responses. Hongru Liang, Jun Wang 0023, Zhenglu Yang |
EMNLP | 4 |
| 2023 | Multi-path Based Self-adaptive Cross-lingual Summarization
Zhongtian Bao, Jun Wang 0023, Zhenglu Yang |
KSEM (3) | 2 |
| 2022 | Multi-Party Empathetic Dialogue Generation: A New Task for Dialog SystemsabstractEmpathetic dialogue assembles emotion understanding, feeling projection, and appropriate response generation.Existing work for empathetic dialogue generation concentrates on the two-party conversation scenario.Multiparty dialogues, however, are pervasive in reality.Furthermore, emotion and sensibility are typically confused; a refined empathy analysis is needed for comprehending fragile and nuanced human feelings.We address these issues by proposing a novel task called Multi-Party Empathetic Dialogue Generation in this study.Additionally, a Static-Dynamic model for Multi-Party Empathetic Dialogue Generation, SDMPED, is introduced as a baseline by exploring the static sensibility and dynamic emotion for the multi-party empathetic dialogue learning, the aspects that help SDMPED achieve the state-of-the-art performance. Lingyu Zhu 0007, Zhengkun Zhang, Jun Wang 0023, Haiying Wu, Zhenglu Yang |
ACL (1) | 3 |
| 2022 | Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine ComprehensionabstractHuibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Huibin Zhang, Zhengkun Zhang, Jun Wang 0023, Zhenglu Yang |
ACL (1) | 4 |
| 2022 | Deep convolutional recurrent model for region recommendation with spatial and temporal contexts
Hengpeng Xu, Wenjian Ding, Wei Shen 0004, Jun Wang 0023, Zhenglu Yang |
Ad Hoc Networks | 4 |
| 2022 | Multi-variable AUC for sifting complementary features and its biomedical applicationabstractAlthough sifting functional genes has been discussed for years, traditional selection methods tend to be ineffective in capturing potential specific genes. First, typical methods focus on finding features (genes) relevant to class while irrelevant to each other. However, the features that can offer rich discriminative information are more likely to be the complementary ones. Next, almost all existing methods assess feature relations in pairs, yielding an inaccurate local estimation and lacking a global exploration. In this paper, we introduce multi-variable Area Under the receiver operating characteristic Curve (AUC) to globally evaluate the complementarity among features by employing Area Above the receiver operating characteristic Curve (AAC). Due to AAC, the class-relevant information newly provided by a candidate feature and that preserved by the selected features can be achieved beyond pairwise computation. Furthermore, we propose an AAC-based feature selection algorithm, named Multi-variable AUC-based Combined Features Complementarity, to screen discriminative complementary feature combinations. Extensive experiments on public datasets demonstrate the effectiveness of the proposed approach. Besides, we provide a gene set about prostate cancer and discuss its potential biological significance from the machine learning aspect and based on the existing biomedical findings of some individual genes. Keyu Du, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040 |
Briefings Bioinform. | 3 |
| 2021 | News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang |
AAAI | 2 |
| 2021 | LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract)abstractMultimodal summarization aims to refine salient information from multiple modalities, among which texts and images are two mostly discussed ones. In recent years, many fantastic works have emerged in this field by modeling image-text interactions; however, they neglect the fact that most of multimodal documents have been elaborately organized by their writers. This means that a critical organized factor has long been short of enough attention, that is, image locations, which may carry illuminating information and imply the key contents of a document. To address this issue, we propose a location-aware approach for multimodal summarization (LAMS) based on Transformer. We investigate image locations for multimodal summarization via a stack of multimodal fusion block, which can formulate the high-order interactions among images and texts. An extensive experimental study on an extended multimodal dataset validates the superior summarization performance of the proposed model. Zhengkun Zhang, Jun Wang 0023, Zhe Sun 0009, Zhenglu Yang |
AAAI | 2 |
| 2021 | Generalized Relation Learning with Semantic Correlation Awareness for Link PredictionabstractDeveloping link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many unseen relations that occur in practical situations. These two problems limit the training effectiveness and practical applications of the existing link prediction models. We advocate a holistic understanding of KGs and we propose in this work a unified Generalized Relation Learning framework GRL to address the above two problems, which can be plugged into existing link prediction models. GRL conducts a generalized relation learning, which is aware of semantic correlations between relations that serve as a bridge to connect semantically similar relations. After training with GRL, the closeness of semantically similar relations in vector space and the discrimination of dissimilar relations are improved. We perform comprehensive experiments on six benchmarks to demonstrate the superior capability of GRL in the link prediction task. In particular, GRL is found to enhance the existing link prediction models making them insensitive to unbalanced relation distributions and capable of learning unseen relations. Jun Wang 0023, Hongru Liang, Wenqiang Lei, Zhe Sun 0009, Adam Jatowt, Zhenglu Yang |
AAAI | 3 |
| 2021 | MM-AVS: A Full-Scale Dataset for Multi-modal SummarizationabstractMultimodal summarization becomes increasingly significant as it is the basis for question answering, Web search, and many other downstream tasks.However, its learning materials have been lacking a holistic organization by integrating resources from various modalities, thereby lagging behind the research progress of this field.In this study, we present a full-scale multimodal dataset comprehensively gathering documents, summaries, images, captions, videos, audios, transcripts, and titles in English from CNN and Daily Mail.To our best knowledge, this is the first collection that spans all modalities and nearly comprises all types of materials available in this community.In addition, we devise a baseline model based on the novel dataset, which employs a newly proposed Jump-Attention mechanism based on transcripts.The experimental results validate the important assistance role of the external information for multimodal summarization. Xiyan Fu, Jun Wang 0023, Zhenglu Yang |
NAACL-HLT | 2 |
| 2021 | Recommending irregular regions using graph attentive networks
Hengpeng Xu, Jun Wang 0023, Jinmao Wei 0001 |
Ad Hoc Networks | 2 |
| 2021 | Unsupervised Cross-View Feature Selection on incomplete data
Yuanyuan Xu 0002, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040, Lina Yao 0001, Wenjie Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Document Summarization with VHTM: Variational Hierarchical Topic-Aware MechanismabstractAutomatic text summarization focuses on distilling summary information from texts. This research field has been considerably explored over the past decades because of its significant role in many natural language processing tasks; however, two challenging issues block its further development: (1) how to yield a summarization model embedding topic inference rather than extending with a pre-trained one and (2) how to merge the latent topics into diverse granularity levels. In this study, we propose a variational hierarchical model to holistically address both issues, dubbed VHTM. Different from the previous work assisted by a pre-trained single-grained topic model, VHTM is the first attempt to jointly accomplish summarization with topic inference via variational encoder-decoder and merge topics into multi-grained levels through topic embedding and attention. Comprehensive experiments validate the superior performance of VHTM compared with the baselines, accompanying with semantically consistent topics. Xiyan Fu, Jun Wang 0023, Jinghan Zhang 0004, Jinmao Wei 0001, Zhenglu Yang |
AAAI | 2 |
| 2020 | To Avoid the Pitfall of Missing Labels in Feature Selection: A Generative Model Gives the AnswerabstractIn multi-label learning, instances have a large number of noisy and irrelevant features, and each instance is associated with a set of class labels wherein label information is generally incomplete. These missing labels possess two sides like a coin; people cannot predict whether their provided information for feature selection is favorable (relevant) or not (irrelevant) during tossing. Existing approaches either superficially consider the missing labels as negative or indiscreetly impute them with some predicted values, which may either overestimate unobserved labels or introduce new noises in selecting discriminative features. To avoid the pitfall of missing labels, a novel unified framework of selecting discriminative features and modeling incomplete label matrix is proposed from a generative point of view in this paper. Concretely, we relax Smoothness Assumption to infer the label observability, which can reveal the positions of unobserved labels, and employ the spike-and-slab prior to perform feature selection by excluding unobserved labels. Using a data-augmentation strategy leads to full local conjugacy in our model, facilitating simple and efficient Expectation Maximization (EM) algorithm for inference. Quantitative and qualitative experimental results demonstrate the superiority of the proposed approach under various evaluation metrics. Yuanyuan Xu 0002, Jun Wang 0023, Jinmao Wei 0001 |
AAAI | 2 |
| 2019 | Probabilistic Margin-Aware Multi-Label Feature Selection by Preserving Spatial ConsistencyabstractMulti-label feature selection focuses on constructing a reduced feature space for discriminating multi-label instances. In consideration of the complex structures of label and feature spaces, a critical issue that explicitly determines selection performance is how to induce consistent information from both spaces to steer feature selection. Existing approaches tackle this issue in various spatial-aware views, without sufficient consideration of the negative effects of irrelevant features and imbalanced neighbors on inferring space structure. Inspired by the superiority of margin theory in assessing reliable space structure, we approach multi-label feature selection in the learning framework of preserving label-feature space consistency through probabilistic margin in this paper. In contrast to existing approaches, our model assesses the weighted margin based on the probabilistic nearest neighbors, and preserves consistent margin information in label and feature spaces. In this manner, label-feature space consistency is elegantly achieved, which conduces to effectively capturing discriminative features suitable for multi-label learning tasks and eliminating noisy features. Experimental results on multi-label data sets demonstrate the encouraging performance of the proposed model. Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001, Jianhua Ruan |
IJCNN | 3 |
| 2019 | A general framework for learning prosodic-enhanced representation of rap lyrics
Hongru Liang, Haozheng Wang, Qian Li 0016, Jun Wang 0023, Guandong Xu, Jinmao Wei 0001, Zhenglu Yang |
World Wide Web | 4 |
| 2018 | Semi-Supervised Multi-Label Feature Selection by Preserving Feature-Label Space ConsistencyabstractSemi-supervised learning and multi-label learning pose different challenges for feature selection, which is one of the core techniques for dimension reduction, and the exploration of reducing feature space for multi-label learning with incomplete label information is far from satisfactory. Existing feature selection approaches devote attention to either of two issues, namely, alleviating negative effects of imperfectly predicted labels and quantitatively evaluating label correlations, exclusively for semi-supervised or multi-label scenarios. A unified framework to extract label correlation information with incomplete prior knowledge and embed this information in feature selection however, is rarely touched. In this paper, we propose a space consistency-based feature selection model to address this issue. Specifically, correlation information in feature space is learned based on the probabilistic neighborhood similarities, and correlation information in label space is optimized by preserving feature-label space consistency. This mechanism contributes to appropriately extracting label information in semi-supervised multi-label learning scenario and effectively employing this information to select discriminative features. An extensive experimental evaluation on real-world data shows the superiority of the proposed approach under various evaluation metrics. Yuanyuan Xu 0002, Jun Wang 0023, Shuai An 0003, Jinmao Wei 0001, Jianhua Ruan |
CIKM | 2 |
| 2018 | JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual FeaturesabstractLearning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we propose to learn social media content by fusing jointly textual, acoustic, and visual information (JTAV). Effective strategies are proposed to extract fine-grained features of each modality, that is, attBiGRU and DCRNN. We also introduce cross-modal fusion and attentive pooling techniques to integrate multi-modal information comprehensively. Extensive experimental evaluation conducted on real-world datasets demonstrate our proposed model outperforms the state-of-the-art approaches by a large margin. Hongru Liang, Haozheng Wang, Jun Wang 0023, Shaodi You, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang |
COLING | 3 |
| 2018 | Probabilistic Topic and Role Model for Information Diffusion in Social Network
Hengpeng Xu, Jinmao Wei 0001, Zhenglu Yang, Jianhua Ruan, Jun Wang 0023 |
PAKDD (2) | 5 |
| 2018 | ANNC: AUC-Based Feature Selection by Maximizing Nearest Neighbor Complementarity
Xuemeng Jiang, Jun Wang 0023, Jinmao Wei 0001, Jianhua Ruan |
PRICAI (1) | 2 |
| 2018 | HAVAE: Learning Prosodic-Enhanced Representations of Rap Lyrics
Hongru Liang, Qian Li 0016, Haozheng Wang, Jun Wang 0023, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang |
PRICAI (1) | 5 |
| 2018 | Local-Nearest-Neighbors-Based Feature Weighting for Gene SelectionabstractSelecting functional genes is essential for analyzing microarray data. Among many available feature (gene) selection approaches, the ones on the basis of the large margin nearest neighbor receive more attention due to their low computational costs and high accuracies in analyzing the high-dimensional data. Yet, there still exist some problems that hamper the existing approaches in sifting real target genes, including selecting erroneous nearest neighbors, high sensitivity to irrelevant genes, and inappropriate evaluation criteria. Previous pioneer works have partly addressed some of the problems, but none of them are capable of solving these problems simultaneously. In this paper, we propose a new local-nearest-neighbors-based feature weighting approach to alleviate the above problems. The proposed approach is based on the trick of locally minimizing the within-class distances and maximizing the between-class distances with the nearest neighbors rule. We further define a feature weight vector, and construct it by minimizing the cost function with a regularization term. The proposed approach can be applied naturally to the multi-class problems and does not require extra modification. Experimental results on the UCI and the open microarray data sets validate the effectiveness and efficiency of the new approach. Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Unsupervised Feature Selection with Joint Clustering AnalysisabstractUnsupervised feature selection has raised considerable interests in the past decade, due to its remarkable performance in reducing dimensionality without any prior class information. Preserving reliable locality information and achieving excellent cluster separation are two critical issues for unsupervised feature selection. However, existing methods cannot tackle two issues simultaneously. To address the problems, we propose a novel unsupervised approach that integrates sparse feature selection and robust joint clustering analysis. The joint clustering analysis seamlessly unifies the spectral clustering and the orthogonal basis clustering. Specifically, a probabilistic neighborhood graph is utilized to preserve reliable locality information in the spectral clustering, and an orthogonal basis matrix is incorporated to achieve excellent cluster separation in the orthogonal basis clustering. A compact and effective iterative algorithm is designed to optimize the proposed selection framework. Extensive experiments on both synthetic data and real-world data validate the effectiveness of our approach under various evaluation indices. Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang |
CIKM | 2 |
| 2017 | AVC: Selecting discriminative features on basis of AUC by maximizing variable complementarityabstractBACKGROUND: The Receiver Operator Characteristic (ROC) curve is well-known in evaluating classification performance in biomedical field. Owing to its superiority in dealing with imbalanced and cost-sensitive data, the ROC curve has been exploited as a popular metric to evaluate and find out disease-related genes (features). The existing ROC-based feature selection approaches are simple and effective in evaluating individual features. However, these approaches may fail to find real target feature subset due to their lack of effective means to reduce the redundancy between features, which is essential in machine learning. RESULTS: In this paper, we propose to assess feature complementarity by a trick of measuring the distances between the misclassified instances and their nearest misses on the dimensions of pairwise features. If a misclassified instance and its nearest miss on one feature dimension are far apart on another feature dimension, the two features are regarded as complementary to each other. Subsequently, we propose a novel filter feature selection approach on the basis of the ROC analysis. The new approach employs an efficient heuristic search strategy to select optimal features with highest complementarities. The experimental results on a broad range of microarray data sets validate that the classifiers built on the feature subset selected by our approach can get the minimal balanced error rate with a small amount of significant features. CONCLUSIONS: Compared with other ROC-based feature selection approaches, our new approach can select fewer features and effectively improve the classification performance. Jun Wang 0023, Jinmao Wei 0001 |
BMC Bioinform. | 2 |
| 2017 | Feature Selection by Maximizing Independent Classification InformationabstractFeature selection approaches based on mutual information can be roughly categorized into two groups. The first group minimizes the redundancy of features between each other. The second group maximizes the new classification information of features providing for the selected subset. A critical issue is that large new information does not signify little redundancy, and vice versa. Features with large new information but with high redundancy may be selected by the second group, and features with low redundancy but with little relevance with classes may be highly scored by the first group. Existing approaches fail to balance the importance of both terms. As such, a new information term denoted as Independent Classification Information is proposed in this paper. It assembles the newly provided information and the preserved information negatively correlated with the redundant information. Redundancy and new information are properly unified and equally treated in the new term. This strategy helps find the predictive features providing large new information and little redundancy. Moreover, independent classification information is proved as a loose upper bound of the total classification information of feature subset. Its maximization is conducive to achieve a high global discriminative performance. Comprehensive experiments demonstrate the effectiveness of the new approach. Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang, Shu-Qin Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Feature Selection via Vectorizing Feature's Discriminative Information
Jun Wang 0023, Hengpeng Xu, Jinmao Wei 0001 |
APWeb (1) | 1 |
| 2016 | Supervised Feature Selection by Preserving Class CorrelationabstractFeature selection is an effective technique for dimension reduction, which assesses the importance of features and constructs an optimal feature subspace suitable for recognition task. Two recognition scenarios, i.e., single-label learning and multi-label learning, pose different challenges for feature selection. For the single-label task, how to accurately measure and reduce feature redundancy is crucial. For the multi-label task, how to effectively exploit class correlation information during selection is critical. However, both issues cannot be simultaneously resolved by any existing selection methods. In this paper, we propose effective supervised feature selection techniques to address the problems. The original class correlation information in the reduced feature space is preserved, and meanwhile the feature redundancy for classification is alleviated. To the best of our knowledge, this study is the first attempt to accomplish both recognition tasks in a unified framework. Comprehensive experimental evaluations on artificial, single-label, and multi-label data sets demonstrate the effectiveness of the new approach. Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang |
CIKM | 1 |
| 2015 | Feature Selection Method Based on Feature's Classification Bias and Performance
Jun Wang 0023, Jinmao Wei 0001 |
ICA3PP (2) | 1 |