VLDB 2026 Research / reviewers in the wild / expert
Huishan Ji
dblp:321/6907
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2024
0009-0002-7815-9508ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
1.4 | 2 | 2024 | Towards Flexible Evaluation for Generative Visual Question Answering · ACM Multimedia 2024 Combo of Thinking and Observing for Outside-Knowledge VQA · ACL (1) 2023 |
Computer vision › Vision and language › visual question answering
visual question answering evaluation |
0.8 | 1 | 2024 | Towards Flexible Evaluation for Generative Visual Question Answering · ACM Multimedia 2024 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation |
0.2 | 1 | 2024 | Towards Flexible Evaluation for Generative Visual Question Answering · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
semantic textual similarity · 0.8large language model evaluator · 0.8multimodal encoder · 0.7cross-modal alignment · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Flexible Evaluation for Generative Visual Question AnsweringabstractThroughout rapid development of multimodal large language models, a crucial ingredient is a fair and accurate evaluation of their multimodal comprehension abilities. Although Visual Question Answering (VQA) could serve as a developed test field, limitations of VQA evaluation, like the inflexible pattern of Exact Match, have hindered MLLMs from demonstrating their real capability and discourage rich responses. Therefore, this paper proposes the use of semantics-based evaluators for assessing unconstrained open-ended responses on VQA datasets. As characteristics of VQA have made such evaluation significantly different from the traditional Semantic Textual Similarity (STS) task, to systematically analyze the behaviour and compare the performance of various evaluators including LLM-based ones, we propose three key properties, i.e., Alignment, Consistency and Generalization, and a corresponding dataset Assessing VQA Evaluators (AVE) to facilitate analysis. In addition, this paper proposes a Semantically Flexible VQA Evaluator (SFVE) with meticulous design based on the unique features of VQA evaluation.Experimental results verify the feasibility of model-based VQA evaluation and effectiveness of the proposed evaluator that surpasses existing semantic evaluators by a large margin. The proposed training scheme generalizes to both the BERT-like encoders and decoder-only LLM. Relaed codes and data available at https://github.com/jihuishan/flexible_evaluation_for_vqa_mm24. Huishan Ji, Qingyi Si, Zheng Lin 0001, Weiping Wang 0005 |
ACM Multimedia | 1 |
| 2023 | Combo of Thinking and Observing for Outside-Knowledge VQAabstractOutside-knowledge visual question answering is a challenging task that requires both the acquisition and the use of open-ended real-world knowledge.Some existing solutions draw external knowledge into the cross-modality space which overlooks the much vaster textual knowledge in natural-language space, while others transform the image into a text that further fuses with the textual knowledge into the natural-language space and completely abandons the use of visual features.In this paper, we are inspired to constrain the cross-modality space into the same space of natural-language space which makes the visual features preserved directly, and the model still benefits from the vast knowledge in natural-language space.To this end, we propose a novel framework consisting of a multimodal encoder, a textual encoder and an answer decoder.Such structure allows us to introduce more types of knowledge including explicit and implicit multimodal and textual knowledge.Extensive experiments validate the superiority of the proposed method which outperforms the state-ofthe-art by 6.17% accuracy.We also conduct comprehensive ablations of each component, and systematically study the roles of varying types of knowledge.Codes and knowledge data can be found at https://github.com/ PhoebusSi/Thinking-while-Observing. Qingyi Si, Yuchen Mo, Zheng Lin 0001, Huishan Ji, Weiping Wang 0005 |
ACL (1) | 4 |
| 2022 | Target Really Matters: Target-aware Contrastive Learning and Consistency Regularization for Few-shot Stance DetectionabstractStance detection aims to identify the attitude from an opinion towards a certain target. Despite the significant progress on this task, it is extremely time-consuming and budget-unfriendly to collect sufficient high-quality labeled data for every new target under fully-supervised learning, whereas unlabeled data can be collected easier. Therefore, this paper is devoted to few-shot stance detection and investigating how to achieve satisfactory results in semi-supervised settings. As a target-oriented task, the core idea of semi-supervised few-shot stance detection is to make better use of target-relevant information from labeled and unlabeled data. Therefore, we develop a novel target-aware semi-supervised framework. Specifically, we propose a target-aware contrastive learning objective to learn more distinguishable representations for different targets. Such an objective can be easily applied with or without unlabeled data. Furthermore, to thoroughly exploit the unlabeled data and facilitate the model to learn target-relevant stance features in the opinion content, we explore a simple but effective target-aware consistency regularization combined with a self-training strategy. The experimental results demonstrate that our approach can achieve state-of-the-art performance on multiple benchmark datasets in the few-shot setting. Rui Liu 0032, Zheng Lin 0001, Huishan Ji, Peng Fu 0008, Weiping Wang 0005 |
COLING | 3 |
| 2022 | Cross-Target Stance Detection Via Refined Meta-LearningabstractCross-target stance detection (CTSD) aims to identify the stance of the text towards a target, where stance annotations are available for (though related but) different targets. Recently, models based on external semantic and emotion knowledge have been proposed for CTSD, achieving promising performance. However, such solutions rely on much external resources and harness only one source target, which is a waste of other available targets. To address the problem above, we propose a many-to-one CTSD model based on meta-learning. To make the most of meta-learning, we further refine it with a balanced and easy-to-hard learning pattern. Specifically, for multiple-target training, we feed the model according to the similarity among targets, and utilize two kinds of re-balanced strategies to deal with the imbalance in data. We conduct experiments on SemEval 2016 task 6, and results demonstrate that our method is effective and establishes a new state-of-the-art macro-f1 score for CTSD. Huishan Ji, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005 |
ICASSP | 1 |