VLDB 2026 Research / reviewers in the wild / expert
Zhihao Fan
dblp:220/0988
· DBLP profile ↗
35ranked-venue papers
13as first author
28since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 11 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and InferenceabstractSiyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li, Zhihao Fan, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Siyuan Wang 0025, Dianyi Wang, Chengxing Zhou, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 5 |
| 2025 | AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorabstractArtificial intelligence has significantly revolutionized healthcare, particularly through large language models (LLMs) that demonstrate superior performance in static medical question answering benchmarks. However, evaluating the potential of LLMs for real-world clinical applications remains challenging due to the intricate nature of doctor-patient interactions. To address this, we introduce AI Hospital, a multi-agent framework emulating dynamic medical interactions between Doctor as player and NPCs including Patient and Examiner. This setup allows for more practical assessments of LLMs in simulated clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and multiple evaluation strategies to quantify the performance of LLM-driven Doctor agents on symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance medical interaction capabilities through iterative discussions. Despite improvements, current LLMs (including GPT-4) still exhibit significant performance gaps in multi-turn interactive scenarios compared to non-interactive scenarios. Our findings highlight the need for further research to bridge these gaps and improve LLMs’ clinical decision-making capabilities. Our data, code, and experimental results are all open-sourced at https://github.com/LibertFan/AI_Hospital. Zhihao Fan, Jialong Tang, Wei Chen 0088, Siyuan Wang 0025, Zhongyu Wei, Fei Huang 0002 |
COLING | 1 |
| 2025 | Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM EvaluationabstractThis paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs). We utilize a multi-agent system to reframe new evolving instances with high confidence that extend existing benchmarks. Towards a more scalable, robust and fine-grained evaluation, we implement six reframing operations to construct evolving instances testing LLMs against diverse queries, shortcut biases and probing their problem-solving sub-abilities. With this framework, we extend datasets across general and specific tasks, through various iterations. Experimental results show a performance decline in most LLMs against their original results under scalable and robust evaluations, offering a more accurate reflection of model capabilities alongside our fine-grained evaluation. Besides, our framework widens performance discrepancies both between different models and within the same model across various tasks, facilitating more informed model selection for specific tasks. We hope this framework contributes the research community for continuously evolving benchmarks alongside LLM development. Siyuan Wang 0025, Zhuohan Long, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei |
COLING | 3 |
| 2025 | Uncover Governing Law of Pathology Propagation Mechanism Through A Mean-Field GameabstractAlzheimer’s disease (AD) is marked by cognitive decline along with the widespread of tau aggregates across the brain cortex. Due to the challenges of imaging pathology spreading flows \textit{in vivo}, however, quantitative analysis on the cortical pathways of tau propagation and its interaction with the cascade of amyloid-beta (A$\beta$) plaques lags behind the experimental insights of underlying pathophysiological mechanisms.
To address this challenge, we present a physics-informed neural network, empowered by mean-field theory, to uncover the biologically meaningful spreading pathways of tau aggregates between two longitudinal snapshots.
Following the notion of `prion-like' mechanism in AD, we first formulate the dynamics of tau propagation as a mean-field game (MFG), where the spread of tau aggregate at each location (aka. agent) depends on the collective behavior of the surrounding agents as well as the potential field formed by amyloid burden. Given the governing equation of propagation dynamics, MFG reaches an equilibrium that allows us to model the evolution of tau aggregates as an optimal transport with the lowest cost in \textit{Wasserstein} space.
By leveraging the variational primal-dual structure in MFG, we propose a \textit{Wasserstein}-1 Lagrangian generative adversarial network (GAN), in which a Lipschitz critic seeks the appropriate transport cost at the population level and a generator parameterizes the flow fields of optimal transport across individuals.
Additionally, we incorporate a symbolic regression module to derive an explicit formulation capturing the A$\beta$-tau crosstalk.
Experimental results on public neuroimaging datasets demonstrate that our explainable deep model not only yields precise and reliable predictions of future tau progression for unseen new subjects but also provides a new window to uncover new understanding of pathology propagation in AD through learning-based approaches. Tingting Dan, Zhihao Fan, Guorong Wu 0001 |
NeurIPS | 2 |
| 2025 | High-Accuracy prediction and efficient adjustment of surface shape distortion in optical elements: Model correction based on uncertainty quantification-driven transfer learning
Zhihao Fan, Xiaokai Mu, Rongxuan Zhao, Kangcheng Yin, Qingchao Sun, Kaike Yang, Wenjing Ma |
Adv. Eng. Informatics | 1 |
| 2024 | DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive LearningabstractVision-and-Language navigation (VLN) requires an agent to navigate in unseen environment by following natural language instruction. For task completion, the agent needs to align and integrate various navigation modalities, including instruction, observation and navigation history. Existing works primarily concentrate on cross-modal attention at the fusion stage to achieve this objective. Nevertheless, modality features generated by disparate uni-encoders reside in their own spaces, leading to a decline in the quality of cross-modal fusion and decision. To address this problem, we propose a Dual-levEL AligNment (DELAN) framework by cross-modal contrastive learning. This framework is designed to align various navigation-related modalities before fusion, thereby enhancing cross-modal interaction and action decision-making. Specifically, we divide the pre-fusion alignment into dual levels: instruction-history level and landmark-observation level according to their semantic correlations. We also reconstruct a dual-level instruction for adaptation to the dual-level alignment. As the training signals for pre-fusion alignment are extremely limited, self-supervised contrastive learning strategies are employed to enforce the matching between different modalities. Our approach seamlessly integrates with the majority of existing models, resulting in improved navigation performance on various VLN benchmarks, including R2R, R4R, RxR and CVDN. Mengfei Du, Binhao Wu, Jiwen Zhang, Zhihao Fan, Ruipu Luo, Xuanjing Huang 0001, Zhongyu Wei |
LREC/COLING | 4 |
| 2024 | From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingabstractThe rapid development of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has exposed vulnerabilities to various adversarial attacks.This paper provides a comprehensive overview of jailbreaking research targeting both LLMs and MLLMs, highlighting recent advancements in evaluation benchmarks, attack techniques and defense strategies.Compared to the more advanced state of unimodal jailbreaking, multimodal domain remains underexplored.We summarize the limitations and potential research directions of multimodal jailbreaking, aiming to inspire future research and further enhance the robustness and security of MLLMs. Zhuohan Long, Zhihao Fan, Zhongyu Wei |
EMNLP | 3 |
| 2024 | ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented BenchmarksabstractRecent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities to perceive visual signals and perform visually grounded reasoning. However, the capabilities of LVLMs have not been comprehensively and quantitatively evaluated. Most existing multi-modal benchmarks require task-oriented input-output formats, posing great challenges to automatically assess the free-form text output of LVLMs. To effectively leverage the annotations available and reduce the manual efforts required for constructing new benchmarks, we propose to re-formulate existing benchmarks into unified LVLM-compatible formats. Through systematic data collection and reformulation, we present ReForm-Eval benchmark, offering substantial data for evaluating various capabilities of LVLMs. Through extensive experiments and analysis in ReForm-Eval, we demonstrate the comprehensiveness and reliability of ReForm-Eval in assessing various LVLMs. Our benchmark and evaluation framework is now available at https://github.com/FudanDISC/ReForm-Eval Mengfei Du, Qingwen Liu 0002, Binhao Wu, Jiwen Zhang, Chengxing Zhou, Zhihao Fan, Jie Fu 0001, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACM Multimedia | 8 |
| 2024 | Graph Interpretation of Image-Text Matching: Link Prediction on Concept-Enhanced Cross-Modal Graph
Zhihao Fan, Zhongyu Wei, Haijun Shan |
NLPCC (3) | 1 |
| 2024 | An Iterative Framework for Document-Level Event Argument Extraction Assisted by Long Short-Term Memory
Tao You, Zhihao Fan, Cunxiang Yin, Yancheng He, Jinhua Fu, Zhongyu Wei |
NLPCC (2) | 3 |
| 2024 | Unifying Structure Reasoning and Language Pre-Training for Complex Reasoning TasksabstractRecent pre-trained language models (PLMs) equipped with foundation reasoning skills have shown remarkable performance on downstream complex tasks. However, the significant structure reasoning skill has been rarely studied, which involves modeling implicit structure information within the text and performing explicit logical reasoning over them to deduce the conclusion. This paper proposes a unified learning framework that combines explicit structure reasoning and language pre-training to endow PLMs with the structure reasoning skill. It first identifies several elementary structures within contexts to construct structured queries and performs step-by-step reasoning along the queries to identify the answer entity. The fusion of textual semantics and structure reasoning is achieved by using contextual representations learned by PLMs to initialize the representation space of structures, and performing stepwise reasoning on this semantic representation space. Experimental results on four datasets demonstrate that the proposed model achieves significant improvements in complex reasoning tasks involving diverse structures, and shows transferability to downstream tasks with limited training data and effectiveness for complex reasoning of KGs modality. Siyuan Wang 0025, Zhongyu Wei, Jiarong Xu, Taishan Li, Zhihao Fan |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Language Pre-trainingabstractMultilingual Vision-Language Pre-training (VLP) is a promising but challenging topic due to the lack of large-scale multilingual imagetext pairs.Existing works address the problem by translating English data into other languages, which is intuitive and the generated data is usually limited in form and scale.In this paper, we explore a more practical and scalable setting: weakly supervised multilingual VLP with only English image-text pairs and multilingual text corpora.We argue that the universal multilingual representation learned from texts allows the cross-modal interaction learned in English to be transferable to other languages.To this end, we propose a framework to effectively unify cross-lingual and cross-modal pre-training.For unified modeling on different data, we design an architecture with flexible modules to learn different interactions.Moreover, two unified tasks are introduced to efficiently guide the unified crosslingual cross-modal learning.Extensive experiments demonstrate that our pre-trained model learns universal multilingual multimodal representations, allowing effective cross-lingual transfer on multimodal tasks.Code and models are available at https://github.com/ FudanDISC/weakly-supervised-mVLP. Zhihao Fan, Jingjing Chen 0001, Qi Zhang 0001, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 2 |
| 2023 | Query Structure Modeling for Inductive Logical Reasoning Over Knowledge GraphsabstractSiyuan Wang, Zhongyu Wei, Meng Han, Zhihao Fan, Haijun Shan, Qi Zhang, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Haijun Shan, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2023 | Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseabstractIn this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pre-trained diffusion language model that consists of an encoder and a diffusion-based decoder, which can generate text by gradually transforming a random noise sequence into a coherent text sequence. To pre-train GENIE on a large-scale language corpus, we design a new continuous paragraph denoise objective, which encourages the diffusion-decoder to reconstruct a clean text paragraph from a corrupted version, while preserving the semantic and syntactic coherence. We evaluate GENIE on four downstream text generation benchmarks, namely XSum, CNN/DailyMail, Gigaword, and CommonGen. Our experimental results show that GENIE achieves comparable performance with the state-of-the-art autoregressive models on these benchmarks, and generates more diverse text samples. The code and models of GENIE are available at https://github.com/microsoft/ProphetNet/tree/master/GENIE. Zhenghao Lin, Yeyun Gong, Yelong Shen, Zhihao Fan, Chen Lin 0001, Nan Duan 0001, Weizhu Chen |
ICML | 5 |
| 2023 | Topic-Aware Modeling for Unsupervised Extractive SummarizationabstractThe recent success of extractive summarization depends on the availability of large-scale annotated datasets. Existing unsupervised approaches are mostly directed graph based by combining location information with centrality computing. These methods tend to generate summaries with two problems, one is low topic coverage of the source document called the facet bias problem, and the other is continuous position distribution of extracted sentences called the position bias problem. To solve these problems, we propose the topic-aware centrality-based sum-marization method (TACSUM). Specifically, we employ clustering techniques to explicitly model the topics of the document and define the metrics for topic consistency and topic coverage to improve the performance of summarization. The metric topic consistency is used to guide the calculation of centrality, which solves the position bias problem and achieves a more general effect in different scenarios. We combine the metric topic coverage with the centrality to enhance the topic awareness of the model, which ensures the selected sentences are important and diverse. Numerical experimental results on four datasets show that our method outperforms previous unsupervised methods, especially in long document domains. Extensive analyses confirm that our method can generate high-quality summaries by eliminating position bias and facet bias problems. Zhihao Fan, Huiyong Li 0005, Shasha Mo, Jianwei Niu 0002 |
IJCNN | 1 |
| 2023 | DANDELION: An ASV Deployed Micro-Profiler Array for Air-Sea ObservationabstractThe air-sea interface is vital in studying heat and energy exchange between the sea and air. The field observation technology of the air-sea interface is an effective way to explore the nature of the air-sea interface. This paper presents an observation system called DANDELION for the air-sea interface environment. The system includes an automatic surface vehicle (ASV), a launching device using a pair of high-speed rotating friction wheels, and eight dandelion-like micro profilers. This system can implement the environmental observation of the air-sea interface in an extensive range through micro profilers' rapid and multi-point placement. The DANDELION system was characterized by establishing the friction wheel launching mechanism model and summarizing the effects of different wings on the profiler performance. A series of experiments were conducted in Qiandao Lake, China, to characterize the DANDELION system. We demonstrate the developed system with data from field experiments, which show very high flexibility and feasibility to observe the air-sea interface, implying potential applications in ocean transient phenomena observation. Zhihao Fan, Chenxin Lyu, Zheng Zeng 0003 |
IROS | 1 |
| 2023 | AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationabstractDiffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently.
However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach.
To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right.
In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion. Zhihao Fan, Xiao Liu 0029, Hai-Tao Zheng 0002, Yeyun Gong, Yelong Shen, Jian Jiao 0007, Zhongyu Wei, Jian Guo 0016, Nan Duan 0001, Weizhu Chen |
NeurIPS | 2 |
| 2022 | Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain ConversationsabstractWei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wei Chen 0088, Yeyun Gong, Can Xu 0002, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan 0001 |
ACL (1) | 7 |
| 2022 | Locate Then Ask: Interpretable Stepwise Reasoning for Multi-hop Question AnsweringabstractMulti-hop reasoning requires aggregating multiple documents to answer a complex question. Existing methods usually decompose the multi-hop question into simpler single-hop questions to solve the problem for illustrating the explainable reasoning process. However, they ignore grounding on the supporting facts of each reasoning step, which tends to generate inaccurate decompositions. In this paper, we propose an interpretable stepwise reasoning framework to incorporate both single-hop supporting sentence identification and single-hop question generation at each intermediate step, and utilize the inference of the current hop for the next until reasoning out the final result. We employ a unified reader model for both intermediate hop reasoning and final hop inference and adopt joint optimization for more accurate and robust multi-hop reasoning. We conduct experiments on two benchmark datasets HotpotQA and 2WikiMultiHopQA. The results show that our method can effectively boost performance and also yields a better interpretable reasoning process without decomposition supervision. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 3 |
| 2022 | DRAGONFLY: a UAV Rapidly Deployed Micro-Profiler Array for Underwater Thermocline ObservationabstractUnderwater thermocline, common in the lakes and ocean, plays a vital role in meteorological forecasting in the ocean and lakes dynamics research. This letter proposes a method for rapid and multipoint observation of thermocline variations with time and space using an airdropped micro-profiler array, named the DRAGONFLY system. It comprises specially designed disposable low-cost micro-profilers, a general unmanned aerial carrier platform, and a ground control system. This system can conduct periodic profile observations at a single point or quickly survey a large area. A series of experiments to characterize the micro-profiler and the DRAGONFLY system were conducted in Qiandao Lake, China. We demonstrate the developed system with data from field experiments, which show very high flexibility, and feasibility to observe the lake thermocline, implying potential applications in ocean transient phenomena observation. Chenxin Lyu, Zhihao Fan, Yuanbo Bi, Zheng Zeng 0003, Lian Lian |
ICRA | 2 |
| 2022 | Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text RetrievalabstractExisting research for image text retrieval mainly relies on sentence-level supervision to distinguish matched and mismatched sentences for a query image. However, semantic mismatch between an image and sentences usually happens in finer grain, i.e., phrase level. In this paper, we explore to introduce additional phrase-level supervision for the better identification of mismatched units in the text. In practice, multi-grained semantic labels are automatically constructed for a query image in both sentence-level and phrase-level. We construct text scene graphs for the matched sentences and extract entities and triples as the phrase-level labels. In order to integrate both supervision of sentence-level and phrase-level, we propose Semantic Structure Aware Multimodal Transformer (SSAMT) for multi-modal representation learning. Inside the SSAMT, we utilize different kinds of attention mechanisms to enforce interactions of multi-grained semantic units in both sides of vision and language. For the training, we propose multi-scale matching from both global and local perspectives, and penalize mismatched phrases. Experimental results on MS-COCO and Flickr30K show the effectiveness of our approach compared to some state-of-the-art models. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001, Jianqing Fan |
ICMR | 1 |
| 2022 | MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningabstractPrevious vision-language pre-training models mainly construct multi-modal inputs with tokens and objects (pixels) followed by performing cross-modality interaction between them. We argue that the input of only tokens and object features limits high-level semantic alignment like phrase-to-region grounding. Meanwhile, multi-level alignments are inherently consistent and able to facilitate the representation learning synergistically. Therefore, in this paper, we propose to learn Multi-level semantic alignment for Vision-language Pre-TRaining (MVPTR). In MVPTR, we follow the nested structure of both modalities to introduce concepts as high-level semantics. To ease the learning from multi-modal multi-level inputs, our framework is split into two stages, the first stage focuses on intra-modality multi-level representation learning, the second enforces interactions across modalities via both coarse-grained and fine-grained semantic alignment tasks. In addition to the commonly used image-text matching and masked language model tasks, we introduce a masked concept recovering task in the first stage to enhance the concept representation learning, and two more tasks in the second stage to explicitly encourage multi-level alignments across modalities. Our model achieves state-of-the-art results on several vision and language tasks. Zhihao Fan, Huaixiao Tou, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACM Multimedia | 2 |
| 2022 | GJTD-LR: A Trainable Grouped Joint Tensor Dictionary With Low-Rank Prior for Single Hyperspectral Image Super-ResolutionabstractReconstructing a high-resolution hyperspectral image (HR-HSI) by using a single low-resolution hyperspectral image (LR-HSI) is a significant technique for increasing the spatial resolution of HSIs and overcoming the physical limitation of the HSI sensor. Most single HSI super-resolution methods have achieved great success recently. However, owning to the difficulty of acquiring an HSI, the available training samples are relatively few, which will inevitably lead to relatively low performance. To address this issue, in the paper, we propose a novel single HSI super-resolution method by combining a trainable grouped joint tensor dictionary and a low-rank prior (GJTD-LR). First, we design a trainable grouped joint tensor dictionary, which can build an accurate mapping relationship between training HR-HSIs and their corresponding LR-HSIs with relatively few training samples. To be specific, the training HR-HSI and LR-HSI pairs are decomposed into a joint tensor dictionary and a set of sparse coefficients by using tensor-tensor product to fully preserve the spectral correlation. In addition, we apply a grouped strategy to divide the training images into several groups and learn a compact joint dictionary for each group. Second, a tensor low-rank model is forced into the reconstruction model to further capture the spatial correlation. At last, GJTD-LR is optimized by employing alternating direction method of multipliers (ADMM), soft threshold algorithm, singular value decomposition and fourier domain transform. The experimental results on both remote sensed HSIs and indoor HSIs show the superiority of GJTD-LR to some other traditional and advanced single HSI super-resolution methods. Cong Liu 0011, Zhihao Fan, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-level Structural Information
Zhongyu Wei, Zhihao Fan, Haijun Shan, Xuanjing Huang 0001 |
AAAI | 3 |
| 2021 | TCIC: Theme Concepts Learning Cross Language and Vision for Image CaptioningabstractExisting research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concepts extended Image Captioning (TCIC) framework that incorporates theme concepts to represent high-level cross-modality semantics. In practice, we model theme concepts as memory vectors and propose Transformer with Theme Nodes (TTN) to incorporate those vectors for image captioning. Considering that theme concepts can be learned from both images and captions, we propose two settings for their representations learning based on TTN. On the vision side, TTN is configured to take both scene graph based features and theme concepts as input for visual representation learning. On the language side, TTN is configured to take both captions and theme concepts as input for text representation re-construction. Both settings aim to generate target captions with the same transformer-based decoder. During the training, we further align representations of theme concepts learned from images and corresponding captions to enforce the cross-modality learning. Experimental results on MS COCO show the effectiveness of our approach compared to some state-of-the-art models. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001 |
IJCAI | 1 |
| 2021 | Mask Attention Networks: Rethinking and Strengthen TransformerabstractZhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001 |
NAACL-HLT | 1 |
| 2021 | Fusion of multi-source retinal fundus images via automatic registration for clinical diagnosis
Tingting Dan, Yu Hu 0004, Chu Han, Zhihao Fan, Zhuobin Huang, Bin Zhang 0050, Guihua Tao, Baoyi Liu, Honghua Yu, Hongmin Cai |
Neurocomputing | 4 |
| 2021 | SGUNet: Style-guided UNet for adversely conditioned fundus image super-resolution
Zhihao Fan, Tingting Dan, Baoyi Liu, Xiaoqi Sheng, Honghua Yu, Hongmin Cai |
Neurocomputing | 1 |
| 2020 | Reconstruction of 3D Retina from Multi-viewed Stereo Fundus Images via Dynamic RegistrationabstractThe human retinal surface resembles to a sphere while it is captured by two-dimensional (2D) planar imaging to have a stereo sequence in clinical practice. Reconstructing its three-dimensional (3D) structure from the 2D planar retinal images is crucial for analyzing the relationship between the topological morphology and clinical implication. In this regard, we propose to reconstruct the 3D retina structure from 2D stereo fundus images via dynamic registration. The fundus images from different viewpoints are first co-registrated by using multi-scale deep convolutional feature and geometric structure feature by building their transformation function. The aligned images are then mosaicked together and a 3D reconstruction is obtained by a learned weighted smoothing project the registered images onto 3D coordinates. We compare the proposed registration method with five state-of-the-art methods. Extensive experimental results demonstrate that the proposed framework achieves superior performances, even with challenging scenarios in which the tested images are severely degraded by illness, large eyeball rotation and low resolutions. Tingting Dan, Zhihao Fan, Yu Hu 0004, Bin Zhang 0050, Guihua Tao, Hongmin Cai |
BIBM | 2 |
| 2020 | An Enhanced Knowledge Injection Model for Commonsense GenerationabstractCommonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics. Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang |
COLING | 1 |
| 2020 | PathQG: Neural Question Generation from FactsabstractExisting research for question generation encodes the input text as a sequence of tokens without explicitly modeling fact information.These models tend to generate irrelevant and uninformative questions.In this paper, we explore to incorporate facts in the text for question generation in a comprehensive way.We present a novel task of question generation given a query path in the knowledge graph constructed from the input text.We divide the task into two steps, namely, query representation learning and query-based question generation.We formulate query representation learning as a sequence labeling problem for identifying the involved facts to form a query and employ an RNN-based generator for question generation.We first train the two modules jointly in an end-to-end fashion, and further enforce the interaction between these two modules in a variational framework.We construct the experimental datasets on top of SQuAD and results show that our model outperforms other state-of-the-art approaches, and the performance margin is larger when target questions are complex.Human evaluation also proves that our model is able to generate relevant and informative questions. 1 Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Zengfeng Huang, Weijian Sun, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 3 |
| 2019 | A Multi-Agent Communication Framework for Question-Worthy Phrase Extraction and Question GenerationabstractQuestion generation aims to produce questions automatically given a piece of text as input. Existing research follows a sequence-to-sequence fashion that constructs a single question based on the input. Considering each question usually focuses on a specific fragment of the input, especially in the scenario of reading comprehension, it is reasonable to identify the corresponding focus before constructing the question. In this paper, we propose to identify question-worthy phrases first and generate questions with the assistance of these phrases. We introduce a multi-agent communication framework, taking phrase extraction and question generation as two agents, and learn these two tasks simultaneously via message passing mechanism. The results of experiments show the effectiveness of our framework: we can extract question-worthy phrases, which are able to improve the performance of question generation. Besides, our system is able to extract more than one question worthy phrases and generate multiple questions accordingly. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Yang Liu 0004, Xuanjing Huang 0001 |
AAAI | 3 |
| 2019 | Bridging by Word: Image Grounded Vocabulary Construction for Visual CaptioningabstractExisting research for visual captioning usually employs a CNN-RNN architecture that combines a CNN for image encoding with a RNN for caption generation, where the vocabulary is constructed from the entire training dataset as the decoding space.Such approaches typically suffer from the problem of generating N-grams which occur frequently in the training set but are irrelevant to the given image.To tackle this problem, we propose to construct an image-grounded vocabulary that leverages image semantics for more effective caption generation.More concretely, a two-step approach is proposed to construct the vocabulary by incorporating both visual information and relationships among words.Two strategies are then explored to utilize the constructed vocabulary for caption generation.One constrains the generator to select words from the image-grounded vocabulary only and the other integrates the vocabulary information into the RNN cell during the caption generation process.Experimental results on two public datasets show the effectiveness of our framework compared to state-of-the-art models.Our code is available on Github 1 . Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2018 | A Reinforcement Learning Framework for Natural Question Generation using Bi-discriminatorsabstractVisual Question Generation (VQG) aims to ask natural questions about an image automatically. Existing research focus on training model to fit the annotated data set that makes it indifferent from other language generation tasks. We argue that natural questions need to have two specific attributes from the perspectives of content and linguistic respectively, namely, natural and human-written. Inspired by the setting of discriminator in adversarial learning, we propose two discriminators, one for each attribute, to enhance the training. We then use the reinforcement learning framework to incorporate scores from the two discriminators as the reward to guide the training of the question generator. Experimental results on a benchmark VQG dataset show the effectiveness and robustness of our model compared to some state-of-the-art models in terms of both automatic and human evaluation metrics. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Yang Liu 0004, Xuanjing Huang 0001 |
COLING | 1 |
| 2018 | A Question Type Driven Framework to Diversify Visual Question GenerationabstractVisual question generation aims at asking questions about an image automatically. Existing research works on this topic usually generate a single question for each given image without considering the issue of diversity. In this paper, we propose a question type driven framework to produce multiple questions for a given image with different focuses. In our framework, each question is constructed following the guidance of a sampled question type in a sequence-to-sequence fashion. To diversify the generated questions, a novel conditional variational auto-encoder is introduced to generate multiple questions with a specific question type. Moreover, we design a strategy to conduct the question type distribution learning for each image to select the final questions. Experimental results on three benchmark datasets show that our framework outperforms the state-of-the-art approaches in terms of both relevance and diversity. Zhihao Fan, Zhongyu Wei, Piji Li, Yanyan Lan, Xuanjing Huang 0001 |
IJCAI | 1 |