VLDB 2026 Research / reviewers in the wild / expert
Yihao Ding
dblp:321/1682
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-5065-6911ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationabstractRadiology report generation from chest X-rays is an important task in artificial intelligence with the potential to greatly reduce radiologists' workload and shorten patient wait times. Despite recent advances, existing approaches often lack sufficient disease-awareness in visual representations and adequate vision-language alignment to meet the specialized requirements of medical image analysis. As a result, these models usually overlook critical pathological features on chest X-rays and struggle to generate clinically accurate reports. To address these limitations, we propose a novel dual-stage disease-aware framework for chest X-ray report generation. In Stage 1, our model learns Disease-Aware Semantic Tokens (DASTs) corresponding to specific pathology categories through cross-attention mechanisms and multi-label classification, while simultaneously aligning vision and language representations via contrastive learning. In Stage 2, we introduce a Disease-Visual Attention Fusion (DVAF) module to integrate disease-aware representations with visual features, along with a Dual-Modal Similarity Retrieval (DMSR) mechanism that combines visual and disease-specific similarities to retrieve relevant exemplars, providing contextual guidance during report generation. Extensive experiments on benchmark datasets (i.e., CheXpert Plus, IU X-ray, and MIMIC-CXR) demonstrate that our disease-aware framework achieves state-of-the-art performance in chest X-ray report generation, with significant improvements in clinical accuracy and linguistic quality. Puzhen Wu, Hexin Dong, Yi Lin 0009, Yihao Ding, Yifan Peng 0002 |
AAAI | 4 |
| 2026 | NarrativeSense: Predicting Affective States in University Students through Smartphone Sensing and Contextual NarrativesabstractMental health challenges are increasingly prevalent among university students, yet often go undetected due to reliance on traditional assessments that are subjective, infrequent, and lack behavioral context. Digital phenotyping through passively collected smartphone data offers a scalable alternative, but existing approaches often fail to integrate predictive accuracy with narrative-based insights. To overcome these limitations, we present NarrativeSense, a novel framework that combines machine learning models with narrative-based descriptions of daily life events inferred from smartphone sensing data to predict weekly affective states. The system incorporates language model components to transform behavioral patterns into contextualized, human-readable narratives that ground affective predictions in everyday experiences. This narrative layer complements structured prediction by offering intuitive, user-centered insights. Applied to longitudinal data from 58 university students over 119 days, NarrativeSense outperforms baseline machine learning models, standalone LLMs, and ensemble methods, while providing richer insights. Our findings demonstrate the potential of narrative-enhanced digital phenotyping for scalable and explainable mental health monitoring in educational and clinical settings. Yan Li 0186, Yihao Ding, Hong Jia, Vassilis Kostakos, Simon D'Alfonso |
ACM Trans. Comput. Heal. | 3 |
| 2026 | LPL3D: LVLM-Driven Pseudo-Labeling for 3D Object DetectionabstractEffective 3D object detection requires large-scale annotated datasets, which are expensive and time-consuming to produce - especially in indoor environments containing dense object arrangements. To address this, we propose an Large Vision-Language Model (LVLM)-driven automatic high-quality pseudo-label generation technique for 3D object detection in single- and multi-view scenarios. We propose an Auto3DLabeler that introduces the first-ever text-to-3D Bounding Box transformation. Its pipeline employs a text-based detector, a segmenter and an LVLM to generate annotation estimates, which are further refined by our IoU-guided iterative Box Aggregator and layout-aware prompt Class Refiner modules. We also introduce a semantic-enhanced multi-modal fusion module that integrates image-level semantic information into point cloud representations for precise detections. Collectively, our contributions provide a remarkable boost to the 3D object detection state-of-the-art. Extensive experiments on SUN RGB-D and ScanNet datasets show our unsupervised detector variant outperforming existing semi-supervised detectors, and our semi-supervised variant achieving up to 28.2% absolute gain in challenging scenarios - all this while maintaining considerable compute advantage over existing label-efficient methods. Our code and models will be made public for the community. Our code and model will be made public after acceptance. Zechuan Li, Hongshan Yu, Yihao Ding, Shuai Yuan 0013, Naveed Akhtar |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | KIEPrompter: Leveraging Lightweight Models' Predictions for Cost-Effective Key Information Extraction using Vision LLMsabstractKey information extraction (KIE) from visually rich documents, such as receipts and forms, involves a deep understanding of textual, visual, and layout feature information. Transformers fine-tuned for KIE achieve state-of-the-art performance but lack generality and portability across different domains. In contrast, vision large language models (VLLMs) offer higher flexibility and zero-shot capability but fall short with domain-specific layout relations unless performing a resource-demanding supervised fine-tuning. To reach the best compromise solution between lightweight models and VLLMs, we propose KIEPrompter, a cost-effective LLM-based KIE approach that leverages the predictions of lightweight models as external knowledge injected into VLLM prompts. By incorporating these auxiliary predictions, VLLMs are guided to attend relevant multimodal content without ad hoc training. The accuracy results achieved by KIEPrompter in three benchmark document collections are superior to those of VLLMs in both zero-shot and layout-sensitive scenarios. We compare various strategies for incorporating lightweight model predictions, ranging from coarse-grained predictions without explicit confidence scores to fine-grained per-element network logits. We also demonstrate that our approach is robust to the absence of specific classes in trained lightweight models, as the VLLMs' pre-training compensates for the limited generality of lightweight models. Lorenzo Vaiani, Yihao Ding, Luca Cagliero, Jean Lee, Paolo Garza, Josiah Poon, Soyeon Caren Han |
CIKM | 2 |
| 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorabstractWe propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet. Zechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao, Basim Azam, Naveed Akhtar |
CVPR | 3 |
| 2025 | VRD-IU: Lessons from Visually Rich Document Intelligence and UnderstandingabstractVisually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges due to their complex layouts, multi-stakeholder involvement, and high structural variability. Addressing these issues, the VRD-IU Competition was introduced, focusing on extracting and localizing key information from multi-format forms within the Form-NLU dataset, which includes digital, printed, and handwritten documents. This paper presents insights from the competition, which featured two tracks: Track A, emphasizing entity-based key information retrieval, and Track B, targeting end-to-end key information localization from raw document images. With over 20 participating teams, the competition showcased various state-of-the-art methodologies, including hierarchical decomposition, transformer-based retrieval, multimodal feature fusion, and advanced object detection techniques. The top-performing models set new benchmarks in VRDU, providing valuable insights into document intelligence. Yihao Ding, Soyeon Caren Han, Yan Li 0186, Josiah Poon |
IJCAI | 1 |
| 2025 | Pseudo-labeling and knowledge-guided contrastive learning for radiology report generation
Yihao Ding |
J. Biomed. Informatics | 3 |
| 2024 | The Language Model Can Have the Personality: Joint Learning for Personality Enhanced Language Model (Student Abstract)abstractWith the introduction of large language models, chatbots are becoming more conversational to communicate effectively and capable of handling increasingly complex tasks. To make a chatbot more relatable and engaging, we propose a new language model idea that maps the human-like personality. In this paper, we propose a systematic Personality-Enhanced Language Model (PELM) approach by using a joint learning mechanism of personality classification and language generation tasks. The proposed PELM leverages a dataset of defined personality typology, Myers-Briggs Type Indicator, and produces a Personality-Enhanced Language Model by using a joint learning and cross-teaching structure consisting of a classification and language modelling to incorporate personalities via both distinctive types and textual information. The results show that PELM can generate better personality-based outputs than baseline models. Feiqi Cao, Yihao Ding, Soyeon Caren Han |
AAAI | 3 |
| 2024 | MMVQA: A Comprehensive Dataset for Investigating Multipage Multimodal Information Retrieval in PDF-based Visual Question Answering
Yihao Ding, Kaixuan Ren, Jiabin Huang 0005, Siwen Luo, Soyeon Caren Han |
IJCAI | 1 |
| 2023 | Workshop on Document Intelligence UnderstandingabstractDocument understanding and information extraction include different tasks to understand a document and extract valuable information automatically. Recently, there has been a rising demand for developing document understanding among different domains, including business, law, and medicine, to boost the efficiency of work that is associated with a large number of documents. This workshop aims to bring together researchers and industry developers in the field of document intelligence and understanding diverse document types to boost automatic document processing and understanding techniques. We also release a data challenge on the recently introduced document-level VQA dataset, PDFVQA. The PDFVQA challenge examines the model's structural and contextual understandings on the natural full document level of multiple consecutive document pages by including questions with a sequence of answers extracted from multi-pages of the full document. This task helps to boost the document understanding step from the single-page level to the full document level understanding. Soyeon Caren Han, Yihao Ding, Siwen Luo, Josiah Poon, Hee-Guen Yoon, Paul Duuring, Eun-Jung Holden |
CIKM | 2 |
| 2023 | Form-NLU: Dataset for the Form Natural Language UnderstandingabstractCompared to general document analysis tasks, form document structure understanding and retrieval are challenging. Form documents are typically made by two types of authors; A form designer, who develops the form structure and keys, and a form user, who fills out form values based on the provided keys. Hence, the form values may not be aligned with the form designer's intention (structure and keys) if a form user gets confused. In this paper, we introduce Form-NLU, the first novel dataset for form structure understanding and its key and value information extraction, interpreting the form designer's intent and the alignment of user-written value on it. It consists of 857 form images, 6k form keys and values, and 4k table keys and values. Our dataset also includes three form types: digital, printed, and handwritten, which cover diverse form appearances and layouts. We propose a robust positional and logical relation-based form key-value information extraction framework. Using this dataset, Form-NLU, we first examine strong object detection models for the form layout understanding, then evaluate the key information extraction task on the dataset, providing fine-grained results for different types of forms and keys. Furthermore, we examine it with the off-the-shelf pdf layout extraction tool and prove its feasibility in real-world cases. Yihao Ding, Siqu Long, Jiabin Huang 0005, Kaixuan Ren, Xingxiang Luo, Hyunsuk Chung, Soyeon Caren Han |
SIGIR | 1 |
| 2022 | Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout AnalysisabstractRecognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent studies in Document Layout Analysis usually rely on visual cues to understand documents while ignoring other information, such as contextual information or the relationships between document layout components, which are vital to boost better layout analysis performance. Our Doc-GCN presents an effective way to harmonize and integrate heterogeneous aspects for Document Layout Analysis. We construct different graphs to capture the four main features aspects of document layout components, including syntactic, semantic, density, and appearance features. Then, we apply graph convolutional networks to enhance each aspect of features and apply the node-level pooling for integration. Finally, we concatenate features of all aspects and feed them into the 2-layer MLPs for document layout component classification. Our Doc-GCN achieves state-of-the-art results on three widely used DLA datasets: PubLayNet, FUNSD, and DocBank. The code will be released at https://github.com/adlnlp/doc_gcn Siwen Luo, Yihao Ding, Siqu Long, Josiah Poon, Soyeon Caren Han |
COLING | 2 |
| 2022 | V-Doc : Visual questions answers with DocumentsabstractWe propose V-Doc, a question-answering tool using document images and PDF, mainly for researchers and general non-deep learning experts looking to generate, process, and understand the document visual question answering tasks. The V-Doc supports generating and using both extractive and abstractive question-answer pairs using documents images. The extractive QA selects a subset of tokens or phrases from the document contents to predict the answers, while the abstractive QA recognises the language in the content and generates the answer based on the trained model. Both aspects are crucial to understanding the documents, especially in an image format. We include a detailed scenario of question generation for the abstractive QA task. V-Doc supports a wide range of datasets and models, and is highly extensible through a declarative, framework-agnostic platform.11Data and demo video: https://github.com/usydnlp/vdoc Yihao Ding, Runlin Wang, Yanhang Zhang, Xianru Chen, Yuzhong Ma, Hyunsuk Chung, Soyeon Caren Han |
CVPR | 1 |