EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Zhang 0021
dblp:58/4582-21
· DBLP profile ↗
15ranked-venue papers
7as first author
15since 2021 · last 2026
0009-0008-7658-0116ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Correspondence Coverage Matters for Multi-Modal Dataset DistillationabstractMulti-modal dataset distillation (DD) condenses large datasets into compact ones that retain task efficacy by capturing correspondence patterns, i.e., shared semantics between paired modalities. However, such patterns rely on cross-modal similarity and cannot be faithfully captured by intra-modal similarity of current unimodal strategies. As a result, current multi-modal DD methods tend to over-concentrate, redundantly encoding similar correspondence patterns and thus limiting generalizability. To this end, we propose a novel multi-modal DD framework to systematically Promote Correspondence coverage, i.e., ProCo. Initially, we develop a correspondence consistency metric based on cross-modal retrieval distributions to cluster correspondence patterns. These clusters capture the underlying correspondence distribution, enabling ProCo to initialize distilled data with representative patterns while regularizing optimization to promote correspondence representativeness and diversity. Moreover, we employ conditional neural fields for efficient distilled data parameterization, enhancing fine-grained pattern capture while allowing more distilled data under a fixed budget to boost correspondence coverage. Extensive experiments verify that our ProCo achieves superior and elastic budget-efficacy trade-offs, surpassing prior methods by over 15% with 10x distillation budget reduction, highlighting its real-world practicality. Zhuohang Dang, Minnan Luo, Chengyou Jia, Hangwei Qian, Xinyu Zhang 0021, Xiaojun Chang, Ivor W. Tsang |
AAAI | 5 |
| 2026 | Encode Geometric Diagram as Geo-Graph in Geometry Problem SolvingabstractGeometry Problem Solving has become a hot topic these years due to its complexity of enabling the machine with geometric abstraction, multi-modal reasoning and mathematical capabilities. Majority of research works place their attention on the fusion of multi-modal data or the synergistic combination of neural and symbolic systems for performance improvement. However, their neglect of the unique characteristics of geometric diagrams, which distinguish them from natural images, impedes the further exploring of critical information in geometric diagrams. In this work, we introduce the novel concept of geo-graph and propose the Geo-Graph Geometry Problem Solving model which encodes the geometric diagram from a new perspective. The geo-graph is designed to include semantic, structural and spatial information in the diagram, which is crucial to subsequent problem reasoning stage. To facilitate the model's comprehension of the actual layout of geometric diagram, spatial and connecting attentions are devised to serve as intrinsic knowledge guidance for feature propagation. An extra cross-modal attention is used as external guidance to instruct the encoding of geo-graph to be related to specific problem target. Fused multi-modal features are then sent into a commonly used encoder-decoder framework for final solution generation. The model is first trained with three carefully designed pre-training tasks to establish its fundamental knowledge of geo-graph, leveraging numerous varied samples generated through a geo-graph-based augmentation method. Experiments on popular geometry problem solving datasets demonstrate the effectiveness and superiority of our model for geometric diagram encoding. Lingling Zhang 0005, Xinyu Zhang 0021, Yaqiang Wu |
AAAI | 5 |
| 2026 | Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem SolvingabstractXinyu Zhang, Yuchen Wan, Boxuan Zhang, Zesheng Yang, Lingling Zhang, Bifan Wei, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinyu Zhang 0021, Yuchen Wan, Zesheng Yang, Lingling Zhang 0005, Bifan Wei, Jun Liu 0002 |
ACL (1) | 1 |
| 2026 | Memory-enriched thought-by-thought framework for complex Diagram Question Answering
Xinyu Zhang 0021, Lingling Zhang 0005, Yanrui Wu, Muye Huang, Qianying Wang 0002, Jun Liu 0002 |
Comput. Vis. Image Underst. | 1 |
| 2025 | EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart UnderstandingabstractChart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension. While existing Visual Language Models (VLMs) have shown progress in chart understanding, the lack of high-quality training data and comprehensive evaluation benchmarks hinders VLM chart comprehension. In this paper, we introduce EvoChart, a novel self-training method for generating synthetic chart data to enhance VLMs' capabilities in real-world chart comprehension. We also propose EvoChart-QA, a noval benchmark for measuring models' chart comprehension abilities in real-world scenarios. Specifically, EvoChart is a unique self-training data synthesis approach that simultaneously produces high-quality training corpus and a high-performance chart understanding model. EvoChart-QA consists of 650 distinct real-world charts collected from 140 different websites and 1,250 expert-curated questions that focus on chart understanding. Experimental results on various open-source and proprietary VLMs tested on EvoChart-QA demonstrate that even the best proprietary model, GPT-4o, achieves only 49.8% accuracy. Moreover, the EvoChart method significantly boosts the performance of open-source VLMs on real-world chart understanding tasks, achieving 54.2% accuracy on EvoChart-QA. Muye Huang, Han Lai, Xinyu Zhang 0021, Jie Ma 0001, Lingling Zhang 0005, Jun Liu 0002 |
AAAI | 3 |
| 2025 | VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution ReasoningabstractCharts are widely used for data visualization across various fields, including education, research, and business. Chart Question Answering (CQA) is an emerging task focused on the automatic interpretation and reasoning of data presented in charts. However, chart images are inherently difficult to interpret, and chart-related questions often involve complex logical and numerical reasoning, which hinders the performance of existing models. This paper introduces VProChart, a novel framework designed to address these challenges in CQA by integrating a lightweight Visual Perception Alignment Agent (VPAgent) and a Programmatic Solution Reasoning approach. VPAgent aligns and models chart elements based on principles of human visual perception, enhancing the understanding of chart context. The Programmatic Solution Reasoning approach leverages large language models (LLMs) to transform natural language reasoning questions into structured solution programs, facilitating precise numerical and logical reasoning. Extensive experiments on benchmark datasets such as ChartQA and PlotQA demonstrate that VProChart significantly outperforms existing methods, highlighting its capability in understanding and reasoning with charts. Muye Huang, Lingling Zhang 0001, Han Lai, Xinyu Zhang 0021, Jun Liu 0002 |
AAAI | 5 |
| 2025 | PhysReason: A Comprehensive Benchmark towards Physics-Based ReasoningabstractLarge language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Xinyu Zhang 0021, Yanrui Wu, Chengyou Jia, Basura Fernando, Zheng Shou 0001, Lingling Zhang 0005, Jun Liu 0036 |
ACL (1) | 1 |
| 2025 | Diagram-Driven Course Questions GenerationabstractXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang, Wenjun Wu, Bo Li, Shaowei Wang, Basura Fernando, Jun Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xinyu Zhang 0021, Lingling Zhang 0005, Yanrui Wu, Muye Huang, Basura Fernando, Jun Liu 0002 |
EMNLP | 1 |
| 2025 | Cognitive Predictive Coding Network: Rethinking the Generalization in Raven's Progressive MatricesabstractAbstract visual reasoning, exemplified by Raven's Progressive Matrices (RPM), remains a significant challenge in artificial intelligence. A critical difficulty lies in disentangling abstract relational rules from image-specific features, as these rules operate independently of visual appearances. To address this challenge, we propose the Cognitive Predictive Coding Network (CPCN), inspired by predictive coding theory from cognitive science. CPCN features a three-component architecture: a Relation Disentangler that separates abstract rules from image-specific features through prediction error minimization; Stacked Free Energy Minimizers (FEMs) that leverage energy minimization principles to progressively reduce uncertainty during hierarchical abstraction; and a classifier for solution identification. Unlike previous approaches, our model employs mutual information constraints to explicitly separate relation-relevant and relation-irrelevant features, enabling more robust pattern recognition. Our novel FEMs provide a principled approach to uncertainty reduction through iterative refinement of pattern understanding. Experiments demonstrate PAH's superior performance across multiple benchmarks (such as 98.9% on RAVEN-Fair) with state-of-the-art 59.7% average accuracy across all PGM subtasks. Xinyu Zhang 0021, Lingling Zhang 0005, Yanrui Wu, Muye Huang, Jun Liu 0002 |
ACM Multimedia | 1 |
| 2025 | CoFFT: Chain of Foresight-Focus Thought for Visual Language ModelsabstractDespite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input.
When images contain large amounts of irrelevant information, VLMs are susceptible to interference, thus generating excessive task-irrelevant reasoning processes or even hallucinations.
This limitation stems from their inability to discover and process the required regions during reasoning precisely.
To address this limitation, we present the Chain of Foresight-Focus Thought (CoFFT), a novel training-free approach that enhances VLMs' visual reasoning by emulating human visual cognition.
Each Foresight-Focus Thought consists of three stages:
(1) Diverse Sample Generation: generates diverse reasoning samples to explore potential reasoning paths, where each sample contains several reasoning steps;
(2) Dual Foresight Decoding: rigorously evaluates these samples based on both visual focus and reasoning progression, adding the first step of optimal sample to the reasoning process;
(3) Visual Focus Adjustment: precisely adjust visual focus toward regions most beneficial for future reasoning, before returning to stage (1) to generate subsequent reasoning samples until reaching the final answer.
These stages function iteratively, creating an interdependent cycle where reasoning guides visual focus and visual focus informs subsequent reasoning.
Empirical results across multiple benchmarks using Qwen2.5-VL, InternVL-2.5, and Llava-Next demonstrate consistent performance improvements of 3.1-5.8\% with controllable increasing computational overhead. Xinyu Zhang 0021, Lingling Zhang 0005, Chengyou Jia, Zhuohang Dang, Basura Fernando, Jun Liu 0036, Zheng Shou 0001 |
NeurIPS | 1 |
| 2025 | Alignment-Guided Self-Supervised Learning for Diagram Question AnsweringabstractDiagram question answering (DQA), which is defined as answering natural language questions according to the visual diagram context, has attracted attention and has recently become a new benchmark for evaluating the complex reasoning ability of models. However, this reasoning task is extremely challenging because of the inclusion of abstract visual objects and specialized textual terms, as well as the complex relationships between them. The rarity of data caused by the high cost of annotation also makes large-scale deep models invalid for the DQA task. To address the above challenges, this paper proposes the cross-modal alignment-guided self-supervised learning model for DQA (CAS-DQA). Unlike previous works, the CAS-DQA model focuses on learning internal visual-textual object relationships, innovatively proposes an attention mechanism module based on object alignment, and effectively integrates cross-modal knowledge units for diagram understanding. In addition, the CAS-DQA model constructs two self-supervised learning (SSL) tasks via intermediate results of visual-textual object alignment. These two tasks exploit the unnoticed objects inside the diagram to fully and completely understand the diagram. They also effectively increase the amount of diagram question-answering data to address the challenge of data scarcity. To the best of our knowledge, the CAS-DQA model is the first to extend SSL strategies to the diagram question-answering task. We evaluate the CAS-DQA model on three different datasets. The results of extensive experiments show that our model significantly outperforms baselines on different scenarios and that the internal object alignment module and self-supervised tasks produce excellent results. Lingling Zhang 0005, Tao Qin 0002, Xinyu Zhang 0021, Jun Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringabstractDiagram Question Answering (DQA) is a challenging task, requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring, technical support, and more practical applications. DQA poses significant challenges, such as the demand for domain-specific knowledge and the scarcity of annotated data, which restrict the applicability of large-scale deep models. Previous approaches have explored external knowledge integration through pretraining, but these methods are costly and can be limited by domain disparities. While Large Language Models (LLMs) show promise in question-answering, there is still a gap in how to cooperate and interact with the diagram parsing process. In this paper, we introduce the Chain-of-Guiding Learning Model for Diagram Question Answering (CoG-DQA), a novel framework that effectively addresses DQA challenges. CoG-DQA leverages LLMs to guide diagram parsing tools (DPTs) through the guiding chains, enhancing the precision of diagram parsing while introducing rich background knowledge. Our experimental findings reveal that CoG-DQA surpasses all comparison models in various DQA scenarios, achieving an average accuracy enhancement exceeding 5% and peaking at 11% across four datasets. These results underscore CoG-DQA's capacity to advance the field of visual question answering and promote the integration of LLMs into specialized domains. Lingling Zhang 0005, Longji Zhu, Tao Qin 0002, Kim-Hui Yap, Xinyu Zhang 0021, Jun Liu 0002 |
CVPR | 6 |
| 2024 | Alignment Relation is What You Need for Diagram ParsingabstractAs a knowledge carrier, the diagram is widely distributed in many aspects of human life, such as textbooks, architectural drawings, and documents. Different from natural images, representations of visual elements in the diagram are sparser, and similar visual representations can reflect dissimilar semantics. Thus, current methods fail to capture the visual elements with precise semantics. To address this issue, regarding the aligned visual and textual elements as pairs is the way to assign the precise semantics of textual elements to visual elements. We build the first diagram dataset named align diagram element (ADE), which includes annotations for alignment relations between visual and textual elements. And we propose a visual-textual alignment model (VTAM) including graph construction and optimal aligning phases. In the graph construction phase, the relational graphs are constructed between different elements with four relational operators. The relational operators are designed to measure the relations between different elements, according to distance, connection line, inclusion, and feature similarity. In the optimal aligning phase, the representation at each visual-textual pair is improved as a weighted sum of the representations on all relational graphs. Experimental results show that our VTAM achieves a significant improvement of 10.9% on mean test folds of the ADE dataset than the current best competitor. In order to explore the role of alignment relations in diagram parsing, we introduce VTAM to diagram-related tasks, such as diagram question answering (DQA). And we achieve 2.8% to 5.9% and 4.6% to 5.1% improvements on AI2D and Foodwebs after adding VTAM. Our dataset and code are released at: https://github.com/ADE-dataset/ADE-dataset. Xinyu Zhang 0021, Lingling Zhang 0005, Jun Liu 0002, Qianying Wang 0002 |
IEEE Trans. Image Process. | 1 |
| 2023 | Diagram Visual Grounding: Learning to See with Gestalt-Perceptual AttentionabstractDiagram visual grounding aims to capture the correlation between language expression and local objects in the diagram, and plays an important role in the applications like textbook question answering and cross-modal retrieval. Most diagrams consist of several colors and simple geometries. This results in sparse low-level visual features, which further aggravates the gap between low-level visual and high-level semantic features of diagrams. The phenomenon brings challenges to the diagram visual grounding. To solve the above issues, we propose a gestalt-perceptual attention model to align the diagram objects and language expressions. For low-level visual features, inspired by the gestalt that simulates human visual system, we build a gestalt-perception graph network to make up the features learned by the traditional backbone network. For high-level semantic features, we design a multi-modal context attention mechanism to facilitate the interaction between diagrams and language expressions, so as to enhance the semantics of diagrams. Finally, guided by diagram features and linguistic embedding, the target query is gradually decoded to generate the coordinates of the referred object. By conducting comprehensive experiments on diagrams and natural images, we demonstrate that the proposed model achieves superior performance over the competitors. Our code will be released at https://github.com/AIProCode/GPA. Lingling Zhang 0005, Jun Liu 0002, Xinyu Zhang 0021, Qianying Wang 0002 |
IJCAI | 4 |
| 2023 | RPMG-FSS: Robust Prior Mask Guided Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) has been developed to perform pixel-level segmentation with only a few dense labeled examples for training, which relieves the expensive annotation problem in traditional segmentation models. Current researches on FSS generally act the labeled masks on the corresponding support images to obtain the class-specific embeddings, and predict the pixel-level masks for query images by matching their pixels to these class-specific embeddings. Their performance is difficult to further break through because of the limited supervision from single-view support images and the neglect of position information from similar pixels between query and support images. To solve these issues, we propose a novel robust prior mask guided model named RPMG-FSS for the challenging FSS task. The core of RPMG-FSS is to produce a robust prior mask with good generalization ability on novel classes to better assist the following query mask prediction. Note that each element in the prior mask corresponds to one pixel in query image. It not only considers the interaction within one view and between multiple views of the support image, but also fuses the top-$k$similarity values to all support pixels and these pixels’ position information. The parameters in RPMG-FSS are optimized with the combination of segmentation loss and multi-view contrastive loss. Comprehensive experiments on two datasets show that our RPMG-FSS achieves outstanding performance comparing with the current popular baselines. The code is released onhttps://github.com/dxzxy12138/RPMG-FSS/tree/master Lingling Zhang 0005, Xinyu Zhang 0021, Qianying Wang 0002, Xiaojun Chang, Jun Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |