EDBT 2026 Demo / reviewers in the wild / expert
Xianfu Cheng
dblp:05/10105
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QAabstractTables serve as a core format for representing structured data on the web, as their two-dimensional layouts effectively encode complex inter-entity relationships. However, real-world web tables often feature heterogeneous structures and rich semantics. Accurately interpreting such tables requires not only spatial layout perception but also multi-step reasoning across rows and columns, posing substantial challenges to web intelligence systems. Multimodal large language models (MLLMs) show promise in table question answering (TableQA) by leveraging visual layouts. However, their performance on complex web tables remains uneven, as existing benchmarks often blur the impact of individual difficulty factors, hindering precise capability analysis. To advance TableQA beyond superficial task difficulty and toward interpretable capability modeling, we introduce MMTableBench, a multi-level benchmark that systematically evaluates MLLMs along two fine-grained dimensions: layout complexity and reasoning complexity. By organizing table-question pairs along these axes, MMTableBench facilitates a detailed evaluation of model performance under varying structural and reasoning challenges, while revealing the respective strengths and limitations of multimodal inputs. Our comprehensive analysis shows that state-of-the-art MLLMs continue to exhibit notable limitations when confronted with complex layouts and deep reasoning tasks, underscoring persistent gaps despite the structural advantages offered by visual inputs. MMTableBench thus provides not only a rigorous evaluation framework but also a diagnostic tool for analyzing and interpreting model behaviors, enabling more transparent and explainable progress in multimodal TableQA development. Xianjie Wu, Xiaohang Xu 0002, Tingyu Jiang, Jian Yang 0030, Di Liang, Xianfu Cheng, Zhenhe Wu, Linzheng Chai, Wei Zhang 0384, Ge Zhang 0009, Bob Simons, Tongliang Li, Zhoujun Li 0001 |
WWW | 6 |
| 2026 | A multi-scale representation and multi-level decision learning network for multimodal sentiment analysis
Xiang Li 0117, Zhiqiang Dong, Xianfu Cheng, Dezhuang Miao, Haijun Zhang 0007, Tianbo Wang 0001, Xiaoming Zhang 0001, Zhoujun Li 0001 |
Expert Syst. Appl. | 3 |
| 2025 | TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringabstractRecent advancements in Large Language Models (LLMs) have markedly enhanced the interpretation and processing of tabular data, introducing previously unimaginable capabilities. Despite these achievements, LLMs still encounter significant challenges when applied in industrial scenarios, particularly due to the increased complexity of reasoning required with real-world tabular data, underscoring a notable disparity between academic benchmarks and practical applications. To address this discrepancy, we conduct a detailed investigation into the application of tabular data in industrial scenarios and propose a comprehensive and complex benchmark TableBench, including 18 fields within four major categories of table question answering (TableQA) capabilities. Furthermore, we introduce TableLLM, trained on our meticulously constructed training set TableInstruct, achieving comparable performance with GPT-3.5. Massive experiments conducted on TableBench indicate that both open-source and proprietary LLMs still have significant room for improvement to meet real-world demands, where the most advanced model, GPT-4, achieves only a modest score compared to humans. Xianjie Wu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Xeron Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Tongliang Li, Zhoujun Li 0001, Guanglin Niu |
AAAI | 9 |
| 2025 | ECLIPSE: Efficient Cross-Lingual Log Intelligence Parser with Semantic Entropy-Enhanced LCS Algorithm
Wei Zhang 0384, Xianfu Cheng, Xiang Li 0117, Jian Yang 0030, Xiangyuan Guan, Zhoujun Li 0001 |
CIKM | 2 |
| 2025 | XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form ParserabstractIn the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task. The advent of pre-trained multimodal models significantly empowers Document AI frameworks to extract key information from form documents in different formats such as PDF, Word, and images. Nonetheless, form parsing is still encumbered by notable challenges like subpar capabilities in multilingual parsing and diminished recall in industrial contexts in rich text and rich visuals. In this work, we introduce a simple but effective Multimodal and Multilingual semi-structured FORM PARSER (XFormParser), which is anchored on a comprehensive Transformer-based pre-trained language model and innovatively amalgamates semantic entity recognition (SER) and relation extraction (RE) into a unified framework. Combined with Bi-LSTM, the performance of multilingual parsing is significantly improved. Furthermore, we develop InDFormSFT, a pioneering supervised fine-tuning (SFT) industrial dataset that specifically addresses the parsing needs of forms in a variety of industrial contexts. Through rigorous testing on established benchmarks, XFormParser has demonstrated its unparalleled effectiveness and robustness. Compared to existing state-of-the-art (SOTA) models, XFormParser notably achieves up to 1.79% F1 score improvement on RE tasks in language-specific settings. It also exhibits exceptional improvements in cross-task performance in both multilingual and zero-shot settings. Xianfu Cheng, Jian Yang 0030, Xiang Li 0117, Weixiao Zhou, Kui Wu 0007, Xiangyuan Guan, Tao Sun 0016, Xianjie Wu, Tongliang Li, Zhoujun Li 0001 |
COLING | 1 |
| 2025 | Breaking Size Barrier: Enhancing Reasoning for Large-Size Table Question Answering
Xianjie Wu, Di Liang, Jian Yang 0037, Xianfu Cheng, Linzheng Chai, Tongliang Li, Liqun Yang, Zhoujun Li 0001 |
DASFAA (2) | 4 |
| 2025 | SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsabstractThe increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work, we introduce SimpleVQA, the first comprehensive multi-modal benchmark to evaluate the factuality ability of MLLMs to answer natural language short questions. SimpleVQA is characterized by six key features: it covers multiple tasks and multiple scenarios, ensures high quality and challenging queries, maintains static and timeless reference answers, and is straightforward to evaluate. Our approach involves categorizing visual question-answering items into 9 different tasks around objective events or common knowledge and situating these within 9 topics. Rigorous quality control processes are implemented to guarantee high-quality, concise, and clear answers, facilitating evaluation with minimal variance via an LLM-as-a-judge scoring system. Using SimpleVQA, we perform a comprehensive assessment of leading 18 MLLMs and 8 text-only LLMs, delving into their image comprehension and text generation abilities by identifying and analyzing error cases. Xianfu Cheng, Wei Zhang 0384, Jian Yang 0030, Xiangyuan Guan, Xianjie Wu, Xiang Li 0117, Ge Zhang 0009, Yuying Mai, Yutao Zeng, Zhoufutu Wen, Baorui Wang, Weixiao Zhou, Yunhong Lu, Hangyuan Ji, Tongliang Li, Wenhao Huang 0001, Zhoujun Li 0001 |
ICCV | 1 |
| 2025 | Learning fine-grained representation with token-level alignment for multimodal sentiment analysis
Xiang Li 0117, Haijun Zhang 0007, Zhiqiang Dong, Xianfu Cheng, Yun Liu 0017, Xiaoming Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2024 | SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
Xianfu Cheng, Weixiao Zhou, Xiang Li 0117, Jian Yang 0030, Tao Sun 0016, Wei Zhang 0384, Yuying Mai, Tongliang Li, Xiaoming Chen 0007, Zhoujun Li 0001 |
CIKM | 1 |
| 2022 | An improved stochastic gradient descent algorithm based on Rényi differential privacyabstractDeep learning techniques based on the neural network have made significant achievements in various fields of artificial intelligence. However, model training requires large-scale data sets, these data sets are crowd-sourced and model parameters will contain the encoding of private information, resulting in the risk of privacy leakage. With the trend toward sharing pretrained models, the risk of stealing training data sets through member inference attacks and model inversion attacks is further heightened. To tackle the privacy-preserving problems in deep learning tasks, we propose an improved Differential Privacy Stochastic Gradient Descent algorithm, using Simulated Annealing algorithm and Laplace Smooth denoising mechanism to optimize the allocation method of privacy loss, replacing the constant clipping method with adaptive gradient clipping method to improve model accuracy. we also analyze privacy cost under random shuffle data batch processing method in detail within the framework of Subsampled Rényi Differential Privacy. Compared with the existing privacy protection training methods with fixed parameters and dynamic privacy parameters in classification tasks, our implementation and experiments show that we can use less privacy budget train deep neural networks with the nonconvex objective function, obtain a higher model evaluation, and have almost zero additional cost in terms of model complexity, training efficiency, and model quality. Xianfu Cheng, Ao Liu 0006, Zhoujun Li 0001 |
Int. J. Intell. Syst. | 1 |
| 2021 | An integrated product modularity method based on transfer network of failure mode-recycling decision for remanufacturingabstractAbstract The product modularity has an important influence on the whole product life cycle. In order to improve product recovery of a used product at the end‐of‐life stage, it is imperative to develop modular product architecture with remanufacturing strategies consideration. In this paper, an integrated modular design method for green remanufacturing considering hierarchical structure of the product and fuzzy logic is proposed, and the modularity of the architecture from the viewpoint of product remanufacturing is assessed. The transfer network of failure mode‐recycling decision using BP neural network optimized by adaptive genetic algorithm is built. Design structure matrix is employed to represent the relationships between components, and the hierarchical structure of product modularity with functionality, physical connection as well as geometric position is constructed. The failure modes of used components are analyzed, and four modular drivers with remanufacturing strategies consideration, namely, recycling mode, component lifetime, material compatibility, and remanufacturing processability, are introduced to reconstruct the modular product architecture for remanufacturing. Finally, an example on the grab is provided to illustrate the proposed method for product modularity taking into account function, structure as well as remanufacturing characteristics. Xianfu Cheng, Minhua You, Zhihu Guo |
Concurr. Comput. Pract. Exp. | 1 |