VLDB 2026 Research / reviewers in the wild / expert
Jingyang Deng
dblp:368/3826
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0008-3487-225XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TopSUMseg: A Topology-Aware Swin Transformer-Mamba Framework for 3D Seismic Fault Image SegmentationabstractSeismic fault image segmentation is crucial for interpreting subsurface geological structures, supporting geologists in resource exploration and structural analysis. However, current deep learning models struggle with single-architecture limitations and the distinctive characteristics of seismic faults, which are distinguished by elongated structures with uneven spatial distributions. To address these challenges, we propose TopSUMseg, a novel Topology-Aware Swin Transformer-Mamba framework for 3D seismic fault image segmentation. Our framework combines Swin Transformer’s local feature extraction with Mamba’s efficient sequence modeling, and boosts 3D spatial modeling in Mamba with a newly designed Global-Local Attention module (GLA). Additionally, we design a Topology-Aware Structural Constraint (TASC) to align predictions with ground-truth structures in the feature space, promoting the modeling of complex fault geometries. Experiments on Thebe, the largest public seismic dataset, demonstrate that TopSUMseg achieves state-of-the-art performance with OIS and ODS scores of 0.879 and 0.875, respectively. Trained entirely from scratch, TopSUMseg nonetheless achieves superior performance compared to extensively pre-trained counterparts. In addition, TopSUMseg maintains a significantly lower parameter count while achieving a favorable trade-off between segmentation performance and time complexity, making it a practical and generalizable solution for real-world seismic fault interpretation. Ran Chen 0002, Jingyang Deng, Zeren Zhang, Ruohua Shi, Jinwen Ma |
ECAI | 2 |
| 2025 | Reframing Multimodal Complex Document Layout Understanding: A Layout-Aware Multi-Source Reasoning Decision FrameworkabstractMultimodal large language models (MLLMs) have achieved significant progress in document understanding. However, complex layout reasoning, characterized by concise answers and cross-page integration, remains a challenge. Unlike conventional semantics-oriented tasks, this task demands accurate visual perception of fine-grained structural elements and logical reasoning across multi-page documents. Existing approaches primarily focus on information extraction and semantic understanding, limiting the capacity of fine-tuned autoregressive models to capture short-answer reasoning signals and generalize to complex layout structures. To address this, we propose the Layout-Aware Multi-Source Reasoning Decision Framework (LAMRD), which reframes complex layout reasoning as a decision-making task over multi-source reasoning paths. In the reasoning path construction stage, LAMRD generates layout-aware reasoning paths by integrating internal visual cues and external knowledge from three complementary perspectives: Visual Structural Awareness (VSA), Logical Reasoning Paths (LRP), and External Knowledge Augmentation (EKA). In the reasoning path decision stage, we employ Group Relative Policy Optimization (GRPO) to train a decision model that produces the final answer based on these paths. We conduct comprehensive evaluations using Qwen2.5-VL-7B-Instruct on the CEP-7K dataset, covering layout structure understanding, information extraction, and logical association. Experimental results demonstrate that LAMRD outperforms advanced MLLMs in accuracy, validating its effectiveness for complex document layout understanding. Ran Chen 0002, Jingyang Deng, Zeren Zhang, Xuefei Tong, Jinwen Ma, Qinghui Shi, Yuanjun Li |
ECAI | 3 |
| 2025 | Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference AlignmentabstractIn handling complex, domain-specific tasks, particularly in the context of state-owned assets and enterprises (SOAEs), general LLMs suffer from the knowledge gap due to insufficient exposure to domain-specific corpora, and the value disagreement, as they are aligned with universal values rather than domain-specific ones. To tackle these challenges, we propose a novel training strategy tailored for the SOAEs domain. This strategy includes a improved domain-adaptive pretraining (DAP) phase with a replay mechanism to mitigate catastrophic forgetting. Following DAP, we utilize a selective portion of domain-specific data for supervised fine-tuning (SFT), and innovatively integrate low-quality data with the remaining SFT data to curate tailored preference datasets, leveraging the Kahneman-Tversky Optimization technique to align our LLMs. Our proposed approach effectively utilizes the data that is often discarded in conventional training procedures, highlighting the substantial improvements in model performance and the importance of training methodologies for domain-specific tasks. Jingyang Deng, Zeren Zhang, Jo-Ku Cheng, Jinwen Ma |
ICASSP | 1 |
| 2025 | Diagram Formalization Enhanced Multi-Modal Geometry Problem SolverabstractMathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems, which require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geometry problem-solving than LLMs that only process text. This limitation is further amplified by the lack of effective methods for representing geometric relationships. To address these issues, we introduce the Diagram Formalization Enhanced Geometry Problem Solver (DFE-GPS), a new framework that integrates visual features, geometric formal language, and natural language representations. Specifically, we propose a novel synthetic data approach and construct a large-scale geometric dataset, SynthGeo228K, annotated with formal and natural language captions, designed to enhance the vision encoder to understand geometric structures better. Our framework improves MLLMs’ ability to process geometric diagrams and extends their application to open-ended tasks on the formalgeo7k dataset. Zeren Zhang, Jo-Ku Cheng, Jingyang Deng, Jinwen Ma, Ziran Qin, Tuo Leng |
ICASSP | 3 |
| 2025 | GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions
Jo-Ku Cheng, Zeren Zhang, Ran Chen 0002, Jingyang Deng, Ziran Qin, Jinwen Ma |
ACM Multimedia | 4 |
| 2024 | FltLM: An Intergrated Long-Context Large Language Model for Effective Context Filtering and UnderstandingabstractThe development of Long-Context Large Language Models (LLMs) has markedly advanced natural language processing by facilitating the process of textual data across long documents and multiple corpora. However, Long-Context LLMs still face two critical challenges: The lost in the middle phenomenon, where crucial middle-context information is likely to be missed, and the distraction issue that the models lose focus due to overly extended contexts. To address these challenges, we propose the Context Filtering Language Model (FltLM), a novel integrated Long-Context LLM which enhances the ability of the model on multi-document question-answering (QA) tasks. Specifically, FltLM innovatively incorporates a context filter with a soft mask mechanism, identifying and dynamically excluding irrelevant content to concentrate on pertinent information for better comprehension and reasoning. Our approach not only mitigates these two challenges, but also enables the model to operate conveniently in a single forward pass. Experimental results demonstrate that FltLM significantly outperforms supervised fine-tuning and retrieval-based methods in complex QA scenarios, suggesting a promising solution for more accurate and reliable long-context natural language understanding applications. Jingyang Deng, Zhengyang Shen, Lixin Su, Suqi Cheng, Ying Nie 0006, Junfeng Wang 0009, Dawei Yin 0001, Jinwen Ma |
ECAI | 1 |
| 2024 | Geometry-Guided Conditional Adaptation for Surrogate Models of Large-Scale 3D PDEs on Arbitrary Geometries
Jingyang Deng, Xingjian Li 0002, Haoyi Xiong, Xiaoguang Hu, Jinwen Ma |
IJCAI | 1 |