EDBT 2026 Demo / reviewers in the wild / expert
Xiaohang Xu 0002
dblp:171/2451-2
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-1266-9943ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 47% Spatial and temporal data management · 27% Data mining · 18% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 32% Language models and text generation · 25% Planning, search and constraint satisfaction · 16% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization |
1.0 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
GUI automation |
1.0 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Natural language and speech › Language models and text generation
instruction tuning |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Information retrieval › evaluation › benchmark
benchmark construction |
1.0 | 1 | 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA · WWW 2026 |
Information retrieval
evaluation |
1.0 | 1 | 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA · WWW 2026 |
Information retrieval
question answering |
1.0 | 1 | 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA · WWW 2026 |
Information retrieval › question answering
table question answering |
1.0 | 1 | 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA · WWW 2026 |
Machine learning › Representation and self-supervised learning
similarity structure learning |
0.8 | 1 | 2024 | SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory Similarity · Proc. VLDB Endow. 2024 |
Data mining › spatiotemporal data mining
human mobility prediction |
0.8 | 1 | 2024 | Taming the Long Tail in Human Mobility Prediction · NeurIPS 2024 |
Spatial and temporal data management › trajectory data management › trajectory similarity
learning-based trajectory similarity |
0.8 | 1 | 2024 | SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory Similarity · Proc. VLDB Endow. 2024 |
Data mining › predictive modeling › classification › imbalanced classification
long-tail learning |
0.8 | 1 | 2024 | Taming the Long Tail in Human Mobility Prediction · NeurIPS 2024 |
Recommender systems › point-of-interest recommendation
next POI recommendation |
0.8 | 1 | 2024 | Taming the Long Tail in Human Mobility Prediction · NeurIPS 2024 |
Spatial and temporal data management
trajectory data management |
0.8 | 1 | 2024 | SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory Similarity · Proc. VLDB Endow. 2024 |
Spatial and temporal data management › trajectory data management
trajectory similarity |
0.8 | 1 | 2024 | SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory Similarity · Proc. VLDB Endow. 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Natural language and speech › Language models and text generation › LLM agents
multimodal large language model agent |
0.3 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
transformer encoder · 1.5representation similarity functions · 1.5self-play · 1.0multimodal large language model · 1.0model instruction weakness value · 1.0importance scoring · 1.0group relative policy optimization · 1.0data distillation · 1.0loss reweighting · 0.8graph neural network · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Importance-Aware Data Selection for Efficient LLM Instruction TuningabstractInstruction tuning plays a critical role in enhancing the performance and efficiency of Large Language Models (LLMs). Its success depends not only on the quality of the instruction data but also on the inherent capabilities of the LLM itself. Some studies suggest that even a small amount of high-quality data can achieve instruction fine-tuning results that are on par with, or even exceed, those from using a full-scale dataset. However, rather than focusing solely on calculating data quality scores to evaluate instruction data, there is a growing need to select high-quality data that maximally enhances the performance of instruction tuning for a given LLM. In this paper, we propose the Model Instruction Weakness Value (MIWV) as a novel metric to quantify the importance of instruction data in enhancing model's capabilities. The MIWV metric is derived from the discrepancies in the model’s responses when using In-Context Learning (ICL), helping identify the most beneficial data for enhancing instruction tuning performance. Our experimental results demonstrate that selecting only the top 1% of data based on MIWV can outperform training on the full dataset. Furthermore, this approach extends beyond existing research that focuses on data quality scoring for data selection, offering strong empirical evidence supporting the effectiveness of our proposed method. Tingyu Jiang, Yiyao Song, Hualei Zhu, Xiaohang Xu 0002, Kenjiro Taura, Hao Henry Wang |
AAAI | 7 |
| 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI AgentsabstractGraphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current methodologies exhibit two fundamental limitations: (1) insufficient exploitation of cross-model synergies, and (2) over-reliance on synthetic data generation without sufficient utilization. To address these challenges, we propose Co-EPG, a self-iterative training framework for Co-Evolution of Planning and Grounding. Co-EPG establishes an iterative positive feedback loop: through this loop, the planning model explores superior strategies under grounding-based reward guidance via Group Relative Policy Optimization (GRPO), generating diverse data to optimize the grounding model. Concurrently, the optimized Grounding model provides more effective rewards for subsequent GRPO training of the planning model, fostering continuous improvement. Co-EPG thus enables iterative enhancement of agent capabilities through self-play optimization and training data distillation. On the Multimodal-Mind2Web and AndroidControl benchmarks, our framework outperforms existing state-of-the-art methods after just three iterations without requiring external data. The agent consistently improves with each iteration, demonstrating robust self-enhancement capabilities. This work establishes a novel training paradigm for GUI agents, shifting from isolated optimization to an integrated, self-driven co-evolution approach. Hualei Zhu, Tingyu Jiang, Xiaohang Xu 0002, Hao Henry Wang |
AAAI | 5 |
| 2026 | MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QAabstractTables serve as a core format for representing structured data on the web, as their two-dimensional layouts effectively encode complex inter-entity relationships. However, real-world web tables often feature heterogeneous structures and rich semantics. Accurately interpreting such tables requires not only spatial layout perception but also multi-step reasoning across rows and columns, posing substantial challenges to web intelligence systems. Multimodal large language models (MLLMs) show promise in table question answering (TableQA) by leveraging visual layouts. However, their performance on complex web tables remains uneven, as existing benchmarks often blur the impact of individual difficulty factors, hindering precise capability analysis. To advance TableQA beyond superficial task difficulty and toward interpretable capability modeling, we introduce MMTableBench, a multi-level benchmark that systematically evaluates MLLMs along two fine-grained dimensions: layout complexity and reasoning complexity. By organizing table-question pairs along these axes, MMTableBench facilitates a detailed evaluation of model performance under varying structural and reasoning challenges, while revealing the respective strengths and limitations of multimodal inputs. Our comprehensive analysis shows that state-of-the-art MLLMs continue to exhibit notable limitations when confronted with complex layouts and deep reasoning tasks, underscoring persistent gaps despite the structural advantages offered by visual inputs. MMTableBench thus provides not only a rigorous evaluation framework but also a diagnostic tool for analyzing and interpreting model behaviors, enabling more transparent and explainable progress in multimodal TableQA development. Xianjie Wu, Xiaohang Xu 0002, Tingyu Jiang, Jian Yang 0030, Di Liang, Xianfu Cheng, Zhenhe Wu, Linzheng Chai, Wei Zhang 0384, Ge Zhang 0009, Bob Simons, Tongliang Li, Zhoujun Li 0001 |
WWW | 2 |
| 2024 | Taming the Long Tail in Human Mobility PredictionabstractWith the popularity of location-based services, human mobility prediction plays a key role in enhancing personalized navigation, optimizing recommendation systems, and facilitating urban mobility and planning. This involves predicting a user's next POI (point-of-interest) visit using their past visit history. However, the uneven distribution of visitations over time and space, namely the long-tail problem in spatial distribution, makes it difficult for AI models to predict those POIs that are less visited by humans. In light of this issue, we propose the $\underline{\bf{Lo}}$ng-$\underline{\bf{T}}$ail Adjusted $\underline{\bf{Next}}$ POI Prediction (LoTNext) framework for mobility prediction, combining a Long-Tailed Graph Adjustment module to reduce the impact of the long-tailed nodes in the user-POI interaction graph and a novel Long-Tailed Loss Adjustment module to adjust loss by logit score and sample weight adjustment strategy. Also, we employ the auxiliary prediction task to enhance generalization and accuracy. Our experiments with two real-world trajectory datasets demonstrate that LoTNext significantly surpasses existing state-of-the-art works. Xiaohang Xu 0002, Renhe Jiang, Chuang Yang 0002, Zipei Fan, Kaoru Sezaki |
NeurIPS | 1 |
| 2024 | SIMformer: Single-Layer Vanilla Transformer Can Learn Free-Space Trajectory SimilarityabstractFree-space trajectory similarity calculation, e.g., DTW, Hausdorff, and Fréchet, often incur quadratic time complexity, thus learning-based methods have been proposed to accelerate the computation. The core idea is to train an encoder to transform trajectories into representation vectors and then compute vector similarity to approximate the ground truth. However, existing methods face dual challenges of effectiveness and efficiency: 1) they all utilize Euclidean distance to compute representation similarity, which leads to the severe curse of dimensionality issue - reducing the distinguishability among representations and significantly affecting the accuracy of subsequent similarity search tasks; 2) most of them are trained in triplets manner and often necessitate additional information which downgrades the efficiency; 3) previous studies, while emphasizing the scalability in terms of efficiency, overlooked the deterioration of effectiveness when the dataset size grows. To cope with these issues, we propose a simple, yet accurate, fast, scalable model that only uses a single-layer vanilla transformer encoder as the feature extractor and employs tailored representation similarity functions to approximate various ground truth similarity measures. Extensive experiments demonstrate our model significantly mitigates the curse of dimensionality issue and outperforms the state-of-the-arts in effectiveness, efficiency, and scalability. Chuang Yang 0002, Renhe Jiang, Xiaohang Xu 0002, Chuan Xiao 0001, Kaoru Sezaki |
Proc. VLDB Endow. | 3 |
| 2023 | Revisiting Mobility Modeling with Graph: A Graph Transformer Model for Next Point-of-Interest RecommendationabstractNext Point-of-Interest (POI) recommendation plays a crucial role in urban mobility applications. Recently, POI recommendation models based on Graph Neural Networks (GNN) have been extensively studied and achieved, however, the effective incorporation of both spatial and temporal information into such GNN-based models remains challenging. Temporal information is extracted from users' trajectories, while spatial information is obtained from POIs. Extracting distinct fine-grained features unique to each piece of information is difficult since temporal information often includes spatial information, as users tend to visit nearby POIs. To address the challenge, we propose Mobility Graph Transformer (MobGT) that enables us to fully leverage graphs to capture both the spatial and temporal features in users' mobility patterns. MobGT combines individual spatial and temporal graph encoders to capture unique features and global user-location relations. Additionally, it incorporates a mobility encoder based on Graph Transformer to extract higher-order information between POIs. To address the long-tailed problem in spatial-temporal data, MobGT introduces a novel loss function, Tail Loss. Experimental results demonstrate that MobGT outperforms state-of-the-art models on various datasets and metrics, achieving 24% improvement on average. Our codes are available at https://github.com/Yukayo/MobGT. Xiaohang Xu 0002, Toyotaro Suzumura, Jiawei Yong, Masatoshi Hanai, Chuang Yang 0002, Hiroki Kanezashi, Renhe Jiang, Shintaro Fukushima |
SIGSPATIAL/GIS | 1 |
| 2022 | Privacy-Preserving Federated Depression Detection From Multisource Mobile Health DataabstractDepression is one of the most common mental illnesses, and the symptoms shown by patients are different, making it difficult to diagnose in the process of clinical practice and pathological research. Although researchers hope that artificial intelligence can contribute to the diagnosis and treatment of depression, the traditional centralized machine learning methods need to aggregate patient data, and the data privacy of patients with mental illness needs to be strictly confidential, which hinders machine learning algorithms’ clinical application. To solve the problem of medical data privacy with depression, in this article, we implement a study of federated learning to analyze and diagnose depression. First, we propose a general multiview federated learning framework using multisource data, which can extend any traditional machine learning model to support federated learning across different institutions or parties. Second, we employ later fusion methods to solve the problem of inconsistent time series of multiview data. Finally, we compare the federated framework with other cooperative learning frameworks in performance and discuss the related results. The experimental results show that in the case of participating in federated learning with enough participants, the prediction accuracy of depression score can reach 85.13%, which is about 15% higher than local training. When the number of participants is small and the amount of data is sufficient, the prediction accuracy of depression score can also reach 84.32%, and the improvement rate is about 9%. Xiaohang Xu 0002, Hao Peng 0001, Md. Zakirul Alam Bhuiyan, Zhifeng Hao 0004, Lianzhong Liu, Lichao Sun 0001, Lifang He 0001 |
IEEE Trans. Ind. Informatics | 1 |