EDBT 2026 Demo / reviewers in the wild / expert
Xichen Ding
dblp:244/9767
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationabstractAerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agents typically adopt mono-granularity frameworks that struggle to balance these two aspects. To address this limitation, this work proposes a History-Enhanced Two-Stage Transformer (HETT) framework, which integrates the two aspects through a coarse-to-fine navigation pipeline. Specifically, HETT first predicts coarse-grained target positions by fusing spatial landmarks and historical context, then refines actions via fine-grained visual analysis. In addition, a historical grid map is designed to dynamically aggregate visual features into a structured spatial memory, enhancing comprehensive scene awareness. Additionally, the CityNav dataset annotations are manually refined to enhance data quality. Experiments on the refined CityNav dataset show that HETT delivers significant performance gains, while extensive ablation studies further verify the effectiveness of each component. Xichen Ding, Jianzhe Gao, Wenguan Wang |
AAAI | 1 |
| 2026 | PRISM: Preference-Guided Semantic Reasoning with Vision-Language Models for Object Goal NavigationabstractObject Goal Navigation (ObjectNav) requires an agent to locate a target object in unseen environments based on partial visual observations. A central challenge lies in reasoning about where the target is likely to appear under semantic uncertainty, sparse observations, and ambiguous scene cues. Existing approaches often rely on extensive training or static semantic priors, limiting adaptability and interpretability in unfamiliar environments. We introduce PRISM, a preference-guided semantic reasoning and mapping framework that integrates vision-language models (VLMs) into the navigation loop to support efficient semantic object search. PRISM constructs a dynamic Preference-guided Semantic Map (PSM) that aggregates observed semantics, predicted unobserved regions, and language-mediated semantic reasoning, enabling the agent to incrementally refine its belief over likely target locations. Based on the PSM, PRISM further employs Adaptive Preference Exploration Strategies (APES) to adjust exploration behavior using contextual semantic cues and historical feedback under semantic ambiguity. Experiments on the Gibson, Matterport3D (MP3D), and Habitat-Matterport3D (HM3D) datasets show that our method achieves consistent improvements over prior training-free approaches, yielding absolute gains of +2.7% SR on Gibson, +2.8% SR on MP3D, and +3.2% SR on HM3D. Cong Pan 0001, Chengjie Fan, Wanjie Cai, Xichen Ding, Jie Qin 0004 |
ICMR | 5 |
| 2024 | An efficient algorithm for optimal route node sensing in smart tourism Urban traffic based on priority constraints
Xichen Ding, Rongju Yao, Edris Khezri |
Wirel. Networks | 1 |
| 2023 | An Unified Search and Recommendation Foundation Model for Cold-Start ScenarioabstractIn modern commercial search engines and recommendation systems, data from multiple domains is available to jointly train the multi-domain model. Traditional methods train multi-domain models in the multi-task setting, with shared parameters to learn the similarity of multiple tasks, and task-specific parameters to learn the divergence of features, labels, and sample distributions of individual tasks. With the development of large language models, LLM can extract global domain-invariant text features that serve both search and recommendation tasks. We propose a novel framework called S&R Multi-Domain Foundation, which uses LLM to extract domain invariant features, and Aspect Gating Fusion to merge the ID feature, domain invariant text features and task-specific heterogeneous sparse features to obtain the representations of query and item. Additionally, samples from multiple search and recommendation scenarios are trained jointly with Domain Adaptive Multi-Task module to obtain the multi-domain foundation model. We apply the S&R Multi-Domain foundation model to cold start scenarios in the pretrain-finetune manner, which achieves better performance than other SOTA transfer learning methods. The S&R Multi-Domain Foundation model has been successfully deployed in Alipay Mobile Application's online services, such as content query recommendation and service card recommendation, etc. Yuqi Gong, Xichen Ding, Yehui Su, Kaiming Shen, Zhongyi Liu 0001 |
CIKM | 2 |
| 2022 | Prototypical Contrastive Learning and Adaptive Interest Selection for Candidate Generation in RecommendationsabstractDeep Candidate Generation plays an important role in large-scale recommender systems. It takes user history behaviors as inputs and learns user and item latent embeddings for candidate generation. In the literature, conventional methods suffer from two problems. First, a user has multiple embeddings to reflect various interests, and such number is fixed. However, taking into account different levels of user activeness, a fixed number of interest embeddings is sub-optimal. For example, for less active users, they may need fewer embeddings to represent their interests compared to active users. Second, the negative samples are often generated by strategies with unobserved supervision, and similar items could have different labels. Such a problem is termed as class collision. In this paper, we aim to advance the typical two-tower DNN candidate generation model. Specifically, an Adaptive Interest Selection Layer is designed to learn the number of user embeddings adaptively in an end-to-end way, according to the level of their activeness. Furthermore, we propose a Prototypical Contrastive Learning Module to tackle the class collision problem introduced by negative sampling. Extensive experimental evaluations show that the proposed scheme remarkably outperforms competitive baselines on multiple benchmarks. Qunwei Li, Xichen Ding, Shaohu Chen, Leon Wenliang Zhong |
CIKM | 3 |
| 2022 | Device-cloud Collaborative Recommendation via Meta ControllerabstractOn-device machine learning enables the lightweight deployment of recommendation models in local clients, which reduces the burden of the cloud-based recommenders and simultaneously incorporates more real-time user features. Nevertheless, the cloud-based recommendation in the industry is still very important considering its powerful model capacity and the efficient candidate generation from the billion-scale item pool. Previous attempts to integrate the merits of both paradigms mainly resort to a sequential mechanism, which builds the on-device recommender on top of the cloud-based recommendation. However, such a design is inflexible when user interests dramatically change: the on-device model is stuck by the limited item cache while the cloud-based recommendation based on the large item pool do not respond without the new re-fresh feedback. To overcome this issue, we propose a meta controller to dynamically manage the collaboration between the on-device recommender and the cloud-based recommender, and introduce a novel efficient sample construction from the causal perspective to solve the dataset absence issue of meta controller. On the basis of the counterfactual samples and the extended training, extensive experiments in the industrial recommendation scenarios show the promise of meta controller in the device-cloud collaboration. Jiangchao Yao, Feng Wang 0072, Xichen Ding, Shaohu Chen, Bo Han 0003, Jingren Zhou 0001, Hongxia Yang |
KDD | 3 |
| 2019 | Infer Implicit Contexts in Real-time Online-to-Offline RecommendationabstractUnderstanding users' context is essential for successful recommendations, especially for Online-to-Offline (O2O) recommendation, such as Yelp, Groupon, and Koubei. Different from traditional recommendation where individual preference is mostly static, O2O recommendation should be dynamic to capture variation of users' purposes across time and location. However, precisely inferring users' real-time contexts information, especially those implicit ones, is extremely difficult, and it is a central challenge for O2O recommendation. In this paper, we propose a new approach, called Mixture Attentional Constrained Denoise AutoEncoder (MACDAE), to infer implicit contexts and consequently, to improve the quality of real-time O2O recommendation. In MACDAE, we first leverage the interaction among users, items, and explicit contexts to infer users' implicit contexts, then combine the learned implicit-context representation into an end-to-end model to make the recommendation. MACDAE works quite well in the real system. We conducted both offline and online evaluations of the proposed approach. Experiments on several real-world datasets (Yelp, Dianping, and Koubei) show our approach could achieve significant improvements over state-of-the-arts. Furthermore, online A/B test suggests a 2.9% increase for click-through rate and 5.6% improvement for conversion rate in real-world traffic. Our model has been deployed in the product of "Guess You Like" recommendation in Koubei. Xichen Ding, Jie Tang 0001, Tracy Xiao Liu, Qixia Jiang |
KDD | 1 |