EDBT 2026 Demo / reviewers in the wild / expert
Jiayi Xie
dblp:231/7875
· DBLP profile ↗
19ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive VideosabstractPerception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-world video data in adverse weather. Existing weather generation approaches struggle to balance visual quality and annotation reusability. We present AutoAWG, a controllable Adverse Weather video Generation framework for Autonomous driving. Our method employs a semantics-guided adaptive fusion of multiple controls to balance strong weather stylization with high-fidelity preservation of safety-critical targets; leverages a vanishing point-anchored temporal synthesis strategy to construct training sequences from static images, thereby reducing reliance on synthetic data; and adopts masked training to enhance long-horizon generation stability. On the nuScenes validation set, AutoAWG significantly outperforms prior state-of-the-art methods: without first-frame conditioning, FID and FVD are relatively reduced by 50.0% and 16.1%; with first-frame conditioning, they are further reduced by 8.7% and 7.2%, respectively. Extensive qualitative and quantitative results demonstrate advantages in style fidelity, temporal consistency, and semantic–structural integrity, underscoring the practical value of AutoAWG for improving downstream perception in autonomous driving. Our code is available at: https://github.com/higherhu/AutoAWG Jiagao Hu, Daiguo Zhou, Danzhen Fu, Fuhao Li, Fei Wang 0032, Wenhua Liao, Jiayi Xie |
ICMR | 8 |
| 2026 | DarkVision: A Benchmark and Study for Low-Light Image/Video AnalysisabstractLow-light image/video analysis is essential for various applications, e.g., night surveillance and photography, high-speed imaging, and autonomous vehicles. Under such conditions, cameras suffer from low signal-to-noise ratio, which degrades image quality severely and poses challenges for downstream tasks such as object detection. Data-driven methods have achieved enormous success for normal-light image/video restoration and high-level vision tasks. However, the lack of a high-quality benchmark dataset with accurate semantic annotations for low-light images and especially videos greatly hinders research progress. In this paper, we contribute the first multi-illuminance, multi-camera, low-light dataset, DarkVision, serving both image/video enhancement and object detection applications. We provide bright and dark pairs with pixel-wise registration, in which the bright counterpart provides a reliable reference for enhancement and annotation. This dataset comprises 13,455 images of 900 static scenes with objects from 15 categories, and 89,411 frames of 32 dynamic scenes with 4 categories of objects. For each scene, images/videos were captured at 5 illuminance levels using three cameras of different quality grades; average photon numbers can be reliably estimated from the calibration curves for quantitative studies. The static images and dynamic videos respectively contain around 7344 and 320,667 object instances in total. With DarkVision, we establish baselines for image/video enhancement and object detection by representative algorithms. To demonstrate an exemplary application of DarkVision, we propose two simple yet effective approaches to improve the performance of video enhancement and object detection respectively by exploiting temporal cues. Furthermore, we study the relationship between image enhancement and object detection. We believe DarkVision can help to advance the state-of-the art in both low-light image/video enhancement and object detection, as well as benefiting cross-task studies. Bo Zhang 0109, Runzhao Yang, Zhihong Zhang 0004, Jiayi Xie, Jin-Li Suo |
Comput. Vis. Media | 5 |
| 2025 | Contrastive Learning-Based Feature Modulation Strategy for Test-Time Adaptation in Medical Image SegmentationabstractMedical image segmentation plays a critical role in various clinical applications, including organ delineation, tumor detection, and surgical planning. However, deploying segmentation models in real-world clinical environments remains challenging due to domain shifts between the training and test data, often leading to performance degradation. To address these challenges, we introduce contrastive learning into Test-Time Adaptation (TTA) for medical image segmentation. By leveraging data augmentation to generate new data sources and calculating NT-Xent loss, we enhance feature representation learning, improving model robustness against complex distribution changes. Furthermore, we propose a robust feature modulation strategy (FMS) comprising Enhanced Feature Optimization (EFO) and Selective Feature Regularization (SFR). This strategy not only improves the model's adaptability but also mitigates the inaccuracies in edge segmentation caused by entropy minimization. We rigorously evaluate our approach to segmentation and classification tasks across two different medical imaging modalities. Experimental results demonstrate the versatility of our method across multiple network architectures, achieving measurable performance improvements and providing a reliable solution for medical image segmentation in dynamic environments. Yunze Bi, Jiayi Xie |
CSCWD | 2 |
| 2025 | Contrastive Learning with Positive Augmentation in Open World Test-Time TrainingabstractTest-time training (TTT) seeks to adapt models trained on a source domain to unseen target domains during testing, addressing challenges of distribution shifts. While previous studies have concentrated on improving model performance against common corruptions and exploring robustness in real-world scenarios, the issue of generalizing pre-trained models to target domains comprising both corrupted and semantically distinct data remains underexplored. This paper introduces a contrastive test-time training approach that incorporates sample identification and discriminability regularization to address these challenges. We utilize entropy as an indicator to selectively filter strong out-of-distribution (OOD) data from test samples, enhancing model adaptation to target corrupted samples. By employing contrastive representation learning, our approach improves feature effectiveness and model robustness. Additionally, we define a feature diversity regularization loss to balance the model's transfer and discriminative capabilities. Experimental results demonstrate that our method surpasses state-of-the-art techniques in three evaluation metrics across multiple benchmark datasets. Jiayi Xie, Yunze Bi |
CSCWD | 1 |
| 2025 | State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect SimulatorabstractIn reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL to jointly learn transferable policies from limited offline data and imperfect simulators. However, due to the unrestricted exploration in the imperfect simulator, the hybrid offline-and-online RL methods inevitably suffer from low sample efficiency and insufficient state-action space coverage during training. To solve this problem, we propose a State Revisit and Re-exploration (SR2) hybrid offline-and-online RL framework. In particular, the proposed algorithm employs a meta-policy and a sub-policy, where the meta-policy aims to find high-quality states in the offline trajectories for online exploration, and the sub-policy learns the robot skill using mixed offline and online data. By introducing the state revisit and explore mechanism, our approach efficiently improves performance on a set of sim-to-real robotic tasks. Through extensive simulation and real-world tasks, we demonstrate the superior performance of our approach against other state-of-the-art methods. Xingyu Chen 0001, Jiayi Xie, Ruixun Liu, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan |
IJCAI | 2 |
| 2025 | SEP: Self-Rewarded Entity-Level Preference for Open Named Entity RecognitionabstractWith the robust capabilities of recent open Named Entity Recognition (NER) model, Direct Preference Optimization (DPO) through self-training emerges as a promising approach for expanding the application of NER tasks. However, the preference annotations from AI feedback do not fully capture high-quality and fine-grained response in the NER task. To address this issue, we propose SEP, a novel preference learning framework that focuses on Self-Rewarded Entity-Level Preference for the open NER task. Our work employs the NER model itself to automatically annotate entity-level pair for preference optimization. Furthermore, self-rewarding strategy via log probability is explored to score the response, aiming to reduce the impact of low-quality annotations in constructing training data. Based on the response and their reward, multiple entity-level preference pairs via a coarse-to-fine mechanism are constructed to pinpoint detailed wrong entities in response. Experimental results show that SEP achieves significant improvements over vanilla DPO across a benchmark of 20 NER datasets. Ablation study underscores the critical roles of self-generating response rewarding and entity-aware preference learning. Further exploration confirms consistent improvements across multiple iterations and reflects the impact of data size and data domain in the self-training NER task. Jiayi Xie, Tingyu Xie |
IJCNN | 2 |
| 2025 | ECS-Net: Extracellular space segmentation with contrastive and shape-aware loss by using cryo-electron microscopy imaging
Chuqiao Yang, Jiayi Xie, Xinrui Huang, Hanbo Tan, Qirun Li, Zeqing Tang, Xinlei Ma, Jiabin Lu, Qingyuan He, Wanyi Fu, Yixing Huang, Junhao Yan, Zhaoheng Xie, Yao Sui, Yanye Lu, Hongbin Han |
Expert Syst. Appl. | 2 |
| 2025 | Disentangling User Interest and Geographical Context for POI RecommendationsabstractPOI recommendation plays an important role in many applications, such as mobility prediction and location-based advertisements. Existing POI recommendation methods mainly capture the observed patterns in user visits for recommendations, without a comprehensive consideration of the underlying reasons behind the visits. Therefore, different causes of a visit, i.e., users’ interest and geographical context, are entangled. When the underlying causes change (e.g., when a user moves to a new place), the robustness of the recommendations cannot be guaranteed. To address the above challenges, we propose DUIG, a novel user interest and geographical influences disentanglement framework for POI recommendations. We first design a personalized disentanglement strategy to divide check-ins through geographical influence. Specifically, the colliding effect of causality is leveraged to the divide cause-specific check-ins, such that user interest and geographical influence can be properly disentangled in user and POI embeddings. Through this mechanism, even if the underlying reasons that affect a user’s preference change, intervention can be conducted upon the causes to make recommendations generalized to the new scenario. In addition, a geographical-aware negative sampling strategy is proposed to utilize hard negatives to regularize the embedding and disentanglement in the latent space, where a larger sampling probability is introduced for negative samples containing more geographic information. Extensive experiments on two real-world POI recommendation datasets demonstrate the superior performance of DUIG. Wenhui Meng, Jiayi Xie, Jing Yi, Yaochen Zhu, Zhenzhong Chen 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | UnifiedSSR: A Unified Framework of Sequential Search and RecommendationabstractIn this work, we propose a Unified framework of Sequential Search and Recommendation (UnifiedSSR) for joint learning of user behavior history in both search and recommendation scenarios. Specifically, we consider user-interacted products in the recommendation scenario, as well as user-interacted products and user-issued queries in the search scenario as three distinct types of user behaviors. We propose a dual-branch network to encode the pair of interacted product history and issued query history in the search scenario in parallel. This allows for cross-scenario modeling by deactivating the query branch for the recommendation scenario. Through the parameter sharing between dual branches, as well as between product branches in two scenarios, we incorporate cross-view and cross-scenario associations of user behaviors, providing a comprehensive understanding of user behavior patterns. To further enhance user behavior modeling by capturing the underlying dynamic intent, an Intent-oriented Session Modeling module is designed for inferring intent-oriented semantic sessions from the contextual information in behavior sequences. In particular, we consider self-supervised learning signals from two perspectives for intent-oriented semantic session locating, which encourage session discrimination within each behavior sequence and session alignment between dual behavior sequences. Extensive experiments on three public datasets demonstrate that UnifiedSSR consistently outperforms state-of-the-art methods for both search and recommendation. Jiayi Xie, Shang Liu 0005, Gao Cong, Zhenzhong Chen 0001 |
WWW | 1 |
| 2024 | CPADS: a web tool for comprehensive pancancer analysis of drug sensitivityabstractDrug therapy is vital in cancer treatment. Accurate analysis of drug sensitivity for specific cancers can guide healthcare professionals in prescribing drugs, leading to improved patient survival and quality of life. However, there is a lack of web-based tools that offer comprehensive visualization and analysis of pancancer drug sensitivity. We gathered cancer drug sensitivity data from publicly available databases (GEO, TCGA and GDSC) and developed a web tool called Comprehensive Pancancer Analysis of Drug Sensitivity (CPADS) using Shiny. CPADS currently includes transcriptomic data from over 29 000 samples, encompassing 44 types of cancer, 288 drugs and more than 9000 gene perturbations. It allows easy execution of various analyses related to cancer drug sensitivity. With its large sample size and diverse drug range, CPADS offers a range of analysis methods, such as differential gene expression, gene correlation, pathway analysis, drug analysis and gene perturbation analysis. Additionally, it provides several visualization approaches. CPADS significantly aids physicians and researchers in exploring primary and secondary drug resistance at both gene and pathway levels. The integration of drug resistance and gene perturbation data also presents novel perspectives for identifying pivotal genes influencing drug resistance. Access CPADS at https://smuonco.shinyapps.io/CPADS/ or https://robinl-lab.com/CPADS. Anqi Lin, Jiayi Xie, Jianguo Zhou, Shamus R. Carr, Zaoqu Liu, Jian Zhang 0104, David S. Schrump, Peng Luo 0005 |
Briefings Bioinform. | 4 |
| 2024 | Meta-path aware dynamic graph learning for friend recommendation with user mobility
Ding Ding 0004, Jing Yi, Jiayi Xie, Zhenzhong Chen 0001 |
Inf. Sci. | 3 |
| 2024 | Hybrid federated learning with brain-region attention network for multi-center Alzheimer's disease detectionabstractIdentifying reproducible and interpretable biomarkers for Alzheimer's disease (AD) detection remains a challenge. AD detection using multi-center datasets can expand the sample size to improve robustness but might lead to a data privacy problem. Moreover, due to the high cost of labeling data, a lot of unlabeled data in each center is not fully utilized. To address this, a hybrid FL (HFL) framework is proposed that not only uses unlabeled data to train deep learning networks, but also achieves data privacy protection. We propose a novel Brain-region Attention Network (BANet), which highlights important regions via attention to represent the region of interest (ROIs).Specifically, we use a brain template to extract ROI signals from the preprocessed structure magnetic resonance imaging (sMRI) data. In addition, we add a self-supervised loss to the current loss to guide the attention map generation to learn the representations from unlabeled data. Finally, we evaluate our method on a multi-center database which is constructed using five AD datasets. The experimental results show that the proposed method performs better than state-of-the-art methods, achieving mean accuracy rates of 85.69 %, 63.34 %, and 69.89 % on the AD vs. NC, MCI vs. NC, and AD vs. MCI respectively. The source code is available for reproducibility at: https://github.com/yuliangCarmelo/HFL . Bai Ying Lei, Jiayi Xie, Enmin Liang, Yong Liu 0018, Peng Yang 0011, Tianfu Wang 0001, Jichen Du, Xiaohua Xiao, Shuqiang Wang |
Pattern Recognit. | 3 |
| 2024 | Deep Causal Reasoning for RecommendationsabstractTraditional recommender systems aim to estimate a user’s rating to an item based on observed ratings from the population. As with all observational studies, hidden confounders, which are factors that affect both item exposures and user ratings, lead to a systematic bias in the estimation. Consequently, causal inference has been introduced in recommendations to address the influence of unobserved confounders. Observing that confounders in recommendations are usually shared among items and are therefore multi-cause confounders, we model the recommendation as a multi-cause multi-outcome (MCMO) inference problem. Specifically, to remedy the confounding bias, we estimate user-specific latent variables that render the item exposures independent Bernoulli trials. The generative distribution is parameterized by a DNN with factorized logistic likelihood and the intractable posteriors are estimated by variational inference. Controlling these factors as substitute confounders, under mild assumptions, can eliminate the bias incurred by multi-cause confounders. Furthermore, we show that MCMO modeling may lead to high variance due to scarce observations associated with the high-dimensional treatment space. Therefore, we theoretically demonstrate that controlling user features as pre-treatment variables can substantially improve sample efficiency and alleviate overfitting. Empirical studies on both simulated and real-world datasets demonstrate that the proposed deep causal recommender shows more robustness to unobserved confounders than state-of-the-art causal recommenders. Codes and datasets are released at https://github.com/yaochenzhu/Deep-Deconf. Yaochen Zhu, Jing Yi, Jiayi Xie, Zhenzhong Chen 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | Hierarchical Transformer with Spatio-temporal Context Aggregation for Next Point-of-interest RecommendationabstractNext point-of-interest (POI) recommendation is a critical task in location-based social networks, yet remains challenging due to a high degree of variation and personalization exhibited in user movements. In this work, we explore the latent hierarchical structure composed of multi-granularity short-term structural patterns in user check-in sequences. We propose a Spatio-Temporal context AggRegated Hierarchical Transformer (STAR-HiT) for next POI recommendation, which employs stacked hierarchical encoders to recursively encode the spatio-temporal context and explicitly locate subsequences of different granularities. More specifically, in each encoder, the global attention layer captures the spatio-temporal context of the sequence, while the local attention layer performed within each subsequence enhances subsequence modeling using the local context. The sequence partition layer infers positions and lengths of subsequences from the global context adaptively, such that semantics in subsequences can be well preserved. Finally, the subsequence aggregation layer fuses representations within each subsequence to form the corresponding subsequence representation, thereby generating a new sequence of higher-level granularity. The stacking of hierarchical encoders captures the latent hierarchical structure of the check-in sequence, which is used to predict the next visiting POI. Extensive experiments on three public datasets demonstrate that the proposed model achieves superior performance while providing explanations for recommendations. Jiayi Xie, Zhenzhong Chen 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2023 | Micro-Video Popularity Prediction Via Multimodal Variational Information BottleneckabstractIn this paper, we propose a Hierarchical Multimodal Variational Encoder-Decoder (HMMVED) to predict the popularity of micro-videos by comprehensively leveraging the user information and the micro-video content in a hierarchical fashion. In particular, the multimodal variational encoder-decoder framework encodes the input modalities to a lower dimensional stochastic embedding, from which the popularity of micro-videos can be decoded. Considering the leading role of the user’s social influence in social media for information dissemination, a user encoder-decoder is designed to learn the prior Gaussian embedding of the micro-video from the user information, which is informative about the coarse-grained popularity. In order to incorporate the fluctuation around the coarse-grained popularity caused by the diverse multimodal content, in the micro-video encoder-decoder, the refined posterior distribution of the micro-video embedding is encoded from the content features while encouraged to be close to the learned prior distribution. The fine-grained popularity of each micro-video is decoded from the posterior embedding of the micro-video. Based on the multimodal extension of variational information bottleneck theory, we show that the learned latent embeddings of micro-videos are maximally expressive about the popularity whilst maximally compressing the information from input modalities. Extensive experiments conducted on two real-world datasets demonstrate the effectiveness of the proposed method. Codes and datasets are available at:https://github.com/JennyXieJiayi/HMMVED. Jiayi Xie, Yaochen Zhu, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Cross-Modal Variational Auto-Encoder for Content-Based Micro-Video Background Music RecommendationabstractIn this paper, we propose a cross-modal variational auto-encoder (CMVAE) for content-based micro-video background music recommendation. CMVAE is a hierarchical Bayesian generative model that matches relevant background music to a micro-video by projecting these two multimodal inputs into a shared low-dimensional latent space, where the alignment of two corresponding embeddings of a matched video-music pair is achieved by cross-generation. Moreover, the multimodal information is fused by the product-of-experts (PoE) principle, where the semantic information in visual and textual modalities of the micro-video are weighted according to their variance estimations such that the modality with a lower noise level is given more weights. Therefore, the micro-video latent variables contain less irrelevant information that results in a more robust model generalization. Furthermore, we establish a large-scale content-based micro-video background music recommendation dataset, TT-150k, composed of extracted features from approximately 3,000 different background music clips associated to 150,000 micro-videos from different users. Extensive experiments on the established TT-150k dataset demonstrate the effectiveness of the proposed method. A qualitative assessment of CMVAE by visualizing some recommendation results is also included. Jing Yi, Yaochen Zhu, Jiayi Xie, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Automated Essay Scoring via Pairwise Contrastive RegressionabstractAutomated essay scoring (AES) involves the prediction of a score relating to the writing quality of an essay. Most existing works in AES utilize regression objectives or ranking objectives respectively. However, the two types of methods are highly complementary. To this end, in this paper we take inspiration from contrastive learning and propose a novel unified Neural Pairwise Contrastive Regression (NPCR) model in which both objectives are optimized simultaneously as a single loss. Specifically, we first design a neural pairwise ranking model to guarantee the global ranking order in a large list of essays, and then we further extend this pairwise ranking model to predict the relative scores between an input essay and several reference essays. Additionally, a multi-sample voting strategy is employed for inference. We use Quadratic Weighted Kappa to evaluate our model on the public Automated Student Assessment Prize (ASAP) dataset, and the experimental results demonstrate that NPCR outperforms previous methods by a large margin, achieving the state-of-the-art average performance for the AES task. Jiayi Xie, Kaiwei Cai, Junsheng Zhou, Weiguang Qu |
COLING | 1 |
| 2020 | User Conditional Hashtag Recommendation for Micro-VideosabstractWhen a user tend to publish a micro-video, hashtag recommendation aims to suggest hashtags that can reflect the theme or contents of the micro-video, and meet the user tagging preference as well. In this paper, we show how user profile and historical hashtags combined with micro-video representations can be used to perform hashtag recommendation. Specifically, a User-guided Hierarchical Multi-head Attention Network (UHMAN) is proposed to attend both image-level and video-level representations of micro-videos with user side information. We evaluate the proposed model on the dataset collected from micro-video sharing platform Musical.ly. The experimental results demonstrate the effectiveness of the proposed method. Jiayi Xie, Cong Zou |
ICME | 2 |
| 2020 | A Multimodal Variational Encoder-Decoder Framework for Micro-video Popularity PredictionabstractPredicting the popularity of a micro-video is a challenging task, due to a number of factors impacting the distribution such as the diversity of the video content and user interests, complex online interactions, etc. In this paper, we propose a multimodal variational encoder-decoder (MMVED) framework that considers the uncertain factors as the randomness for the mapping from the multimodal features to the popularity. Specifically, the MMVED first encodes features from multiple modalities in the observation space into latent representations and learns their probability distributions based on variational inference, where only relevant features in the input modalities can be extracted into the latent representations. Then, the modality-specific hidden representations are fused through Bayesian reasoning such that the complementary information from all modalities is well utilized. Finally, a temporal decoder implemented as a recurrent neural network is designed to predict the popularity sequence of a certain micro-video. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed model in the micro-video popularity prediction task. Jiayi Xie, Yaochen Zhu, Jing Yi, Yaosi Hu, Hongyi Liu 0003, Zhenzhong Chen 0001 |
WWW | 1 |