EDBT 2026 Demo / reviewers in the wild / expert
Yuying Chen
dblp:207/7692
· DBLP profile ↗
20ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 9 since 2021Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised Diffusion-Based Degradation Modeling for Real-World Super-ResolutionabstractSingle image super-solution (SR) aims to restore a high-resolution (HR) image from a degraded low-resolution (LR) image. However, existing SR models still face a significant domain gap between synthetic and real-world datasets due to the mismatched degradation distributions, hindering SR models from achieving optimal results. In this paper, we propose an unsupervised diffusion-based degradation modeling framework (UDDM) to effectively capture real-world degradation distributions. Specifically, given unpaired LR and HR images, a diffusion-based degradation module (DDM) first models the degradation distribution by diffusing real-world LR images to downsampled LR images, which does not require HR images. It then applies reverse diffusion to generate real-world LR images from extremely downsampled HR images. This approach allows DDM to model and generate real-world degradation distributions without requiring paired data, by using extreme downsampling to link unpaired LR and HR images. Additionally, we introduce a physics-based dynamic degradation module (P-DDM) that adaptively models content-aware degradation, ensuring both content and structural accuracy. Finally, the LR images generated by DDM and P-DDM are adaptively weighted to produce the final LR images, which are paired with the given HR images for training the SR network. Extensive experiments across multiple real-world datasets demonstrate that our framework achieves state-of-the-art performance in both qualitative and quantitative comparison. Yuying Chen, Mingde Yao, Renjing Pei, Jinjing Zhao, Wenqi Ren |
AAAI | 1 |
| 2025 | Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel LearningabstractWith the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating high-quality UHD images on a resource-constrained device. To alleviate the limitations of extremely long-sequence inputs, inspired by the Large Language Model (LLM) for processing infinitely long texts, we propose a novel learning paradigm to achieve UHD multi-exposure dynamic scene image fusion on a single consumer-grade GPU, named Infinite Pixel Learning (IPL). The design of our approach comes from three key components: The first step is to slice the input sequences to relieve the pressure generated by the model processing the data stream; Second, we develop an attention cache technique, which is similar to the KV cache for infinite data stream processing; Finally, we design a method for attention cache compression to alleviate the storage burden of the cache on the device. In addition, we provide a new UHD benchmark to evaluate the effectiveness of our method. Extensive experimental results show that our method maintains high-quality visual performance while fusing UHD dynamic multi-exposure images in real-time (>40fps) on a single consumer-grade GPU. Xingchi Chen, Zhuoran Zheng, Xuerui Li, Yuying Chen, Wenqi Ren |
AAAI | 4 |
| 2025 | Enhancing the Flexibility of a Quadruped Robot with a 2-DOF Active Spine Using Nonlinear Model Predictive ControlabstractFor quadrupeds, a flexible spine allows them to traverse space and make quick turns. From the perspective of mechanical design in quadruped robots, an active spine with 2 degrees of freedom (2-DOF) can achieve dynamic posture adjustment similar to biological organisms which allows for pitch and yaw control. In this work, we present a novel approach to enhance the flexibility of a quadruped robot, Yatsen Lion II, by incorporating a 2-DOF active spine, which is mechanically designed as a linkage-driven parallelogram mechanism. To optimize its motion, we utilize nonlinear model predictive control (NMPC), which combines centroidal dynamics with full kinematics. By incorporating the two extra DOFs of the spinal joint into the generalized coordinates and velocities, we represent the robot as a hybrid dynamic system, capturing the intricate interplay between the legs and spine. Centroidal dynamics act as a crucial bridge between joint movements and the robot’s overall momentum, enabling the controller to synchronize the quadruped’s movements with dynamic spinal adjustments and adaptive gait patterns. We validate our approach through both simulation and real-world experiments. We compare spinal quadruped robot to their rigid-spine counterparts across key locomotion metrics, including in-place turning, straight-line speed, and turning radius. The results indicate that the spined quadrupedal robot outperforms its rigid counterpart by up to 26%, highlighting its flexibility. Zeyi Yang, Haoming Rong, Shaolin Mo, Yuying Chen, Zujian Chen |
IROS | 5 |
| 2025 | PlanVWM: Autonomous Vehicle Planning Method Based on Vectorized World Model ModelingabstractWith the increasing demand for handling complex scenarios in autonomous driving, data-driven planning methods based on imitation learning have attracted significant attention. In this context, this work proposes the PlanVMN method, aiming to enhance the model's robustness to data and its causal inference ability in planning. Based on vectorized scene inputs, this method integrates the world model into the core architecture of the planning model and deduces the temporal features of the scene based on the actions of the ego vehicle. Meanwhile, we introduce an attention-based history enhancement component within the world model, which remarkably improves the planning performance of the ego vehicle. Experiments show that the model has achieved outstanding results in the open-loop and closed-loop tests of nuPlan. Notably, thanks to the generative ability endowed by the unique world model architecture, the model can still exhibit excellent performance when facing blind area, sensor misfuctioning or detection instability, providing strong support for planning in complex scenarios of autonomous driving. Yuying Chen, Ziqing Gu, Siyuan Cheng 0012, Xueqian Wang 0001 |
IV | 2 |
| 2025 | Stream Normalization for CTR Prediction
Yizhou Sang, Yuying Chen, Zhiwei Fang, Changping Peng, Zhangang Lin, Ching Law, Jingping Shao |
RecSys | 3 |
| 2025 | Post-event Modeling via Causal Optimal Transport for CTR PredictionabstractAccurate click-through rate (CTR) prediction is critical for online advertising, relying on regular features like browsing history and demographics and post-event features such as exposed position and detailed page behaviors. However, post-event features, unavailable during inference, often face training-inference inconsistency and low coverage issues, especially post-click features like dwell time that are available only for clicked items. To address these challenges, we propose Causal Optimal Transport (COT), a novel framework that (1) generates pseudo post-click features via semi-supervised pseudo-labeling (2) causally generates accurate feature distributions using a Causal Distribution Shaper (CDS), and (3) refines generated features through optimal transport to minimize distributional divergence, facilitating further knowledge transfer. Experiments on real-world data confirm COT's superiority and practical efficacy in enhancing CTR prediction via improved user interest modeling and bias mitigation. Theoretical guarantees underpin the framework's robustness. Yizhou Sang, Yuying Chen, Zhiwei Fang, Changping Peng, Zhangang Lin, Ching Law, Jingping Shao |
SIGIR | 3 |
| 2025 | An adaptive microbiome regression-based kernel association test using generalized estimating equations for longitudinal microbiome studies
Ranxin Yan, Yuying Chen, Yongrui Liu, Pengyan Wang |
Expert Syst. Appl. | 3 |
| 2023 | Improving Vehicle Trajectory Prediction with Online LearningabstractIn autonomous driving systems, predicting the trajectory of surrounding vehicles facilitates the decision-making and trajectory planning of ego cars. Most previous works train the trajectory predictor based on the offline dataset. While the offline general predictors capture the average distribution in the dataset, they commonly neglect the low-probability corner cases that outside the general distribution of dataset. However, the corner cases might cause inevitable displacement errors and goals missing in the evaluation. Offline data augmentation and model refinement could essentially improve the accuracy while we can leverage the historical trajectories to specialize the predictor by self-supervision. In this paper, we propose an online learning framework that updates the general trajectory predictor with sequential history data at corner cases. We design the temporal data-generation method and training scheme for online learning, which can generally accommodate learning-based predictors. Since online learning costs expensive online computation resources, we devise a switching discriminator to distinguish the corner cases by evaluating the metric improvement of online learning over the offline general predictor. We validate that the online learning method can effectively reduce the displacement errors and miss rate caused by vehicle velocity variation and promote the intention convergence at multi-modal intersections. Experiments also show that the switching discriminator limits the switching-on rate of online learning cases to 2.45% while reducing the miss rate by 33.4%. The code for generating sequentially arranged online learning dataset(Argoverse) is in https://gitee.com/mindspore/models/tree/master/research/cv/TraPred_OnlineLearning. Ce Hao, Yuying Chen, Siyuan Cheng 0012 |
IV | 2 |
| 2022 | HGCN-GJS: Hierarchical Graph Convolutional Network with Groupwise Joint Sampling for Trajectory PredictionabstractPedestrian trajectory prediction is of great importance for downstream tasks, such as autonomous driving and mobile robot navigation. Realistic models of the social interactions within the crowd is crucial for accurate pedestrian trajectory prediction. However, most existing methods do not capture group level interactions well, focusing only on pairwise interactions and neglecting group-wise interactions. In this work, we propose a hierarchical graph convolutional network, HGCN-GJS, for trajectory prediction which well leverages group level interactions within the crowd. Furthermore, we introduce a joint sampling scheme that captures co-dependencies between pedestrian trajectories during trajectory generation. Based on group information, this scheme ensures that generated trajectories within each group are consistent with each other, but enables different groups to act more independently. We demonstrate that our proposed network achieves state of the art performance on all datasets we have considered. Yuying Chen, Xiaodong Mei 0001, Bertram E. Shi, Ming Liu 0001 |
IROS | 1 |
| 2021 | AVGCN: Trajectory Prediction using Graph Convolutional Networks Guided by Human AttentionabstractPedestrian trajectory prediction is a critical yet challenging task especially for crowded scenes. We suggest that introducing an attention mechanism to infer the importance of different neighbors is critical for accurate trajectory prediction in scenes with varying crowd size. In this work, we propose a novel method, AVGCN, for trajectory prediction utilizing graph convolutional networks (GCN) based on human attention (A denotes attention, V denotes visual field constraints). First, we train an attention network that estimates the importance of neighboring pedestrians, using gaze data collected as subjects perform a bird’s eye view crowd navigation task. Then, we incorporate the learned attention weights modulated by constraints on the pedestrian’s visual field into a trajectory prediction network that uses a GCN to aggregate information from neighbors efficiently. AVGCN also considers the stochastic nature of pedestrian trajectories by taking advantage of variational trajectory prediction. Our approach achieves state-of-the-art performance on several trajectory prediction benchmarks, and the lowest average prediction error over all considered benchmarks. Yuying Chen, Ming Liu 0002, Bertram E. Shi |
ICRA | 2 |
| 2021 | Using Eye Gaze to Enhance Generalization of Imitation Networks to Unseen EnvironmentsabstractVision-based autonomous driving through imitation learning mimics the behavior of human drivers by mapping driver view images to driving actions. This article shows that performance can be enhanced via the use of eye gaze. Previous research has shown that observing an expert's gaze patterns can be beneficial for novice human learners. We show here that neural networks can also benefit. We trained a conditional generative adversarial network to estimate human gaze maps accurately from driver-view images. We describe two approaches to integrating gaze information into imitation networks: eye gaze as an additional input and gaze modulated dropout. Both significantly enhance generalization to unseen environments in comparison with a baseline vanilla network without gaze, but gaze-modulated dropout performs better. We evaluated performance quantitatively on both single images and in closed-loop tests, showing that gaze modulated dropout yields the lowest prediction error, the highest success rate in overtaking cars, the longest distance between infractions, lowest epistemic uncertainty, and improved data efficiency. Using Grad-CAM, we show that gaze modulated dropout enables the network to concentrate on task-relevant areas of the image. Yuying Chen, Ming Liu 0001, Bertram E. Shi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Neural Cognitive Diagnosis for Intelligent Education SystemsabstractCognitive diagnosis is a fundamental issue in intelligent education, which aims to discover the proficiency level of students on specific knowledge concepts. Existing approaches usually mine linear interactions of student exercising process by manual-designed function (e.g., logistic function), which is not sufficient for capturing complex relations between students and exercises. In this paper, we propose a general Neural Cognitive Diagnosis (NeuralCD) framework, which incorporates neural networks to learn the complex exercising interactions, for getting both accurate and interpretable diagnosis results. Specifically, we project students and exercises to factor vectors and leverage multi neural layers for modeling their interactions, where the monotonicity assumption is applied to ensure the interpretability of both factors. Furthermore, we propose two implementations of NeuralCD by specializing the required concepts of each exercise, i.e., the NeuralCDM with traditional Q-matrix and the improved NeuralCDM+ exploring the rich text content. Extensive experimental results on real-world datasets show the effectiveness of NeuralCD framework with both accuracy and interpretability. Fei Wang 0063, Qi Liu 0003, Enhong Chen, Zhenya Huang, Yuying Chen, Yu Yin 0002, Zai Huang, Shijin Wang 0001 |
AAAI | 5 |
| 2020 | CoMoGCN: Coherent Motion Aware Trajectory Prediction with Graph Representation
Yuying Chen, Bertram E. Shi, Ming Liu 0001 |
BMVC | 1 |
| 2020 | Learning or Forgetting? A Dynamic Approach for Tracking the Knowledge Proficiency of StudentsabstractThe rapid development of the technologies for online learning provides students with extensive resources for self-learning and brings new opportunities for data-driven research on educational management. An important issue of online learning is to diagnose the knowledge proficiency (i.e., the mastery level of a certain knowledge concept) of each student. Considering that it is a common case that students inevitably learn and forget knowledge from time to time, it is necessary to track the change of their knowledge proficiency during the learning process. Existing approaches either relied on static scenarios or ignored the interpretability of diagnosis results. To address these problems, in this article, we present a focused study on diagnosing the knowledge proficiency of students, where the goal is to track and explain their evolutions simultaneously. Specifically, we first devise an explanatory probabilistic matrix factorization model, Knowledge Proficiency Tracing (KPT), by leveraging educational priors. KPT model first associates each exercise with a knowledge vector in which each element represents a specific knowledge concept with the help of Q -matrix. Correspondingly, at each time, each student can be represented as a proficiency vector in the same knowledge space. Then, our KPT model jointly applies two classical educational theories (i.e., learning curve and forgetting curve ) to capture the change of students’ proficiency level on concepts over time. Furthermore, for improving the predictive performance, we develop an improved version of KPT, named Exercise-correlated Knowledge Proficiency Tracing (EKPT), by considering the connectivity among exercises with the same knowledge concepts. Finally, we apply our KPT and EKPT models to three important diagnostic tasks, including knowledge estimation, score prediction, and diagnosis result visualization. Extensive experiments on four real-world datasets demonstrate that both of our models could track the knowledge proficiency of students effectively and interpretatively. Zhenya Huang, Qi Liu 0003, Yuying Chen, Le Wu 0001, Keli Xiao, Enhong Chen, Haiping Ma |
ACM Trans. Inf. Syst. | 3 |
| 2019 | Hierarchical Multi-label Text Classification: An Attention-based Recurrent Network ApproachabstractHierarchical multi-label text classification (HMTC) is a fundamental but challenging task of numerous applications (e.g., patent annotation), where documents are assigned to multiple categories stored in a hierarchical structure. Categories at different levels of a document tend to have dependencies. However, the majority of prior studies for the HMTC task employ classifiers to either deal with all categories simultaneously or decompose the original problem into a set of flat multi-label classification subproblems, ignoring the associations between texts and the hierarchical structure and the dependencies among different levels of the hierarchical structure. To that end, in this paper, we propose a novel framework called Hierarchical Attention-based Recurrent Neural Network (HARNN) for classifying documents into the most relevant categories level by level via integrating texts and the hierarchical category structure. Specifically, we first apply a documentation representing layer for obtaining the representation of texts and the hierarchical structure. Then, we develop an hierarchical attention-based recurrent layer to model the dependencies among different levels of the hierarchical structure in a top-down fashion. Here, a hierarchical attention strategy is proposed to capture the associations between texts and the hierarchical structure. Finally, we design a hybrid method which is capable of predicting the categories of each level while classifying all categories in the entire hierarchical structure precisely. Extensive experimental results on two real-world datasets demonstrate the effectiveness and explanatory power of HARNN. Wei Huang 0002, Enhong Chen, Qi Liu 0003, Yuying Chen, Zai Huang, Yang Liu 0278, Zhou Zhao 0001, Shijin Wang 0001 |
CIKM | 4 |
| 2019 | A gaze model improves autonomous drivingabstractEnd-to-end behavioral cloning trained by human demonstration is now a popular approach for vision-based autonomous driving. A deep neural network maps drive-view images directly to steering commands. However, the images contain much task-irrelevant data. Humans attend to behaviorally relevant information using saccades that direct gaze towards important areas. We demonstrate that behavioral cloning also benefits from active control of gaze. We trained a conditional generative adversarial network (GAN) that accurately predicts human gaze maps while driving in both familiar and unseen environments. We incorporated the predicted gaze maps into end-to-end networks for two behaviors: following and overtaking. Incorporating gaze information significantly improves generalization to unseen environments. We hypothesize that incorporating gaze information enables the network to focus on task critical objects, which vary little between environments, and ignore irrelevant elements in the background, which vary greatly. Yuying Chen, Lei Tai, Haoyang Ye, Ming Liu 0001, Bertram E. Shi |
ETRA | 2 |
| 2019 | Tightly Coupled 3D Lidar Inertial Odometry and MappingabstractEgo-motion estimation is a fundamental requirement for most mobile robotic applications. By sensor fusion, we can compensate the deficiencies of stand-alone sensors and provide more reliable estimations. We introduce a tightly coupled lidar-IMU fusion method in this paper. By jointly minimizing the cost derived from lidar and IMU measurements, the lidarIMU odometry (LIO) can perform well with considerable drifts after long-term experiment, even in challenging cases where the lidar measurement can be degraded. Besides, to obtain more reliable estimations of the lidar poses, a rotation-constrained refinement algorithm (LIO-mapping) is proposed to further align the lidar poses with the global map. The experiment results demonstrate that the proposed method can estimate the poses of the sensor pair at the IMU update rate with high precision, even under fast motion conditions or with insufficient features. Haoyang Ye, Yuying Chen, Ming Liu 0001 |
ICRA | 2 |
| 2019 | Gaze Training by Modulated Dropout Improves Imitation LearningabstractImitation learning by behavioral cloning is a prevalent method that has achieved some success in vision-based autonomous driving. The basic idea behind behavioral cloning is to have the neural network learn from observing a human expert's behavior. Typically, a convolutional neural network learns to predict the steering commands from raw driver-view images by mimicking the behaviors of human drivers. However, there are other cues, such as gaze behavior, available from human drivers that have yet to be exploited. Previous researches have shown that novice human learners can benefit from observing experts' gaze patterns. We present here that deep neural networks can also profit from this. We propose a method, gaze-modulated dropout, for integrating this gaze information into a deep driving network implicitly rather than as an additional input. Our experimental results demonstrate that gaze-modulated dropout enhances the generalization capability of the network to unseen scenes. Prediction error in steering commands is reduced by 23.5% compared to uniform dropout. Running closed loop in the simulator, the gaze-modulated dropout net increased the average distance travelled between infractions by 58.5%. Consistent with these results, the gazemodulated dropout net shows lower model uncertainty. Yuying Chen, Lei Tai, Ming Liu 0001, Bertram E. Shi |
IROS | 1 |
| 2019 | Visual-based Autonomous Driving Deployment from a Stochastic and Uncertainty-aware PerspectiveabstractEnd-to-end visual-based imitation learning has been widely applied in autonomous driving. When deploying the trained visual-based driving policy, a deterministic command is usually directly applied without considering the uncertainty of the input data. Such kind of policies may bring dramatical damage when applied in the real world. In this paper, we follow the recent real-to-sim pipeline by translating the testing world image back to the training domain when using the trained policy. In the translating process, a stochastic generator is used to generate various images stylized under the training domain randomly or directionally. Based on those translated images, the trained uncertainty-aware imitation learning policy would output both the predicted action and the data uncertainty motivated by the aleatoric loss function. Through the uncertainty-aware imitation learning policy, we can easily choose the safest one with the lowest uncertainty among the generated images. Experiments in the Carla navigation benchmark show that our strategy outperforms previous methods, especially in dynamic environments. Lei Tai, Peng Yun, Yuying Chen, Haoyang Ye, Ming Liu 0001 |
IROS | 3 |
| 2017 | Tracking Knowledge Proficiency of Students with Educational PriorsabstractDiagnosing students' knowledge proficiency, i.e., the mastery degrees of a particular knowledge point in exercises, is a crucial issue for numerous educational applications, e.g., targeted knowledge training and exercise recommendation. Educational theories have converged that students learn and forget knowledge from time to time. Thus, it is necessary to track their mastery of knowledge over time. However, traditional methods in this area either ignored the explanatory power of the diagnosis results on knowledge points or relied on a static assumption. To this end, in this paper, we devise an explanatory probabilistic approach to track the knowledge proficiency of students over time by leveraging educational priors. Specifically, we first associate each exercise with a knowledge vector in which each element represents an explicit knowledge point by leveraging educational priors (i.e., Q-matrix ). Correspondingly, each student is represented as a knowledge vector at each time in a same knowledge space. Second, given the student knowledge vector over time, we borrow two classical educational theories (i.e., Learning curve and Forgetting curve ) as priors to capture the change of each student's proficiency over time. After that, we design a probabilistic matrix factorization framework by combining student and exercise priors for tracking student knowledge proficiency. Extensive experiments on three real-world datasets demonstrate both the effectiveness and explanatory power of our proposed model. Yuying Chen, Qi Liu 0003, Zhenya Huang, Le Wu 0001, Enhong Chen, Runze Wu 0001, Yu Su 0002 |
CIKM | 1 |