VLDB 2026 Research / reviewers in the wild / expert
Kexin Zheng
dblp:248/9615
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsabstractYunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Runhui Xu, Kexin Zheng, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun |
ACL (1) | 3 |
| 2025 | Supportive Negatives Spectral Augmentation for Source-Free Cross-Domain SegmentationabstractSource-free domain adaptation (SFDA) aims to transfer knowledge from the well-trained source model and optimize it to adapt target data distribution. SFDA methods are suitable for medical image segmentation task due to its data-privacy protection and achieve promising performances. However, cross-domain distribution shift makes it difficult for the adapted model to provide accurate decisions on several hard instances and negatively affects model generalization. To overcome this limitation, a novel method `supportive negatives spectral augmentation' (SNSA) is presented in this work. Concretely, SNSA includes the instance selection mechanism to automatically discover a few hard samples for which source model produces incorrect predictions. And, active learning strategy is adopted to re-calibrate their predictive masks. Moreover, SNSA deploys the spectral augmentation between hard instances and others to encourage source model to gradually capture and adapt the attributions of target distribution. Considerable experimental studies demonstrate that annotating merely 4%~5% of negative instances from the target domain significantly improves segmentation performance over previous methods. Kexin Zheng, Haifeng Xia, Si-Yu Xia, Ming Shao, Zhengming Ding |
AAAI | 1 |
| 2025 | Diffusion-Based Planning for Autonomous Driving with Flexible GuidanceabstractAchieving human-like driving behaviors in complex open-world environments is a critical challenge in autonomous driving. Contemporary learning-based planning approaches such as imitation learning methods often struggle to balance competing objectives and lack of safety assurance,due to limited adaptability and inadequacy in learning complex multi-modal behaviors commonly exhibited in human planning, not to mention their strong reliance on the fallback strategy with predefined rules. We propose a novel transformer-based Diffusion Planner for closed-loop planning, which can effectively model multi-modal driving behavior and ensure trajectory quality without any rule-based refinement. Our model supports joint modeling of both prediction and planning tasks under the same architecture, enabling cooperative behaviors between vehicles. Moreover, by learning the gradient of the trajectory score function and employing a flexible classifier guidance mechanism, Diffusion Planner effectively achieves safe and adaptable planning behaviors. Evaluations on the large-scale real-world autonomous planning benchmark nuPlan and our newly collected 200-hour delivery-vehicle driving dataset demonstrate that Diffusion Planner achieves state-of-the-art closed-loop performance with robust transferability in diverse driving styles. Yinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng, Liyuan Mao, Weihao Gu, Rui Ai 0001, Shengbo Eben Li, Xianyuan Zhan |
ICLR | 3 |
| 2025 | Contact Map Transfer with Conditional Diffusion Model for Generalizable Dexterous Grasp GenerationabstractDexterous grasp generation is a fundamental challenge in robotics, requiring both grasp stability and adaptability across diverse objects and tasks. Analytical methods ensure stable grasps but are inefficient and lack task adaptability, while generative approaches improve efficiency and task integration but generalize poorly to unseen objects and tasks due to data limitations. In this paper, we propose a transfer-based framework for dexterous grasp generation, leveraging a conditional diffusion model to transfer high-quality grasps from shape templates to novel objects within the same category. Specifically, we reformulate the grasp transfer problem as the generation of an object contact map, incorporating object shape similarity and task specifications into the diffusion process. To handle complex shape variations, we introduce a dual mapping mechanism, capturing intricate geometric relationship between shape templates and novel objects. Beyond the contact map, we derive two additional object-centric maps, the part map and direction map, to encode finer contact details for more stable grasps. We then develop a cascaded conditional diffusion model framework to jointly transfer these three maps, ensuring their intra-consistency. Finally, we introduce a robust grasp recovery mechanism, identifying reliable contact points and optimizing grasp configurations efficiently. Extensive experiments demonstrate the superiority of our proposed method. Our approach effectively balances grasp quality, generation efficiency, and generalization performance across various tasks. Project homepage: https://cmtdiffusion.github.io/ Yiyao Ma, Kai Chen 0028, Kexin Zheng, Qi Dou 0001 |
NeurIPS | 3 |
| 2025 | Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior ModelingabstractModeling interactive driving behaviors in complex scenarios remains a fundamental challenge for autonomous driving planning. Learning-based approaches attempt to address this challenge with advanced generative models, removing the dependency on over-engineered architectures for representation fusion. However, brute-force implementation by simply stacking transformer blocks lacks a dedicated mechanism for modeling interactive behaviors that is common in real driving scenarios. The scarcity of interactive driving data further exacerbates this problem, leaving conventional imitation learning methods ill-equipped to capture high-value interactive behaviors. We propose Flow Planner, which tackles these problems through coordinated innovations in data modeling, model architecture, and learning scheme. Specifically, we first introduce fine-grained trajectory tokenization, which decomposes the trajectory into overlapping segments to decrease the complexity of whole trajectory modeling. With a sophisticatedly designed architecture, we achieve efficient temporal and spatial fusion of planning and scene information, to better capture interactive behaviors. In addition, the framework incorporates flow matching with classifier-free guidance for multi-modal behavior generation, which dynamically reweights agent interactions during inference to maintain coherent response strategies, providing a critical boost for interactive scenario understanding. Experimental results on the large-scale nuPlan dataset demonstrate that Flow Planner achieves state-of-the-art performance among learning-based approaches while effectively modeling interactive behaviors in complex driving scenarios. Tianyi Tan, Yinan Zheng, Ruiming Liang, Zexu Wang, Kexin Zheng, Jinliang Zheng, Xianyuan Zhan |
NeurIPS | 5 |
| 2025 | Towards Robust Zero-Shot Reinforcement LearningabstractThe recent development of zero-shot reinforcement learning (RL) has opened a new avenue for learning pre-trained generalist policies that can adapt to arbitrary new tasks in a zero-shot manner. While the popular Forward-Backward representations (FB) and related methods have shown promise in zero-shot RL, we empirically found that their modeling lacks expressivity and that extrapolation errors caused by out-of-distribution (OOD) actions during offline learning sometimes lead to biased representations, ultimately resulting in suboptimal performance. To address these issues, we propose Behavior-REgularizEd Zero-shot RL with Expressivity enhancement (BREEZE), an upgraded FB-based framework that simultaneously enhances learning stability, policy extraction capability, and representation learning quality. BREEZE introduces behavioral regularization in zero-shot RL policy learning, transforming policy optimization into a stable in-sample learning paradigm. Additionally, BREEZE extracts the policy using a task-conditioned diffusion model, enabling the generation of high-quality and multimodal action distributions in zero-shot RL settings. Moreover, BREEZE employs expressive attention-based architectures for representation modeling to capture the complex relationships between environmental dynamics. Extensive experiments on ExORL and D4RL Kitchen demonstrate that BREEZE achieves the best or near-the-best performance while exhibiting superior robustness compared to prior offline zero-shot RL methods. The official implementation is available at: https://github.com/Whiterrrrr/BREEZE. Kexin Zheng, Lauriane Teyssier, Yinan Zheng, Yu Luo 0021, Xianyuan Zhan |
NeurIPS | 1 |
| 2024 | FastMapSVM/FastMapSVR for Predictive Tasks on CSPs, SAT, and Weighted CSPsabstractPredictive tasks on Constraint Satisfaction Problems (CSPs), Satisfiability (SAT) problems, and Weighted CSPs (WC-SPs) are usually NP-hard but can also be modeled as classification or regression problems suitable for Machine Learning (ML) algorithms. While most existing ML algorithms have had only limited success on such tasks, a newly developed ML frame-work, called FastMapSVM, has been shown to be successful for predicting CSP satisfiability. FastMapSVM leverages a distance function between pairs of CSP instances instead of trying to characterize individual CSP instances. In this paper, we advance FastMapSVM in various ways. For predicting the satisfiability of CSP and SAT instances, we design a distance function that utilizes maxflow computations and strong path-consistency (or a truncated version of it). For predicting the optimal cost of WCSP instances, we design a distance function that also utilizes maxflow computations and replaces the Support Vector Machine (SVM) component of FastMapSVM by a Support Vector-based Regression (SVR) component. We demonstrate the success of our FastMapSVM/FastMapSVR approach over competing state-of-the-art ML algorithms in all three domains: In the CSP domain, we demonstrate our success on several CSP benchmark suites; in the SAT domain, we demonstrate our success on hard 3-SAT instances drawn from the phase transition region; and in the WCSP domain, we demonstrate our success on a wide range of randomly generated WCSP instances. Kexin Zheng, T. K. Satish Kumar |
ICMLA | 1 |
| 2024 | M$^{3}$3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing SystemabstractFace presentation attacks (FPA), also known as face spoofing, have brought increasing concerns to the public through various malicious applications, such as financial fraud and privacy leakage. Therefore, safeguarding face recognition systems against FPA is of utmost importance. Although existing learning-based face anti-spoofing (FAS) models can achieve outstanding detection performance, they lack generalization capability and suffer significant performance drops in unforeseen environments. Many methodologies seek to use auxiliary modality data (e.g., depth and infrared maps) during the presentation attack detection (PAD) to address this limitation. However, these methods can be limited since (1) they require specific sensors such as depth and infrared cameras for data capture, which are rarely available on commodity mobile devices, and (2) they cannot work properly in practical scenarios when either modality is missing or of poor quality. In this paper, we devise an accurate and robustMultiModalMobileFaceAnti-Spoofing system namedM$^{3}$FASto overcome the issues above. The primary innovation of this work lies in the following aspects: (1) To achieve robust PAD, our system combines visual and auditory modalities using three commonly available sensors: camera, speaker, and microphone; (2) We design a novel two-branch neural network with three hierarchical feature aggregation modules to perform cross-modal feature fusion; (3). We propose a multi-head training strategy, allowing the model to output predictions from the vision, acoustic, and fusion heads, resulting in a more flexible PAD. Extensive experiments have demonstrated the accuracy, robustness, and flexibility of M$^{3}$FAS under various challenging experimental settings. The source code and dataset are available at:https://github.com/ChenqiKONG/M3FAS/. Chenqi Kong, Kexin Zheng, Yibing Liu, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | FastMapSVM for Predicting CSP Satisfiability
Kexin Zheng, Han Zhang 0018, T. K. Satish Kumar |
CP | 1 |
| 2022 | An Unsupervised Hyperspectral Image Fusion Method Based on Spectral Unmixing and Deep LearningabstractDue to the limitations of various hardware conditions, in practice, only high resolution multispectral and low resolution hyperspectral images are usually captured. In order to apply hyperspectral images in various fields, better quality hyperspectral images have become a problem to be solved. In this paper, we propose an image fusion method based on spectral unmixing, which effectively combines the advantages of multispectral images and low-resolution hyperspectral images to generate high-resolution hyperspectral images. To be specific, a deep learning model based on spectral decomposition is constructed, using multiplication iterative rules based on the traditional gradient descent algorithm to get initial high-resolution abundance and define degeneration networks to describe the spatial and spectral downsampling operations. Experiments show that this method can get fusion images better quality than other methods. Kexin Zheng, Abdolraheem Khader, Liang Xiao 0001 |
IGARSS | 1 |
| 2022 | Beyond the Pixel World: A Novel Acoustic-Based Face Anti-Spoofing System for Smartphonesabstract2D face presentation attacks are one of the most notorious and pervasive face spoofing types, which have caused pressing security issues to facial authentication systems. While RGB-based face anti-spoofing (FAS) models have proven to counter the face spoofing attack effectively, most existing FAS models suffer from the overfitting problem (i.e., lack generalization capability to data collected from an unseen environment). Recently, many models have been devoted to capturing auxiliary information (e.g., depth and infrared maps) to achieve a more robust face liveness detection performance. However, these methods require expensive sensors and cost extra hardware to capture the specific modality information, limiting their applications in practical scenarios. To tackle these problems, we devise a novel and cost-effective FAS system based on the acoustic modality, named Echo-FAS, which employs the crafted acoustic signal as the probe to perform face liveness detection. We first propose to build a large-scale, high-diversity, and acoustic-based FAS database, Echo-Spoof. Then, based upon Echo-Spoof, we propose designing a novel two-branch framework that combines the global and local frequency clues of input signals to distinguish inputs, live vs. spoofing faces accurately. The devised Echo-FAS comprises the following three merits: (1) It only needs one available speaker and microphone as sensors while not requiring any expensive hardware; (2) It can successfully capture the 3D geometrical information of input queries and achieve a remarkable face anti-spoofing performance; and (3) It can be handily allied with other RGB-based FAS models to mitigate the overfitting problem in the RGB modality and make the FAS model more accurate and robust. Our proposed Echo-FAS provides new insights regarding the development of FAS systems for mobile devices. Chenqi Kong, Kexin Zheng, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | ThermEarhook: Investigating Spatial Thermal Haptic Feedback on the Auricular Skin AreaabstractHaptic feedbacks are widely adopted in mobile and wearable devices to convey various types of notifications to the users. This paper investigates the design and the evaluation of thermal haptic feedback on an earable form factor with multiple thermoelectric (i.e. Peltier) modules. We propose ThermEarhook, a wearable device that can provide hot and cold stimuli at multiple points on the auricular skin area. To investigate users’ thermal perception on the auricular area, we develop a series of ThermEarhook prototypes with 3, 4, and 5 Peltier modules. While most existing research utilized the constant level of haptic signal for different users, our pilot study with ThermEarhook shows that the auricular thermohaptic threshold varies across the feedback locations and the users. With the user-customized thermohaptic signals around the ear, our first study with 12 participants reports on the selection of the auricular configuration with four TEC modules on each side, considering the users’ identification accuracy (averagely 99.3%) and preference. We then conduct three follow-up studies and a total of 36 participants to further evaluate users’ perception of spatial thermal patterns with ThermEarhook, and finalize a set of multi-points auricular thermal patterns that can be reliably perceived by the users with the average accuracy of 85.3%. Lastly, we discuss the user-proposed potential applications of the thermal haptic feedback with ThermEarhook. Arshad Nasser, Kexin Zheng, Kening Zhu |
ICMI | 2 |