VLDB 2026 Research / reviewers in the wild / expert
Kaixiang Chen
dblp:227/9058
· DBLP profile ↗
16ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Security and privacy · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution DetectionabstractRecent vision-language model (VLM)-based methods have achieved promising results in zero-shot out-of-distribution (OOD) detection by effectively leveraging the local patch features. However, the zero-shot nature inherently comes with two limitations: 1) imperfect local feature prototypes; 2) lack of OOD prototypes. In this paper, we propose Intra-Image Mining (IIM), a lightweight framework designed to overcome these limitations in a few-shot manner. IIM is motivated by the fact that local patches within an image often exhibit diverse semantics, with some patches deviating from the main class concept. Therefore, for each image, we first select the top-k class prototype-related patches as positive samples and leverage them to refine and optimize the local feature prototype. Then, the next top-k among the remaining patches are selected as negatives—serving as OOD signals to construct OOD prototypes. This process yields coherent local positives and challenging negatives, effectively enhancing the model’s local feature discrimination. Besides, we propose a novel inference strategy named Symmetric Maximum Concept Matching (S-MCM). While existing approaches typically adopt an image-to-text scheme—comparing the image features to textual class prototypes—S-MCM further incorporate a text-to-image perspective, leading to more reliable OOD detection. We also propose two benchmarks to analyze the impact of semantic diversity within ID dataset. Built on a frozen VLM, IIM, in conjunction with S-MCM, achieves consistent gains in OOD detection on ImageNet-1k and other benchmarks, outperforming prior methods in FPR95 and AUROC across various few-shot settings. Kaixiang Chen, Pengfei Fang, Hui Xue 0002 |
AAAI | 1 |
| 2025 | Multi-Modal Interactive Agent Layer for Few-Shot Universal Cross-Domain Retrieval and BeyondabstractThis paper firstly addresses the challenge of few-shot universal cross-domain retrieval (FS-UCDR), enabling machines trained with limited data to generalize to novel retrieval scenarios, with queries from entirely unknown domains and categories. To achieve this, we first formally define the FS-UCDR task and propose the Multi-Modal Interactive Agent Layer (MAIL), which enhances the cross-modal interaction in vision-language models (VLMs) by aligning the parameter updates of target layer pairs across modalities.
Specifically, MAIL freezes the selected target layer pair and introduces a trainable agent layer pair to approximate localized parameter updates. A bridge function is then introduced to couple the agent layer pair, enabling gradient communication across modalities to facilitate update alignment. The proposed MAIL offers four key advantages: 1) its cross-modal interaction mechanism improves knowledge acquisition from limited data, making it highly effective in low-data scenarios; 2) during inference, MAIL integrates seamlessly into the VLM via reparameterization, preserving inference complexity; 3) extensive experiments validate the superiority of MAIL, which achieves substantial performance gains over data-efficient UCDR methods while requiring significantly fewer training samples; 4) beyond UCDR, MAIL also performs competitively on few-shot classification tasks, underscoring its strong generalization ability. Code. Kaixiang Chen, Pengfei Fang |
NeurIPS | 1 |
| 2025 | DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalabstractThis paper investigates the potential of vision-language models (VLMs) in addressing the challenges of universal cross-domain retrieval (UCDR), where queries originate from unseen domains or classes. A common approach to adapting VLMs for downstream tasks involves prompt tuning, which alleviates the computational burden of full fine-tuning. However, this approach often struggles with the domain and semantic shifts inherent in UCDR. To overcome these limitations, we propose a novel prompt decoupling strategy that separates prompts into universal domain prompts (UDPs) and class prompts (CPs). Specifically, UDPs are designed to unify features from both seen and unseen domains into a cohesive universal domain, while CPs are tailored to capture class-specific visual characteristics, enabling robust retrieval across both known and unknown classes. To ensure effective decoupling, we introduce a dedicated decoupling loss that enforces the domain-agnostic nature of CPs. Additionally, we employ a regulation loss to align features from the frozen CLIP domain with those of the universal domain by selectively integrating or excluding UDPs. This mechanism fosters a synergistic domain ensemble effect, enhancing retrieval generalization across diverse domains. Finally, we propose the domain-aware triplet-hard (DaTri) loss to mitigate overfitting by reducing the risk of class collapse. The proposed framework, referred to as Domain Ensemble using Decoupled Prompts (DePro), demonstrates state-of-the-art performance and effectively enhances the model's generalization capacity across unseen domains and classes, as validated through extensive experiments. Code is here. Kaixiang Chen, Pengfei Fang, Hui Xue 0002 |
SIGIR | 1 |
| 2024 | Multi-Scale Explicit Matching and Mutual Subject Teacher Learning for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification (DG-ReID) stands out as the most challenging task and practically important branch in the ReID field, which enables the direct deployment of pre-trained models in unseen and real scenarios. Recent works have made significant efforts in this task via the image-matching paradigm, which searches for the local correspondences in the feature maps. A common practice of employing pixel-wise matching is typically used to ensure efficient matching. This, however, makes the matching susceptible to deviations caused by identity-irrelevant pixel features. On the other hand, patch-wise matching also demonstrates that it will disregard the spatial orientation of pedestrians and amplify the impact of noise. To address the mentioned issues, this paper proposes the Multi-Scale Query-Adaptive Convolution (QAConv-MS) framework, which encodes patches in the feature maps to pixels using template kernels of various scales. This enables the matching process to enjoy broader receptive fields and robustness to orientations and noises. To stabilize the matching process and facilitate the independent learning of each sub-kernel within the template kernels to capture diverse local patterns, we propose the OrthoGonal Norm (OGNorm), which consists of two orthogonal normalizations. We also present Mutual Subject Teacher Learning (MSTL) to address the potential issues of overconfidence and overfitting in the model. MSTL allows two models to individually select the most challenging data for training, resulting in more dependable soft labels that can provide mutual supervision. Extensive experiments conducted in both single-source and multi-source setups offer compelling evidence of our framework’s generalization and competitiveness. Kaixiang Chen, Pengfei Fang, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Camera-Aware Recurrent Learning and Earth Mover's Test-Time Adaption for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification (ReID) aims to design a generalizable model, which is trained under the supervision of a set of labeled source domains and can be directly deployed on unknown domains. Existing approaches simply treat each identity as a distinct class and ignore the differences among cameras. We argue that the camera information is crucial for learning discriminative representations, as people’s behavior usually varies between cameras. In this paper, we present Multi-Centroid Memory (MCM) to capture different camera information for each identity and Soft Triple Hard (ST-Hard) loss to align the information of the same identity across cameras. Furthermore, in contrast to the traditional approaches of training a single model using a parallel training mechanism, we propose the Recurrent Implicit Lifelong Learning (RILL) that feeds the source domains into the model in a continuous loop to train an expert for each domain. To make each expert further generalized to other source domains, during the training on the current domain, RILL adopts a style replay-based method to simulate the training of the previous domain, encouraging each domain’s expert to extract generalizable features. We also present Earth Mover’s Test-time Adaption (EMTA) to be used in conjunction with RILL, which enables source domains that are more similar to the test domain to play a more significant role in the test. This is achieved by our proposed Earth Mover’s Similarity (EMS), which helps model the similarities between the source and test domains. Extensive experiments on two evaluation protocols fully demonstrate our framework’s generalization and competitiveness. Kaixiang Chen, Tiantian Gong, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-Scale Query-Adaptive Convolution for Generalizable Person Re-IdentificationabstractDomain Generalization in person re-identification (ReID) aims to learn a generalizable model from a single or multi-source domain that can be directly deployed to an unseen domain without fine-tuning. In this paper, we investigate the problem of single-source domain generalization in ReID. Recent research has gained remarkable progress by treating image matching as a search for local correspondences in feature maps. However, to ensure efficient matching, they usually adopt a pixel-wise matching approach, which is prone to be deviated by the identity-irrelevant patch features in the image, such as background patches. To address this problem, we propose the Multi-scale Query-Adaptive Convolution (QAConv-MS) framework. Specifically, we adopt a group of template kernels with different scales to extract local features of different receptive fields from the original feature maps and accordingly perform the local matching process. We also introduce a self-attention branch to extract global features from the feature map as complementary information for local features. Our approach achieves state-of-the-art performances on four large-scale datasets. Kaixiang Chen, Tiantian Gong |
ICME | 1 |
| 2023 | Dynamically Adaptive Instance Normalization and Attention-Aware Incremental Meta-Learning for Generalizable Person Re-identificationabstractDomain generalization person re-identification (DG-ReID) aims to train a generalizable model over several source domains that can perform well on unseen target domains, which makes the DG-ReID task challenging since the model is not allowed to access any target data during training. The classic meta-learning method, which simulates the train-test process of DG-ReID task, is a popular and effective way for DG-ReID. Nevertheless, the method still suffers from the unstable meta-optimization problem. We thus propose a novel Dynamically Adaptive Instance Normalization and Attention-Aware Incremental Meta-Learning (DAIML) optimization method to effectively address this issue. In addition, most existing DG-ReID studies generally utilize Instance Normalization (IN) to learn domain-irrelevant features for eliminating domain shift, which may lead to the removal of effective domain-specific discriminative information that is usually useful to enhance the discriminability of a particular source. Therefore, we propose a dynamically adaptive IN (DAIN) that can balance well the learning of domain-invariant representations and domain-specific discriminative features. Furthermore, we apply channel attention and spatial attention to the proposed DAIN for further improving the domain-irrelevant discriminative information. Extensive experiments demonstrate the effectiveness of our proposed method on several DG-ReID datasets. Tiantian Gong, Kaixiang Chen |
ICME | 2 |
| 2023 | 1dFuzz: Reproduce 1-Day Vulnerabilities with Directed Differential Fuzzingabstract1-day vulnerabilities are common in practice and have posed severe threats to end users, as adversaries could learn from released patches to find them and exploit them. Reproducing 1-day vulnerabilities is also crucial for defenders, e.g., to block attack traffic against 1-day vulnerabilities. A core question that affects the effectiveness of recognizing and triggering 1-day vulnerabilities is what is the unique feature of a security patch. After conducting a large-scale empirical study, we point out that a common and unique feature of patches is the trailing call sequence (TCS) and present a novel directed differential fuzzing solution 1dFuzz to efficiently reproduce 1-day vulnerabilities in this paper. Based on the TCS feature, we present a locator 1dLoc able to find candidate patch locations via static analysis, a novel TCS-based distance metric for directed fuzzing, and a novel sanitizer 1dSan able to catch PoCs for 1-day vulnerabilities during fuzzing. We have systematically evaluated 1dFuzz on a set of real-world software vulnerabilities in 11 different settings. Results show that 1dFuzz significantly outperforms state-of-the-art (SOTA) baselines and could find up to 2.26x more 1-day vulnerabilities with a 43% shorter time. Songtao Yang 0001, Yubo He, Kaixiang Chen, Zheyu Ma, Xiapu Luo, Jianjun Chen 0005, Chao Zhang 0008 |
ISSTA | 3 |
| 2023 | AlphaEXP: An Expert System for Identifying Security-Sensitive Kernel Objects
Kaixiang Chen, Chao Zhang 0008, Zulie Pan, Qianyu Li 0001, Siliang Qin, Shenglin Xu, Min Zhang 0054, Yang Li 0215 |
USENIX Security Symposium | 2 |
| 2023 | Tunter: Assessing Exploitability of Vulnerabilities with Taint-Guided Exploitable States Exploration
Kaixiang Chen, Zulie Pan, Yuwei Li 0002, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Chao Zhang 0008 |
Comput. Secur. | 2 |
| 2023 | Debiased Contrastive Curriculum Learning for Progressive Generalizable Person Re-IdentificationabstractDomain generalization (DG) in person re-identification (ReID) is an extremely challenging but essential task, which aims to learn a generalizable model over multiple labeled source domains that can perform well on unseen target domains. Most existing DG strategies in ReID directly aggregate multiple source data together for training, incurring a large inter-domain bias and unstable model optimization that lead the model apt to overfitting domain bias and the model training more time-consuming respectively, thus hampering the generalization and convergence speed of the model. To tackle the aforementioned issues, inspired by Curriculum Learning that mimics the process of human lifelong DG learning (from easy to hard), we put forward a novel Debiased Contrastive Curriculum Learning (DCCL) strategy for DG ReID, which is designed to incrementally enhance generalization in an easy-to-hard training way that can continuously accumulate learning experience to make learning in unknown domains easier and effectively eliminate the domain bias to help the model learn rich domain-invariant discriminative features, thereby strengthening generalization and accelerating convergence for the model. In addition, to simultaneously learn class-level and instance-level discriminative representations, we raise a non-parametric hybrid contrastive loss to equip the DCCL model. We also particularly design an inter-domain mix module to variegate the features of the newly added source domain at each stage of DCCL, further establishing the advantages of DCCL. Extensive experimental results on four public ReID benchmarks fully demonstrate that our DCCL can effectively strengthen the generalization capacity of the model to unseen domains and outperform the state-of-the-art methods. Tiantian Gong, Kaixiang Chen, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Weighted Supervised Contrastive Learning and Domain Mixture for Generalized Person Re-IdentificationabstractDomain generalization(DG) in person re-identification (ReID) attracts increasing attention due to its practical applications. It aims to learn a model that, after training on multiple source domains, can be applied directly to the unseen domains without further training. In order to develop a domain-robust model for unseen domains, we propose the Memory-based Anti-Hard-instances Domain-Mix(MAD) framework. Specifically, a memory-based non-parametric contrastive loss is adopted to replace the traditional parametric cross-entropy loss. To prevent overfitting on the source domains, we present an Anti-Hard-instances module to mitigate the effect of hard instances. We also introduce a Domain-Mix module to diversify the features in the source domains, further enhancing the generalization ability of our model. Extensive experiments on four large-scale ReID datasets fully demonstrate the strong generalization and competitiveness of our framework. Kaixiang Chen, Tiantian Gong |
ICIP | 1 |
| 2022 | Distributed Ranging SLAM for Multiple Robots with Ultra-WideBand and Odometry MeasurementsabstractTo accomplish task efficiently in a multiple robots system, a problem that has to be addressed is Simultaneous Localization and Mapping (SLAM). LiDAR (Light Detection and Ranging) has been used for many SLAM solutions due to its superb accuracy, but its performance degrades in featureless environments, like tunnels or long corridors. Centralized SLAM solves the problem with a cloud server, which requires a huge amount of computational resources and lacks robustness against central node failure. To address these issues, we present a distributed SLAM solution to estimate the trajectory of a group of robots using Ultra-WideBand (UWB) ranging and odometry measurements. The proposed approach distributes the processing among the robot team and significantly mitigates the computation concern emerged from the centralized SLAM. Our solution determines the relative pose (also known as loop closure) between two robots by minimizing the UWB ranging measurements taken at different positions when the robots are in close proximity. UWB provides a good distance measure in line-of-sight conditions, but retrieving a precise pose estimation remains a challenge, due to ranging noise and unpredictable path traveled by the robot. To deal with the suspicious loop closures, we use Pairwise Consistency Maximization (PCM) to examine the quality of loop closures and perform outlier rejections. The filtered loop closures are then fused with odometry in a distributed pose graph optimization (DPGO) module to recover the full trajectory of the robot team. Extensive experiments are conducted to validate the effectiveness of the proposed approach. Ran Liu 0007, Zhongyuan Deng, Zhiqiang Cao 0004, Muhammad Shalihan, Billy Pik Lik Lau, Kaixiang Chen, Kaushik Bhowmik, Chau Yuen, U-Xuan Tan |
IROS | 6 |
| 2021 | iDEV: exploring and exploiting semantic deviations in ARM instruction processingabstractARM has become the most competitive processor architecture. Many platforms or tools are developed to execute or analyze ARM instructions, including various commercial CPUs, emulators, and binary analysis tools. However, they have deviations when processing the same ARM instructions, and little attention has been paid to systematically analyze such semantic deviations, not to mention the security implications of such deviations. In this paper, we conduct an empirical study on the ARM Instruction Semantic Deviation (ISDev) issue. First, we classify this issue into several categories and analyze the security implications behind them. Then, we further demonstrate several novel attacks which utilize the ISDev issue, including stealthy targeted attacks and targeted defense evasion. Such attacks could exploit the semantic deviations to generate malware that is specific to certain platforms or able to detect and bypass certain detection solutions. We have developed a framework iDEV to systematically explore the ISDev issue in existing ARM instructions processing tools and platforms via differential testing. We have evaluated iDEV on four hardware devices, the QEMU emulator, and five disassemblers which could process the ARMv7-A instruction set. The evaluation results show that, over six million instructions could cause dynamic executors (i.e., CPUs and QEMU) to present different runtime behaviors, and over eight million instructions could cause static disassemblers yielding different decoding results, and over one million instructions cause inconsistency between dynamic executors and static disassemblers. After analyzing the root causes of each type of deviation, we point out they are mostly due to ARM unpredictable instructions and program defects. Shisong Qin, Chao Zhang 0008, Kaixiang Chen, Zheming Li |
ISSTA | 3 |
| 2021 | VScape: Assessing and Escaping Virtual Call Protections
Kaixiang Chen, Chao Zhang 0008, Tingting Yin, Xingman Chen |
USENIX Security Symposium | 1 |
| 2018 | Revery: From Proof-of-Concept to ExploitableabstractAutomatic exploit generation is an open challenge. Existing solutions usually explore in depth the crashing paths, i.e., paths taken by proof-of-concept (POC) inputs triggering vulnerabilities, and generate exploits when exploitable states are found along the paths. However, exploitable states do not always exist in crashing paths. Moreover, existing solutions heavily rely on symbolic execution and are not scalable in path exploration and exploit generation. In addition, few solutions could exploit heap-based vulnerabilities. In this paper, we propose a new solution revery to search for exploitable states in paths diverging from crashing paths, and generate control-flow hijacking exploits for heap-based vulnerabilities. It adopts three novel techniques:(1) a digraph to characterize a vulnerability's memory layout and its contributor instructions;(2) a fuzz solution to explore diverging paths, which have similar memory layouts as the crashing paths, in order to search more exploitable states and generate corresponding diverging inputs;(3) a stitch solution to stitch crashing paths and diverging paths together, and synthesize EXP inputs able to trigger both vulnerabilities and exploitable states. We have developed a prototype of revery based on the binary analysis engine angr, and evaluated it on a set of 19 real world CTF (capture the flag) challenges. Experiment results showed that it could generate exploits for 9 (47%) of them, and generate EXP inputs able to trigger exploitable states for another 5 (26%) of them. Chao Zhang 0008, Xiaobo Xiang, Wenjie Li 0006, Xiaorui Gong, Bingchang Liu, Kaixiang Chen |
CCS | 8 |