Junxi Chen

dblp:188/1312 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mixture of knowledge from multiple pre-trained foundation models for few-shot recognition
Junxi Chen, Zhiyong Gan, Guangxing Wu
Neurocomputing1
2025 Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
abstract
Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, we reveal that T2I-DMs equipped with the IP-Adapter (T2I-IP-DMs) enable a new jailbreak attack named the hijacking attack. We demonstrate that, by uploading imperceptible image-space adversarial examples (AEs), the adversary can hijack massive benign users to jailbreak an Image Generation Service (IGS) driven by T2I-IP-DMs and mislead the public to discredit the service provider. Worse still, the IP-Adapter’s dependency on open-source image encoders reduces the knowledge required to craft AEs. Extensive experiments verify the technical feasibility of the hijacking attack. In light of the revealed threat, we investigate several existing defenses and explore combining the IP-Adapter with adversarially trained models to overcome existing defenses’ limitations. Our code is available at https://github.com/fhdnskfbeuv/attackIPA.
Junxi Chen, Junhao Dong 0001, Xiaohua Xie
CVPR1
2025 Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supervised paradigm. To address these limitations, we propose a novel paradigm: Single-Frame supervised VAD (SF-VAD), which uses a single annotated abnormal frame per abnormal video. SF-VAD ensures annotation efficiency while offering precise anomaly reference, facilitating robust anomaly modeling, and enhancing the detection of subtle anomalies in complex visual contexts. To validate its effectiveness, we construct three SF-VAD benchmarks by manually re-annotating the ShanghaiTech, UCF-Crime, and XD-Violence datasets in a practical procedure. Further, we devise Frame-guided Progressive Learning (FPL), to generalize sparse frame supervision to event-level anomaly understanding. FPL first leverages evidential learning to estimate anomaly relevance guided by annotated frames. Then it extends anomaly supervision by mining discrete abnormal events based on anomaly relevance and feature similarity. Meanwhile, FPL decouples normal patterns by isolating distinct normal frames outside abnormal events, reducing false alarms. Extensive experiments show SF-VAD achieves state-of-the-art detection results while offering a favorable trade-off between performance and annotation cost.
Junxi Chen, Liang Li 0003, Yunbin Tu, Li Su 0003, Zhe Xue, Qingming Huang
NeurIPS1
2025 Piecewise Physics-Constrained Neural Networks for solving second-order impulsive differential equations
Junxi Chen, Liangcai Mei, Jiarong Han, Wulong Yao
Neurocomputing1
2025 EEMD-MST-Resnet: a hybrid deep learning approach for predicting passenger flow in urban transportation hubs
Junxi Chen, Zhenlin Wei
Neural Comput. Appl.1
2025 DAT: Dual-Branch Adapter-Tuning for Few-Shot Recognition
abstract
Parameter-Efficient Fine-Tuning methods based on vision-language models (such as CLIP) for few-shot learning have recently received considerable attention. However, previous works only fine-tune either the image or text branch, breaking the alignment of the original two branches, meanwhile fine-tuning both branches of the CLIP would inevitably introduce more trainable parameters and likely cause more severe over-fitting due to the limited training data. In this study, we propose a novel Dual-branch Adapter-Tuning framework (DAT), which collaboratively trains the visual adapter and textual adapter added to the two branches of the original CLIP with multiple consistency constraints. By effectively utilizing the semantically detailed class-specific prompts and outputs of the original CLIP to guide the fine-tuning of both branches, our method gains exceptional adaptation ability to the downstream few-shot learning tasks and alleviates the over-fitting issue, meanwhile maximally preserving the generalization ability of the original CLIP model. Our proposed framework has achieved superior performance on diverse datasets under various few-shot learning settings compared to the existing approaches. The source code is available athttps://github.com/SandyXi/DAT.
Junxi Chen, Guangxing Wu, Hongxiang Li 0004, Jiankang Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Releasing Inequality Phenomenon in ℓ∞-Norm Adversarial Training via Input Gradient Distillation
abstract
Adversarial training (AT) is considered the most effective defense against adversarial attacks. However, a recent study revealed that ℓ∞-norm adversarial training ( ℓ∞-AT) will also induce unevenly distributed input gradients, which is called the inequality phenomenon. This phenomenon makes the ℓ∞ -norm adversarially trained model more vulnerable than the standard-trained model when high-attribution or randomly selected pixels are perturbed, enabling robust and practical closed-box attacks against ℓ∞ -adversarially trained models. In this paper, we propose a simple yet effective method called Input Gradient Distillation (IGD) to release the inequality phenomenon in ℓ∞-AT. IGD distills the standard-trained teacher model’s equal decision pattern into the ℓ∞-adversarially trained student model by aligning input gradients of the student model and the standard-trained model with the Cosine Similarity. Experiments show that IGD can mitigate the inequality phenomenon and its threats while preserving adversarial robustness. Compared to vanilla ℓ∞-AT, IGD reduces error rates against inductive noise, inductive occlusion, random noise, and noisy images in ImageNet-C by up to 60%, 16%, 50%, and 21%, respectively. Other than empirical experiments, we also conduct a theoretical analysis to explain why releasing the inequality phenomenon can improve such robustness and discuss why the severity of the inequality phenomenon varies according to the dataset’s image resolution.
Junxi Chen, Junhao Dong 0001, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Inf. Forensics Secur.1
2025 Leveraging Spatial-Temporal Heterogeneity and Cross-Mode Interactions: A Meta-Learning Approach for Multimodal Transportation Demand Prediction
abstract
Accurately and jointly predicting multimodal transportation demand is crucial for pre-allocating transport resources, enhancing the resilience of traffic systems. However, current approaches insufficiently explore inter- and intra-mode heterogeneity, resulting in undifferentiated dependency extraction. Moreover, existing research struggles to model cross-mode interactions among three or more transportation modes and adapt to dynamic relations in multimodal demand. To address these limitations, we propose a novel multimodal demand prediction model based on a meta-parameter learning network (MMDNet), centered on characterizing multimodal traffic spatial-temporal heterogeneity and unifying the modeling of cross-mode interactions. Our model features: 1) a spatial-temporal heterogeneity meta-parameter learning method, capturing both inter- and intra-mode heterogeneity to steer more targeted dependency extraction than previous studies; 2) a spatial-temporal evolving unified graph generator, transcending prior studies’ limitations in unifying dynamic interactions across three or more modes by creating dynamic unified graphs. Extensive experiments on three real-world datasets (New York, Beijing and Chicago) covering four different traffic modes are carried out to evaluate the MMDNet. The model achieves a 6.65% performance gain over advanced baselines and demonstrates strong cross-city adaptability. Abundant interpretability analyses show our model can semantically encode explainable cross-mode interactions and differences between modes. Source codes are available athttps://github.com/zhjiang1/MMDNet
Zhihuan Jiang, Ailing Huang, Renhe Jiang, Junxi Chen, Yoshihide Sekimoto
IEEE Trans. Intell. Transp. Syst.4
2024 Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly Detection
abstract
Weakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels, Multi-Instance Learning (MIL) is prevailing in wVAD. However, MIL suffers from insufficiency of binary supervision to model diverse abnormal patterns. Besides, the coupling between abnormality and its context hinders the learning of clear abnormal event boundary. In this paper, we propose prompt-enhanced MIL to detect various abnormal events while ensuring clear event boundaries. Concretely, we design the abnormal-aware prompts by using abnormal class annotations together with learnable prompt, which can incorporate semantic priors into video features dynamically. The detector can utilize the semantic-rich features to capture diverse abnormal patterns. In addition, normal context prompt is introduced to amplify the distinction between abnormality and its context, facilitating the generation of clear boundary. With the mutual enhancement of abnormal-aware and normal context prompt, the model can construct discriminative representations to detect divergent anomalies without ambiguous event boundaries. Extensive experiments demonstrate our method achieves SOTA performance on three public benchmarks. The code is available at https://github.com/Junxi-Chen/PE-MIL.
Junxi Chen, Liang Li 0003, Li Su 0003, Zhengjun Zha, Qingming Huang
CVPR1
2024 Robust Distillation via Untargeted and Targeted Intermediate Adversarial Samples
abstract
Adversarially robust knowledge distillation aims to com-press large-scale models into lightweight models while preserving adversarial robustness and natural performance on a given dataset. Existing methods typically align probability distributions of natural and adversarial samples between teacher and student models, but they overlook intermediate adversarial samples along the “adversarial path” formed by the multi-step gradient ascent of a sample towards the decision boundary. Such paths capture rich information about the decision boundary. In this paper, we propose a novel adversarially robust knowledge distillation approach by incorporating such adversarial paths into the alignment process. Recognizing the diverse impacts of intermediate adversarial samples (ranging from benign to noisy), we propose an adaptive weighting strategy to selectively em-phasize informative adversarial samples, thus ensuring efficient utilization of lightweight model capacity. Moreover, we propose a dual-branch mechanism exploiting two following insights: (i) complementary dynamics of adversar-ial paths obtained by targeted and untargeted adversarial learning, and (ii) inherent differences between the gradient ascent path from class$c_{i}$towards the nearest class bound-ary and the gradient descent path from a specific class$c_{j}$towards the decision region of$c_{i}(i\neq j)$. Comprehensive experiments demonstrate the effectiveness of our method on lightweight models under various settings.
Junhao Dong 0001, Piotr Koniusz, Junxi Chen, Z. Jane Wang 0001, Yew-Soon Ong
CVPR3
2024 Adversarially Robust Few-shot Learning via Parameter Co-distillation of Similarity and Class Concept Learners
abstract
Few-shot learning (FSL) facilitates a variety of computer vision tasks yet remains vulnerable to adversarial attacks. Existing adversarially robust FSL methods rely on either visual similarity learning or class concept learning. Our analysis reveals that these two learning paradigms are complementary, exhibiting distinct robustness due to their unique decision boundary types (concepts clustering by the visual similarity label vs. classification by the class labels). To bridge this gap, we propose a novel framework unifying adversarially robust similarity learning and class concept learning. Specifically, we distill parameters from both network branches into a “unified embedding model” during robust optimization and redistribute them to individual network branches periodically. To capture generalizable robustness across diverse branches, we initialize adversaries in each episode with cross-branch class-wise “global adversarial perturbations” instead of less informative random initialization. We also propose a branch robustness harmonization to modulate the optimization of similarity and class concept learners via their relative adversarial robustness. Extensive experiments demonstrate the state-of-the-art performance of our method in diverse few-shot scenarios.
Junhao Dong 0001, Piotr Koniusz, Junxi Chen, Xiaohua Xie, Yew-Soon Ong
CVPR3
2024 Adversarially Robust Distillation by Reducing the Student-Teacher Variance Gap
Junhao Dong 0001, Piotr Koniusz, Junxi Chen, Yew-Soon Ong
ECCV (4)3
2024 Region Attention Fine-tuning with CLIP for Few-shot Classification
abstract
With the advancements in visual language models such as CLIP and their strong performance in zero-shot recognition, numerous CLIP-based methods have emerged in the field of few-shot classification. However, many of them do not fully leverage the abundant feature information within the CLIP visual encoder and overlook the issue of varying region-specific importance for image classification across different datasets. To address these limitations, we present an attention pooling-based framework for few-shot fine-tuning. Our framework enables the model to learn task-specific attention weights for image regions, while also incorporating background features and a consistency constraint to enhance training. As a result, our approach outperforms the state-of-the-art approaches on 11 benchmarks, demonstrating its effectiveness.
Guangxing Wu, Junxi Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001
ICME2
2024 DFCG: A Dual-Frequency Cascade Graph model for semi-supervised ultrasound image segmentation with diffusion model
Yifeng Yao, Xingxing Duan, Aiping Qu, Mingzhi Chen 0002, Junxi Chen, Lingna Chen
Knowl. Based Syst.5
2023 Feature Adaptation with CLIP for Few-shot Classification
abstract
Large Vision-Language models such as CLIP have demonstrated impressive capabilities in zero-shot recognition. To apply CLIP to few-shot classification tasks, several methods have been proposed based on CLIP, achieving significant improvements. However, these methods either insufficiently leverage CLIP’s prior knowledge during training or neglect the impact of feature adaptation. In this paper, we propose FAR, a novel approach that balances distribution-altered Feature Adaptation with pRior knowledge of CLIP to further improve the performance of CLIP in few-shot classification tasks. Firstly, we introduce an adapter that enhances the effectiveness of CLIP adaptation by amplifying the differences between the fine-tuned CLIP features and the original CLIP features. Secondly, we leverage the prior knowledge of CLIP to mitigate the risk of overfitting. Through this framework, a good trade-off between feature adaptation and preserving prior knowledge is achieved, enabling effective utilization of both components to enhance performance on downstream tasks. We evaluate our method on over 10 datasets for classification, and our approach consistently outperforms existing methods, demonstrating its effectiveness and robustness.
Guangxing Wu, Junxi Chen, Wentao Zhang 0005
MMAsia2
2023 Effective hybrid attention network based on pseudo-color enhancement in ultrasound image segmentation
Xuping Huang, Junxi Chen, Lingna Chen
Image Vis. Comput.3
2020 Reversible Data Hiding in JPEG Images Based on Negative Influence Models
abstract
Reversible data hiding (RDH) can be used to imperceptibly embed data into images in a reversible manner. Many RDH schemes have been developed for uncompressed images. However, JPEG compressed images are more widely used in our daily lives. The existing RDH techniques for JPEG images may cause significant distortion or a large increase in the file size of marked images. In this paper, a novel RDH scheme for JPEG images is proposed. First, the negative influence models of data embedding, including image visual distortion model and file size change model, are mathematically established. Then, a negative index for each frequency is defined as the weighted sum of the normalized average image visual distortion and the normalized average file size change per 1-bit hidden data, and the frequencies with small negative indices will be used for data embedding with a high priority. The weighting factor can be adjusted according to the user's preference for less image distortion or smaller file size. Lastly, secret data is embedded into non-zero quantized AC coefficients of the selected frequencies in ascending order of zero-run length. Extensive experiments conducted on typical images and a well-known image database show that the presented negative influence models are effective, and the proposed RDH scheme based on the models can achieve low image distortion and small increase in the file size of marked images.
Junxi Chen, Shaohua Tang
IEEE Trans. Inf. Forensics Secur.2
2019 A Novel High-Capacity Reversible Data Hiding Scheme for Encrypted JPEG Bitstreams
abstract
As cloud storage becomes more common, concerns about the invasion of privacy are increasing. When images are stored in an encrypted form in the public cloud, reversible data hiding in the encrypted domain can be applied to embed additional data within the encrypted images for ease of management. Most existing works focus on uncompressed images and are not applicable to JPEG images, which are widely used throughout the Internet. Therefore, in this paper, a novel reversible data hiding scheme for encrypted JPEG bitstreams is proposed. First, an effective method of bitstream-based JPEG image encryption is employed to encrypt plaintext JPEG images. Then, we present a reversible data hiding technique for encrypted JPEG images based on invariant zero-run length in the zero-run value pairs. In the cloud, additional data, such as labels, timestamps, origins, and authentication messages, are directly embedded into the encrypted JPEG images with our proposed data hiding technique. From the marked encrypted JPEG images, the extraction of the hidden data and the recovery of the original images can be done independently. Extensive experiments performed on several typical images and three well-known image databases show that the proposed scheme can achieve much higher embedding capacity than that of most recent schemes, and the file sizes of the marked encrypted JPEG images are well preserved compared to those of the related methods. In addition, we provide further analysis to show that the proposed scheme has good format compatibility and low computational complexity.
Junxi Chen, Weiqi Luo 0001, Shaohua Tang, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Novel AMR-WB Speech Steganography Based on Diameter-Neighbor Codebook Partition
abstract
Steganography is a means of covert communication without revealing the occurrence and the real purpose of communication. The adaptive multirate wideband (AMR-WB) is a widely adapted format in mobile handsets and is also the recommended speech codec for VoLTE. In this paper, a novel AMR-WB speech steganography is proposed based on diameter-neighbor codebook partition algorithm. Different embedding capacity may be achieved by adjusting the iterative parameters during codebook division. The experimental results prove that the presented AMR-WB steganography may provide higher and flexible embedding capacity without inducing perceptible distortion compared with the state-of-the-art methods. With 48 iterations of cluster merging, twice the embedding capacity of complementary-neighbor-vertices-based embedding method may be obtained with a decrease of only around 2% in speech quality and much the same undetectability. Moreover, both the quality of stego speech and the security regarding statistical steganalysis are better than the recent speech steganography based on neighbor-index-division codebook partition.
Junxi Chen, Shichang Xiao, Xiao-Yu Huang, Shaohua Tang
Secur. Commun. Networks2
2016 The design of an efficient swap mechanism for hybrid DRAM-NVM systems
abstract
Non-Volatile Memory (NVM) is becoming an attractive candidate to be the swap area in embedded systems for its near-DRAM speed, low energy consumption, high density, and byte-addressability. Swapping data from DRAM out to NVM, however, can cause large performance/energy penalty and deplete the lifetime of NVM. Traditional swap mechanisms may need to be re-studied. Even through there are several swap mechanisms proposed for the hybrid DRAM-NVM systems, most of them have limited performance without considering the data access features of applications.
Xianzhang Chen, Edwin H.-M. Sha, Weiwen Jiang, Qingfeng Zhuge, Junxi Chen, Jiejie Qin, Yuansong Zeng
EMSOFT5
2016 A unified framework for designing high performance in-memory and hybrid memory file systems
Xianzhang Chen, Edwin H.-M. Sha, Qingfeng Zhuge, Weiwen Jiang, Junxi Chen
J. Syst. Archit.5