EDBT 2026 Demo / reviewers in the wild / expert
Jingxuan Zhou
dblp:283/3717
· DBLP profile ↗
15ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsabstractDownstream fine-tuning of Multimodal Large Language Models (MLLMs) is advancing rapidly, allowing general models to achieve superior performance on domain-specific tasks. Yet most prior research focuses on performance gains and overlooks the vulnerability of the fine-tuning pipeline: attackers can easily poison the dataset to implant backdoors into MLLMs. We conduct an in-depth investigation of backdoor attacks on MLLMs and reveal the phenomenon of Attention Hijacking and its Hierarchical Mechanism. Guided by this insight, we propose PurMM, a test-time backdoor purification framework that removes visual tokens exhibiting anomalous attention, thereby avoiding targeted outputs while restoring correct answers. PurMM contains three stages: (1) locating tokens with abnormal attention, (2) filtering them using deep-layer cues, and (3) zeroing out their corresponding components in the visual embeddings. Unlike existing defences, PurMM dispenses with retraining and training-process modifications, operating at test-time to restore model performance while eliminating the backdoor. Extensive experiments across multiple MLLMs and datasets show that PurMM maintains normal performance, sharply reduces attack success rates, and consistently converts backdoor outputs to benign ones, offering a new perspective for safeguarding MLLMs. Wenzheng Jiang, Ke Liang 0006, Xuankun Rong, Jingxuan Zhou, Zhengyi Zhong, Guancheng Wan, Ji Wang 0002 |
AAAI | 4 |
| 2025 | Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent DetectionabstractZero-shot multi-intent detection is capable of capturing multiple intents within a single utterance without any training data, which gains increasing attention. Building on the success of large language models (LLM), dominant approaches in the literature explore prompting techniques to enable zero-shot multi-intent detection. While significant advancements have been witnessed, the existing prompting approaches still face two major issues: lacking explicit reasoning and lacking interpretability. Therefore, in this paper, we introduce a Divide-Solve-Combine Prompting (DSCP) to address the above issues. Specifically, DSCP explicitly decomposes multi-intent detection into three components including (1) single-intent division prompting is utilized to decompose an input query into distinct sub-sentences, each containing a single intent; (2) intent-by-intent solution prompting is applied to solve each sub-sentence recurrently; and (3) multi-intent combination prompting is employed for combining each sub-sentence result to obtain the final multi-intent result. By decomposition, DSCP allows the model to track the explicit reasoning process and improve the interpretability. In addition, we propose an interactive divide-solve-combine prompting (Inter-DSCP) to naturally capture the interaction capabilities of large language models. Experimental results on two standard multi-intent benchmarks (i.e., MixATIS and MixSNIPS) reveal that both DSCP and Inter-DSCP obtain substantial improvements over baselines, achieving superior performance and higher interpretability. Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Hao Fei 0003, Wanxiang Che, Min Li 0007 |
AAAI | 3 |
| 2025 | IPAU: Integrating Prototype, Affinity, and Uncertainty for Weakly-Supervised Histopathology SegmentationabstractWeakly supervised semantic segmentation (WSSS) reduces annotation burden by using only image-level labels for histopathology image segmentation. Current WSSS methods face the challenge of bridging the information gap between weak labels and dense prediction tasks, which often results in insufficient class activation maps (CAMs) and increased false positives. Most approaches address this by mining additional object-related information. Following this direction, we propose IPAU, a framework that integrates Prototype, Affinity, and Uncertainty to enhance WSSS. In our IPAU, Prototype-based Information Enhancement (PIE) that uses class-wise prototypes to enrich CAM generation. Affinity-based Self-Refinement (ASR) that refines CAMs into pseudo-masks using affinity correlations without extra training. And Uncertainty-Aware PseudoSupervision (UAPS) that mitigates noise by focusing learning on reliable regions. Experiments on two public histopathology WSSS datasets demonstrate that our IPAU achieves state-of-theart performance. Jiatai Lin, Jingxuan Zhou, Zhenwei Shi 0002, Zaiyi Liu, Xiao-jing Guo, Chu Han |
BIBM | 2 |
| 2025 | Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse AdapterabstractFederated Learning is a promising paradigm for privacy-preserving collaborative model training. In practice, it is essential not only to continuously train the model to acquire new knowledge but also to guarantee old knowledge the right to be forgotten (i.e., federated unlearning), especially for privacy-sensitive information or harmful knowledge. However, current federated unlearning methods face several challenges, including indiscriminate unlearning of cross-client knowledge, irreversibility of unlearning, and significant unlearning costs. To this end, we propose a method named FUSED, which first identifies critical layers by analyzing each layer’s sensitivity to knowledge and constructs sparse unlearning adapters for sensitive ones. Then, the adapters are trained without altering the original parameters, overwriting the unlearning knowledge with the remaining knowledge. This knowledge overwriting process enables FUSED to mitigate the effects of indiscriminate unlearning. Moreover, the introduction of independent adapters makes unlearning reversible and significantly reduces the unlearning costs. Finally, extensive experiments on three datasets across various unlearning scenarios demonstrate that FUSED’s effectiveness is comparable to Retraining, surpassing all other baselines while greatly reducing unlearning costs. Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Shuai Zhang 0004, Jingxuan Zhou, Lingjuan Lyu, Wei Yang Bryan Lim |
CVPR | 5 |
| 2025 | CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language UnderstandingabstractSlot filling and intent detection are two highly correlated tasks in spoken language understanding (SLU). Recent SLU research attempts to explore zero-shot prompting techniques in large language models to alleviate the data scarcity problem. Nevertheless, the existing prompting work ignores the cross-task interaction information for SLU, which leads to sub-optimal performance. To solve this problem, we present the pioneering work of Cross-task Interactive Prompting (CroPrompt) for SLU, which enables the model to interactively leverage the information exchange across the correlated tasks in SLU. Additionally, we further introduce a multi-task self-consistency mechanism to mitigate the error propagation caused by the intent information injection. We conduct extensive experiments on the standard SLU benchmark and the results reveal that CroPrompt consistently outperforms the existing prompting approaches. In addition, the multi-task self-consistency mechanism can effectively ease the error propagation issue, thereby enhancing the performance. We hope this work can inspire more research on cross-task prompting for SLU. Libo Qin 0001, Fuxuan Wei, Qiguang Chen, Jingxuan Zhou, Shijue Huang, Jiasheng Si, Wenpeng Lu, Wanxiang Che |
ICASSP | 4 |
| 2025 | Frequency-Domain Guided Multiple Parallel Kernels Network for Low-Light Remote Sensing Image EnhancementabstractDue to dark environments, optical aberrations, etc, the remote sensing images are often submerged under low contrast degradation, which greatly hinders their practical applications for agricultural management and other related tasks. The surface features of remote sensing images are often continuously distributed in space, thus, the sizes of the network’s receptive fields and its ability to learn long-range dependencies are crucial for restoring low-light remote sensing images. Existing methods based on CNN provide limited receptive fields, while Transformer-based methods are constrained by their quadratic computational complexity. To cope with these issues, we propose a novel low-light remote sensing image enhancement network that combines multi-scale receptive fields with frequency-domain attention. Specifically, this network employs multiple parallel kernels of varying sizes to learn multi-scale local features in the spatial domain and complements frequency-domain information to learn global long-range correlations, which achieves local-global feature extraction and further facilitates subsequent degraded images enhancement. We have conducted extensive experiments to demonstrate that our network outperforms existing methods quantitatively and achieves exceptional visual performance, which fully highlights the effectiveness and superiority of our method in enhancing low-light remote sensing images. Jingxuan Zhou, Xiongxin Tang, Fanjiang Xu |
ICASSP | 1 |
| 2025 | CUT: Pruning Pre-trained Multi-task Models into Compact Models for Edge Devices
Jingxuan Zhou, Weidong Bao 0001, Ji Wang 0002, Zhengyi Zhong |
ICIC (17) | 1 |
| 2025 | Learning Adaptive High-Frequency Semantic Guidance for Low-light Image EnhancementabstractThe low-light image enhancement has always been an important yet challenging task, which attracts significant attention in many fields. However, prior methods either ignore integrating semantic priors or depend on the masks generated by the pre-trained segmentation model. This way is complex and inevitably leads to inaccurate masks when facing unseen scenarios, which may be incompatible with the original feature and result in suboptimal performance. To address this issue, we first consider the high-frequency physical prior is more related to structural and textural properties, which embrace the rich semantic clues and can adaptively assist the learning process under various scenarios. Inspired by this, we propose the high-frequency semantic-aware guidance framework (HighFreNet) to leverage the guidance of semantic information tailored for enhancing low-light images. Specifically, the core parts are the novel Frequency-based Semantic Embedding Module (FSEM) and the Spatial-based Semantic Embedding Module (SSEM), which are separately designed to fully exploit the structure knowledge to modulate the original representation from frequency and spatial perspectives. Extensive experiments showcase that our method significantly outperforms the state-of-the-art methods on five benchmark datasets both in natural and remote sensing environments. Jingxuan Zhou, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu |
ICME | 2 |
| 2025 | Efficient Multi-Task Modeling through Automated Fusion of Trained ModelsabstractAlthough multi-task learning is widely applied in intelligent services, traditional multi-task modeling methods often require customized designs based on specific task combinations, resulting in a cumbersome modeling process. Inspired by the rapid development and excellent performance of single-task models, this paper proposes an efficient multi-task modeling method that can automatically fuse trained single-task models with different structures and tasks to form a multi-task model. As a general framework, this method allows modelers to simply prepare trained models for the required tasks, simplifying the modeling process while fully utilizing the knowledge contained in the trained models. This eliminates the need for excessive focus on task relationships and model structure design. To achieve this goal, we consider the structural differences among various trained models and employ model decomposition techniques to hierarchically decompose them into multiple operable model components. Furthermore, we design an Adaptive Knowledge Fusion (AKF) module based on Transformer, which adaptively integrates intra-task and inter-task knowledge based on model components. Through the proposed method, we achieve efficient and automated construction of multi-task models, and its effectiveness is verified through extensive experiments on three datasets. Our code and related baseline methods can be found at: https://github.com/zxccvdql/EMM. Jingxuan Zhou, Weidong Bao 0001, Ji Wang 0002, Dayu Zhang, Zhengyi Zhong |
SMC | 1 |
| 2025 | All on board: Efficient reinforcement learning with milestone aggregation in asynchronous distributed training for RTS games
Dayu Zhang, Weidong Bao 0001, Ji Wang 0002, Xiongtao Zhang, Jingxuan Zhou, Yaohong Zhang |
Neurocomputing | 5 |
| 2024 | SS-WSSS: Small-Scale Weakly Supervised Semantic Segmentation for Histopathology ImageabstractSemantic segmentation for histopathology images is one of the fundamental tasks in computational pathology. Due to the high cost of pixel-level annotation acquisition, the weakly supervised semantic segmentation (WSSS) attempts to achieve information-intensive segmentation task for histopathology images to reduce the labeling effort of pathologists by leveraging image-level labels. However, traditional WSSS requires a large-scale training set with image-level labels, which still imposes considerable labeling costs on pathologists. To this end, this work proposes a Small-Scale Weakly Supervised Semantic Segmentation (SS-WSSS) approach to achieve the comparable performance only with small-scale weakly-labeled data to further reduce pathologist’s labeling effort. Since histopathology images can easily generate massive unlabeled data, our SS-WSSS aims to learn with the unlabeled data to bridge the information gap. First, we propose a Single-to-Multi Prototype Similarity (S2M-PS) method to generate reliable pseudo-labels for unlabeled data by measuring the similarity between single-label prototypes and multi-label feature maps. Then, we introduce a Cross-Task CoTraining (CT2) method for pseudo-supervision of models with pseudo-labels self-refinement to avoid overfitting to noisy labels. We conduct the experiment on two public datasets to demonstrate the effectiveness of our SS-WSSS. In the experiment, our method achieves comparable performance with SOTA methods only using 30% labeled data. Jiatai Lin, Guoqiang Han 0002, Jingxuan Zhou, Zhenwei Shi 0002, Zaiyi Liu, Chu Han |
BIBM | 3 |
| 2024 | LabCLIP: Label-Enhanced Clip for Improving Zero-Shot Text ClassificationabstractZero-shot text classification aims to handle the text classification task without any annotated training data, which can greatly alleviate the data scarcity problem. Current dominant approaches follow a novel text-image matching paradigm, reformulating zero-shot text classification into a text-image matching problem, which can capture the visual image information and show promising performance. Nevertheless, existing text-image matching approaches solely focus on the visual image information, ignoring the semantic knowledge embedded in the text labels. To address the challenge, in the work, we present a label-enhanced CLIP framework (Lab-CLIP) for zero-shot text classification to consider both the visual image and text label semantic information simultaneously. Specifically, LabCLIP first converts the label into the corresponding image, and then injects the text label into the corresponding label image to explicitly capture the label semantic knowledge. We conduct experiments on 8 publicly available zero-shot text classification datasets and experimental results indicate that LabCLIP outperforms previous approaches on all datasets (with 4.3% improvement on average). In addition, we provide extensive analysis on exploring how to effectively incorporate the text label information. Yongheng Zhang 0001, Peng Wang 0168, Qiguang Chen, Jingxuan Zhou, Yongmei Michelle Wang, Min Li 0007, Libo Qin 0001 |
ICASSP | 4 |
| 2024 | Decoupling Breaks Data Barriers: A Decoupled Pre-training Framework for Multi-intent Spoken Language Understanding
Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Qinzheng Li, Chunlin Lu, Wanxiang Che |
IJCAI | 3 |
| 2024 | Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-ThoughtabstractChain-of-Thought (CoT) reasoning has emerged as a promising approach for enhancing the performance of large language models (LLMs) on complex reasoning tasks. Recently, a series of studies attempt to explain the mechanisms underlying CoT, aiming to deepen the understanding of its efficacy. Nevertheless, the existing research faces two major challenges: (1) a lack of quantitative metrics to assess CoT capabilities and (2) a dearth of guidance on optimizing CoT performance. Motivated by this, in this work, we introduce a novel reasoning boundary framework (RBF) to address these challenges. To solve the lack of quantification, we first define a reasoning boundary (RB) to quantify the upper-bound of CoT and establish a combination law for RB, enabling a practical quantitative approach applicable to various real-world CoT tasks. To address the lack of optimization, we propose three categories of RBs. We further optimize these categories with combination laws focused on RB promotion and reasoning path optimization for CoT improvement. Through extensive experiments on 27 models and 5 tasks, the study validates the existence and rationality of the proposed framework. Furthermore, it explains the effectiveness of 10 CoT strategies and guides optimization from two perspectives. We hope this work can provide a comprehensive understanding of the boundaries and optimization strategies for reasoning in LLMs. Our code and data are available at https://github.com/LightChen233/reasoning-boundary. Qiguang Chen, Libo Qin 0001, Jiaqi Wang 0012, Jingxuan Zhou, Wanxiang Che |
NeurIPS | 4 |
| 2020 | Image-based early predictions of functional properties in cell manufacturingabstractEffective cell manufacturing is essential to realizing the full potential of cell-based therapies but faces a multitude of challenges. One of the major challenges is the identification of critical quality attributes (CQAs), especially ones that enable early predictions of functional properties of the final products. The main goal of this study is to develop machine learning models for early predictions of the functional properties of mesenchymal stromal/stem cells(MSCs) in cell manufacturing. Deep learning models are trained and tested for image-based prediction of functional property-Collagen II expression after chondrogenic differentiation-of MSCs cells. During the MSC expansion, images of culturing wells were collected daily in the first six days, and the Collagen II level was assayed at the end of differentiation, following expansion. For each day, a deep learning model was trained with images from a specific experimental condition, and each model was tested with images from the same condition and also from other conditions. The trained neural network models showed 70-90 percent accuracy. Most of the models across different days and conditions show high consistency, especially models trained with images past day 2 of cell culture. Such consistency suggests that models are picking up similar features in predicting chondrogenesis capability. Our study highlighted the potential of deep neural network models used for early predictions of the functional properties of MSCs in cell manufacturing. Hong Seo Lim, Madeline E. Smerchansky, Jingxuan Zhou, Paramita Chatterjee, Angela C. Jimenez, Krishnendu Roy, Peng Qiu |
BIBM | 3 |