Xiu Su

dblp:189/3416 · DBLP profile ↗
← Back
48ranked-venue papers
7as first author
46since 2021 · last 2027
0000-0002-9863-5404ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Capacity-bounded expansion: Transferability-driven mixture of experts for continual process monitoring
Qinzhe Wang, Keying Ding, Keke Huang, Xiu Su, Chang Xu 0002, Chunhua Yang 0001
Expert Syst. Appl.6
2026 Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization
abstract
Federated Domain Generalization (FDG) aims to collaboratively train a global model across distributed clients that can generalize well on unseen domains. However, existing FDG methods typically struggle with cross-client data heterogeneity and incur significant communication and computation overhead. To address these challenges, this paper presents a new FDG framework, dubbed FaST-PT, which facilitates local feature augmentation and efficient unseen domain adaptation in a distributed manner. First, we propose a lightweight Multi-Modal Style Transfer (MST) method to transform image embedding under text supervision, which could expand the training data distribution and mitigate domain shift. We then design a dual-prompt module that decomposes the prompt into global and domain prompts. Specifically, global prompts capture general knowledge from augmented embedding across clients, while domain prompts capture domain-specific knowledge from local data. Besides, Domain-aware Prompt Generation (DPG) is introduced to adaptively generate suitable prompts for each sample, which facilitates unseen domain adaptation through knowledge fusion. Extensive experiments on four cross-domain benchmark datasets, e.g., PACS and DomainNet, demonstrate the superior performance of FaST-PT over SOTA FDG methods such as FedDG-GA and DiPrompt. Ablation studies further validate the effectiveness and efficiency of FaST-PT.
Yuliang Chen, Xi Lin 0003, Jun Wu 0001, Xiangrui Cai, Qiaolun Zhang, Xichun Fan, Jiapeng Xu, Xiu Su
AAAI8
2026 ROVER: Robust Generative Continual Identity Unlearning Against Relearning Attacks
abstract
Recent generative unlearning models synthesize high quality samples while protecting private information by unlearning the identity. However, existing generative identity unlearning methods face two challenges in multi-identity unlearning: 1) identity conflicts, which cause conflicts of model parameters in the continuous erasure of multiple identities; 2) fragile unlearning, where the model's unlearning ability deteriorates or fails under malicious attacks. In this paper, we introduce a critical yet under-explored task called robust multi-identity unlearning, with the goals of resolving identity conflicts to achieve interference-free unlearning and protecting against malicious attacks to achieve robust unlearning. To satisfy these goals, we propose a novel framework, RObust generatiVE continual identity unlearning against Relearning attacks (ROVER). By filtering unlearning requests with latent similarity, our method effectively isolates benign unlearning from malicious attacks to preserve identity removal integrity. Meanwhile, residual orthogonal resonator resolves identity conflicts in the continuous erasure of multiple identities, preserving stability in benign continual unlearning. Moreover, we introduce the phantom guard network to block malicious attacks by absorbing adversarial gradients, ensuring irreversible identity unlearning. The extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on the task of robust multi-identity unlearning against relearning attacks.
Tairan Huang 0001, Qiang Chen 0016, Beibei Hu, Yunlong Zhao 0003, Hongyan Xu 0002, Xiu Su
AAAI8
2026 Injection Without Distortion: Geometrically Constrained Knowledge Enhancement for Vision-Language Models
abstract
Vision-Language Models (VLMs) are widely used in tasks like Open-Vocabulary Object Detection and zero-shot Classification, owing to their powerful generalization. However, recent research reveals that VLMs exhibit significant performance instability when tasked with recognizing concepts at varying granularities (e.g., ``animal'' vs. ``dog''). Prevailing methods inject external knowledge from Large Language Models, but this unconstrained approach distorts the VLM's inherent hierarchical orthogonal geometry, leading to performance collapse on general concepts. To address this, we introduce GeCoin, an innovative Geometrically Constrained framework that safely enhances existing VLMs with external knowledge for improved hierarchical understanding, without additional training. By projecting knowledge into the null-space of a query concept's feature space, GeCoin mathematically guarantees the preservation of general knowledge while integrating specialized information. Extensive experiments across large-scale benchmarks, diverse VLMs, and knowledge from various LLMs (e.g., GPT-3.5, Claude-3, Gemini-Pro) show that GeCoin boosts performance by an average of 3.9% over the strongest baseline—crucially eradicating performance collapse on general concepts.
Zhongze Wu, Xiu Su, Shan You, Yueyi Luo
AAAI2
2026 STAG: Biologically guided spatial transcriptomics prediction via hypergraph learning
abstract
Spatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG.
Mingcheng Qu, Yuchuan Zhao, Donglin Di, Xiu Su, Hongyan Xu 0002, Yang Song 0001, Lei Fan 0007
Medical Image Anal.5
2026 Robot Few-Shot Manipulation Skills Learning Based on Meta Imitation Learning and Mixture of Experts Model
Jiahe Zhao, Xiu Su, Bin He 0003
IEEE Trans Autom. Sci. Eng.3
2025 Perturbating, Tuning, and Collaborating: Harnessing Vision Foundation Models for Single Domain Generalization on Medical Imaging
abstract
Single Domain Generalization (SDG) is critical in medical imaging applications. Recently, Vision Foundation Models (VFMs) have spearheaded a trend in AI development due to their robust generalizability and versatility. This work aims to fully explore the generalization capabilities of VFMs alongside the domain-specific expertise of specialized models, thoroughly investigating the boundaries of their respective capabilities, thereby collaboratively addressing SDG challenges within medical imaging. We propose a framework for Collaborative reasoning between Specialized and Universal models for Single Domain Generalization (CollaSU-SDG) in medical imaging. Specifically, we first design a model-aware perturbation injection method from the perspective of single-source domain data, enabling differentiated and adaptive perturbation injection for two different scales of models. Then, a domain expansion adapter is designed for the VFM to adapt to the augmented single-source domain medical data. Lastly, we introduce an adaptive hierarchical transfer and dynamic dense prompting method that facilitate collaborative reasoning between the specialized and universal models, eliminating the need for explicit prompts. Through these designs, CollaSU-SDG fully leverages the strengths of both specialized and universal models, achieving robust out-of-distribution generalization capabilities on single-source domain data. Experimental results demonstrate that CollaSU-SDG significantly advances the state-of-the-art performance across a wide range of medical datasets. All the code will be publicly available.
Yichao Cao, YingYing Zhang, Xiu Su, Haogang Zhu
AAAI4
2025 Seeing Beyond Noise: Joint Graph Structure Evaluation and Denoising for Multimodal Recommendation
abstract
Multimodal Recommendation Systems (MRSs) boost traditional user-item interaction-based methods by incorporating multimodal information. However, existing methods ignore the inherent noise brought by (1) noisy semantic priors in multimodal content, and (2) noisy user interactions in history records, therefore diminishing model performance. To fill this gap, we propose to denoise MRSs by jointly EValuating structure Effectiveness and mitigating Noisy links (EVEN). Firstly, for semantic prior noise in multimodal content, EVEN builds item homogeneous consistency and denoises it by evaluating behavior-driven confidence. Secondly, for noise in user interactions, EVEN updates user feedback by denoising observed interactions following implicit contribution evaluation of high-order representations. Thirdly, EVEN performs cross-modal alignment through self-guided structure learning, reinforcing task-specific inter-modal dependency modeling and cross-modal fusion. Through extensive experiments on three widely-used datasets, EVEN achieves an average improvement of 8.95% and 5.90% in recommendation accuracy compared with LGMRec and FREEDOM, respectively, without extending the total training time.
Yuxin Qi 0001, Xi Lin 0003, Xiu Su, Jiani Zhu, Jingyu Wang 0005, Jianhua Li 0001
AAAI4
2025 VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
abstract
The advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual annotations with insufficient granularity or automatic construction methods with redundant frame-by-frame analysis, limiting their scalability and effectiveness for complex reasoning. To address these challenges, we introduce VideoEspresso, a novel dataset that features VideoQA pairs preserving essential spatial details and temporal coherence, along with multimodal annotations of intermediate reasoning steps. Our construction pipeline employs a semantic-aware method to reduce redundancy, followed by generating QA pairs using GPT-4o. We further develop video Chain-of-Thought (CoT) annotations to enrich reasoning processes, guiding GPT-4o in extracting logical relationships from QA pairs and video content. To exploit the potential of high-quality VideoQA pairs, we propose a Hybrid LVLMs Collaboration framework, featuring a Frame Selector and a two-stage instruction fine-tuned reasoning LVLM. This framework adaptively selects core frames and performs CoT reasoning using multimodal evidence. Evaluated on our proposed benchmark with 14 tasks against 9 popular LVLMs, our method outperforms existing baselines on most tasks, demonstrating superior video reasoning capabilities. Our code and dataset have been released at: https://github.com/hshjerry/VideoEspresso
Songhao Han, Wei Huang 0042, Hairong Shi, Le Zhuo, Xiu Su, Xiaojuan Qi 0001, Yue Liao, Si Liu 0001
CVPR5
2025 HieClip: Hierarchical CLIP with Explicit Alignment for Zero-Shot Anomaly Detection
abstract
Large image-language models(LLM) have made significant progress in zero-shot anomaly detection(ZSAD), however, the semantic gap between images and text limits their performance in hierarchical learning. In this paper, we propose the hierarchical alignment clip(HieClip) framework, to achieve hierarchical alignment between images and text. Specifically, we introduce learnable hierarchical textual(LHT) to reduce the representation differences between various levels of images and text, while performing multi-level comprehensive discrimination. Additionally, the dynamically adjusting the weights of features at different levels, improving the model’s ability to capture both global and local information. Experiments on public industrial datasets demonstrate HieClip’s effectiveness, showing significant accuracy improvement, and its strong generalization capabilities were further validated on medical datasets. Compared to existing methods, HieClip excels in anomaly detection tasks, particularly in industrial inspection and medical diagnosis scenarios.
Liujie Hua, Xiu Su, Yueyi Luo, Shan You
ICASSP2
2025 Harmonizing for defect visibility with Fine-Grained Hierarchical Interaction Learning
abstract
Defect detection is a fundamental task in industrial image analysis, crucial for identifying and delineating defect regions. However, existing models, often struggle to learn critical features effectively under conditions of noisy interference. In this study, we introduce the Fine-Grained Hierarchical Interaction Learning (FINet) framework, designed to enhance the learning process by harmonizing feature interactions at multiple scales. Specifically, FINet incorporates the Adaptive Tensor Interaction (ATI) to facilitate high-order feature interactions amidst noise in a high-dimensional frequency space. Additionally, the FlexiFocus network is developed to dynamically balance feature focus across scales, further enhancing defect feature visibility and providing an effective trade-off between computational speed and performance. Extensive experiments on the PVEL-AD dataset show FINet’s superior accuracies (90.50% mAP50, 62.80% mAP50:5:95 ), surpassing DDQ-DETR by 8.3% and 3.6%, respectively. The code is available at https://github.com/zhongzee/FINet-master.
Zhongze Wu, Yitian Long, Xiu Su, Yueyi Luo, Shan You
ICASSP3
2025 CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point Clouds
Yichao Cao, Xiu Su, Dan Niu, Xuanpeng Li
ICCV3
2025 Adaptive Training Meets Progressive Scaling: Elevating Efficiency in Diffusion Models
abstract
Diffusion models have demonstrated remarkable efficacy in various generative tasks with the predictive prowess of denoising model. Currently, diffusion models employ a uniform denoising model across all timesteps. However, the inherent variations in data distributions at different timesteps lead to conflicts during training, constraining the potential of diffusion models. To address this challenge, we propose a novel two-stage divide-and-conquer training strategy termed TDC Training. It groups timesteps based on task similarity and difficulty, assigning highly customized denoising models to each group, thereby enhancing the performance of diffusion models. While two-stage training avoids the need to train each model separately, the total training cost is even lower than training a single unified denoising model. Additionally, we introduce Proxy-based Pruning to further customize the denoising models. This method transforms the pruning problem of diffusion models into a multi-round decision-making problem, enabling precise pruning of diffusion models. Our experiments validate the effectiveness of TDC Training, demonstrating improvements in FID of 1.5 on ImageNet64 compared to original IDDPM, while saving about 20% of computational resources.
Xiu Su, Shan You, Tao Huang 0020, Chang Xu 0002
ICME2
2025 Stable Fair Graph Representation Learning with Lipschitz Constraint
abstract
Group fairness based on adversarial training has gained significant attention on graph data, which was implemented by masking sensitive attributes to generate fair feature views. However, existing models suffer from training instability due to uncertainty of the generated masks and the trade-off between fairness and utility. In this work, we propose a stable fair Graph Neural Network (SFG) to maintain training stability while preserving accuracy and fairness performance. Specifically, we first theoretically derive a tight upper Lipschitz bound to control the stability of existing adversarial-based models and employ a stochastic projected subgradient algorithm to constrain the bound, which operates in a block-coordinate manner. Additionally, we construct the uncertainty set to train the model, which can prevent unstable training by dropping some overfitting nodes caused by chasing fairness. Extensive experiments conducted on three real-world datasets demonstrate that SFG is stable and outperforms other state-of-the-art adversarial-based methods in terms of both fairness and utility performance. Codes are available at https://github.com/sh-qiangchen/SFG.
Qiang Chen 0016, Zhongze Wu, Xiu Su, Xi Lin 0003, Shan You, Shuo Yang 0006, Chang Xu 0002
ICML3
2025 TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical Imaging
abstract
Medical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vision foundation models to medical imaging SDG. TinyMIG aims to enable lightweight specialized models to mimic the strong generalization capabilities of foundation models in terms of both global feature distribution and local fine-grained details during training. Specifically, for global feature distribution, we propose a Global Distribution Consistency Learning strategy that mimics the prior distributions of the foundation model layer by layer. For local fine-grained details, we further design a Localized Representation Alignment method, which promotes semantic alignment and generalization distillation between the specialized model and the foundation model. These mechanisms collectively enable the specialized model to achieve robust performance in diverse medical imaging scenarios. Extensive experiments on large-scale benchmarks demonstrate that TinyMIG, with extremely low computational cost, significantly outperforms state-of-the-art models, showcasing its superior SDG capabilities. All the code and model weights will be publicly available.
Hongyan Xu 0002, Yichao Cao, Xiu Su, Tianfa Li, Shan An, Haogang Zhu
ICML4
2025 Addressing Granularity-induced Semantic Drift in OvOD via Graph-guided semantically consistent representation
abstract
Open-vocabulary object detection (OvOD) uses Vision-Language Models (VLMs) to detect arbitrary categories specified by natural language. However, existing methods often struggle with performance instability caused by granularity-induced semantic drift, which arises from misaligned label embeddings across varying levels of specificity. In this paper, we propose GraSecon, a Graph-guided Semantically Consistent representation framework that enhances zero-shot detection robustness without requiring additional training. We construct a hierarchical Fine-grained Semantic Graph enriched with visually grounded attributes from large language models (LLMs). This graph captures hierarchical, sibling and cross-level relations, enabling controlled Laplacian refinement to harmonize the embedding space and improve visual-semantic alignment. To strengthen fine-grained discriminability, we introduce a Key Semantic Node Mining module that identifies and anchors semantically sensitive nodes, ensuring robust feature representation. Furthermore, our Semantic Relevance-Driven Laplacian Propagation adaptively propagates information, promoting coherent and context-aware embedding alignment across granularities. Extensive experiments on the iNatLoc and FSOD datasets demonstrate that GraSecon outperforms prior SOTA methods, achieving average mAP50 improvements of 6.5% and 5.4%. Code is publicly available at: https://github.com/minoslab-csu/GraSecon.
Hongyan Xu 0002, Zhongze Wu, Ang He, Xi Lin 0003, Xiu Su
ACM Multimedia6
2025 CaDGS: Modeling Inter-Gaussian Mutual Information for Dynamic Novel View Synthesis
abstract
Dynamic novel view synthesis (NVS) aims to render time-varying scenes from arbitrary viewpoints, balancing rendering quality and computational efficiency. While recent 4D Gaussian Splatting approaches offer promising real-time performance, they fundamentally overlook critical interdependence between Gaussians by modeling deformations independently. Our information-theoretic analysis reveals substantial mutual information across the Gaussian field, manifesting as appearance-preserving radiance coherence and motion-consistent deformation propagation. This finding establishes that rendering quality emerges from coordinated transformation rather than independent processing. We propose Correlation-aware Dynamic Gaussian Splatting (CaDGS) with our novel Gaussian Correlation Tensor Projection (GCTP) method, which efficiently transforms the complex O(n3) mutual information tensor into a dual-channel O(n2) spatial matrix, preserving the critical topological structure of Gaussian interactions. Combined with our Spatio-Temporal Deformation Consistency (STDC) learning, which enforces volumetric coherence through tensor-guided regularization across multiple scales, CaDGS prevents geometric distortions and texture inconsistencies common in previous approaches. Experimental results demonstrate state-of-the-art performance, achieving 32.4 PSNR on the Neu3D dataset with fewer Gaussians while maintaining rendering speeds of 323 FPS at 1353 × 1014 resolution.
Yunlong Zhao 0003, Xiaoheng Deng, Zhuohua Qiu, Chang Xu 0002, Xiangjian He, Shan You, Xiu Su
ACM Multimedia8
2025 DualFPT: Handling Data Heterogeneity in Federated Prompt Tuning from both Generalized and Personalized Perspective
abstract
Federated Prompt Tuning (FPT) integrates large pre-trained Vision Transformers (ViT) into Federated Learning (FL) by leveraging Visual Prompt Tuning (VPT), achieving state-of-the-art performance with enhanced efficiency across various visual downstream tasks. However, data heterogeneity, such as feature shift and class imbalance, limits prompts' transferability and robustness in FPT. Existing methods primarily focus on generalized FPT (GFPT) or personalized FPT (PFPT), while only a few methods make initial attempts to integrate both approaches. In this paper, we propose a new FPT framework, dubbed DualFPT, which handles data heterogeneity from both generalized and personalized perspectives. Specifically, DualFPT divides the learnable prompts into global and local prompts to jointly capture general and client-specific information, achieving the harmonization of GFPT and PFPT. The generalization of DualFPT is realized by the Feature Sharing (FS) mechanism, which effectively narrows the distribution gap by allowing clients to securely share a portion of sensitive features. The key feature for improving personalization is the Prompt Composition Scheme (PCS), which weights local prompts with distribution similarity to generate composite prompts, thus achieving automatic distribution adaptation. Extensive experiments under feature shift and class imbalance scenarios demonstrate the superior performance of DualFPT. On DomainNet and CIFAR-100 (D (0.1)), DualFPT surpasses SGPT by 4.35% and 5.89% for generalization along with 7.89% and 9.09% for personalization. Ablation studies further validate the effectiveness, efficiency, and security of DualFPT.
Yuliang Chen, Xi Lin 0003, Chao Sang, Xiu Su
ACM Multimedia4
2025 Graph Unlearning Meets Influence-aware Negative Preference Optimization
abstract
Recent advancements in graph unlearning models have enhanced model utility by preserving the node representation essentially invariant, while using gradient ascent on the forget set to achieve unlearning. However, this approach causes a drastic degradation in model utility during the unlearning process due to the rapid divergence speed of gradient ascent. In this paper, we introduce INPO, an Influence-aware Negative Preference Optimization framework that focuses on slowing the divergence speed and improving the robustness of the model utility to the unlearning process. Specifically, we first analyze that NPO has slower divergence speed and theoretically propose that unlearning high-influence edges can reduce impact of unlearning. We design an influence-aware message function to amplify the influence of unlearned edges and mitigate the tight topological coupling between the forget set and the retain set. The influence of each edge is quickly estimated by a removal-based method. Additionally, we propose a topological entropy loss from the perspective of topology to avoid excessive information loss in the local structure during unlearning. Extensive experiments conducted on five real-world datasets demonstrate that INPO-based model achieves state-of-the-art performance on all forget quality metrics while maintaining the model's utility. Codes are available at https://github.com/sh-qiangchen/INPO.
Qiang Chen 0016, Zhongze Wu, Ang He, Xi Lin 0003, Shan You, Chang Xu 0002, Xiu Su
ACM Multimedia9
2025 Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable advancements in numerous areas such as multimedia. However, hallucination issues significantly limit their credibility and application potential. Existing mitigation methods typically rely on external tools or the comparison of multi-round inference, which significantly increase inference time. In this paper, we propose SElf-Evolving Distillation (SEED), which identifies hallucinations within the inner knowledge of LVLMs, isolates and purges them, and then distills the purified knowledge back into the model, enabling self-evolution. Furthermore, we identified that traditional distillation methods are prone to inducing void spaces in the output space of LVLMs. To address this issue, we propose a Mode-Seeking Evolving approach, which performs distillation to capture the dominant modes of the purified knowledge distribution, thereby avoiding the chaotic results that could emerge from void spaces. Moreover, we introduce a Hallucination Elimination Adapter, which corrects the dark knowledge of the original model by learning purified knowledge. Extensive experiments on multiple benchmarks validate the superiority of our SEED, demonstrating substantial improvements in mitigating hallucinations for representative LVLM models such as LLaVA-1.5 and InternVL2. Remarkably, the F1 score of LLaVA-1.5 on the hallucination evaluation metric POPE-Random improved from 81.3 to 88.3.
Xiu Su, Yang Liu 0246, Shan You, Chang Xu 0002
ACM Multimedia2
2025 On the Stability and Generalization of Meta-Learning: the Impact of Inner-Levels
abstract
Meta-learning has achieved significant advancements, with generalization emerging as a key metric for evaluating meta-learning algorithms. While recent studies have mainly focused on training strategies, data-split methods, and tightening generalization bounds, they often ignore the impact of inner-levels on generalization. To bridge this gap, this paper focuses on several prominent meta-learning algorithms and establishes two generalization analytical frameworks for them based on their inner-processes: the Gradient Descent Framework (GDF) and the Proximal Descent Framework (PDF). Within these frameworks, we introduce two novel algorithmic stability definitions and derive the corresponding generalization bounds. Our findings reveal a trade-off of inner-levels under GDF, whereas PDF exhibits a beneficial relationship. Moreover, we highlight the critical role of the meta-objective function in minimizing generalization error. Inspired by this, we propose a new, simplified meta-objective function definition to enhance generalization performance. Many real-world experiments support our findings and show the improvement of the new meta-objective function.
Wenjun Ding, Jingling Liu, Lixing Chen, Xiu Su
NeurIPS4
2025 UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
abstract
Data augmentation using generative models has emerged as a powerful paradigm for enhancing performance in computer vision tasks. However, most existing augmentation approaches primarily focus on optimizing intrinsic data attributes -- such as fidelity and diversity -- to generate visually high-quality synthetic data, while often neglecting task-specific requirements. Yet, it is essential for data generators to account for the needs of downstream tasks, as training data requirements can vary significantly across different tasks and network architectures. To address these limitations, we propose UtilGen, a novel utility-centric data augmentation framework that adaptively optimizes the data generation process to produce task-specific, high-utility training data via downstream task feedback. Specifically, we first introduce a weight allocation network to evaluate the task-specific utility of each synthetic sample. Guided by these evaluations, UtilGen iteratively refines the data generation process using a dual-level optimization strategy to maximize the synthetic data utility: (1) model-level optimization tailors the generative model to the downstream task, and (2) instance-level optimization adjusts generation policies -- such as prompt embeddings and initial noise -- at each generation round. Extensive experiments on eight benchmark datasets of varying complexity and granularity demonstrate that UtilGen consistently achieves superior performance, with an average accuracy improvement of 3.87\% over previous SOTA. Further analysis of data influence and distribution reveals that UtilGen produces more impactful and task-relevant synthetic data, validating the effectiveness of the paradigm shift from visual characteristics-centric to task utility-centric data augmentation.
Jiyu Guo, Shuo Yang 0006, Yiming Huang 0001, Yancheng Long, Xiaobo Xia, Xiu Su, Bo Zhao 0038, Zeke Xie, Liqiang Nie
NeurIPS6
2025 L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
abstract
Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained in both contextual coverage and inference efficiency due to its inherently sequential process. To overcome these challenges, we propose leap multi-token prediction~(L-MTP), an innovative token prediction method that extends the capabilities of multi-token prediction (MTP) by introducing a leap-based mechanism. Unlike conventional MTP, which generates multiple tokens at adjacent positions, L-MTP strategically skips over intermediate tokens, predicting non-sequential ones in a single forward pass. This structured leap not only enhances the model's ability to capture long-range dependencies but also enables a decoding strategy specially optimized for non-sequential leap token generation, effectively accelerating inference. We theoretically demonstrate the benefit of L-MTP in improving inference efficiency. Experiments across diverse benchmarks validate its merit in boosting both LLM performance and inference speed. The source code is available at https://github.com/Xiaohao-Liu/L-MTP.
Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang, Xianzhi Yu, Xiu Su, Shuo Yang 0006, See-Kiong Ng, Tat-Seng Chua
NeurIPS6
2025 FairMoE: Decoupled Expert Learning for Unbiased Customized Face Generation
Shan You, Chang Xu 0002, Xiu Su
PRCV (6)7
2025 FeCoGraph: Label-Aware Federated Graph Contrastive Learning for Few-Shot Network Intrusion Detection
abstract
With increasing cyber attacks over the Internet, network intrusion detection systems (NIDS) have been an indispensable barrier to protecting network security. Taking advantage of automatically capturing topology connections, recent deep graph learning approaches have achieved remarkable performance in distinguishing different types of malicious flows. However, there remain some critical challenges. 1) previous supervised learning methods rely heavily on abundant and high-quality annotated samples, while label annotation requires abundant time and expert knowledge. 2) Centralized methods require all data to be uploaded to a server for learning behavior patterns, which results in high detection latency and critical privacy leakage. 3) Diverse attack scenarios exhibit highly imbalanced distribution, making it hard to characterize abnormal behaviors. To address these issues, we proposed FeCoGraph, a label-aware federated graph contrastive learning framework for intrusion detection in few-shot scenarios. The line graph is introduced to directly process flow embeddings, which are compatible with diverse GNNs. Furthermore, We formulate a graph contrastive learning task to effectively leverage label information, allowing intra-class embeddings more compact than inter-class embeddings. To improve the scalability of NIDS, we utilize federated learning to cover more attack scenarios while protecting data privacy. Experiment results show that FeCoGraph surpass E-graphSAGE with an average 8.36% accuracy on binary classification and 6.77% accuracy on multiclass classification, demonstrating the efficiency of our approach.
Qinghua Mao, Xi Lin 0003, Wenchao Xu 0001, Yuxin Qi 0001, Xiu Su, Gaolei Li, Jianhua Li 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language Models
abstract
Large language models (LLMs) offer impressive performance across diverse fields, but their increasing complexity raises both design costs and the need for specialized expertise. These challenges are intensified for Neural Architecture Search (NAS) methods reliant on weight-sharing techniques. This paper introduces GNAS, a new NAS method that boosts the search process with the aid of LLMs for efficient model discovery. With insights from existing architectures, GNAS swiftly identifies superior models that can adapt to changing resource constraints. We provide a mathematical framework to facilitate the transfer of knowledge across different model sizes, thereby improving search efficiency. Our experiments conducted on ImageNet, NAS-Bench-Macro, and ChannelBench-Macro confirm the effectiveness of GNAS across both CNN and Transformer architectures.
Xiu Su, Shan You, Hongyan Xu 0002, Xiuxing Li, Chang Xu 0002
ICASSP1
2024 TCNAS: Transformer Architecture Evolving in Code Clone Detection
abstract
Code clone detection aims at finding code fragments with syntactic or semantic similarity. Most of current approaches mainly focus on detecting syntactic similarity while ignoring semantic long-term context alignment, and these detection methods encode the source code using human-designed models, a process which requires both expert input and a significant cost of time for experimentation and refinement. To address these challenges, we introduce the Transformer Code Neural Architecture Search (TCNAS), an approach designed to optimize transformer-based architectures for detection. In TCNAS, all channels are trained and evaluated equitably to enhance search efficiency. Besides, we introduce the dataflow of the code by extracting the semantic information from the code fragments. TCNAS facilitates the discovery of an optimal model structure geared towards the detection, eliminating the need for manual design. The searched optimal architecture is utilized to detect the code pairs. We conduct various empirical experiments on the benchmark, which covering all four types of code clone detection. The results demonstrate our approach consistently yields competitive detection scores across a range of evaluations.
Hongyan Xu 0002, Xiaohuan Pei, Xiu Su, Shan You, Chang Xu 0002
ICASSP3
2024 DomainVoyager: Embracing The Unknown Domain by Prompting for Automatic Augmentation
abstract
For medical image analysis, domain generalization (DG) faces significant challenges due to variances in data across medical imaging devices. Addressing this, we introduce Prompt Guided Domain Aligning Augmentation (PGDAA), a novel approach that harnesses Large Language Models (LLMs) to iteratively refine data augmentation sequences in DG for medical segmentations. Specifically, by harnessing the LLM’s advanced capabilities in interpreting and responding to specific prompts, PGDAA iteratively voyages the parameter searching space for data augmentation, identifying optimal augmentation sequences. In each searching round, the LLM leverages provided prior data augmentation methods, their parameters, and corresponding evaluation results to acquire a sufficient performance memory bank, thereby proposing more effective augmentation sequences to significantly narrowing inter-domain gaps. Notably, this integration of LLMs in DG represents a pioneering application in the field, which adds minor training parameters and can be easily combined with other DG benchmarks for further improvements. Comprehensive experiments reveal our method outperforms the baseline by 4.47% and 5.08% on Fundus and Prostate datasets, achieving 90.10% and 89.28% accuracy, respectively.
Haogang Zhu, Xiu Su
ICME3
2024 SCD-NAS: Towards Zero-Cost Training in Melanoma Diagnosis
abstract
Diagnosing melanoma remains challenging despite advances in Convolutional Neural Networks (CNNs) for skin cancer detection. Their application in clinical settings is often limited by differences between natural and clinical images. To address this, we introduce the Skin Cancer Detection Neural Architecture Search (SCD-NAS) framework. In our method, Large Language Model (LLM) is leveraged as a proxy, which helps SCD-NAS achieve cost-free training. Additionally, to maximize the benefits of various architectural design spaces, we introduce a Search Space Expansion (SSE) methodology. This effectively combines the merits of diverse architectural configurations, thereby enhancing model performance. We conducted experiments on the ISIC 2020, MedMNISTv2, CIFAR-10 and CIFAR-100 datasets. Our SCD-NAS-derived ResNet50 model achieved an Area Under the Curve (AUC) of 91.23% on the ISIC 2020 dataset, improving the baseline by 5.93%. It also exceeded the CIFAR-10 benchmark by 2.45% in accuracy.
Hongyan Xu 0002, Xiu Su, Arcot Sowmya, Ian Katz, Dadong Wang
ICME2
2024 Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQA
abstract
Medical Visual Question Answering (Med-VQA) interprets complex medical imagery using user instructions for precise diagnostics, yet faces challenges due to diverse, inadequately annotated images. In this paper, we introduce the Universal Instruction-Vision Navigator (Uni-Med) framework for extracting instruction-to-answer relationships, facilitating the understanding of visual evidence behind responses. Specifically, we design the Instruct-to-Answer Clues Interpreter (IAI) to generate visual explanations based on the answers and mark the core part of instructions with "real intent" labels. The IAI-Med VQA dataset, produced using IAI, is now publicly available to advance Med-VQA research. Additionally, our Token-Level Cut-Mix module dynamically aligns visual explanations with image patches, ensuring answers are traceable and learnable. We also implement intention-guided attention to minimize non-core instruction interference, sharpening focus on ’real intent’. Extensive experiments on SLAKE datasets show Uni-Med’s superior accuracies (87.52% closed, 86.12% overall), outperforming MedVInT-PMC-VQA by 1.22% and 0.92%. Code and dataset are available at: https://github.com/zhongzee/Uni-Med-master.
Zhongze Wu, Hongyan Xu 0002, Yitian Long, Shan You, Xiu Su, Yueyi Luo, Chang Xu 0002
ICML5
2024 Image Anomaly Detection Based on Controllable Self-Augmentation
abstract
Based on data synthesis, anomaly detection (AD) methods often rely on external data for data synthesis. However, most external abnormal data exhibits strong randomness, which may lead to a reduced range of diversity among the synthesized data. In order to achieve a broader diversity in data synthesis, it is necessary to not only have highly diverse data but also to incorporate low-diversity noise data. To enhance the diversity range of the synthesized data, this study proposes a diversity measurement assisted by image self-representation: measuring the distance between noise data and normal data and quantitatively synthesizing diversified data by selecting diverse noise data for synthesis, namely, Diversified Synthesis (DS). Diversified Synthesis introduces patch measurement and a controllable enhancement module to establish controllable diversified enhanced data. The contribution of this study lies in proposing a novel diversified synthesis method, which achieves a broader diversity synthesis through the introduction of image self-representation-assisted diversity measurement and quantitative synthesis. Furthermore, through the self-enhancement data augmentation method, the use of image intrinsic features for enhancement achieves diversity and multi-scale characteristics in the synthesized data, thereby improving the training performance of the discriminative model. This provides an effective optimization solution for comprehensive anomaly detection methods.
Liujie Hua, Yichao Cao, Yitian Long, Shan You, Xiu Su, Yueyi Luo, Chang Xu 0002
IJCNN5
2024 Universal Frequency Domain Perturbation for Single-Source Domain Generalization
abstract
In this work, we introduce a novel approach to single-source domain generalization (SDG) in medical imaging, focusing on overcoming the challenge of style variation in out-of-distribution (OOD) domains without requiring domain labels or additional generative models. We propose a Universal Frequency Perturbation framework for SDG termed as UniFreqSDG, that performs hierarchical feature-level frequency domain perturbations, facilitating the model's ability to handle diverse OOD styles. Specifically, we design a learnable spectral perturbation module that adaptively learns the frequency distribution range of samples, allowing for precise low-frequency (LF) perturbation. This adaptive approach not only generates stylistically diverse samples but also preserves domain-invariant anatomical features without the need for manual hyperparameter tuning. Then, the frequency features before and after perturbation are decoupled and recombined through the Content Preservation Reconstruction operation, effectively preventing the loss of discriminative content information. Furthermore, we introduce the Active Domain-variance Inducement Loss to encourage effective perturbation in the frequency domain while ensuring the sufficient decoupling of domain-invariant and domain-style features. Extensive experiments demonstrate that UniFreqSDG increases the dice score by an average of 7.47% (from 77.98% to 85.45%) on the fundus dataset and 4.99% (from 71.42% to 76.73%) on the prostate dataset compared to the state-of-the-art approaches.
Yichao Cao, Xiu Su, Haogang Zhu
ACM Multimedia3
2023 Neural Architecture Search for Wide Spectrum Adversarial Robustness
abstract
One major limitation of CNNs is that they are vulnerable to adversarial attacks. Currently, adversarial robustness in neural networks is commonly optimized with respect to a small pre-selected adversarial noise strength, causing them to have potentially limited performance when under attack by larger adversarial noises in real-world scenarios. In this research, we aim to find Neural Architectures that have improved robustness on a wide range of adversarial noise strengths through Neural Architecture Search. In detail, we propose a lightweight Adversarial Noise Estimator to reduce the high cost of generating adversarial noise with respect to different strengths. Besides, we construct an Efficient Wide Spectrum Searcher to reduce the cost of adjusting network architecture with the large adversarial validation set during the search. With the two components proposed, the number of adversarial noise strengths searched can be increased significantly while having a limited increase in search time. Extensive experiments on benchmark datasets such as CIFAR and ImageNet demonstrate that with a significantly richer search signal in robustness, our method can find architectures with improved overall robustness while having a limited impact on natural accuracy and around 40% reduction in search time compared with the naive approach of searching. Codes available at: https://github.com/zhicheng2T0/Wsr-NAS.git
Zhi Cheng, Yanxi Li 0001, Minjing Dong, Xiu Su, Shan You, Chang Xu 0002
AAAI4
2023 Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detection
abstract
Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predicttriplets. Despite the challenges posed by the numerous interaction combinations, they also offer opportunities for multi-modal learning of visual texts. In this paper, we present a systematic and unified framework (RmLR) that enhances HOI detection by incorporating structured text knowledge. Firstly, we qualitatively and quantitatively analyze the loss of interaction information in the two-stage HOI detector and propose a re-mining strategy to generate more comprehensive visual representation. Secondly, we design more fine-grained sentence- and word-level alignment and knowledge transfer strategies to effectively address the many-to-many matching problem between multiple interactions and multiple texts. These strategies alleviate the matching confusion problem that arises when multiple interactions occur simultaneously, thereby improving the effectiveness of the alignment process. Finally, HOI reasoning by visual features augmented with textual knowledge substantially improves the understanding of interactions. Experimental results illustrate the effectiveness of our approach, where state-of-the-art performance is achieved on public benchmarks.
Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002
ICCV4
2023 DiffNAS: Bootstrapping Diffusion Models by Prompting for Better Architectures
abstract
Diffusion models have recently exhibited remarkable performance on synthetic data. After a diffusion path is selected, a base model, such as UNet, operates as a denoising autoencoder, primarily predicting noises that need to be eliminated step by step. Consequently, it is crucial to employ a model that aligns with the expected budgets to facilitate superior synthetic performance. In this paper, we meticulously analyze the diffusion model and engineer a base model search approach, denoted "DiffNAS". Specifically, we leverage GPT-4 as a supernet to expedite the search, supplemented with a search memory to enhance the results. Moreover, we employ RFID as a proxy to promptly rank the experimental outcomes produced by GPT-4. We also adopt a rapid-convergence training strategy to boost search efficiency. Rigorous experimentation corroborates that our algorithm can augment the search efficiency by $2 \times$ under GPT-based scenarios, while also attaining a performance of 2.82 with 0.37 improvement in FID on CIFAR10 relative to the benchmark IDDPM algorithm.
Xiu Su, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002
ICDM2
2023 Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models
abstract
Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting <human, action, object> triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as UniHOI. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (i.e. GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights will be made publicly available.
Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002
NeurIPS3
2023 Searching for Network Width With Bilaterally Coupled Network
abstract
Searching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfil the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this article, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we propose to reduce the redundant search space and present the BCNetV2 as the enhanced supernet to ensure rigorous training fairness over channels. Furthermore, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. We also propose a new open-source width search benchmark on macro structures named Channel-Bench-Macro for the better comparisons of the width search algorithms with MobileNet- and ResNet-like architectures. Extensive experiments on the benchmark datasets demonstrate that our method can achieve state-of-the-art performance.
Xiu Su, Shan You, Jiyang Xie 0001, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 ViTAS: Vision Transformer Architecture Search
Xiu Su, Shan You, Jiyang Xie 0001, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002
ECCV (21)1
2022 ScaleNet: Searching for the Model to Scale
Jiyang Xie 0001, Xiu Su, Shan You, Zhanyu Ma, Fei Wang 0032, Chen Qian 0006
ECCV (21)2
2022 Data Agnostic Filter Gating For Efficient Deep Networks
abstract
Filter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art.
Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya
ICASSP2
2022 Sufficient Vision Transformer
abstract
Currently, Vision Transformer (ViT) and its variants have demonstrated promising performance on various computer vision tasks. Nevertheless, task-irrelevant information such as background nuisance and noise in patch tokens would damage the performance of ViT-based models. In this paper, we develop Sufficient Vision Transformer (Suf-ViT) as a new solution to address this issue. In our research, we propose the Sufficiency-Blocks (S-Blocks) to be applied across the depth of Suf-ViT to disentangle and discard task-irrelevant information accurately. Besides, to boost the training of Suf-ViT, we formulate a Sufficient-Reduction Loss (SRLoss) leveraging the concept of Mutual Information (MI) that enables Suf-ViT to extract more reliable sufficient representations by removing task-irrelevant information. Extensive experiments on benchmark datasets such as ImageNet, ImageNet-C, and CIFAR-10 indicate that our method can achieve state-of-the-art or competing performance over other baseline methods. Codes are available at: https://github.com/zhicheng2T0/Sufficient-Vision-Transformer.git
Zhi Cheng, Xiu Su, Xueyu Wang, Shan You, Chang Xu 0002
KDD2
2022 Searching for Better Spatio-temporal Alignment in Few-Shot Action Recognition
abstract
Spatio-Temporal feature matching and alignment are essential for few-shot action recognition as they determine the coherence and effectiveness of the temporal patterns. Nevertheless, this process could be not reliable, especially when dealing with complex video scenarios. In this paper, we propose to improve the performance of matching and alignment from the end-to-end design of models. Our solution comes at two-folds. First, we encourage to enhance the extracted Spatio-Temporal representations from few-shot videos in the perspective of architectures. With this aim, we propose a specialized transformer search method for videos, thus the spatial and temporal attention can be well-organized and optimized for stronger feature representations. Second, we also design an efficient non-parametric spatio-temporal prototype alignment strategy to better handle the high variability of motion. In particular, a query-specific class prototype will be generated for each query sample and category, which can better match query sequences against all support sequences. By doing so, our method SST enjoys significant superiority over the benchmark UCF101 and HMDB51 datasets. For example, with no pretraining, our method achieves 17.1\% Top-1 accuracy improvement than the baseline TRX on UCF101 5-way 1-shot setting but with only 3x fewer FLOPs.
Yichao Cao, Xiu Su, Qingfei Tang, Shan You, Xiaobo Lu, Chang Xu 0002
NeurIPS2
2021 Prioritized Architecture Sampling With Monto-Carlo Tree Search
abstract
One-shot neural architecture search (NAS) methods significantly reduce the search cost by considering the whole search space as one network, which only needs to be trained once. However, current methods select each operation independently without considering previous layers. Besides, the historical information obtained with huge computation costs is usually used only once and then discarded. In this paper, we introduce a sampling strategy based on Monte Carlo tree search (MCTS) with the search space modeled as a Monte Carlo tree (MCT), which captures the dependency among layers. Furthermore, intermediate results are stored in the MCT for future decisions and a better exploration-exploitation balance. Concretely, MCT is updated using the training loss as a reward to the architecture performance; for accurately evaluating the numerous nodes, we propose node communication and hierarchical node selection methods in the training and search stages, respectively, making better uses of the operation rewards and hierarchical information. Moreover, for a fair comparison of different NAS methods, we construct an open-source NAS benchmark of a macro search space evaluated on CIFAR-10, namely NAS-Bench-Macro. Extensive experiments on NAS-Bench-Macro and ImageNet demonstrate that our method significantly improves search efficiency and performance. For example, by only searching 20 architectures, our obtained architecture achieves 78.0% top-1 accuracy with 442M FLOPs on ImageNet. Code (Benchmark) is available at: https://github.com/xiusu/NAS-Bench-Macro.
Xiu Su, Tao Huang 0020, Yanxi Li 0001, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
CVPR1
2021 BCNet: Searching for Network Width With Bilaterally Coupled Network
abstract
Searching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfill the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this paper, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. Extensive experiments on benchmark CIFAR-10 and ImageNet datasets indicate that our method can achieve state-of-the-art or competing performance over other baseline methods. Moreover, our method turns out to further boost the performance of NAS models by refining their network widths. For example, with the same FLOPs budget, our obtained EfficientNet-B0 achieves 77.36% Top-1 accuracy on ImageNet dataset, surpassing the performance of original setting by 0.48%.
Xiu Su, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
CVPR1
2021 Locally Free Weight Sharing for Network Width Search
Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
ICLR1
2021 K-shot NAS: Learnable Weight-Sharing for NAS with K-shot Supernets
abstract
In one-shot weight sharing for NAS, the weights of each operation (at each layer) are supposed to be identical for all architectures (paths) in the supernet. However, this rules out the possibility of adjusting operation weights to cater for different paths, which limits the reliability of the evaluation results. In this paper, instead of counting on a single supernet, we introduce $K$-shot supernets and take their weights for each operation as a dictionary. The operation weight for each path is represented as a convex combination of items in a dictionary with a simplex code. This enables a matrix approximation of the stand-alone weight matrix with a higher rank ($K>1$). A \textit{simplex-net} is introduced to produce architecture-customized code for each path. As a result, all paths can adaptively learn how to share weights in the $K$-shot supernets and acquire corresponding weights for better evaluation. $K$-shot supernets and simplex-net can be iteratively trained, and we further extend the search to the channel dimension. Extensive experiments on benchmark datasets validate that K-shot NAS significantly improves the evaluation accuracy of paths and thus brings in impressive performance improvements.
Xiu Su, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002
ICML1
2017 Monitoring the thermal discharge of hongyanhe nuclear power plant with aerial remote sensing technology using a UAV platform
abstract
Existing monitoring approaches are not effective in deal with routine thermal discharge monitoring requirements of nuclear power plants. This paper describes a monitoring methodology using an aerial remote sensing monitoring system based on an unmanned aerial vehicle (UAV) platform by taking the monitoring of the thermal discharge of the Hongyanhe Nuclear Power Plant as an example, and conducts a study on remote sensing extraction of thermal diffusion information of the thermal discharge. In this study, quartic polynomial fitting and parallel real data correction are used to correct the wide-angle distortion and acquire the water body surface temperature information, respectively. Synchronized measured data validation of independent samples indicates that the system can acquire diffusion information of the thermal discharge accurately, and the retrieval error of the surface water temperature is within 0.4°C. Results analysis shows that this efficient and convenient aerial remote sensing monitoring system based on a UAV platform can effectively make up inadequacies of existing monitoring technical measures and offers high precision. This system is expected to be further adopted and applied for post-assessment of the environmental impact of nuclear power plant thermal discharge and service monitoring.
Jianchao Fan, Xiu Su, Dejun Zou
IGARSS5
2016 Comparison of different spatial resolution thermal infrared data in monitoring thermal plume from the Hongyanhe nuclear power plant
abstract
Nuclear power industry had a great development in China in recent years and the environmental problems, such as thermal plume, has caused wide public concern. Taking the Hongyanhe nuclear power plant as example, this paper computed and achieved the sea surface temperature distribution with MODIS, HJ-1B and Landsat-8 thermal infrared data separately. These data were imaged in same time phase but different spatial resolution. Based on verified method of average gulf temperature, thermal plume distribution in three kinds of data was achieved. The comparison showed that the Landsat-8 data with higher monitoring precision and better details in thermal plume was more suitable for small area monitoring than the other two, which were affected by mixed pixels caused by low spatial resolution. Considering the different hydrogeological conditions of various nuclear power plants, it's smart to monitoring thermal plume with different temporal and spatial resolution satellite data. It can be foreseen that unmanned plane with infrared scanner and remote sensing technique will be complementary for each other in thermal plume monitoring.
Jianchao Fan, Shiyong Wen, Xiu Su
IGARSS6