EDBT 2026 Demo / reviewers in the wild / expert
Zhichao Lu
dblp:144/1417
· DBLP profile ↗
63ranked-venue papers
15as first author
47since 2021 · last 2026
0000-0002-4618-3573ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 13 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 20 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic BanditsabstractIn multi-objective decision-making with hierarchical preferences, lexicographic bandits provide a natural framework for optimizing multiple objectives in a prioritized order. In this setting, a learner repeatedly selects arms and observes reward vectors, aiming to maximize the reward for the highest-priority objective, then the next, and so on. While previous studies have primarily focused on regret minimization, this work bridges the gap between regret minimization and best arm identification under lexicographic preferences. We propose two elimination-based algorithms to address this joint objective. The first algorithm eliminates suboptimal arms sequentially, layer by layer, in accordance with the objective priorities, and achieves sample complexity and regret bounds comparable to those of the best single-objective algorithms. The second algorithm simultaneously leverages reward information from all objectives in each round, effectively exploiting cross-objective dependencies. Remarkably, it outperforms the known lower bound for the single-objective bandit problem, highlighting the benefit of cross-objective information sharing in the multi-objective setting. Empirical results further validate their superior performance over baselines. Bo Xue 0004, Yuanyu Wan, Zhichao Lu, Qingfu Zhang 0001 |
AAAI | 3 |
| 2026 | CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious AttacksabstractMultimodal Large Language Models (MLLMs) achieve strong reasoning and perception capabilities but are increasingly vulnerable to jailbreak attacks. While existing work focuses on explicit attacks, where malicious content resides in a single modality, recent studies reveal implicit attacks, in which benign text and image inputs jointly express unsafe intent. Such joint-modal threats are difficult to detect and remain underexplored, largely due to the scarcity of high-quality implicit data. We propose ImpForge, an automated red-teaming pipeline that leverages reinforcement learning with tailored reward modules to generate diverse implicit samples across 14 domains. Building on this dataset, we further develop CrossGuard, an intent-aware safeguard providing robust and comprehensive defense against both explicit and implicit threats. Extensive experiments across safe and unsafe benchmarks, implicit and explicit attacks, and multiple out-of-domain settings demonstrate that CrossGuard significantly outperforms existing defenses, including advanced MLLMs and guardrails, achieving stronger security while maintaining high utility. This offers a balanced and practical solution for enhancing MLLM robustness against real-world multimodal threats. Zhichao Lu |
ACL (1) | 3 |
| 2026 | Functional consistency of LLM code embeddings: A self-evolving data synthesis framework for benchmarking
Zhuohao Li, Wenqing Chen, Jianxing Yu, Zhichao Lu |
Expert Syst. Appl. | 4 |
| 2026 | BinaryAD: Efficient image anomaly detection via binarized representations
Bingyang Guo, Hanzhe Liang, LinLin Shen, Jinbao Wang 0001, Zhichao Lu |
Pattern Recognit. | 8 |
| 2026 | I2EKD: Efficient and Versatile Image-to-Event Knowledge DistillationabstractRecently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%. Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Toward Zero-Shot Point Cloud Anomaly Detection: A Multiview Projection FrameworkabstractDetecting anomalies within point clouds is crucial for various industrial applications, but traditional unsupervised methods face challenges due to data acquisition costs, early stage production constraints, and limited generalization across product categories. To overcome these challenges, we introduce the multiview projection (MVP) framework, leveraging pretrained vision-language models (VLMs) to detect anomalies. Specifically, MVP projects point cloud data into multiview depth images, thereby translating point cloud anomaly detection into image anomaly detection. Following zero-shot image anomaly detection methods, pretrained VLMs are utilized to detect anomalies on these depth images. Given that pretrained VLMs are not inherently tailored for zero-shot point cloud anomaly detection and may lack specificity, we propose the integration of learnable visual and adaptive text prompting techniques to fine-tune these VLMs, thereby enhancing their detection performance. Extensive experiments on the MVTec 3-D-AD and Real3D-AD demonstrate our proposed MVP framework’s superior zero-shot anomaly detection performance and the prompting techniques’ effectiveness. Real-world evaluations on automotive plastic part inspection further showcase that the proposed method can also be generalized to practical, unseen scenarios. Yunkang Cao, Guoyang Xie, Zhichao Lu, Weiming Shen 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural PerspectiveabstractExisting efforts to boost multimodal fusion of 3D anomaly detection (3D-AD) primarily concentrate on devising more effective multimodal fusion strategies. However, little attention was devoted to analyzing the role of multimodal fusion architecture (topology) design in contributing to 3D-AD. In this paper, we aim to bridge this gap and present a systematic study on the impact of multimodal fusion architecture design on 3D-AD. This work considers the multimodal fusion architecture design at the intra-module fusion level, i.e., independent modality-specific modules, involving early, middle or late multimodal features with specific fusion operations, and also at the inter-module fusion level, i.e., the strategies to fuse those modules. In both cases, we first derive insights through theoretically and experimentally exploring how architectural designs influence 3D-AD. Then, we extend SOTA neural architecture search (NAS) paradigm and propose 3D-ADNAS to simultaneously search across multimodal fusion strategies and modality-specific modules for the first time. Extensive experiments show that 3D-ADNAS obtains consistent improvements in 3D-AD across various model capacities in terms of accuracy, frame rate, and memory usage, and it exhibits great potential in dealing with few-shot 3D-AD tasks. Kaifang Long, Guoyang Xie, Lianbo Ma 0004, Zhichao Lu |
AAAI | 5 |
| 2025 | SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space ModelsabstractKnown as low energy consumption networks, spiking neural networks (SNNs) have gained a lot of attention within the past decades. While SNNs are increasing competitive with artificial neural networks (ANNs) for vision tasks, they are rarely used for long sequence tasks, despite their intrinsic temporal dynamics. In this work, we develop spiking state space models (SpikingSSMs) for long sequence learning by leveraging on the sequence learning abilities of state space models (SSMs). Inspired by dendritic neuron structure, we hierarchically integrate neuronal dynamics with the original SSM block, meanwhile realizing sparse synaptic computation. Furthermore, to solve the conflict of event-driven neuronal dynamics with parallel computing, we propose a light-weight surrogate dynamic network which accurately predicts the after-reset membrane potential and compatible to learnable thresholds, enabling orders of acceleration in training speed compared with conventional iterative methods. On the long range arena benchmark task, SpikingSSM achieves competitive performance to state-of-the-art SSMs meanwhile realizing on average 90% of network sparsity. On language modeling, our network significantly surpasses existing spiking large language models (spikingLLMs) on the WikiText-103 dataset with only a third of the model size, demonstrating its potential as backbone architecture for low computation cost LLMs. Shuaijie Shen, Renzhuo Huang, Yan Zhong 0001, Qinghai Guo, Zhichao Lu, Jianguo Zhang 0001, Luziwei Leng |
AAAI | 6 |
| 2025 | Mitigating Social Bias in Large Language Models: A Multi-Objective Approach Within a Multi-Agent FrameworkabstractNatural language processing (NLP) has seen remarkable advancements with the development of large language models (LLMs). Despite these advancements, LLMs often produce socially biased outputs. Recent studies have mainly addressed this problem by prompting LLMs to behave ethically, but this approach results in unacceptable performance degradation. In this paper, we propose a multi-objective approach within a multi-agent framework (MOMA) to mitigate social bias in LLMs without significantly compromising their performance. The key idea of MOMA involves deploying multiple agents to perform causal interventions on bias-related contents of the input questions, breaking the shortcut connection between these contents and the corresponding answers. Unlike traditional debiasing techniques leading to performance degradation, MOMA substantially reduces bias while maintaining accuracy in downstream tasks. Our experiments conducted in two datasets and two models demonstrate that MOMA reduces bias scores by up to 87.7%, with only a marginal performance degradation of up to 6.8% in the BBQ dataset. Additionally, it significantly enhances the multi-objective metric icat in the StereoSet dataset by up to 58.1%. Zhenjie Xu, Wenqing Chen, Xuanying Li, Zhixuan Chu, Kui Ren 0001, Zibin Zheng, Zhichao Lu |
AAAI | 9 |
| 2025 | Multi-Objective Evolution of Heuristic Using Large Language ModelabstractHeuristics are commonly used to tackle various search and optimization problems. Design heuristics usually require tedious manual crafting with domain knowledge. Recent works have incorporated Large Language Models (LLMs) into automatic heuristic search, leveraging their powerful language and coding capacity. However, existing research focuses on the optimal performance on the target problem as the sole objective, neglecting other criteria such as efficiency and scalability, which are vital in practice. To tackle this challenge, we propose to model the heuristic search as a multi-objective optimization problem and consider introducing additional practical criteria beyond optimal performance. Due to the complexity of the search space, conventional multi-objective optimization methods struggle to effectively handle LLM-based multi-objective heuristic search. We propose the first LLM-based multi-objective heuristic search framework, Multi-objective Evolution of Heuristic (MEoH), which integrates LLMs in a zero-shot manner to generate a non-dominated set of heuristics to meet multiple design criteria. We design a new dominance-dissimilarity mechanism for effective population management and selection, which incorporates both code dissimilarity in the search space and dominance in the objective space. MEoH is demonstrated in two well-known combinatorial optimization problems: the online Bin Packing Problem (BPP) and the Traveling Salesman Problem (TSP). The results indicate that a variety of elite heuristics are automatically generated in a single run, offering more trade-off options than the existing methods. It successfully achieves competitive or superior performance while improving efficiency up to 10 times. Moreover, we also observe that the multi-objective search introduces novel insights into heuristic design and leads to the discovery of diverse heuristics. Shunyu Yao 0002, Fei Liu 0044, Xi Lin 0001, Zhichao Lu, Zhenkun Wang 0001, Qingfu Zhang 0001 |
AAAI | 4 |
| 2025 | Design Principle Transfer in Neural Architecture Search via Large Language ModelsabstractTransferable neural architecture search (TNAS) has been introduced to design efficient neural architectures for multiple tasks, to enhance the practical applicability of NAS in real-world scenarios. In TNAS, architectural knowledge accumulated in previous search processes is reused to warm up the architecture search for new tasks. However, existing TNAS methods still search in an extensive search space, necessitating the evaluation of numerous architectures. To overcome this challenge, this work proposes a novel transfer paradigm, i.e., design principle transfer. In this work, the linguistic description of various structural components' effects on architectural performance is termed design principles. They are learned from established architectures and then can be reused to reduce the search space tasks by discarding unpromising architectures. Searching in the refined search space can boost both the search performance and efficiency for new NAS tasks. To this end, a large language model (LLM)-assisted design principle transfer (LAPT) framework is devised. In LAPT, LLM is applied to automatically reason the design principles from a set of given architectures, and then a principle adaptation method is applied to refine these principles progressively based on the search results. Experimental results demonstrate that LAPT can beat the state-of-the-art TNAS methods on most tasks and achieve comparable performance on the remainder. Liang Feng 0001, Zhichao Lu, Kay Chen Tan |
AAAI | 4 |
| 2025 | MOS-Attack: A Scalable Multi-objective Adversarial Attack FrameworkabstractCrafting adversarial examples is crucial for evaluating and enhancing the robustness of Deep Neural Networks (DNNs), presenting a challenge equivalent to maximizing a non-differentiable 0-1 loss function. However, existing single objective methods, namely adversarial attacks focus on a surrogate loss function, do not fully harness the benefits of engaging multiple loss functions, as a result of insufficient understanding of their synergistic and conflicting nature. To overcome these limitations, we propose the Multi-Objective Set-Based Attack (MOS Attack), a novel adversarial attack framework leveraging multiple loss functions and automatically uncovering their interrelations. The MOS Attack adopts a set-based multi-objective optimization strategy, enabling the incorporation of numerous loss functions without additional parameters. It also automatically mines synergistic patterns among various losses, facilitating the generation of potent adversarial attacks with fewer objectives. Extensive experiments have shown that our MOS Attack outperforms single-objective attacks. Furthermore, by harnessing the identified synergistic patterns, MOS Attack continues to show superior results with a reduced number of loss functions. Our code is available at https://github.com/pgg3/MOS-Attack. Ping Guo 0007, Xi Lin 0001, Fei Liu 0044, Zhichao Lu, Qingfu Zhang 0001, Zhenkun Wang 0001 |
CVPR | 5 |
| 2025 | DEIM: DETR with Improved Matching for Fast ConvergenceabstractWe introduce DEIM, an innovative and efficient training framework designed to accelerate convergence in real-time object detection with Transformer-based architectures (DETR). To mitigate the sparse supervision inherent in one-to-one (O2O) matching in DETR models, DEIM employs a Dense O2O matching strategy. This approach increases the number of positive samples per image by incorporating additional targets, using standard data augmentation techniques. While Dense O2O matching speeds up convergence, it also introduces numerous low-quality matches that could affect performance. To address this, we propose the Matchability-Aware Loss (MAL), a novel loss function that optimizes matches across various quality levels, enhancing the effectiveness of Dense O2O. Extensive experiments on the COCO dataset validate the efficacy of DEIM. When integrated with RT-DETR and D-FINE, it consistently boosts performance while reducing training time by 50%. Notably, paired with RT-DETRv2, DEIM achieves 53.2% AP in a single day of training on an NVIDIA 4090 GPU. Additionally, DEIM-trained real-time models outperform leading real-time object detectors, with DEIM-D-FINE-L and DEIM-D-FINE-X achieving 54.7% and 56.5% AP at 124 and 78 FPS on an NVIDIA T4 GPU, respectively, without the need for additional data. We believe DEIM sets a new baseline for advancements in real-time object detection. Our code and pre-trained models are available at https://www.shihuahuang.cn/DEIM/. Shihua Huang, Zhichao Lu, Xiaodong Cun, Yongjun Yu, Xi Shen 0001 |
CVPR | 2 |
| 2025 | Trade-Offs in Image Generation: How Do Different Dimensions Interact?
Binzhu Xie, Zhonghao Yan, Shi Qiu 0001, Guoyang Xie, Zhichao Lu |
ICCV | 10 |
| 2025 | Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient AligningabstractMixed Precision Quantization (MPQ) has become an essential technique for optimizing neural network by determining the optimal bitwidth per layer. Existing MPQ methods, however, face a major hurdle: they require a computationally expensive search for quantization strategies on large-scale datasets. To resolve this issue, we introduce a novel approach that first searches for quantization strategies on small datasets and then generalizes them to large-scale datasets. This approach simplifies the process, eliminating the need for large-scale quantization fine-tuning and only necessitating model weight adjustment. Our method is characterized by three key techniques: sharpness-aware minimization for enhanced quantized model generalization, implicit gradient direction alignment to handle gradient conflicts among different optimization objectives, and an adaptive perturbation radius to accelerate optimization. It offers advantages such as no intricate computation of feature maps and high search efficiency. Both theoretical analysis and experimental results validate our approach. Using the CIFAR10 dataset (just 0.5\% the size of ImageNet training data) for MPQ policy search, we achieved equivalent accuracy on ImageNet with a significantly lower computational cost, while improving efficiency by up to 150\% over the baselines. Lianbo Ma 0004, Jianlun Ma, Yuee Zhou, Guoyang Xie, Qiang He 0002, Zhichao Lu |
ICML | 6 |
| 2025 | Uncertainty-Guided Enhancement on Driving Perception System Via Foundation ModelsabstractMultimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models to refine predictions from existing driving perception modelssuch as enhancing object classification accuracy-while minimizing the frequency of using these resource-intensive models. The method quantitatively characterizes uncertainties in the perception model's predictions and engages the foundation model only when these uncertainties exceed a pre-specified threshold. Specifically, it characterizes uncertainty by calibrating the perception model's confidence scores into theoretical lower bounds on the probability of correct predictions using conformal prediction. Then, it sends images to the foundation model and queries for refining the predictions only if the theoretical bound of the perception model's outcome is below the threshold. Additionally, we propose a temporal inference mechanism that enhances prediction accuracy by integrating historical predictions, leading to tighter theoretical bounds. The method demonstrates a 10 to 15 percent improvement in prediction accuracy and reduces the number of queries to the foundation model by 50 percent, based on quantitative evaluations from driving datasets. Yunhao Yang, Zaiwei Zhang, Zhichao Lu, Ufuk Topcu, Ben Snyder |
ICRA | 5 |
| 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly SynthesisabstractIndustrial anomaly segmentation relies heavily on pixel-level annotations, yet real-world anomalies are often scarce, diverse, and costly to label. Segmentation-oriented industrial anomaly synthesis (SIAS) has emerged as a promising alternative; however, existing methods struggle to balance sampling efficiency and generation quality. Moreover, most approaches treat all spatial regions uniformly, overlooking the distinct statistical differences between anomaly and background areas. This uniform treatment hinders the synthesis of controllable, structure-specific anomalies tailored for segmentation tasks. In this paper, we propose FAST, a foreground-aware diffusion framework featuring two novel modules: the Anomaly-Informed Accelerated Sampling (AIAS) and the Foreground-Aware Reconstruction Module (FARM). AIAS is a training-free sampling algorithm specifically designed for segmentation-oriented industrial anomaly synthesis, which accelerates the reverse process through coarse-to-fine aggregation and enables the synthesis of state-of-the-art segmentation-oriented anomalies in as few as 10 steps. Meanwhile, FARM adaptively adjusts the anomaly-aware noise within the masked foreground regions at each sampling step, preserving localized anomaly signals throughout the denoising trajectory. Extensive experiments on multiple industrial benchmarks demonstrate that FAST consistently outperforms existing anomaly synthesis methods in downstream segmentation tasks. We release the code in https://github.com/Chhro123/fast-foreground-aware-anomaly-synthesis. Xichen Xu, Yanshu Wang, Jinbao Wang 0001, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu |
NeurIPS | 7 |
| 2025 | Enhancing Multimodal Learning via Hierarchical Fusion Architecture Search With Inconsistency MitigationabstractThe design of effective multimodal feature fusion strategies is the key task for multimodal learning, which often requires huge computational costs with extensive expertise. In this paper, we seek to enhance multimodal learning via hierarchical fusion architecture search with inconsistency mitigation. Different from previous works, our Hierarchical Fusion Multimodal Neural Architecture Search (HF-MNAS) considers the inconsistency in modalities and labels, and fine-grained exploitation in multi-level fusion architectures. Specifically, it disentangles the hierarchical fusion problem into two-level (macro- and micro-level) search spaces. In the macro-level search space, the high-level and low-level features are extracted and then connected in a fine-grained way, where the inconsistency mitigation module is designed to minimize discrepancies between modalities and labels in cell outputs. In the micro-level search space, we find that different intermediate nodes in the cells exhibit different importance degrees. Then, we propose an importance-based node selection mechanism to form the optimal cells for feature fusion. We evaluate HF-MNAS on a series of multimodal classification tasks. Empirical evidence shows that HF-MNAS achieves competitive trade-off performance across accuracy, search time, and inference speed. In particular, HF-MNAS consumes minimal computational cost compared with state-of-the-art MNASs. Furthermore, we theoretically and experimentally verify that the modality-label inconsistency deteriorates the overall fusion performance of models such as accuracy and F1 score, and demonstrate that the proposed inconsistency mitigation module could effectively mitigate this phenomenon. Kaifang Long, Guoyang Xie, Lianbo Ma 0004, Qing Li 0006, Min Huang 0001, Jianhui Lv, Zhichao Lu |
IEEE Trans. Image Process. | 7 |
| 2025 | NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNNabstractDeploying quantized deep neural network (DNN) models with resource adaptation capabilities on ubiquitous Internet of Things (IoT) devices to provide high-quality AI services can leverage the benefits of compression and meet multi-scenario resource requirements. However, existing dynamic/mixed precision quantization requires retraining or special hardware, whereas post-training quantization (PTQ) has two limitations for resource adaptation: (i) The state-of-the-art PTQ methods only provide one fixed bitwidth model, which makes it challenging to adapt to the dynamic resources of IoT devices; (ii) Deploying multiple PTQ models with diverse bitwidths consumes large storage resources and switching overheads. To this end, this paper introduces a resource-friendly post-training integer-nesting quantization, i.e., NestQuant, for on-device quantized model switching on IoT devices. The proposed NestQuant incorporates the integer weight decomposition, which bit-wise splits quantized weights into higher-bit and lower-bit weights of integer data types. It also contains a decomposed weights nesting mechanism to optimize the higher-bit weights by adaptive rounding and nest them into the original quantized weights. In deployment, we can send and store only one NestQuant model and switch between the full-bit/part-bit model by paging in/out lower-bit weights to adapt to resource changes and reduce consumption. Experimental results on the ImageNet-1K pretrained DNNs demonstrated that the NestQuant model can achieve high performance in top-1 accuracy, and reduce in terms of data transmission, storage consumption, and switching overheads. In particular, the ResNet-101 with INT8 nesting INT6 can achieve 78.1% and 77.9% accuracy for full-bit and part-bit models, respectively, and reduce switching overheads by approximately 78.1% compared with diverse bitwidths PTQ models. Code:https://github.com/jianhayes/NESTQUANT. Jianhang Xie, Chuntao Ding, Xiaqing Li, Shenyuan Ren, Yidong Li, Zhichao Lu |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Rethinking Unsupervised Outlier Detection via Multiple Thresholding
Zhonghang Liu, Panzhong Lu, Guoyang Xie, Zhichao Lu, Wen-Yan Lin |
ECCV (18) | 4 |
| 2024 | Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language ModelabstractHeuristics are widely used for dealing with complex search and optimization problems. However, manual design of heuristics can be often very labour extensive and requires rich working experience and knowledge. This paper proposes Evolution of Heuristic (EoH), a novel evolutionary paradigm that leverages both Large Language Models (LLMs) and Evolutionary Computation (EC) methods for Automatic Heuristic Design (AHD). EoH represents the ideas of heuristics in natural language, termed thoughts. They are then translated into executable codes by LLMs. The evolution of both thoughts and codes in an evolutionary search framework makes it very effective and efficient for generating high-performance heuristics. Experiments on three widely studied combinatorial optimization benchmark problems demonstrate that EoH outperforms commonly used handcrafted heuristics and other recent AHD methods including FunSearch. Particularly, the heuristic produced by EoH with a low computational budget (in terms of the number of queries to LLMs) significantly outperforms widely-used human hand-crafted baseline algorithms for the online bin packing problem. Fei Liu 0044, Xialiang Tong, Mingxuan Yuan, Xi Lin 0001, Fu Luo, Zhenkun Wang 0001, Zhichao Lu, Qingfu Zhang 0001 |
ICML | 7 |
| 2024 | Understanding the Importance of Evolutionary Search in Automated Heuristic Design with Large Language Models
Rui Zhang 0042, Fei Liu 0044, Xi Lin 0001, Zhenkun Wang 0001, Zhichao Lu, Qingfu Zhang 0001 |
PPSN (2) | 5 |
| 2024 | Neural Architecture Search as Multiobjective Optimization Benchmarks: Problem Formulation and Performance AssessmentabstractThe ongoing advancements in network architecture design have led to remarkable achievements in deep learning across various challenging computer vision tasks. Meanwhile, the development of neural architecture search (NAS) has provided promising approaches to automating the design of network architectures for lower prediction error. Recently, the emerging application scenarios of deep learning (e.g., autonomous driving) have raised higher demands for network architectures considering multiple design criteria: number of parameters/weights, number of floating-point operations, inference latency, among others. From an optimization point of view, the NAS tasks involving multiple design criteria are intrinsically multiobjective optimization problems; hence, it is reasonable to adopt evolutionary multiobjective optimization (EMO) algorithms for tackling them. Nonetheless, there is still a clear gap confining the related research along this pathway: on the one hand, there is a lack of a general problem formulation of NAS tasks from an optimization point of view; on the other hand, there are challenges in conducting benchmark assessments of EMO algorithms on NAS tasks. To bridge the gap: 1) we formulate NAS tasks into general multiobjective optimization problems and analyze the complex characteristics from an optimization point of view; 2) we present an end-to-end pipeline, dubbedEvoXBench, to generate benchmark test problems for EMO algorithms to run efficiently—without the requirement of GPUs or Pytorch/Tensorflow; and 3) we instantiate two test suites comprehensively covering two datasets, seven search spaces, and three hardware devices, involving up to eight objectives. Based on the above, we validate the proposed test suites using six representative EMO algorithms and provide some empirical analyses. The code ofEvoXBenchis available athttps://github.com/EMI-Group/EvoXBench. Zhichao Lu, Ran Cheng 0004, Yaochu Jin, Kay Chen Tan, Kalyanmoy Deb |
IEEE Trans. Evol. Comput. | 1 |
| 2024 | A Resource-Efficient Feature Extraction Framework for Image Processing in IoT DevicesabstractExtracting features from image data on Internet of Things (IoT) devices to reduce the amount of data that needs to be uploaded to cloud/edge servers has received increasing attention. However, most of the existing related approaches suffer from two major limitations, (i) low performance and high network traffic, and (ii) a lot of storage resource consumption. To this end, we propose a resource-efficient feature extraction framework for image processing in IoT devices. The proposed framework consists of the edge-assisted extractor generation method and the NestE method. The extractor generated by the edge-assisted extractor generation method can extract the features required by the application, which can not only avoid the IoT device uploading useless feature data but also improve application performance. The proposed NestE generates a nonredundant subextractor by splitting the extractor into multiple subextractors, removing redundant subextractors, and nesting small-capacity subextractors in large-capacity subextractors in a parameter-sharing manner. Compared with deploying multiple independent subextractors on IoT devices, deploying the nonredundant multifunctional extractor can save considerable storage resources and switching overhead. Extensive experimental results show that the proposed framework reduces the storage footprint by approximately 90.7% and switching overhead by approximately 92.4% compared with deploying independent subextractors when using the classical principal component analysis algorithm. Chuntao Ding, Yidong Li, Zhichao Lu, Shangguang Wang, Song Guo 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Mitigating Task Interference in Multi-Task Learning via Explicit Task Routing with Non-Learnable PrimitivesabstractMulti-task learning (MTL) seeks to learn a single model to accomplish multiple tasks by leveraging shared information among the tasks. Existing MTL models, however, have been known to suffer from negative interference among tasks. Efforts to mitigate task interference have focused on either loss/gradient balancing or implicit parameter partitioning with partial overlaps among the tasks. In this paper, we propose ETR-NLP to mitigate task interference through a synergistic combination of non-learnable primitives (NLPs) and explicit task routing (ETR). Our key idea is to employ non-learnable primitives to extract a diverse set of task-agnostic features and recombine them into a shared branch common to all tasks and explicit task-specific branches reserved for each task. The non-learnable primitives and the explicit decoupling of learnable parameters into shared and task-specific ones afford the flexibility needed for minimizing task interference. We evaluate the efficacy of ETR-NLP networks for both image-level classification and pixel-level dense prediction MTL problems. Experimental results indicate that ETR-NLP significantly outperforms state-of-the-art baselines with fewer learnable parameters and similar FLOPs across all datasets. Code is available at this URL. Chuntao Ding, Zhichao Lu, Shangguang Wang, Ran Cheng 0004, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2023 | Revisiting Residual Networks for Adversarial RobustnessabstractEfforts to improve the adversarial robustness of convolutional neural networks have primarily focused on developing more effective adversarial training methods. In contrast, little attention was devoted to analyzing the role of architectural elements (e.g., topology, depth, and width) on adversarial robustness. This paper seeks to bridge this gap and present a holistic study on the impact of architectural design on adversarial robustness. We focus on residual networks and consider architecture design at the block level as well as at the network scaling level. In both cases, we first derive insights through systematic experiments. Then we design a robust residual block, dubbed RobustResBlock, and a compound scaling rule, dubbed RobustScaling, to distribute depth and width at the desired FLOP count. Finally, we combine RobustResBlock and RobustScaling and present a portfolio of adversarially robust residual networks, RobustResNets, spanning a broad spectrum of model capacities. Experimental validation across multiple datasets and adversarial attacks demonstrate that RobustResNets consistently outperform both the standard WRNs and other existing robust architectures, achieving state-of-the-art AutoAttack robust accuracy 63.7% with 500K external data while being 2× more compact in terms of parameters. Code is available at this URL. Shihua Huang, Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2023 | Seed Feature Maps-based CNN Models for LEO Satellite Remote Sensing ServicesabstractDeploying high-performance convolutional neural network (CNN) models on low-earth orbit (LEO) satellites for rapid remote sensing image processing has attracted significant interest from industry and academia. However, the limited resources available on LEO satellites contrast with the demands of resource-intensive CNN models, necessitating the adoption of ground-station server assistance for training and updating these models. Existing approaches often require large floating-point operations (FLOPs) and substantial model parameter transmissions, presenting considerable challenges. To address these issues, this paper introduces a ground-station server-assisted framework. With the proposed framework, each layer of the CNN model contains only one learnable feature map (called the seed feature map) from which other feature maps are generated based on specific rules. The hyperparameters of these rules are randomly generated instead of being trained, thus enabling the generation of multiple feature maps from the seed feature map and significantly reducing FLOPs. Furthermore, since the random hyperparameters can be saved using a few random seeds, the ground station server assistance can be facilitated in updating the CNN model deployed on the LEO satellite. Experimental results on the ISPRS Vaihingen, ISPRS Potsdam, UAVid, and LoveDA datasets for semantic segmentation services demonstrate that the proposed framework outperforms existing state-of-the-art approaches. In particular, the SineFM-based model achieves a higher mIoU than the UNetFormer on the UAVid dataset, with 3.3 × fewer parameters and 2.2 × fewer FLOPs. Zhichao Lu, Chuntao Ding, Shangguang Wang, Ran Cheng 0004, Felix Juefei-Xu, Vishnu Naresh Boddeti |
ICWS | 1 |
| 2023 | GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 17 |
| 2023 | A general framework for enhancing relaxed Pareto dominance methods in evolutionary many-objective optimization
Shuwei Zhu, Lihong Xu, Erik D. Goodman, Kalyanmoy Deb, Zhichao Lu |
Nat. Comput. | 5 |
| 2023 | Minimizing Expected Deviation in Upper Level Outcomes Due to Lower Level Decision Making in Hierarchical Multiobjective ProblemsabstractMany societal and industrial problem-solving tasks involving search, optimization, design, and management are conveniently decomposed into hierarchical subproblems. While this process allows a systematic procedure to have a multistakeholder solution, the independent decision-making process for the lower level problem causes a deviation in the expected outcome of the upper level problem. In this article, we provide a new and computationally efficient evolutionary approach allowing upper level decision makers to analyze the vagaries of lower level decision making when choosing a preferred solution with the minimum deviation from their expectations. This concept is novel and pragmatic. We demonstrate the concept through a search for optimistic–pessimistic tradeoff solutions found by an evolutionary multiobjective optimization approach first on two difficult test problems, then on a watershed management problem and a telecommunication management problem. The approach is generic and can be applied to similar hierarchical management problems to achieve minimum deviation with a more predictive and reliable outcome. The proposed solution procedure is found to choose an optimistic solution that has approximately 31%–65% reduced deviation compared to another optimistic solution chosen at random in the test problems and approximately 85%–95% reduced deviation in the two practical problems, making the method of this study applicable to practical hierarchical problems. Kalyanmoy Deb, Zhichao Lu, Ian Kropp, Juan Sebastian Hernandez-Suarez, Rayan Hussein, A. Pouyan Nejadhashemi |
IEEE Trans. Evol. Comput. | 2 |
| 2023 | Towards Transmission-Friendly and Robust CNN Models over Cloud and DeviceabstractDeploying deep convolutional neural network (CNN) models on ubiquitous Internet of Things (IoT) devices has attracted much attention from industry and academia since it greatly facilitates our lives by providing various rapid-response services. Due to the limited resources of IoT devices, cloud-assisted training of CNN models has become the mainstream. However, most existing related works suffer froma large amount of model parameter transmission and weak model robustness. To this end, this paper proposes a cloud-assisted CNN training framework with low model parameter transmission and strong model robustness. In the proposed framework, we first introduce MonoCNN, which contains only a few learnable filters, and other filters are nonlearnable. These nonlearnable filter parameters are generated according to certain rules, i.e., the filter generation function (FGF), and can be saved and reproduced by a few random seeds. Thus, the cloud server only needs to send these learnable filters and a few seeds to the IoT device. Compared to transmitting all model parameters, sending several learnable filter parameters and seeds can significantly reduce parameter transmission. Then, we investigate multiple FGFs and enable the IoT device to use the FGF to generate multiple filters and combine them into MonoCNN. Thus, MonoCNN is affected not only by the training data but also by the FGF. The rules of the FGF play a role in regularizing the MonoCNN, thereby improving its robustness. Experimental results show that compared to state-of-the-art methods, our proposed framework can reduce a large amount of model parameter transfer between the cloud server and the IoT device while improving the performance by approximately 2.2% when dealing with corrupted data. Chuntao Ding, Zhichao Lu, Felix Juefei-Xu, Vishnu Naresh Boddeti, Yidong Li, Jiannong Cao 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | TFormer: A Transmission-Friendly ViT Model for IoT DevicesabstractDeploying high-performance vision transformer (ViT) models on ubiquitous Internet of Things (IoT) devices to provide high-quality vision services will revolutionize the way we live, work, and interact with the world. Due to the contradiction between the limited resources of IoT devices and resource-intensive ViT models, the use of cloud servers to assist ViT model training has become mainstream. However, due to the larger number of parameters and floating-point operations (FLOPs) of the existing ViT models, the model parameters transmitted by cloud servers are large and difficult to run on resource-constrained IoT devices. To this end, this article proposes a transmission-friendly ViT model, TFormer, for deployment on resource-constrained IoT devices with the assistance of a cloud server. The high performance and small number of model parameters and FLOPs of TFormer are attributed to the proposed hybrid layer and the proposed partially connected feed-forward network (PCS-FFN). The hybrid layer consists of nonlearnable modules and a pointwise convolution, which can obtain multitype and multiscale features with only a few parameters and FLOPs to improve the TFormer performance. The PCS-FFN adopts group convolution to reduce the number of parameters. The key idea of this article is to propose TFormer with few model parameters and FLOPs to facilitate applications running on resource-constrained IoT devices to benefit from the high performance of the ViT models. Experimental results on the ImageNet-1K, MS COCO, and ADE20K datasets for image classification, object detection, and semantic segmentation tasks demonstrate that the proposed model outperforms other state-of-the-art models. Specifically, TFormer-S achieves 5% higher accuracy on ImageNet-1K than ResNet18 with 1.4× fewer parameters and FLOPs. Zhichao Lu, Chuntao Ding, Felix Juefei-Xu, Vishnu Naresh Boddeti, Shangguang Wang, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-conceptsabstractLearning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this challenge, we introduce a new method for pre-training video action recognition models using queried web videos. Instead of trying to filter out potential noises, we propose to provide fine-grained supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations and improves data utilization efficiency of untrimmed videos. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset. Experiments show that SPL outperforms several existing pre-training strategies and the learned representations lead to competitive results on several benchmarks. Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu 0001, Tomas Pfister |
AAAI | 6 |
| 2022 | Multiview Transformers for Video RecognitionabstractVideo understanding requires reasoning at multiple spatiotemporal resolutions – from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal resolutions. To this end, we present Multiview Transformers for Video Recognition (MTV). Our model consists of separate encoders to represent different views of the input video with lateral connections to fuse information across views. We present thorough ablation studies of our model and show that MTV consistently performs better than single-view counterparts in terms of accuracy and computational cost across a range of model sizes. Furthermore, we achieve state-of-the-art results on six standard datasets, and improve even further with large-scale pretraining. Code and checkpoints are available at: https://github.com/google-research/scenic. Shen Yan 0008, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang 0002, Chen Sun 0002, Cordelia Schmid |
CVPR | 4 |
| 2022 | Tailor: removing redundant operations in memristive analog neural network acceleratorsabstractAnalog in-situ computation based on memristive circuits has been regarded as a promising approach for designing high-performance and low-power neural network accelerators. However, despite the low-cost and highly parallel memristive crossbars, the peripheral circuits especially analog-digital-converters (ADCs) induce significant overhead. Quantitative analysis shows that ADCs can contribute up to 91% energy consumption and 72% chip area, which significantly offset the advantages of memristive NN accelerators. Zhihang Yuan, Guangyu Sun 0003, Zhichao Lu |
DAC | 5 |
| 2022 | Enabling High-Quality Uncertainty Quantification in a PIM Designed for Bayesian Neural NetworkabstractUncertainty quantification measures the prediction uncertainty of a neural network facing out-of-training-distribution samples. Bayesian Neural Networks (BNNs) can provide high-quality uncertainty quantification by introducing specific noise to the weights during inference. To accelerate BNN inference, ReRAM processing-in-memory (PIM) architecture is a competitive solution to provide both high-efficient computing and in-situ noise generation at the same time. However, there normally exists a huge gap between the generated noise in PIM hardware and that required by a BNN model. We demonstrate that the quality of uncertainty quantification is substantially degraded due to this gap. To solve this problem, we propose a holistic framework called W2W-PIM. We first introduce an efficient method to generate noise in ReRAM PIM design according to the demand of a BNN model. In addition, the PIM architecture is carefully modified to enable the noise generation and evaluate uncertainty quality. Moreover, a calibration unit is further introduced to reduce the noise gap caused by imperfection of the noise model. Comprehensive evaluation results demonstrate that W2W-PIM framework can achieve high-quality uncertainty quantification and high energy-efficiency at the same time. Bingzhe Wu, Guangyu Sun 0003, Zhe Zhang 0006, Zhihang Yuan, Runsheng Wang, Ru Huang 0001, Dimin Niu, Hongzhong Zheng, Zhichao Lu, Meng-Fan Chang, Tianchan Guan, Xin Si |
HPCA | 10 |
| 2022 | VLMixer: Unpaired Vision-Language Pre-training via Cross-Modal CutMixabstractExisting vision-language pre-training (VLP) methods primarily rely on paired image-text datasets, which are either annotated by enormous human labors or crawled from the internet followed by elaborate data cleaning techniques. To reduce the dependency on well-aligned image-text pairs, it is promising to directly leverage the large-scale text-only and image-only corpora. This paper proposes a data augmentation method, namely cross-modal CutMix (CMC), for implicit cross-modal alignment learning in unpaired VLP. Specifically, CMC transforms natural sentences in the textual view into a multi-modal view, where visually-grounded words in a sentence are randomly replaced by diverse image patches with similar semantics. There are several appealing proprieties of the proposed CMC. First, it enhances the data diversity while keeping the semantic meaning intact for tackling problems where the aligned data are scarce; Second, by attaching cross-modal noise on uni-modal data, it guides models to learn token-level interactions across modalities for better denoising. Furthermore, we present a new unpaired VLP method, dubbed as VLMixer, that integrates CMC with contrastive learning to pull together the uni-modal and multi-modal views for better instance-level alignments among different modalities. Extensive experiments on five downstream tasks show that VLMixer could surpass previous state-of-the-art unpaired VLP methods. Teng Wang 0007, Zhichao Lu, Feng Zheng 0001, Ran Cheng 0004, Chengguo Yin, Ping Luo 0002 |
ICML | 3 |
| 2022 | PERF-Net: Pose Empowered RGB-Flow NetabstractIn recent years, many works in the video action recognition literature have shown that two stream models (combining spatial and temporal input streams) are necessary for achieving state-of-the-art performance. In this paper we show the benefits of including yet another stream based on human pose estimated from each frame — specifically by rendering pose on input RGB frames. At first blush, this additional stream may seem redundant given that human pose is fully determined by RGB pixel values — however we show (perhaps surprisingly) that this simple and flexible addition can provide complementary gains. Using this insight, we propose a new model, which we dub PERF-Net (short for Pose Empowered RGB-Flow Net), which combines this new pose stream with the standard RGB and flow based input streams via distillation techniques and show that our model outperforms the state-of-the-art by a large margin in a number of human action recognition datasets while not requiring flow or pose to be explicitly computed at inference time. The proposed pose stream is also part of the winner solution of the ActivityNet Kinetics Challenge 2020 [1]. Yinxiao Li, Zhichao Lu, Xuehan Xiong, Jonathan Huang |
WACV | 2 |
| 2022 | A New Many-Objective Evolutionary Algorithm Based on Generalized Pareto DominanceabstractIn the past several years, it has become apparent that the effectiveness of Pareto-dominance-based multiobjective evolutionary algorithms deteriorates progressively as the number of objectives in the problem, given by M , grows. This is mainly due to the poor discriminability of Pareto optimality in many-objective spaces (typically M ≥ 4 ). As a consequence, research efforts have been driven in the general direction of developing solution ranking methods that do not rely on Pareto dominance (e.g., decomposition-based techniques), which can provide sufficient selection pressure. However, it is still a nontrivial issue for many existing non-Pareto-dominance-based evolutionary algorithms to deal with unknown irregular Pareto front shapes. In this article, a new many-objective evolutionary algorithm based on the generalization of Pareto optimality (GPO) is proposed, which is simple, yet effective, in addressing many-objective optimization problems. The proposed algorithm used an "( M-1 ) + 1" framework of GPO dominance, ( M-1 )-GPD for short, to rank solutions in the environmental selection step, in order to promote convergence and diversity simultaneously. To be specific, we apply M symmetrical cases of ( M-1 )-GPD, where each enhances the selection pressure of M-1 objectives by expanding the dominance area of solutions, while remaining unchanged for the one objective left out of that process. Experiments demonstrate that the proposed algorithm is very competitive with the state-of-the-art methods to which it is compared, on a variety of scalable benchmark problems. Moreover, experiments on three real-world problems have verified that the proposed algorithm can outperform the others on each of these problems. Shuwei Zhu, Lihong Xu, Erik D. Goodman, Zhichao Lu |
IEEE Trans. Cybern. | 4 |
| 2021 | A Compute-in-Memory Architecture Compatible with 3D NAND Flash that Parallelly Activates Multi-LayersabstractCompute-In-Memory (CIM) architectures based on emerging non-volatile memories have demonstrated great potential in accelerating neural network computation for AI applications. However, the reliability challenges associated with multi-level cells and the lack of mature 3D-integration scheme have limited the model size and energy efficiency of these architectures. In this work, we propose a novel NAND-based architecture to efficiently accelerate the vector-matrix multiplication for deep neural networks. The proposed approach is fully compatible with 3D-NAND and allows multiple layers of wordline (WL) planes to be activated in parallel, as opposed to the previous layer-by-layer activation. The revolutionary linear-VTcorrection and positive-negative weights techniques help to achieve multilevel weight storage and better computing precision. The feasibility and accuracy of the proposed architecture have been verified using TCAD, SPICE and system-level simulations based on commercial 3D-NAND parameters. Major advantages of the approach include $16 \sim32\mathrm{x}$ increase of array utilization and $64 \sim128\mathrm{x}$ reduction of read power consumption. Chu Yan, Fan Yang 0054, Shifan Gao, Gabriel Rosca, Dan Manea, Zhichao Lu, Yi Zhao 0015 |
DAC | 7 |
| 2021 | Multi-objective Neural Architecture Search with Almost No Training
Shengran Hu, Ran Cheng 0004, Cheng He 0001, Zhichao Lu |
EMO | 4 |
| 2021 | The (M-1)+1 Framework of Relaxed Pareto Dominance for Evolutionary Many-Objective Optimization
Shuwei Zhu, Lihong Xu, Erik D. Goodman, Kalyanmoy Deb, Zhichao Lu |
EMO | 5 |
| 2021 | The surprising impact of mask-head architecture on novel class segmentationabstractInstance segmentation models today are very accurate when trained on large annotated datasets, but collecting mask annotations at scale is prohibitively expensive. We address the partially supervised instance segmentation problem in which one can train on (significantly cheaper) bounding boxes for all categories but use masks only for a subset of categories. In this work, we focus on a popular family of models which apply differentiable cropping to a feature map and predict a mask based on the resulting crop. Under this family, we study Mask R-CNN and discover that instead of its default strategy of training the mask-head with a combination of proposals and groundtruth boxes, training the mask-head with only groundtruth boxes dramatically improves its performance on novel classes. This training strategy also allows us to take advantage of alternative mask-head architectures, which we exploit by replacing the typical mask-head of 2-4 layers with significantly deeper off-the-shelf architectures (e.g. ResNet, Hourglass models). While many of these architectures perform similarly when trained in fully supervised mode, our main finding is that they can generalize to novel classes in dramatically different ways. We call this ability of mask-heads to generalize to unseen classes the strong mask generalization effect and show that without any specialty modules or losses, we can achieve state-of-the-art results in the partially supervised COCO instance segmentation benchmark. Finally, we demonstrate that our effect is general, holding across underlying detection methodologies (including anchor-based, anchor-free or no detector at all) and across different backbone networks. Code and pre-trained models are available at https://git.io/deepmac. Vighnesh Birodkar, Zhichao Lu, Siyang Li 0002, Vivek Rathod, Jonathan Huang |
ICCV | 2 |
| 2021 | FaPN: Feature-aligned Pyramid Network for Dense Image PredictionabstractRecent advancements in deep neural networks have made remarkable leap-forwards in dense image prediction. However, the issue of feature alignment remains as neglected by most existing approaches for simplicity. Direct pixel addition between upsampled and local features leads to feature maps with misaligned contexts that, in turn, translate to mis-classifications in prediction, especially on object boundaries. In this paper, we propose a feature alignment module that learns transformation offsets of pixels to contextually align upsampled higher-level features; and another feature selection module to emphasize the lower-level features with rich spatial details. We then integrate these two modules in a top-down pyramidal architecture and present the Feature-aligned Pyramid Network (FaPN). Extensive experimental evaluations on four dense prediction tasks and four datasets have demonstrated the efficacy of FaPN, yielding an overall improvement of 1.2 - 2.6 points in AP / mIoU over FPN when paired with Faster / Mask R-CNN. In particular, our FaPN achieves the state-of-the-art of 56.7% mIoU on ADE20K when integrated within Mask-Former. The code is available from https://github.com/EMI-Group/FaPN. Shihua Huang, Zhichao Lu, Ran Cheng 0004, Cheng He 0001 |
ICCV | 2 |
| 2021 | End-to-End Dense Video Captioning with Parallel DecodingabstractDense video captioning aims to generate multiple associated captions with their temporal locations from the video. Previous methods follow a sophisticated "localizethen-describe" scheme, which heavily relies on numerous hand-crafted components. In this paper, we proposed a simple yet effective framework for end-to-end dense video captioning with parallel decoding (PDVC), by formulating the dense caption generation as a set prediction task. In practice, through stacking a newly proposed event counter on the top of a transformer decoder, the PDVC precisely segments the video into a number of event pieces under the holistic understanding of the video content, which effectively increases the coherence and readability of predicted captions. Compared with prior arts, the PDVC has several appealing advantages: (1) Without relying on heuristic non-maximum suppression or a recurrent event sequence selection network to remove redundancy, PDVC directly produces an event set with an appropriate size; (2) In contrast to adopting the two-stage scheme, we feed the enhanced representations of event queries into the localization head and caption head in parallel, making these two sub-tasks deeply interrelated and mutually promoted through the optimization; (3) Without bells and whistles, extensive experiments on ActivityNet Captions and YouCook2 show that PDVC is capable of producing high-quality captioning results, surpassing the state-of-the-art two-stage methods when its localization accuracy is on par with them. Code is available at https://github.com/ttengwang/PDVC. Teng Wang 0007, Ruimao Zhang, Zhichao Lu, Feng Zheng 0001, Ran Cheng 0004, Ping Luo 0002 |
ICCV | 3 |
| 2021 | Neural Architecture TransferabstractNeural architecture search (NAS) has emerged as a promising avenue for automatically designing task-specific neural networks. Existing NAS approaches require one complete search for each deployment specification of hardware or objective. This is a computationally impractical endeavor given the potentially large number of application scenarios. In this paper, we propose Neural Architecture Transfer (NAT) to overcome this limitation. NAT is designed to efficiently generate task-specific custom models that are competitive under multiple conflicting objectives. To realize this goal we learn task-specific supernets from which specialized subnets can be sampled without any additional training. The key to our approach is an integrated online transfer learning and many-objective evolutionary search procedure. A pre-trained supernet is iteratively adapted while simultaneously searching for task-specific subnets. We demonstrate the efficacy of NAT on 11 benchmark image classification tasks ranging from large-scale multi-class to small-scale fine-grained datasets. In all cases, including ImageNet, NATNets improve upon the state-of-the-art under mobile settings ( ≤ 600M Multiply-Adds). Surprisingly, small-scale fine-grained datasets benefit the most from NAT. At the same time, the architecture search and transfer is orders of magnitude more efficient than existing NAS methods. Overall, experimental evaluation indicates that, across diverse image classification tasks and computational objectives, NAT is an appreciably more effective alternative to conventional transfer learning of fine-tuning weights of an existing network architecture learned on standard datasets. Code is available at https://github.com/human-analysis/neural-architecture-transfer. Zhichao Lu, Gautam Sreekumar, Erik D. Goodman, Wolfgang Banzhaf, Kalyanmoy Deb, Vishnu Naresh Boddeti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Multiobjective Evolutionary Design of Deep Convolutional Neural Networks for Image ClassificationabstractConvolutional neural networks (CNNs) are the backbones of deep learning paradigms for numerous vision tasks. Early advancements in CNN architectures are primarily driven by human expertise and by elaborate design processes. Recently, neural architecture search was proposed with the aim of automating the network design process and generating task-dependent architectures. While existing approaches have achieved competitive performance in image classification, they are not well suited to problems where the computational budget is limited for two reasons: 1) the obtained architectures are either solely optimized for classification performance, or only for one deployment scenario and 2) the search process requires vast computational resources in most approaches. To overcome these limitations, we propose an evolutionary algorithm for searching neural architectures under multiple objectives, such as classification performance and floating point operations (FLOPs). The proposed method addresses the first shortcoming by populating a set of architectures to approximate the entire Pareto frontier through genetic operations that recombine and modify architectural components progressively. Our approach improves computational efficiency by carefully down-scaling the architectures during the search as well as reinforcing the patterns commonly shared among past successful architectures through Bayesian model learning. The integration of these two main contributions allows an efficient design of architectures that are competitive and in most cases outperform both manually and automatically designed architectures on benchmark image classification datasets: CIFAR, ImageNet, and human chest X-ray. The flexibility provided from simultaneously obtaining multiple architecture choices for different compute requirements further differentiates our approach from other methods in the literature. Zhichao Lu, Ian Whalen, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
IEEE Trans. Evol. Comput. | 1 |
| 2020 | MUXConv: Information Multiplexing in Convolutional Neural NetworksabstractConvolutional neural networks have witnessed remarkable improvements in computational efficiency in recent years. A key driving force has been the idea of trading-off model expressivity and efficiency through a combination of 1x1 and depth-wise separable convolutions in lieu of a standard convolutional layer. The price of the efficiency, however, is the sub-optimal flow of information across space and channels in the network. To overcome this limitation, we present MUXConv, a layer that is designed to increase the flow of information by progressively multiplexing channel and spatial information in the network, while mitigating computational complexity. Furthermore, to demonstrate the effectiveness of MUXConv, we integrate it within an efficient multi-objective evolutionary algorithm to search for the optimal model hyper-parameters while simultaneously optimizing accuracy, compactness, and computational efficiency. On ImageNet, the resulting models, dubbed MUXNets, match the performance (75.3% top-1 accuracy) and multiply-add operations (218M) of MobileNetV3 while being 1.6x more compact, and outperform other mobile models in all the three criteria. MUXNet also performs well under transfer learning and when adapted to object detection. On the ChestX-Ray 14 benchmark, its accuracy is comparable to the state-of-the-art while being 3.3x more compact and 14x more efficient. Similarly, detection on PASCAL VOC 2007 is 1.2% more accurate, 28% faster and 6% more compact compared to MobileNetV2. Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh Boddeti |
CVPR | 1 |
| 2020 | RetinaTrack: Online Single Stage Joint Detection and TrackingabstractTraditionally multi-object tracking and object detection are performed using separate systems with most prior works focusing exclusively on one of these aspects over the other. Tracking systems clearly benefit from having access to accurate detections, however and there is ample evidence in literature that detectors can benefit from tracking which, for example, can help to smooth predictions over time. In this paper we focus on the tracking-by-detection paradigm for autonomous driving where both tasks are mission critical. We propose a conceptually simple and efficient joint model of detection and tracking, called RetinaTrack, which modifies the popular single stage RetinaNet approach such that it is amenable to instance-level embedding training. We show, via evaluations on the Waymo Open Dataset, that we outperform a recent state of the art tracking algorithm while requiring significantly less computation. We believe that our simple yet effective approach can serve as a strong baseline for future work in this area. Zhichao Lu, Vivek Rathod, Ronny Votel, Jonathan Huang |
CVPR | 1 |
| 2020 | DOPS: Learning to Detect 3D Objects and Predict Their 3D ShapesabstractWe propose DOPS, a fast single-stage 3D object detection method for LIDAR data. Previous methods often make domain-specific design decisions, for example projecting points into a bird-eye view image in autonomous driving scenarios. In contrast, we propose a general-purpose method that works on both indoor and outdoor scenes. The core novelty of our method is a fast, single-pass architecture that both detects objects in 3D and estimates their shapes. 3D bounding box parameters are estimated in one pass for every point, aggregated through graph convolutions, and fed into a branch of the network that predicts latent codes representing the shape of each detected object. The latent shape space and shape decoder are learned on a synthetic dataset and then used as supervision for the end-to-end training of the 3D object detection pipeline. Thus our model is able to extract shapes without access to ground-truth shape information in the target dataset. During experiments, we find that our proposed method achieves state-of-the-art results by~5% on object detection in ScanNet scenes, and it gets top results by 3.4% in the Waymo Open Dataset, while reproducing the shapes of detected cars. Mahyar Najibi, Guangda Lai, Abhijit Kundu, Zhichao Lu, Vivek Rathod, Thomas A. Funkhouser, Caroline Pantofaru, David A. Ross, Larry Davis 0001, Alireza Fathi |
CVPR | 4 |
| 2020 | NSGANetV2: Evolutionary Multi-objective Surrogate-Assisted Neural Architecture Search
Zhichao Lu, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
ECCV (1) | 1 |
| 2020 | NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm (Extended Abstract)abstractConvolutional neural networks (CNNs) are the backbones of deep learning paradigms for numerous vision tasks. Early advancements in CNN architectures are primarily driven by human expertise and elaborate design. Recently, neural architecture search (NAS) was proposed with the aim of automating the network design process and generating task-dependent architectures. This paper introduces NSGA-Net -- an evolutionary search algorithm that explores a space of potential neural network architectures in three steps, namely, a population initialization step that is based on prior-knowledge from hand-crafted architectures, an exploration step comprising crossover and mutation of architectures, and finally an exploitation step that utilizes the hidden useful knowledge stored in the entire history of evaluated neural architectures in the form of a Bayesian Network. The integration of these components allows an efficient design of architectures that are competitive and in many cases outperform both manually and automatically designed architectures on CIFAR-10 classification task. The flexibility provided from simultaneously obtaining multiple architecture choices for different compute requirements further differentiates our approach from other methods in the literature. Zhichao Lu, Ian Whalen, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
IJCAI | 1 |
| 2019 | NSGA-Net: neural architecture search using multi-objective genetic algorithmabstractThis paper introduces NSGA-Net --- an evolutionary approach for neural architecture search (NAS). NSGA-Net is designed with three goals in mind: (1) a procedure considering multiple and conflicting objectives, (2) an efficient procedure balancing exploration and exploitation of the space of potential neural network architectures, and (3) a procedure finding a diverse set of trade-off network architectures achieved in a single run. NSGA-Net is a population-based search algorithm that explores a space of potential neural network architectures in three steps, namely, a population initialization step that is based on prior-knowledge from hand-crafted architectures, an exploration step comprising crossover and mutation of architectures, and finally an exploitation step that utilizes the hidden useful knowledge stored in the entire history of evaluated neural architectures in the form of a Bayesian Network. Experimental results suggest that combining the dual objectives of minimizing an error metric and computational complexity, as measured by FLOPs, allows NSGA-Net to find competitive neural architectures. Moreover, NSGA-Net achieves error rate on the CIFAR-10 dataset on par with other state-of-the-art NAS methods while using orders of magnitude less computational resources. These results are encouraging and shows the promise to further use of EC methods in various deep-learning paradigms. Zhichao Lu, Ian Whalen, Vishnu Naresh Boddeti, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf |
GECCO | 1 |
| 2018 | Balancing Survival of Feasible and Infeasible Solutions in Constraint Evolutionary Optimization AlgorithmsabstractReal-world optimization problems often involve constraints that relate to viability of implementing a solution. To solve such problems efficiently, a good constraint handling method is indispensable for an optimization algorithm. Population-based optimization algorithms allow a flexible way to handle constraints by making a careful comparison between feasible and infeasible solutions present in the population. A previous approach, which emphasized feasible solutions infinitely more than the infeasible solutions, has been popularly applied for more than one-and-half decade, mostly with real-parameter genetic algorithms (RGAs). Despite its popular use, the idea was criticized for its extreme selection pressure against infeasible solutions. Since optimal solutions often lie on the constraint boundaries, survival of certain infeasible solutions close to critical constraint boundaries should help RGA's recombination and mutation operators to produce near-optimal solutions. In this paper, we extend the earlier parameter-less constraint handling approach so as to strike a balance between survival of feasible and infeasible solutions in a GA population. The balance is controlled through an additional parameter that could be pre-specified or adaptively updated as the algorithm progresses. A parametric study is conducted to determine an appropriate value which works the best on most problems of this study. A significant improvement in performance is observed for the commonly-used g-series test problem suite and a real-world application problem (welded beam design). The approach is generic and can be easily extended to other real-parameter evolutionary algorithms, multi-objective and other advanced optimization tasks. Zhichao Lu, Kalyanmoy Deb, Hemant K. Singh |
CEC | 1 |
| 2018 | Uncertainty Handling in Bilevel Optimization for Robust and Reliable SolutionsabstractUncertainties in variables and parameters cause optimization problems to move away from globally-optimal and uncertain solutions. Practitioners resort to finding robust and reliable solutions in such situations. Bilevel optimization problems involving a hierarchy of two nested optimization problems have received a growing attention in the recent past due to their relevance in practice. While a number of studies on bilevel solution methodologies and applications are available for a deterministic setup, but studies on uncertainties in bilevel optimization are rare. In this paper, we suggest methodologies for handling uncertainty in both lower and upper level variables that may occur from different practicalities. For the first time, we perform a systematic study demonstrating the effect of uncertainties in each level along with the definition of robustness and reliability in the context of bilevel optimization. The issues and complexities introduced due to such uncertainties are then studied through a number of test cases, for brevity, we only show results on three test cases. Finally, two real-world bilevel problems involving uncertainties in their variables are solved. The study provides foundations and demon- strates viable directions for further research in uncertainty-based bilevel optimization problems. Zhichao Lu, Kalyanmoy Deb, Ankur Sinha 0001 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2017 | Handling practicalities in agricultural policy optimization for water quality improvementsabstractBilevel and multi-objective optimization methods are often useful to spatially target agri-environmental policy throughout a watershed. This type of problem is complex and is comprised of a number of practicalities: (i) a large number of decision variables, (ii) at least two inter-dependent levels of optimization between policy makers and policy followers, and (iii) uncertainty in decision variables and problem parameters. Given agricultural and economic data from the Raccoon watershed in central Iowa, we formulate a bilevel multi-objective optimization problem that accommodates objectives of both policy makers and farmers. The solution procedure then explicitly accounts for the nested nature of farm-level management decisions in response to agri-environmental policy incentives constructed by policy makers. We specifically examine the spatial targeting of a fertilizer-reduction incentive policy while seeking to maximize farm-level productivity while generating mandated water quality improvements using this framework. We test three different evolutionary optimization algorithms - m-BLEAQ, NSGA-II, and SPEA2 - and show that m-BLEAQ is well suited for handling the bilevel optimization problems and the considered practicalities. Bradley L. Barnhart, Zhichao Lu, Moriah Bostian, Ankur Sinha 0001, Kalyanmoy Deb, Luba Kurkalova, Manoj Jha, Gerald Whittaker |
GECCO | 2 |
| 2017 | Solving a supply-chain management problem using a bilevel approachabstractSupply-chain management problems are common to most industries and they involve a hierarchy of subtasks, which must be coordinated well to arrive at an overall optimal solution. Such problems involve a hierarchy of decision-makers, each having its own objectives and constraints, but importantly requiring a coordination of their actions to make the overall supply chain process optimal from cost and quality considerations. In this paper, we consider a specific supply-chain management problem from a company, which involves two levels of coordination: (i) yearly strategic planning in which a decision on establishing an association of every destination point with a supply point must be made so as to minimize the yearly transportation cost, and (ii) weekly operational planning in which, given the association between a supply and a destination point, a decision on the preference of available transport carriers must be made for multiple objectives: minimization of transport cost and maximization of service quality and satisfaction of demand at each destination point. We propose a customized multi-objective bilevel evolutionary algorithm, which is computationally tractable. We then present results on state-level and ZIP-level accuracy (involving about 40,000 upper level variables) of destination points over the mainland USA. We compare our proposed method with current non-optimization based practices and report a considerable cost saving. Zhichao Lu, Kalyanmoy Deb, Erik D. Goodman, John M. Wassick |
GECCO | 1 |
| 2017 | Resistive Random Access Memory for Future Information Processing SystemabstractResistive random access memory (RRAM) is regarded as one of the most promising emerging memory technologies for next-generation embedded, standalone nonvolatile memory (NVM), and storage class memory (SCM) due to its speed, density, cost, and scalability. Considerable progress has been made in recent years on the manufacturability of RRAM, with low-density RRAM products now in production and the path to higher density parts becoming clearer. This review updates the learning on the fundamental materials and process integration needed for high-volume manufacturing and summarizes very recent progress on array level performance improvement methodology using novel techniques, and circuit level contributions for different applications. The device performance, array integration, and device/circuit codesign for memory systems are discussed. Novel applications besides embedded memory and standalone memory are addressed, including hardware security, neuromorphic computing, and nonvolatile logic systems. Huaqiang Wu, Xiao Hu Wang, Bin Gao 0006, Ning Deng 0008, Zhichao Lu, Brent Haukness, Gary Bronner, He Qian |
Proc. IEEE | 5 |
| 2016 | Finding Reliable Solutions in Bilevel Optimization Problems Under UncertaintiesabstractBilevel optimization problems are referred to as having a nested inner optimization problem as a constraint to a outer optimization problem in the domain of mathematical programming. It is also known as Stackelberg problems in game theory. In the recent past, bilevel optimization problems have received a growing attention because of its relevance in practice applications. However, the hierarchical structure makes these problems difficult to handle and they are commonly optimized with a deterministic setup. With presence of constrains, bilevel optimization problems are considered for finding reliable solutions which are subjected to a possess a minimum reliability requirement under decision variable uncertainties. Definition of reliable bilevel solution, the effect of lower and upper level uncertainties on reliable bilevel solution, development of efficient reliable bilevel evolutionary algorithm, and supporting simulation results on test and engineering design problems amply demonstrate their further use in other practical bilevel problems. Zhichao Lu, Kalyanmoy Deb, Ankur Sinha 0001 |
GECCO | 1 |
| 2015 | Continuous Symmetric Stereo with Adaptive Outlier HandlingabstractWe present a method for symmetric stereo matching in which outliers from occlusions, texture-less regions, and repeated patterns are handled in a soft and adaptive manner. Rather than making binary outlier decisions, our model incorporates continuous-valued confidence weights that account for outlier likelihood, to promote robustness in disparity estimation. In contrast to previous outlier labeling techniques that fix the labels at the start of optimization, our method iteratively updates our outlier confidence weights as the matching results are gradually refined. By doing this, errors in an initial labeling can be rectified in the matching process. Our model is optimized in an Expectation-Maximization framework that efficiently produces continuous disparity estimates. This approach provides a good combination of accuracy and speed. Experiments show that our method compares favorably to prior outlier labeling techniques on the Middlebury benchmark, and that it can generate high-quality reconstruction for outdoor images with much more complex occlusions. Chen Li 0031, Lap-Fai Yu, Zhichao Lu, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
3DV | 3 |
| 2015 | Towards optimal ship design and valuable knowledge discovery under uncertain conditionsabstractShip design is a complex engineering activity which requires a multidisciplinary consideration in arriving at design objectives and constraints. An optimal design of such problems require a multi-objective optimization method that is capable of finding multiple trade-off solutions, not only to choose a preferred solution for implementation, but also to have a deeper understanding of the interactions among design variables. In this paper, we consider two ship design models involving uncertainties in design variables, and demonstrate the usefulness of an evolutionary multiobjective optimization (EMO) method and subsequent data analysis procedures in arriving at valuable design principles that enhance the knowledge of a designer. The study is pedagogical yet provide key insights of ship design issues and importantly outlines the systematic procedure for applying the technology to other more complex design problems. Kalyanmoy Deb, Zhichao Lu, Chris B. McKesson, Cherie Courseault Trumbach, Larry DeCan |
CEC | 2 |
| 2015 | Handling decision variable uncertainty in bilevel optimization problemsabstractBilevel optimization problems have received a growing attention in the recent past. In this paper, we suggest methodologies for handling uncertainty in both lower and upper level decision variables that may occur from different practicalities. For the first time, we discuss and demonstrate the effect of uncertainties in each level on the overall definition of a robust bilevel solution and present simulation results on a number of test problems. Finally, the robust solutions of a bilevel circuit design problem are found using a previously suggested fast bilevel evolutionary algorithm (BLEAQ). Definition of robust bilevel solutions, effect of lower and upper level uncertainties in robust bilevel solutions, development of a robust bilevel evolutionary algorithm and simulation results on test and engineering design problems are contributions of this study. Zhichao Lu, Kalyanmoy Deb, Ankur Sinha 0001 |
CEC | 1 |
| 2014 | Similarity-Aware Patchwork Assembly for Depth Image Super-resolutionabstractThis paper describes a patchwork assembly algorithm for depth image super-resolution. An input low resolution depth image is disassembled into parts by matching similar regions on a set of high resolution training images, and a super-resolution image is then assembled using these corresponding matched counterparts. We convert the super resolution problem into a Markov Random Field (MRF) labeling problem, and propose a unified formulation embedding (1) the consistency between the resolution enhanced image and the original input, (2) the similarity of disassembled parts with the corresponding regions on training images, (3) the depth smoothness in local neighborhoods, (4) the additional geometric constraints from self-similar structures in the scene, and (5) the boundary coincidence between the resolution enhanced depth image and an optional aligned high resolution intensity image. Experimental results on both synthetic and real-world data demonstrate that the proposed algorithm is capable of recovering high quality depth images with X4 resolution enhancement along each coordinate direction, and that it outperforms state-of-the-arts [14] in both qualitative and quantitative evaluations. Zhichao Lu, Rui Gan, Hongbin Zha |
CVPR | 2 |