Chengzhi Mao

dblp:189/4529 · DBLP profile ↗
← Back
30ranked-venue papers
11as first author
25since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Learning to Rewrite: Generalized LLM-Generated Text Detection
abstract
Detecting text generated by Large Language Models (LLMs) is crucial, yet current detectors often struggle to generalize in open-world settings.We introduce Learning2Rewrite, a novel framework to detect LLM-generated text with exceptional generalization to unseen domains.Capitalized on the finding that LLMs inherently modify LLM-generated content less than human-written text when rewriting, we train an LLM to amplify this disparity, yielding a more distinguishable and generalizable edit distance across diverse text distributions.Extensive experiments on data from 21 independent domains and four major LLMs (GPT-3.5,GPT-4, Gemini, and Llama-3) demonstrate that our detector outperforms state-of-the-art detection methods by up to 23.04% in AU-ROC for in-distribution tests, 35.10% for outof-distribution tests, and 48.66% under adversarial attacks.Our unique training objective ensures better generalizability compared to directly training for classification, even when leveraging the same amount of tunable parameters.Our findings suggest that reinforcing LLMs' inherent rewriting tendencies offers a robust and scalable solution for detecting LLMgenerated text.
Weiliang Zhao, Chengzhi Mao
ACL (1)5
2025 Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
abstract
Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about their fair usage. We investigate the mechanisms behind this discrepancy and find that alignment surprisingly amplifies implicit bias in model outputs. Specifically, we show that aligned LMs, unlike their unaligned counterparts, overlook racial concepts in early internal representations when the context is ambiguous. Not representing race likely fails to activate safety guardrails, leading to unintended biases. Inspired by this insight, we propose a new bias mitigation strategy that works by incentivizing the representation of racial concepts in the early model layers. In contrast to conventional mitigation methods of machine unlearning, our interventions find that steering the model to be more aware of racial concepts effectively mitigates implicit bias. Similar to race blindness in humans, ignoring racial nuances can inadvertently perpetuate subtle biases in LMs.
Lihao Sun, Chengzhi Mao, Valentin Hofmann, Xuechunzi Bai
ACL (1)2
2025 Nazar: Monitoring and Adapting ML Models on Mobile Devices
Lauren Hong, Nader Karayanni, AnMei Dasbach-Prisk, Chengzhi Mao, Asaf Cidon
ASPLOS (1)7
2025 I Can Hear You: Selective Robust Training for Deepfake Audio Detection
abstract
Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named DeepFakeVox-HQ, comprising 1.3 million samples, including 270,000 high-quality deepfake samples from 14 diverse sources. Despite previously reported high accuracy, existing deepfake voice detectors struggle with our diversely collected dataset, and their detection success rates drop even further under realistic corruptions and adversarial attacks. We conduct a holistic investigation into factors that enhance model robustness and show that incorporating a diversified set of voice augmentations is beneficial. Moreover, we find that the best detection models often rely on high-frequency features, which are imperceptible to humans and can be easily manipulated by an attacker. To address this, we propose the F-SAT: Frequency-Selective Adversarial Training method focusing on high-frequency components. Empirical results demonstrate that using our training dataset boosts baseline model performance (without robust training) by 33%, and our robust training further improves accuracy by 7.7% on clean samples and by 29.3% on corrupted and attacked samples, over the state-of-the-art RawNet3 model.
Aroon Sankoh, William Lin, Emanuel Mendiola-Ortiz, Chengzhi Mao
ICLR7
2025 EditLord: Learning Code Transformation Rules for Code Editing
abstract
Code editing is a foundational task in software development, where its effectiveness depends on whether it introduces desired code property changes without changing the original code's intended functionality. Existing approaches often formulate code editing as an implicit end-to-end task, omitting the fact that code-editing procedures inherently consist of discrete and explicit steps, and thus suffer from suboptimal performance and lack of robustness and generalization. We introduce EditLord, a code editing framework that makes the code transformation steps explicit. Our key insight is to employ a language model (LM) as an inductive learner to extract code editing rules from the training code pairs as concise meta-rule sets. Such rule sets will be manifested for each training sample to augment them for finetuning or assist in prompting- and iterative-based code editing. EditLord outperforms the state-of-the-art by an average of 22.7% in editing performance and 58.1% in robustness while achieving 20.2% higher functional correctness, across critical software engineering and security applications, LM models, and editing modes.
Weichen Li 0001, Albert Jan, Baishakhi Ray, Chengzhi Mao, Kexin Pei
ICML5
2025 Diversity Helps Jailbreak Large Language Models
abstract
Weiliang Zhao, Daniel Ben-Levi, Wei Hao, Junfeng Yang, Chengzhi Mao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Weiliang Zhao, Daniel Ben-Levi, Chengzhi Mao
NAACL (Long Papers)5
2025 LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
abstract
Efficient red-teaming method to uncover vulnerabilities in Large Language Models (LLMs) is crucial. While recent attacks often use LLMs as optimizers, the discrete language space make gradient-based methods struggle. We introduce LARGO (Latent Adversarial Reflection through Gradient Optimization), a novel latent self-reflection attack that reasserts the power of gradient-based optimization for generating fluent jailbreaking prompts. By operating within the LLM's continuous latent space, LARGO first optimizes an adversarial latent vector and then recursively call the same LLM to decode the latent into natural language. This methodology yields a fast, effective, and transferable attack that produces fluent and stealthy prompts. On standard benchmarks like AdvBench and JailbreakBench, LARGO surpasses leading jailbreaking techniques, including AutoDAN, by 44 points in attack success rate. Our findings demonstrate a potent alternative to agentic LLM prompting, highlighting the efficacy of interpreting and attacking LLM internals through gradient optimization.
Chengzhi Mao
NeurIPS3
2025 Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
abstract
Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suitable for tracking without task-specific training. This ability arises because their denoising process isolates motion in early, high-noise stages, distinct from later appearance refinement. Capitalizing on this discovery, our self-supervised tracker significantly improves performance in distinguishing visually similar objects, an underexplored failure point for existing methods. Our method achieves up to a 6-point improvement over recent self-supervised approaches on established benchmarks and our newly introduced tests focused on tracking visually similar items. Visualizations confirm that these diffusion-derived motion representations enable robust tracking of even identical objects across challenging viewpoint changes and deformations. Project page: \small{\url{https://chenshuang-zhang.github.io/projects/ted}}.
Chenshuang Zhang, Kang Zhang 0008, Joon Son Chung, In-So Kweon, Junmo Kim 0002, Chengzhi Mao
NeurIPS6
2024 ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
abstract
We establish rigorous benchmarks for visual perception robustness. Synthetic images such as ImageNet-C, ImageNet-9, and Stylized ImageNet provide specific type of evaluation over synthetic corruptions, backgrounds, and textures, yet those robustness benchmarks are restricted in specified variations and have low synthetic quality. In this work, we introduce generative model as a data source for synthesizing hard images that benchmark deep models' robustness. Leveraging diffusion models, we are able to generate images with more diversified backgrounds, textures, and materials than any prior work, where we term this benchmark as ImageNet-D. Experimental results show that ImageNet-D results in a significant accuracy drop to a range of vision models, from the standard ResNet visual classifier to the latest foundation models like CLIP and MiniGPT-4, significantly reducing their accuracy by up to 60%. Our work suggests that diffusion models can be an effective source to test vision models. The code and dataset are available at https://github.com/chenshuang-zhang/imagenet_d.
Chenshuang Zhang, Junmo Kim 0002, In-So Kweon, Chengzhi Mao
CVPR5
2024 RAFT: Realistic Attacks to Fool Text Detectors
abstract
Large language models (LLMs) have exhibited remarkable fluency across various tasks.However, their unethical applications, such as disseminating disinformation, have become a growing concern.Although recent works have proposed a number of LLM detection methods, their robustness and reliability remain unclear.In this paper, we present RAFT: a grammar error-free black-box attack against existing LLM detectors.In contrast to previous attacks for language models, our method exploits the transferability of LLM embeddings at the wordlevel while preserving the original text quality.We leverage an auxiliary embedding to greedily select candidate words to perturb against the target detector.Experiments reveal that our attack effectively compromises all detectors in the study across various domains by up to 99%, and are transferable across source models.Manual human evaluation studies show our attacks are realistic and indistinguishable from original human-written text.We also show that examples generated by RAFT can be used to train adversarially robust detectors.Our work shows that current LLM detectors are not adversarially robust, underscoring the urgent need for more resilient detection mechanisms. Original GPT3.5Maj Richard Scott, 40, is accused of driving at speeds of up to 95mph (153km/h) in bad weather before the fatal crash that claimed the lives of two young children.The incident occurred on the A34 motorway near Newbury, Berkshire, last Saturday.DetectGPT LLM Likelihood: 0.7100 Perplexity: 12.14 Red-Teaming (Shi et al.)Maj Richard Scott, Two score, is accused of driving at quickens of up to lightning tempo (swift pace) in bad Sunny before the fatal crash that claimed the lives of two young children.The circumstance betided on the path motorway near Newbury, Berkshire, last Saturday.DetectGPT LLM Likelihood: 0.0500 Perplexity: 137.72 Ours Maj Richard Scott, 40, is reproached of steering at speeds of up to 95mph (153km/h) under bad weather before the fatal crash that claimed the deaths of two young minors.The mishap occurred near the A34 motorway near Newbury, UK, previous Saturday.
James Wang, Chengzhi Mao
EMNLP4
2024 INViTE: INterpret and Control Vision-Language Models with Text Explanations
abstract
Large-scale pre-trained vision foundation models, such as CLIP, have become de facto backbones for various vision tasks. However, due to their black-box nature, understanding the underlying rules behind these models’ predictions and controlling model behaviors have remained open challenges. We present INViTE: a framework for INterpreting Vision Transformer’s latent tokens with Text Explanations. Given a latent token, INViTE retains its semantic information to the final layer using transformer’s local operations and retrieves the closest text for explanation. INViTE enables understanding of model visual reasoning procedure without needing additional model training or data collection. Based on the obtained interpretations, INViTE allows for model editing that controls model reasoning behaviors and improves model robustness against biases and spurious correlations. Our code is available at https://github.com/tonychenxyz/vit-interpret.
Carl Vondrick, Chengzhi Mao
ICLR4
2024 Raidar: geneRative AI Detection viA Rewriting
abstract
We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer modifications. We introduce a method to detect AI-generated content by prompting LLMs to rewrite text and calculating the editing distance of the output. We dubbed our geneRative AI Detection viA Rewriting method Raidar. Raidar significantly improves the F1 detection scores of existing AI content detection models -- both academic and commercial -- across various domains, including News, creative writing, student essays, code, Yelp reviews, and arXiv papers, with gains of up to 29 points. Operating solely on word symbols without high-dimensional features, our method is compatible with black box LLMs, and is inherently robust on new content. Our results illustrate the unique imprint of machine-generated text through the lens of the machines themselves.
Chengzhi Mao, Carl Vondrick, Hao Wang 0014
ICLR1
2024 SelfIE: Self-Interpretation of Large Language Model Embeddings
abstract
How do large language models (LLMs) obtain their answers? The ability to explain and control an LLM’s reasoning process is key for reliability, transparency, and future model developments. We propose SelfIE (Self-Interpretation of Embeddings), a framework that enables LLMs to interpret their own embeddings in natural language by leveraging their ability to respond to inquiries about a given passage. Capable of interpreting open-world concepts in the hidden embeddings, SelfIE reveals LLM internal reasoning in cases such as making ethical decisions, internalizing prompt injection, and recalling harmful knowledge. SelfIE’s text descriptions on hidden embeddings open avenues to control LLM reasoning. We propose Supervised Control, which allows editing open-ended concepts while only requiring gradient computation of individual layer. We extend RLHF to hidden embeddings and propose Reinforcement Control that erases harmful knowledge in LLM without supervision targets.
Carl Vondrick, Chengzhi Mao
ICML3
2024 Towards Causal Deep Learning for Vulnerability Detection
abstract
Deep learning vulnerability detection has shown promising results in recent years. However, an important challenge that still blocks it from being very useful in practice is that the model is not robust under perturbation and it cannot generalize well over the out-of-distribution (OOD) data, e.g., applying a trained model to unseen projects in real world. We hypothesize that this is because the model learned non-robust features, e.g., variable names, that have spurious correlations with labels. When the perturbed and OOD datasets no longer have the same spurious features, the model prediction fails. To address the challenge, in this paper, we introduced causality into deep learning vulnerability detection. Our approach CausalVul consists of two phases. First, we designed novel perturbations to discover spurious features that the model may use to make predictions. Second, we applied the causal learning algorithms, specifically, do-calculus, on top of existing deep learning models to systematically remove the use of spurious features and thus promote causal based prediction. Our results show that CausalVul consistently improved the model accuracy, robustness and OOD performance for all the state-of-the-art models and datasets we experimented. To the best of our knowledge, this is the first work that introduces do calculus based causal learning to software engineering models and shows it's indeed useful for improving the model accuracy, robustness and generalization. Our replication package is located at https://figshare.com/s/0ffda320dcb96c249ef2.
Ira Ceka, Chengzhi Mao, Saikat Chakraborty 0001, Baishakhi Ray, Wei Le
ICSE3
2023 What You Can Reconstruct from a Shadow
abstract
3D reconstruction is a fundamental problem in computer vision, and the task is especially challenging when the object to reconstruct is partially or fully occluded. We introduce a method that uses the shadows cast by an unobserved object in order to infer the possible 3D volumes under occlusion. We create a differentiable image formation model that allows us to jointly infer the 3D shape of an object, its pose, and the position of a light source. Since the approach is end-to-end differentiable, we are able to integrate learned priors of object geometry in order to generate realistic 3D shapes of different object categories. Experiments and visualizations show that the method is able to generate multiple possible solutions that are consistent with the observation of the shadow. Our approach works even when the position of the light source and object pose are both unknown. Our approach is also robust to real-world images where ground-truth shadow mask is unknown.
Ruoshi Liu, Sachit Menon, Chengzhi Mao, Dennis Park, Simon Stent, Carl Vondrick
CVPR3
2023 Doubly Right Object Recognition: A Why Prompt for Visual Rationales
abstract
Many visual recognition models are evaluated only on their classification accuracy, a metric for which they obtain strong performance. In this paper, we investigate whether computer vision models can also provide correct rationales for their predictions. We propose a “doubly right” object recognition benchmark, where the metric requires the model to simultaneously produce both the right labels as well as the right rationales. We find that state-of-the-art visual models, such as CLIP, often provide incorrect rationales for their categorical predictions. However, by transferring the rationales from language models into visual representations through a tailored dataset, we show that we can learn a “why prompt,” which adapts large visual representations to produce correct rationales. Visualizations and empirical experiments show that our prompts significantly improve performance on doubly right object recognition, in addition to zero-shot transfer to unseen tasks and datasets.
Chengzhi Mao, Revant Teotia, Amrutha Sundar, Sachit Menon, Xin Wang 0066, Carl Vondrick
CVPR1
2023 Landscape Learning for Neural Network Inversion
abstract
Many machine learning methods operate by inverting a neural network at inference time, which has become a popular technique for solving inverse problems in computer vision, robotics, and graphics. However, these methods often involve gradient descent through a highly non-convex loss landscape, causing the optimization process to be unstable and slow. We introduce a method that learns a loss landscape where gradient descent is efficient, bringing massive improvement and acceleration to the inversion process. We demonstrate this advantage on a number of methods for both generative and discriminative tasks, including GAN inversion, adversarial defense, and 3D human pose reconstruction.
Ruoshi Liu, Chengzhi Mao, Purva Tendulkar, Hao Wang 0014, Carl Vondrick
ICCV2
2023 Understanding Zero-shot Adversarial Robustness for Large-Scale Models
Chengzhi Mao, Scott Geng, Xin Wang 0066, Carl Vondrick
ICLR1
2023 Robust Perception through Equivariance
abstract
Deep networks for computer vision are not reliable when they encounter adversarial examples. In this paper, we introduce a framework that uses the dense intrinsic constraints in natural images to robustify inference. By introducing constraints at inference time, we can shift the burden of robustness from training to testing, thereby allowing the model to dynamically adjust to each individual image's unique and potentially novel characteristics at inference time. Our theoretical results show the importance of having dense constraints at inference time. In contrast to existing single-constraint methods, we propose to use equivariance, which naturally allows dense constraints at a fine-grained level in the feature space. Our empirical experiments show that restoring feature equivariance at inference time defends against worst-case adversarial perturbations. The method obtains improved adversarial robustness on four datasets (ImageNet, Cityscapes, PASCAL VOC, and MS-COCO) on image recognition, semantic segmentation, and instance segmentation tasks.
Chengzhi Mao, Abhishek Vaibhav Joshi, Hao Wang 0014, Carl Vondrick
ICML1
2023 Convolutional Visual Prompt for Robust Visual Perception
abstract
Vision models are often vulnerable to out-of-distribution (OOD) samples without adapting. While visual prompts offer a lightweight method of input-space adaptation for large-scale vision models, they rely on a high-dimensional additive vector and labeled data. This leads to overfitting when adapting models in a self-supervised test-time setting without labels. We introduce convolutional visual prompts (CVP) for label-free test-time adaptation for robust visual perception. The structured nature of CVP demands fewer trainable parameters, less than 1\% compared to standard visual prompts, combating overfitting. Extensive experiments and analysis on a wide variety of OOD visual perception tasks show that our approach is effective, improving robustness by up to 5.87\% over several large-scale models.
Yun-Yun Tsai, Chengzhi Mao
NeurIPS2
2022 Causal Transportability for Visual Recognition
abstract
Visual representations underlie object recognition tasks, but they often contain both robust and non-robust features. Our main observation is that image classifiers may perform poorly on out-of-distribution samples because spurious correlations between non-robust features and labels can be changed in a new environment. By analyzing procedures for out-of-distribution generalization with a causal graph, we show that standard classifiers fail because the association between images and labels is not transportable across settings. However, we then show that the causal effect, which severs all sources of confounding, remains invariant across domains. This motivates us to develop an algorithm to estimate the causal effect for image classification, which is transportable (i.e., invariant) across source and target environments. Without observing additional variables, we show that we can derive an estimand for the causal effect under empirical assumptions using representations in deep models as proxies. Theoretical analysis, empirical results, and visualizations show that our approach captures causal invariances and improves overall generalization.
Chengzhi Mao, Kevin Xia 0001, James Wang, Hao Wang 0014, Elias Bareinboim, Carl Vondrick
CVPR1
2022 Real-Time Neural Voice Camouflage
Mia Chiquier, Chengzhi Mao, Carl Vondrick
ICLR2
2022 Discrete Representations Strengthen Vision Transformer Robustness
Chengzhi Mao, Lu Jiang 0004, Mostafa Dehghani 0001, Carl Vondrick, Rahul Sukthankar, Irfan A. Essa
ICLR1
2021 Generative Interventions for Causal Learning
abstract
We introduce a framework for learning robust visual representations that generalize to new viewpoints, backgrounds, and scene contexts. Discriminative models often learn naturally occurring spurious correlations, which cause them to fail on images outside of the training distribution. In this paper, we show that we can steer generative models to manufacture interventions on features caused by confounding factors. Experiments, visualizations, and theoretical results show this method learns robust representations more consistent with the underlying causal relationships. Our approach improves performance on multiple datasets demanding out-of-distribution generalization, and we demonstrate state-of-the-art performance generalizing from ImageNet to ObjectNet dataset.
Chengzhi Mao, Augustine Cha, Amogh Gupta, Hao Wang 0014, Carl Vondrick
CVPR1
2021 Adversarial Attacks are Reversible with Natural Supervision
abstract
We find that images contain intrinsic structure that enables the reversal of many adversarial attacks. Attack vectors cause not only image classifiers to fail, but also collaterally disrupt incidental structure in the image. We demonstrate that modifying the attacked image to restore the natural structure will reverse many types of attacks, providing a defense. Experiments demonstrate significantly improved robustness for several state-of-the-art models across the CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets. Our results show that our defense is still effective even if the attacker is aware of the defense mechanism. Since our defense is deployed during inference instead of training, it is compatible with pre-trained networks as well as most other defenses. Our results suggest deep networks are vulnerable to adversarial examples partly because their representations do not enforce the natural structure of images.
Chengzhi Mao, Mia Chiquier, Hao Wang 0014, Carl Vondrick
ICCV1
2020 Multitask Learning Strengthens Adversarial Robustness
Chengzhi Mao, Amogh Gupta, Vikram Nitin, Baishakhi Ray, Shuran Song, Carl Vondrick
ECCV (2)1
2019 Bidirectional Inference Networks: A Class of Deep Bayesian Networks for Health Profiling
abstract
We consider the problem of inferring the values of an arbitrary set of variables (e.g., risk of diseases) given other observed variables (e.g., symptoms and diagnosed diseases) and high-dimensional signals (e.g., MRI images or EEG). This is a common problem in healthcare since variables of interest often differ for different patients. Existing methods including Bayesian networks and structured prediction either do not incorporate high-dimensional signals or fail to model conditional dependencies among variables. To address these issues, we propose bidirectional inference networks (BIN), which stich together multiple probabilistic neural networks, each modeling a conditional dependency. Predictions are then made via iteratively updating variables using backpropagation (BP) to maximize corresponding posterior probability. Furthermore, we extend BIN to composite BIN (CBIN), which involves the iterative prediction process in the training stage and improves both accuracy and computational efficiency by adaptively smoothing the optimization landscape. Experiments on synthetic and real-world datasets (a sleep study and a dermatology dataset) show that CBIN is a single model that can achieve state-of-the-art performance and obtain better accuracy in most inference tasks than multiple models each specifically trained for a different task.
Hao Wang 0014, Chengzhi Mao, Hao He 0011, Mingmin Zhao, Tommi S. Jaakkola, Dina Katabi
AAAI2
2019 Metric Learning for Adversarial Robustness
abstract
Deep networks are well-known to be fragile to adversarial attacks. We conduct an empirical analysis of deep representations under the state-of-the-art attack method called PGD, and find that the attack causes the internal representation to shift closer to the ``false'' class. Motivated by this observation, we propose to regularize the representation space under attack with metric learning to produce more robust classifiers. By carefully sampling examples for metric learning, our learned representation not only increases robustness, but also detects previously unseen adversarial samples. Quantitative experiments show improvement of robustness accuracy by up to 4% and detection efficiency by up to 6% according to Area Under Curve score over prior work. The code of our work is available at https://github.com/columbia/MetricLearningAdversarial_Robustness.
Chengzhi Mao, Ziyuan Zhong, Carl Vondrick, Baishakhi Ray
NeurIPS1
2018 A Probabilistic Learning Approach to UWB Ranging Error Mitigation
abstract
Ultra-Wide Band (UWB) radio is capable of providing sufficient information for high accuracy localization. However, its actual performance is degraded due to the non-line-of-sight (NLOS) propagation. This paper introduces a probabilistic learning approach to mitigate the ranging error and yield uncertainties which correlate with the mitigation results. By combining variational inference with probabilistic neural networks, we propose a new probabilistic deep learning architecture, which can improve the accuracy significantly especially when the training data is limited. Results show that the proposed model can reduce the root mean square error (RMSE) of UWB ranging by 16%~56% compared with existing support vector machine approach in practical environment.
Chengzhi Mao, Kangbo Lin, Tiancheng Yu, Yuan Shen 0001
GLOBECOM1
2016 Deep self-organizing reservoir computing model for visual object recognition
abstract
Reservoir computing becomes increasingly a hot spot in recent years. In this paper, we propose a deep self-organizing reservoir computing model for visual object recognition. First, through combination of Kohonen's self-organizing map and SHESN network, we present a self-organizing SHESN (SO-SHESN). In the new model, we adopt the same mechanism of generating reservoir as SHESN, but McCulloch-Pitts type reservoir neuron is replaced with radial basis function neuron. Correspondingly, unsupervised competitive learning is exploited to train both input weights and reservoir weights of SO-SHESN. Second, we propose a deep SO-SHESN model through a stack of well-trained reservoir layers. In such a stacked structure, a novel trial-and-readout learning algorithm is used for pre-training of layer-wise reservoir, in which each layer is trained independently from each other. Finally, the experimental results obtained on MNIST benchmark dataset show that our SO-SHESN achieves the test recognition error rate of 5.66%, which improves classical ESN and SHESN by 6.44% and 1.74%, respectively. Furthermore, the test error rate of our deep SO-SHESN could reach up to 1.39%, which outperforms SO-SHESN with single reservoir layer by 4.27% and approximately approaches the state-of-the-art result of 1% among existing traditional machine learning approaches with non-CNN features.
Zhidong Deng, Chengzhi Mao
IJCNN2