Sang (Peter) Chin

dblp:133/4494 · also Peter Chin 0001, Sang Chin 0001, Sang Peter Chin · DBLP profile ↗
← Back
46ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-1913-4223ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 AsymPuzl: A minimal puzzle testbed for LLM-based two agent communication with information asymmetry
abstract
Large Language Model (LLM) agents are increasingly studied in multi-turn, multi-agent scenarios, yet most existing setups emphasize open-ended roleplay rather than controlled evaluation.We introduce AsymPuzl, a minimal but expressive two-agent puzzle environment isolating communication under information asymmetry.Each agent observes complementary but incomplete views of a puzzle and must exchange messages to solve it.Using contemporary LLMs, we show that (i) models such as GPT-5 and Claude-4.0reliably solve puzzles of different sizes by sharing complete information in few turns, (ii) feedback design in multi-agent LLM systems is non-trivial, more information can degrade performance.
Xavier Cadet, Edward Koh, Sang (Peter) Chin
ESANN3
2026 FORMULA: FORmation MPC with neUral barrier Learning for safety Assurance
Qintong Xie, Weishu Zhan, Sang (Peter) Chin
IV3
2025 ZipNN: Lossless Compression for AI Models
abstract
With the growth of model sizes and the scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these. While there is a vast model compression literature deleting parts of the model weights for faster inference, we investigate a more traditional type of compression - one that represents the model in a compact form and is coupled with a decompression algorithm that returns it to its original form and size - namely lossless compression. We present ZipNN, a lossless compression tailored to neural networks. Somewhat surprisingly, we show that specific lossless compression can gain significant network and storage reduction on popular models, often saving 33% and at times reducing over 50% of the model size. We investigate the source of model compressibility and introduce specialized compression variants tailored for models that further increase the effectiveness of compression. On popular models (e.g. Llama 3) ZipNN shows space savings that are over 17% better than vanilla compression while also improving compression and decompression speeds by 62%. Using multiple workers and threads, ZipNN can achieve decompression speeds of up to 80GB/s and compression speed of up to 13GB/s. We estimate that these methods could save over an ExaByte per year of network traffic downloaded from a large model hub like Hugging Face.
Moshe Hershcovitch, Andrew Wood, Leshem Choshen, Guy Girmonsky, Roy Leibovitz, Or Ozeri, Ilias Ennmouri, Michal Malka, Sang (Peter) Chin, Swaminathan Sundararaman, Danny Harnik
CLOUD9
2025 Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
Aditya Vikram Singh 0002, Ethan Rathbun, Emma Graham, Lisa Oakley, Simona Boboila, Sang (Peter) Chin, Alina Oprea
AAMAS6
2024 Weisfeiler and Lehman Go Paths: Learning Topological Features via Path Complexes
abstract
Graph Neural Networks (GNNs), despite achieving remarkable performance across different tasks, are theoretically bounded by the 1-Weisfeiler-Lehman test, resulting in limitations in terms of graph expressivity. Even though prior works on topological higher-order GNNs overcome that boundary, these models often depend on assumptions about sub-structures of graphs. Specifically, topological GNNs leverage the prevalence of cliques, cycles, and rings to enhance the message-passing procedure. Our study presents a novel perspective by focusing on simple paths within graphs during the topological message-passing process, thus liberating the model from restrictive inductive biases. We prove that by lifting graphs to path complexes, our model can generalize the existing works on topology while inheriting several theoretical results on simplicial complexes and regular cell complexes. Without making prior assumptions about graph sub-structures, our method outperforms earlier works in other topological domains and achieves state-of-the-art results on various benchmarks.
Quang Truong, Sang (Peter) Chin
AAAI2
2024 Detecting Continuous Gravitational Waves Using Generated Training Data
abstract
Detecting continuous gravitational waves using machine learning approaches is an active research topic. With signal strengths between 0.1% and 2%, this classification task is very difficult. The presence of noise makes it impossible even for humans to distinguish between data with and without traces of continuous gravitational waves. The European Gravitational Observatory (EGO) formulated this problem as a Kaggle challenge. As participants, we present our approach in this paper. In particular, we focus on our innovative data generation solution, which provides great flexibility while maintaining efficiency and training accuracy. Our generated training data is fully compatible with state-of-the-art image classifiers.
Judith Herrmann, Raphael Kunert, Ron Hachmon, Aviv Markus, Allison Gunby-Mann, Sarel Cohen, Tobias Friedrich 0001, Sang (Peter) Chin
ICASSP8
2024 SocioDojo: Building Lifelong Analytical Agents with Real-world Text and Time Series
abstract
We introduce SocioDojo, an open-ended lifelong learning environment for developing ready-to-deploy autonomous agents capable of performing human-like analysis and decision-making on societal topics such as economics, finance, politics, and culture. It consists of (1) information sources from news, social media, reports, etc., (2) a knowledge base built from books, journals, and encyclopedias, plus a toolbox of Internet and knowledge graph search interfaces, (3) 30K high-quality time series in finance, economy, society, and polls, which support a novel task called "hyperportfolio", that can reliably and scalably evaluate societal analysis and decision-making power of agents, inspired by portfolio optimization with time series as assets to "invest". We also propose a novel Analyst-Assistant-Actuator architecture for the hyperportfolio task, and a Hypothesis & Proof prompting for producing in-depth analyses on input news, articles, etc. to assist decision-making. We perform experiments and ablation studies to explore the factors that impact performance. The results show that our proposed method achieves improvements of 32.4% and 30.4% compared to the state-of-the-art method in the two experimental settings.
Junyan Cheng, Sang (Peter) Chin
ICLR2
2024 Bridging Neural and Symbolic Representations with Transitional Dictionary Learning
abstract
This paper introduces a novel Transitional Dictionary Learning (TDL) framework that can implicitly learn symbolic knowledge, such as visual parts and relations, by reconstructing the input as a combination of parts with implicit relations. We propose a game-theoretic diffusion model to decompose the input into visual parts using the dictionaries learned by the Expectation Maximization (EM) algorithm, implemented as the online prototype clustering, based on the decomposition results. Additionally, two metrics, clustering information gain, and heuristic shape score are proposed to evaluate the model. Experiments are conducted on three abstract compositional visual object datasets, which require the model to utilize the compositionality of data instead of simply exploiting visual features. Then, three tasks on symbol grounding to predefined classes of parts and relations, as well as transfer learning to unseen classes, followed by a human evaluation, were carried out on these datasets. The results show that the proposed method discovers compositional patterns, which significantly outperforms the state-of-the-art unsupervised part segmentation methods that rely on visual features from pre-trained backbones. Furthermore, the proposed metrics are consistent with human evaluations.
Junyan Cheng, Sang (Peter) Chin
ICLR2
2024 Safe multi-agent reinforcement learning for bimanual dexterous manipulation
abstract
Bimanual dexterous manipulation in robotics, essential for a wide range of applications, addresses the critical challenge of balancing intricate operational capabilities with assured safety and reliability. While Safe Reinforcement Learning is integral to the robustness of robotic systems, the area of safe multi-agent reinforcement learning (MARL), cooperative control of multiple robots has been scarcely studied. In this study, we explore MARL for safe cooperative control with multiple robot hands. Each robot must follow individual and collective safety guidelines to ensure safe team actions. However, the non-stationarity inherent in current algorithms hinders the precise updating of strategies to satisfy these safety constraints effectively. In this paper, we propose Multi-Agent Constrained Proximal Advantage Optimization (MAC-PAO), which considers the sequence of agent updates and integrates non-stationarity into sequential update schemes. This algorithm ensures consistent improvement in both rewards and adherence to safety constraints in each iteration. We tested MACPAO on various tasks with safety constraints and demonstrated that it outperforms other MARL algorithms in balancing reward enhancement and safety compliance. Supplementary materials and code are available at the provided link https://github.com/YONEX4090/MultiSafeHand.git.
Weishu Zhan, Sang (Peter) Chin
IROS2
2024 TopoX: A Suite of Python Packages for Machine Learning on Topological Domains
abstract
We introduce TopoX, a Python software suite that provides reliable and user-friendly building blocks for computing and machine learning on topological domains that extend graphs: hypergraphs, simplicial, cellular, path and combinatorial complexes. TopoX consists of three packages: TopoNetX facilitates constructing and computing on these domains, including working with nodes, edges and higher-order cells; TopoEmbedX provides methods to embed topological domains into vector spaces, akin to popular graph-based embedding algorithms such as node2vec; TopoModelX is built on top of PyTorch and offers a comprehensive toolbox of higher-order message passing functions for neural networks on topological domains. The extensively documented and unit-tested source code of TopoX is available under MIT license at https://pyt-team.github.io.
Mustafa Hajij, Mathilde Papillon, Florian Frantzen, Jens Agerberg, Ibrahem AlJabea, Rubén Ballester, Claudio Battiloro, Guillermo Bernárdez, Tolga Birdal, Aiden Brent, Sang (Peter) Chin, Sergio Escalera, Simone Fiorellino, Odin Hoff Gardaa, Gurusankar Gopalakrishnan, Devendra Govil, Josef Hoppe, Maneel Reddy Karri, Jude Khouja, Manuel Lecha, Neal Livesay, Jan Meißner, Alexander Nikitin 0002, Theodore Papamarkou, Jaro Prílepok, Karthikeyan Natesan Ramamurthy, Paul Rosen 0001, Aldo Guzmán-Sáenz, Alessandro Salatiello, Shreyas N. Samaga, Simone Scardapane, Michael T. Schaub, Luca Scofano, Indro Spinelli, Lev Telyatnikov, Quang Truong, Robin Walters 0001, Maosheng Yang, Olga Zaghen, Ghada Zamzmi, Ali Zia, Nina Miolane
J. Mach. Learn. Res.11
2022 Training Robust Zero-Shot Voice Conversion Models with Self-Supervised Features
abstract
Unsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation has been shown to produce useful linguistic units without using transcripts, which can be directly passed to a VC model. In this paper, we showed that high-quality audio samples can be achieved by using a length resampling decoder, which enables the VC model to work in conjunction with different linguistic feature extractors and vocoders without requiring them to operate on the same sequence length. We showed that our method can outperform many baselines on the VCTK dataset. Without modifying the architecture, we further demonstrated that a) using pairs of different audio segments from the same speaker, b) adding a cycle consistency loss, and c) adding a speaker classification loss can help to learn a better speaker embedding. Our model trained on LibriTTS using these techniques achieves the best performance, producing audio samples transferred well to the target speaker’s voice, while preserving the linguistic content that is comparable with actual human utterances in terms of Character Error Rate.
Trung Dang 0002, Dung N. Tran, Sang (Peter) Chin, Kazuhito Koishida
ICASSP3
2022 A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter IT
abstract
End-to-end Automatic Speech Recognition (ASR) models are commonly trained over spoken utterances using optimization methods like Stochastic Gradient Descent (SGD). In distributed settings like Federated Learning, model training requires transmission of gradients over a network. In this work, we design the first method for revealing the identity of the speaker of a training utterance with access only to a gradient. We propose Hessian-Free Gradients Matching, an input reconstruction technique that operates without second derivatives of the loss function (required in prior works), which can be expensive to compute. We show the effectiveness of our method using the DeepSpeech model architecture, demonstrating that it is possible to reveal the speaker’s identity with 34% top-1 accuracy (51% top-5 accuracy) on the LibriSpeech dataset. Further, we study the effect of Dropout on the success of our method. We show that a dropout rate of 0.2 can reduce the speaker identity accuracy to 0% top-1 (0.5% top-5).
Trung Dang 0002, Om Thakkar 0001, Swaroop Ramaswamy, Rajiv Mathews, Sang (Peter) Chin, Françoise Beaufays
ICASSP5
2022 Substitutional Neural Image Compression
abstract
We describe Substitutional Neural Image Compression (SNIC), a general approach for enhancing any neural image compression model, that requires no data or additional tuning of the trained model. It boosts compression performance toward a flexible distortion metric and enables bit-rate control using a single model instance. The key idea is to replace the image to be compressed with a substitutional one that outperforms the original one in a desired way. Finding such a substitute is inherently difficult for conventional codecs, yet surprisingly favorable for neural compression models thanks to their fully differentiable structures. With gradients of a particular loss back-propogated to the input, a desired substitute can be efficiently crafted iteratively. We demonstrate the effectiveness of SNIC, when combined with various neural compression models and target metrics, in improving compression quality and performing bit-rate control measured by rate-distortion curves.
Xiao Wang 0028, Ding Ding 0004, Wei Jiang 0001, Wei Wang 0311, Xiaozhong Xu, Shan Liu 0001, Brian Kulis, Sang (Peter) Chin
PCS8
2021 Intrinsic Examples: Robust Fingerprinting of Deep Neural Networks
Siyue Wang, Pu Zhao 0001, Xiao Wang 0028, Sang (Peter) Chin, Thomas Wahl, Yunsi Fei, Qi Alfred Chen, Xue Lin 0001
BMVC4
2021 Training Many-to-Many Recurrent Neural Networks with Target Propagation
Peilun Dai, Sang (Peter) Chin
ICANN (4)2
2021 A Scale Invariant Measure of Flatness for Deep Network Minima
abstract
It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of flatness are not invariant to rescaling of the network parameters. This means that the measure of flatness can be made as small or as large as possible through rescaling, rendering the quantitative measures meaningless. In this paper we show that for deep networks with positively homogenous activations, these rescalings constitute equivalence relations, and that these equivalence relations induce a quotient manifold structure in the parameter space. Using an appropriate Riemannian metric, we propose a Hessian-based measure for flatness that is invariant to rescaling and perform simulations to empirically verify our claim. Finally we perform experiments to verify that our flatness measure correlates with generalization by using minibatch stochastic gradient descent with different batch sizes to find deep network minima with different generalization properties.
Akshay Rangamani, Nam H. Nguyen, Dzung T. Phan, Sang (Peter) Chin, Trac D. Tran
ICASSP5
2021 How Dense Autoencoders can still Achieve the State-of-the-art in Time-Series Anomaly Detection
abstract
Time series data has become ubiquitous in the modern era of data collection. With the increase of these time series data streams, the demand for automatic time series anomaly detection has also increased. Automatic monitoring of data allows engineers to investigate only unusual behavior in their data streams. Despite this increase in demand for automatic time series anomaly detection, many popular methods fail to offer a general purpose solution. Some demand expensive labelling of anomalies, others require the data to follow certain assumed patterns, some have long and unstable training, and many suffer from high rates of false alarms. In this paper we demonstrate that simpler is often better, showing that a fully unsupervised multilayer perceptron autoencoder is able to outperform much more complicated models with only a few critical improvements. We offer improvements to help distinguish anomalous subsequences near to each other, and to distinguish anomalies even in the midst of changing distributions of data. We compare our model with state-of-the-art competitors on benchmark datasets sourced from NASA, Yahoo, and Numenta, achieving improvements beyond competitive models in all three datasets.
Louis Jensen, Jayme Fosa, Ben Teitelbaum, Sang (Peter) Chin
ICMLA4
2021 Mixed Spatio-Temporal Neural Networks on Real-time Prediction of Crimes
abstract
Forecasting the crime rate in real-time is always an important task to public safety. However, there are no known models that provide satisfactory approximation to this complex spatio-temporal problem until recently. The crime rate may be affected by various factors, such as local education, public events, weather, etc. Such factors make the prediction of crimes more complex and challenging than other problems that are less influenced by outer factors. In this paper, we propose a deep-learning-based approach, which combines various methods in neural networks to handle the spatial temporal prediction problem. Some optimization techniques, such as Bayesian optimization, are applied for finding the optimal hyper-parameters as well as dealing with noises in the dataset. The model is trained on a dataset about crime information in Los Angeles at a scale of hours in block-divided areas, released by the LA Police Department (LAPD). The results of experiments on this dataset demonstrates the proposed model’s ability in predicting potential crimes in real time.
Xiao Zhou 0016, Xiao Wang 0028, Gavin Brown 0003, Chengchen Wang, Sang (Peter) Chin
ICMLA5
2021 PATCHCOMM: Using Commonsense Knowledge to Guide Syntactic Parsers
abstract
Syntactic parsing technologies have become significantly more robust thanks to advancements in their underlying statistical and Deep Neural Network (DNN) techniques: most modern syntactic parsers can produce a syntactic parse tree for almost any sentence, including ones that may not be strictly grammatical. Despite improved robustness, such parsers still do not reflect the alternatives in parsing that are intrinsic in syntactic ambiguities. Two most notable such ambiguities are prepositional phrase (PP) attachment ambiguities and pronoun coreference ambiguities. In this paper, we discuss PatchComm, which uses commonsense knowledge to help resolve both kinds of ambiguities. To the best of our knowledge, we are the first to propose the general-purpose approach of using external commonsense knowledge bases to guide syntactic parsers. We evaluated PatchComm against the state-of-the-art (SOTA) spaCy parser on a PP attachment task and against the SOTA NeuralCoref module on a coreference task. Results show that PatchComm is successful at detecting syntactic ambiguities and using commonsense knowledge to help resolve them.
Yida Xin, Henry Lieberman, Sang (Peter) Chin
KR3
2021 Revealing and Protecting Labels in Distributed Training
abstract
Distributed learning paradigms such as federated learning often involve transmission of model updates, or gradients, over a network, thereby avoiding transmission of private data. However, it is possible for sensitive information about the training data to be revealed from such gradients. Prior works have demonstrated that labels can be revealed analytically from the last layer of certain models (e.g., ResNet), or they can be reconstructed jointly with model inputs by using Gradients Matching [Zhu et al.] with additional knowledge about the current state of the model. In this work, we propose a method to discover the set of labels of training samples from only the gradient of the last layer and the id to label mapping. Our method is applicable to a wide variety of model architectures across multiple domains. We demonstrate the effectiveness of our method for model training in two domains - image classification, and automatic speech recognition. Furthermore, we show that existing reconstruction techniques improve their efficacy when used in conjunction with our method. Conversely, we demonstrate that gradient quantization and sparsification can significantly reduce the success of the attack.
Trung Dang 0002, Om Thakkar 0001, Swaroop Ramaswamy, Rajiv Mathews, Sang (Peter) Chin, Françoise Beaufays
NeurIPS5
2020 AdvMS: A Multi-Source Multi-Cost Defense Against Adversarial Attacks
abstract
Designing effective defense against adversarial attacks is a crucial topic as deep neural networks have been proliferated rapidly in many security-critical domains such as malware detection and self-driving cars. Conventional defense methods, although shown to be promising, are largely limited by their single-source single-cost nature: The robustness promotion tends to plateau when the defenses are made increasingly stronger while the cost tends to amplify. In this paper, we study principles of designing multi-source and multi-cost schemes where defense performance is boosted from multiple defending components. Based on this motivation, we propose a multi-source and multi-cost defense scheme, Adversarially Trained Model Switching (AdvMS), that inherits advantages from two leading schemes: adversarial training and random model switching. We show that the multi-source nature of AdvMS mitigates the performance plateauing issue and the multi-cost nature enables improving robustness at a flexible and adjustable combination of costs over different factors which can better suit specific restrictions and needs in practice.
Xiao Wang 0028, Siyue Wang, Xue Lin 0001, Sang (Peter) Chin
ICASSP5
2020 NodeDrop: A Method for Finding Sufficient Network Architecture Size
abstract
Determining an appropriate number of features for each layer in a neural network is an important and difficult task. This task is especially important in applications on systems with limited memory or processing power. Many current approaches to reduce network size either utilize iterative procedures, which can extend training time significantly, or require very careful tuning of algorithm parameters to achieve reasonable results. In this paper we propose NodeDrop, a new method for eliminating features in a network. With NodeDrop, we define a condition to identify and guarantee which nodes carry no information, and then use regularization to encourage nodes to meet this condition. We find that NodeDrop drastically reduces the number of features in a network while maintaining high performance. NodeDrop reduces the number of parameters by a factor of 114x for a VGG like network on CIFAR10 without a drop in accuracy.
Louis Jensen, Jacob Harer, Sang (Peter) Chin
IJCNN3
2019 Protecting Neural Networks with Hierarchical Random Switching: Towards Better Robustness-Accuracy Trade-off for Stochastic Defenses
abstract
Despite achieving remarkable success in various domains, recent studies have uncovered the vulnerability of deep neural networks to adversarial perturbations, creating concerns on model generalizability and new threats such as prediction-evasive misclassification or stealthy reprogramming. Among different defense proposals, stochastic network defenses such as random neuron activation pruning or random perturbation to layer inputs are shown to be promising for attack mitigation. However, one critical drawback of current defenses is that the robustness enhancement is at the cost of noticeable performance degradation on legitimate data, e.g., large drop in test accuracy.This paper is motivated by pursuing for a better trade-off between adversarial robustness and test accuracy for stochastic network defenses. We propose Defense Efficiency Score (DES), a comprehensive metric that measures the gain in unsuccessful attack attempts at the cost of drop in test accuracy of any defense. To achieve a better DES, we propose hierarchical random switching (HRS), which protects neural networks through a novel randomization scheme. A HRS-protected model contains several blocks of randomly switching channels to prevent adversaries from exploiting fixed model structures and parameters for their malicious purposes. Extensive experiments show that HRS is superior in defending against state-of-the-art white-box and adaptive adversarial misclassification attacks. We also demonstrate the effectiveness of HRS in defending adversarial reprogramming, which is the first defense against adversarial programs. Moreover, in most settings the average DES of HRS is at least 5X higher than current stochastic network defenses, validating its significantly improved robustness-accuracy trade-off.
Xiao Wang 0028, Siyue Wang, Yanzhi Wang 0001, Brian Kulis, Xue Lin 0001, Sang (Peter) Chin
IJCAI7
2019 JOBS: Joint-Sparse Optimization from Bootstrap Samples
abstract
Classical sparse regression based on ℓ1minimization solves the least squares problem with all available measurements via sparsity-promoting regularization. In challenging practical applications with high levels of noise and missing or adversarial samples, solving the problem using all measurements simultaneously may fail. In this paper, we propose a robust global sparse recovery strategy, named JOBS, which uses bootstrap samples of measurements to improve sparse regression in difficult cases. K measurement vectors are generated from the original pool of m measurements via bootstrapping, with each bootstrap sample containing L elements, and then a joint-sparse constraint is enforced to ensure support consistency among multiple predictors. The final estimate is obtained by averaging over K estimators. The performance limits associated with finite bootstrap sampling ratio L/m and number of estimates K is analyzed theoretically. Simulation results validate the theoretical analysis of proper choice of (L,K) and show that the proposed method yields state-of-the-art recovery performance, outperforming ℓ1minimization and other existing bootstrap-based techniques, especially when the number of measurements are limited. With a proper choice of bootstrap sampling ratio (0.3-0.5) and a reasonably large number of estimates K (≥ 30), the SNR improvement over the baseline ℓ1-minimization algorithm can reach up to 336%.
Luoluo Liu, Sang (Peter) Chin, Trac D. Tran
ISIT2
2018 A Greedy Pursuit Algorithm for Separating Signals from Nonlinear Compressive Observations
abstract
In this paper we study the unmixing problem which aims to separate a set of structured signals from their superposition. In this paper, we consider the scenario in which the mixture is observed via nonlinear compressive measurements. We present a fast, robust, greedy algorithm called Unmixing Matching Pursuit (UnmixMP) to solve this problem. We prove rigorously that the algorithm can recover the constituents from their noisy nonlinear compressive measurements with arbitrarily small error. We compare our algorithm to the Demixing with Hard Thresholding (DHT) algorithm [1], in a number of experiments on synthetic and real data.
Sang (Peter) Chin, Trac D. Tran, Dung N. Tran, Akshay Rangamani
ICASSP1
2018 Defensive dropout for hardening deep neural networks under adversarial attacks
Siyue Wang, Xiao Wang 0028, Pu Zhao 0001, Wujie Wen, David R. Kaeli, Sang (Peter) Chin, Xue Lin 0001
ICCAD6
2018 Using Deep Learning to Extract Scenery Information in Real Time Spatiotemporal Compressed Sensing
abstract
One of the problems of real time compressed sensing system is the computational cost of the reconstruction algorithms. It is especially problematic for close loop sensory applications where the sensory parameters needs to be constantly adjust to adapt to a dynamic scene. Through a preliminary experiment with MNIST dataset, we showed that we can extract some scene information (object recognition, scene movement direction and speed) based on the compressed samples using a deep convolutional neural network. It achieves 100% percent accuracy in distinguishing moving velocity, 96.22% in recognizing the digit and 90.04% in detecting moving direction after the code images are re-centered. Even though the classification accuracy drops slightly compared to using original videos, the computational speed is two time faster than classification on videos directly. This method also eliminates the need for sparse reconstruction prior to classification.
Xiao Wang 0028, Jie Zhang 0063, Trac D. Tran, Sang (Peter) Chin, Ralph Etienne-Cummings
ISCAS5
2018 Sparse Coding and Autoencoders
abstract
In this work we study the landscape of squared loss of an Autoencoder when the data generative model is that of “Sparse Coding”/“Dictionary Learning”. The neural net considered is an$\mathbb{R}^{n}\rightarrow \mathbb{R}^{n}$mapping and has a single ReLU activation layer of size$h > n$. The net has access to vectors$y\in \mathbb{R}^{n}$obtained as$y=A^{\ast}x^{\ast}$where$x^{\ast}\in \mathbb{R}^{h}$are sparse high dimensional vectors and$A^{\ast}\in \mathbb{R}^{n\times h}$is an overcomplete incoherent matrix. Under very mild distributional assumptions on$x^{\ast}$, we prove that the norm of the expected gradient of the squared loss function is asymptotically (in sparse code dimension) negligible for all points in a small neighborhood of$A^{\ast}$. This is supported with experimental evidence using synthetic data. We conduct experiments to suggest that$A^{\ast}$sits at the bottom of a well in the landscape and we also give experiments showing that gradient descent on this loss function gets columnwise very close to the original dictionary even with far enough initialization. Along the way we prove that a layer of ReLU gates can be set up to automatically recover the support of the sparse codes. Since this property holds independent of the loss function we believe that it could be of independent interest. A full version of this paper is accessible at: https://arxiv.org/abs/1708.03735
Akshay Rangamani, Anirbit Mukherjee, Amitabh Basu, Ashish Arora, Tejaswini Ganapathi, Sang (Peter) Chin, Trac D. Tran
ISIT6
2018 Sound Signal Processing with Seq2Tree Network
Kai Cao 0005, Zhaoheng Ni, Sang (Peter) Chin, Xiang Li 0066
LREC4
2018 Learning to Repair Software Vulnerabilities with Generative Adversarial Networks
abstract
Motivated by the problem of automated repair of software vulnerabilities, we propose an adversarial learning approach that maps from one discrete source domain to another target domain without requiring paired labeled examples or source and target domains to be bijections. We demonstrate that the proposed adversarial learning approach is an effective technique for repairing software vulnerabilities, performing close to seq2seq approaches that require labeled pairs. The proposed Generative Adversarial Network approach is application-agnostic in that it can be applied to other problems similar to code repair, such as grammar correction or sentiment translation.
Jacob Harer, Onur Ozdemir, Tomo Lazovich, Christopher P. Reale, Rebecca L. Russell, Louis Y. Kim, Sang (Peter) Chin
NeurIPS7
2017 A provable nonconvex model for factoring nonnegative matrices
abstract
We study the Nonnegative Matrix Factorization problem which approximates a nonnegative matrix by a low-rank factorization. This problem is particularly important in Machine Learning, and finds itself in a large number of applications. Unfortunately, the original formulation is ill-posed and NP-hard. In this paper, we propose a row sparse model based on Row Entropy Minimization to solve the NMF problem under separable assumption which states that each data point is a convex combination of a few distinct data columns. We utilize the concentration of the entropy function and the ℓ∞norm to concentrate the energy on the least number of latent variables. We prove that under the separability assumption, our proposed model robustly recovers data columns that generate the dataset, even when the data is corrupted by noise. We empirically justify the robustness of the proposed model and show that it is significantly more robust than the state-of-the-art separable NMF algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP2
2017 Live demonstration: A compact all-CMOS spatiotemporal compressed sensing video camera
abstract
A compact all-CMOS spatiotemporal compressed sensing (CS) video camera is demonstrated. This CS-based framework [1], implemented on integrated circuits, is able to achieve 20-fold reduction in the readout speed and consumes only 14μW to provide 100 fps videos. Taking advantage of dictionary learning and sparse recovery, this prototype image sensor (127×90 pixels) can reconstruct 100 fps videos from the coded images sampled at 5 fps.
Jie Zhang 0063, Chetan Singh Thakur, John M. Rattray, Sang (Peter) Chin, Trac D. Tran, Ralph Etienne-Cummings
ISCAS5
2017 Deep Image-to-Image Recurrent Network with Shape Basis Learning for Automatic Vertebra Labeling in Large-Scale 3D CT Volumes
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Zhoubing Xu, Mingqing Chen, Jin Hyeong Park, Sasa Grbic, Trac D. Tran, Sang (Peter) Chin, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)10
2017 A Mathematical Analysis of Network Controllability Through Driver Nodes
abstract
A May 2011 Nature article by Liu, Slotine, and Barabasi laid a mathematical foundation for analyzing network controllability of self-organizing networks and how to identify the minimum number of nodes needed to control a network, or driver nodes. In this paper, we continue to explore this topic, beginning with a look at how Laplacian eigenvalues relate to the percentage of nodes required to control a network. Next, we define and analyze super driver nodes, or those driver nodes that survive graph randomization. Finally, we examine node properties to differentiate super driver nodes from other types of nodes in a graph.
Sang (Peter) Chin, Alison Albin, Mykola Hayvanovych, Elizabeth Reilly, Gavin Brown 0003, Jacob Harer
IEEE Trans. Comput. Soc. Syst.1
2016 Partial face recognition: A sparse representation-based approach
abstract
Partial face recognition is a problem that often arises in practical settings and applications. We propose a sparse representation-based algorithm for this problem. Our method firstly trains a dictionary and the classifier parameters in a supervised dictionary learning framework and then aligns the partially observed test image and seeks for the sparse representation with respect to the training data alternatively to obtain its label. We also analyze the performance limit of sparse representation-based classification algorithms on partial observations. Finally, face recognition experiments on the popular AR data-set are conducted to validate the effectiveness of the proposed method.
Luoluo Liu, Trac D. Tran, Sang (Peter) Chin
ICASSP3
2016 Low-rank matrices recovery via entropy function
abstract
The low-rank matrix recovery problem consists of reconstructing an unknown low-rank matrix from a few linear measurements, possibly corrupted by noise. One of the most popular method in low-rank matrix recovery is based on nuclear-norm minimization, which seeks to simultaneously estimate the most significant singular values of the target low-rank matrix by adding a penalizing term on its nuclear norm. In this paper, we introduce a new method that requires substantially fewer measurements needed for exact matrix recovery compared to nuclear norm minimization. The proposed optimization program utilizes a sparsity promoting regularization in the form of the entropy function of the singular values. Numerical experiments on synthetic and real data demonstrates that the proposed method outperforms stage-of-the-art nuclear norm minimization algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP3
2015 Targeted Dot Product Representation for Friend Recommendation in Online Social Networks
abstract
In this paper, we develop Targeted Dot Product Representation (TarDPR), a DPR-based feature selection and combination framework for friend recommendation in online social networks (OSNs). Our approach modifies conventional DPR techniques and makes itself applicable to OSNs by focusing on computing a consistent representation while minimizing unnecessary suggestions made outside these interested regions. A notable property of TarDPR is its ability to effectively incorporate different types of social features and produce new meaningful features that help competitive approaches to significantly improve their recommendation quality. We derive an iterative algorithm for TarDPR that is supported by mathematical analysis, and is efficient on large social traces. To certify the usability of our approach, we conduct empirical experiments on real social traces including Facebook and Foursquare social networks. The competitive experimental results show that TarDPR achieves up to 15% improvement in comparison with other competitive methods. These results consequently confirm the efficacy of our suggested framework.
Minh Dao, Akshay Rangamani, Sang (Peter) Chin, Nam P. Nguyen, Trac D. Tran
ASONAM3
2015 Stochastic Block Model and Community Detection in Sparse Graphs: A spectral algorithm with optimal rate of recovery
abstract
In this paper, we present and analyze a simple and robust spectral algorithm for the stochastic block model with k blocks, for any k fixed. Our algorithm works with graphs having constant edge density, under an optimal condition on the gap between the density inside a block and the density between the blocks. As a co-product, we settle an open question posed by Abbe et. al. concerning censor block models.
Sang (Peter) Chin, Anup B. Rao, Van H. Vu
COLT1
2015 Nonnegative matrix factorization with gradient vertex pursuit
abstract
Nonnegative Matrix Factorization (NMF), defined as factorizing a nonnegative matrix into two nonnegative factor matrices, is a particularly important problem in machine learning. Unfortunately, it is also ill-posed and NP-hard. We propose a fast, robust, and provably correct algorithm, namely Gradient Vertex Pursuit (GVP), for solving a well-defined instance of the problem which results in a unique solution: there exists a polytope, whose vertices consist of a few columns of the original matrix, covering the entire set of remaining columns. Our algorithm is greedy: it detects, at each iteration, a correct vertex until the entire polytope is identified. We evaluate the proposed algorithm on both synthetic and real hyperspectral data, and show its superior performance compared with other state-of-the-art greedy pursuit algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP3
2015 Local sensing with global recovery
abstract
In this paper, we study Locally Compressed Sensing for images, where sampling process is allowed to be performed on arbitrary local regions of the images. We propose a fast and efficient reconstruction algorithm which utilizes local structures of images. Several numerical experiments on real images demonstrates that our algorithm yields better reconstruction quality than existing techniques at much lower computational complexity and memory requirement.
Dung N. Tran, Duyet N. Tran, Sang (Peter) Chin, Trac D. Tran
ICIP3
2015 Randomized Minmax Regret for Combinatorial Optimization Under Uncertainty
Andrew Mastin, Patrick Jaillet, Sang (Peter) Chin
ISAAC3
2015 An unsupervised dictionary learning algorithm for neural recordings
abstract
To meet the growing demand of wireless and power efficient neural recordings systems, we demonstrate an unsupervised dictionary learning algorithm in Compressed Sensing (CS) framework which can be implemented in VLSI systems. Without prior label information of neural spikes, we extend our previous work to unsupervised learning and construct a dictionary with discriminative structures for spike sorting. To further improve the reconstruction and classification performance, we proposed a joint prediction to determine the class of neural spikes in dictionary learning. When the neural spikes is compressed 50 times, our approach can achieve an average gain of 2 dB and 15 percentage units over state-of-the-art of CS approaches in terms of the reconstruction quality and classification accuracy respectively.
Jie Zhang 0063, Yuanming Suo, Dung N. Tran, Ralph Etienne-Cummings, Sang (Peter) Chin, Trac D. Tran
ISCAS6
2014 Computing Diffusion State Distance Using Green's Function and Heat Kernel on Graphs
Edward Boehnlein, Sang (Peter) Chin, Amit Sinha, Linyuan Lu
WAW2
2013 Video frame interpolation via weighted robust principal component analysis
abstract
In this paper, we propose a new video frame interpolation technique by a locally-adaptive robust principal component analysis (RPCA) with weight priors. The proposed algorithm relies on two main steps: 1. the pre-processing step initializes the new frame by a simplified motion-compensated frame interpolation and assigns each pixel a confident weight based on both the difference of motion estimation and local consistency; and 2. the refinement step updates the frame by a proposed weighted robust principal component analysis (WRPCA) algorithm. Experiments demonstrate that the proposed method outperforms the state-of-the-art algorithms, both in visual quality and PSNR performance.
Minh Dao, Yuanming Suo, Sang (Peter) Chin, Trac D. Tran
ICASSP3
2013 Reconstruction of neural action potentials using signal dependent sparse representations
abstract
We demonstrate a method to build signal dependent sparse representation dictionary for neural action potentials using K-SVD algorithm and Discrete Wavelets Transform. We also show a method to utilize this dictionary to recover the neural signal in the Compressive Sensing (CS) framework. Comparing against the non-signal dependent CS recovery algorithms, this new recovery algorithm can achieve same reconstruction quality with 2.5 times less compressed sensing measurements. For the same compression ratio, the purposed approach can increase recovery signal's signal to noise and distortion ratio (SNDR) by around 6 dB compare to non-signal dependent recovery method. We also evaluated the recovered signal using spike sorting techniques. The results have shown that the spikes clusters still maintain clear separation even when the compression ratio is at 15-20% of the Nyquist rate. This work also implies that any hardware implementation of compressed sensing could be scaled down in term of power and chip area by the same order if this signal dependent framework is used to recover the signal.
Jie Zhang 0063, Yuanming Suo, Srinjoy Mitra, Sang (Peter) Chin, Trac D. Tran, Refet Firat Yazicioglu, Ralph Etienne-Cummings
ISCAS4
2012 A fast multiscale framework for data in high-dimensions: Measure estimation, anomaly detection, and compressive measurements
abstract
Data sets are often modeled as samples from some probability distribution lying in a very high dimensional space. In practice, they tend to exhibit low intrinsic dimensionality, which enables both fast construction of efficient data representations and solving statistical tasks such as regression of functions on the data, or even estimation of the probability distribution from which the data is generated. In this paper we introduce a novel multiscale density estimator for high dimensional data and apply it to the problem of detecting changes in the distribution of dynamic data, or in a time series of data sets. We also show that our data representations, which are not standard sparse linear expansions, are amenable to compressed measurements. Finally, we test our algorithms on both synthetic data and a real data set consisting of a times series of hyperspectral images, and demonstrate their high accuracy in the detection of anomalies.
Guangliang Chen, Mark A. Iwen, Sang (Peter) Chin, Mauro Maggioni
VCIP3