Yilin Shen

dblp:30/383 · DBLP profile ↗
← Back
93ranked-venue papers
21as first author
41since 2021 · last 2025
0000-0002-1955-1529ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 10 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 11 · 7 first-authorComputer networks · 9 · 4 first-authorSecurity and privacy · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Theory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter Network
abstract
Multichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottleneck in denoising speech within one stage. In contrast, conventional signal processing algorithms often feature a two-stage design, where the first stage focuses on spatially aligning the received signals with respect to the speech source, followed by the second stage to filter out noise. In this paper, we introduce Align-and-Filter network (AFnet) for deep learning-based SE that decouples the primal denoising problem into two sub-problems, which imitates the alignment-followed-by-filtering wisdom from signal processing. The key is to leverage the relative transfer functions (RTFs) that encode meaningful spatial information via a tactically designed alignment strategy. Experimental results show that by leveraging the proposed RTF-based spatial alignment supervision, AFnet learns interpretable directional features to better exploit spatial separability of sound sources for improved SE performance.
Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin
ICASSP5
2025 MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword Spotting
abstract
Deep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation of existing State-of-the-Art (SOTA) keyword spotting models on OOD data w.r.t. in-domain testing data, and propose a training mechanism to improve performance on OOD data. Specifically, we propose a novel combination of Mixup and Information Bottleneck, called MIB, to achieve SOTA performance on OOD data. Considering on-device applications, we show across multiple models ranging from sizes of 12.5K parameters to 350K parameters, that MIB achieves as much as 2.5% (absolute) improvement in performance over OOD data. Further, in the more realistic case where OOD keywords are uttered in the presence of OOD noise, MIB achieves as much as 10% (absolute) performance improvement over SOTA models. The proposed MIB is model-agnostic, i.e., it can be applied to enhance the training of any deep keyword spotting model.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP6
2025 MoDeGPT: Modular Decomposition for Large Language Model Compression
abstract
Large Language Models (LLMs) have significantly advanced AI with their exceptional performance across a wide range of tasks. However, their extensive computational requirements restrict their use on devices with limited resources. While recent compression methods based on low-rank matrices show potential solutions, they often suffer from significant loss of accuracy or introduce substantial overhead in parameters and inference time. In this paper, we introduce Modular De- composition (MoDeGPT), a new, efficient, and structured compression framework that overcomes these limitations. MoDeGPT jointly decomposes pairs of consecu- tive subcomponents within Transformer blocks, reduces hidden dimensions through output reconstruction on a larger structural scale than conventional low-rank meth- ods, and repurposes three classical matrix decomposition algorithms—Nyström approximation, CR decomposition, and SVD—to ensure bounded errors in our novel decomposition approach. Our experiments show that MoDeGPT, without relying on backward propagation, consistently matches or surpasses the performance of prior techniques that depend on gradient information, while achieving a 98% reduction in compute costs when compressing a 13B-parameter model. On LLaMA-2/3 and OPT models, MoDeGPT retains 90-95% of zero-shot performance with compression rates of 25-30%. The compression process can be completed on a single GPU in a few hours, boosting inference throughput by up to 46%.
Chi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel, Shikhar Tuli, Yilin Shen, Hongxia Jin, Yen-Chang Hsu
ICLR6
2025 RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
abstract
Denoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass certain useful information about the target clean data distribution. However, existing adoption of the standard Gaussian as the prior distribution in turn discards such information when shaping the prior, resulting in sub-optimal performance. In this paper, we propose to improve conditional DDPMs for signal restoration by leveraging a more informative prior that is jointly learned with the diffusion model. The proposed framework, called RestoreGrad, seamlessly integrates DDPMs into the variational autoencoder (VAE) framework, taking advantage of the correlation between the degraded and clean signals to encode a better diffusion prior. On speech and image restoration tasks, we show that RestoreGrad demonstrates faster convergence (5-10 times fewer training steps) to achieve better quality of restored signals over existing DDPM baselines and improved robustness to using fewer sampling steps in inference time (2-2.5 times fewer), advocating the advantages of leveraging jointly learned prior for efficiency improvements in the diffusion process.
Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin
ICML6
2025 FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
abstract
James Seale Smith, Chi-Heng Lin, Shikhar Tuli, Haris Jeelani, Shangqian Gao, Yilin Shen, Hongxia Jin, Yen-Chang Hsu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
James Seale Smith, Chi-Heng Lin, Shikhar Tuli, Haris Jeelani, Shangqian Gao, Yilin Shen, Hongxia Jin, Yen-Chang Hsu
NAACL (Long Papers)6
2024 Unified Srgb Real Noise Synthesizing with Adaptive Feature Modulation
abstract
Recently, the Neighboring Correlation-Aware (NeCA) noise model has achieved impressive performance on both noise synthesis and the downstream image denoising task. However, its design regarding noise-level prediction requires training NeCA separately for each camera type. To this end, by making use of an adaptive feature modulation technique, we improve NeCA’s noise-level prediction model to be unified for different camera types and thus enable a unified sRGB real noise synthesis method. We also find out that in the neigh-boring correlation network of NeCA, there is no mechanism to maintain the signal dependency of the synthesized noise. Therefore, we introduce another adaptive feature modulation technique to the neighboring correlation network to maintain the signal dependency of the noise.
Wenbo Li 0001, Zhipeng Mo, Yilin Shen, Hongxia Jin
ICASSP3
2024 End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG Signals
abstract
Cuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this paper, we introduce a novel end-to-end model based on transformers. Further, we also introduce a novel contrastive loss-based loss function for robust training. To study the limits of performance for our proposed ideas, we first study personalized models trained on large subject-specific datasets, and achieve an average mean absolute error of 1.08/0.68 mmHg for systolic (SBP) and diastolic BP (DBP) across all subjects while achieving a best case of 0.29/0.19 mmHg. Further, in the case where subject-specific data is scarce, we leverage transfer learning using multi-subject data, and show that our model outperforms State-of-the-Art (SOTA) methods across varying amounts of subject-specific data.
Suhas BN, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP7
2024 Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language Model
abstract
Zero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that addresses an issue in the original CL which treats non-corresponding pairs indiscriminately. We confirm that SSCL outperforms CL through experiments. Then, we explore how including text or speech in-domain data during the SSCL training affects the out-of-domain intent classification.During the zero-shot classification, embeddings for a set of classes in the new domain are generated to calculate the similarities between each class embedding and an input utterance embedding, after which the most similar class is predicted for the utterance’s intent. Although manually-collected text sentences per class can be used to generate the class embedding, the data collection can be costly. Thus, we explore how to generate better class embeddings without human-collected text data in the target domain. The best proposed method employing an instruction-tuned Llama2, a public large language model, shows the performance comparable to the case where the human-collected text data was used, implying the importance of accurate class embedding generation.
Rakshith Sharma Srinivasa, Ching Hua Lee, Yashas Malur Saidutta, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP6
2024 An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech Enhancement
abstract
In multichannel speech enhancement (SE) systems, deep neural networks (DNNs) are often utilized to directly estimate the clean speech for effective beamforming. This approach, however, may not generalize adequately to new acoustic or noise conditions. Alternatively, DNNs can indirectly perform SE by predicting the time-frequency masks of speech and noise patterns to assist classic statistical beamformers. Despite being robust, its effectiveness is constrained by the later statistical component relying on certain modeling assumptions, e.g., covariance-based modeling in the minimum-variance-distortionless-response (MVDR) beamformer. In this paper, we propose a novel integration of the two types of methodology, by introducing an intra-MVDR module embedded in the U-Net beamformer, that encompasses the merits of both, i.e., effectiveness and robustness. Experiments show that intra-MVDR leads to improvements that are not achievable by simply enlarging the baseline SE network.
Ching Hua Lee, Kashyap Patel, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP4
2024 Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech Enhancement
abstract
Deep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perform adequately in another environment, the target (new) domain of unseen acoustic or noise conditions. Even though we can improve the target domain performance by leveraging paired data in that domain, in reality, noisy data is more straightforward to collect. Therefore, it is worth studying unsupervised domain adaptation techniques for SE that utilize only noisy data from the target domain, together with exploiting the knowledge available from the source domain paired data, for improved SE in the new domain. In this paper, we present a novel adaptation framework for SE by leveraging self-supervised learning (SSL) based speech models. SSL models are pre-trained with large amount of raw speech data to extract representations rich in phonetic and acoustics information. We explore the potential of leveraging SSL representations for effective SE adaptation to new domains. To our knowledge, it is the first attempt to apply SSL models for domain adaptation in SE.
Ching Hua Lee, Chouchang Yang, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Yilin Shen, Hongxia Jin
ICASSP6
2024 Enabling Device Control Planning Capabilities of Small Language Model
abstract
Smart home device control is a difficult task if the instruction is abstract and the planner needs to adjust dynamic home configurations. With the increasing capability of Large Language Model (LLM), they have become the customary model for zero-shot planning tasks similar to smart home device control. Although cloud supported large language models can seamlessly do device control tasks, on-device small language models show limited capabilities. In this work, we show how we can leverage large language models to enable small language models for device control task. Towards this goal, we develop an automated system to generate device control planning data leveraging large language model and use the generated data to finetune the small language models. We empirically validate the improvement of small language models’ performance for device control task.
Sudipta Paul 0011, Yilin Shen, Hongxia Jin
ICASSP3
2024 Adaptive Rank Selections for Low-Rank Approximation of Language Models
abstract
Shangqian Gao, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shangqian Gao, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin
NAACL-HLT4
2024 DynaMo: Accelerating Language Model Inference with Dynamic Multi-Token Sampling
abstract
Shikhar Tuli, Chi-Heng Lin, Yen-Chang Hsu, Niraj Jha, Yilin Shen, Hongxia Jin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shikhar Tuli, Chi-Heng Lin, Yen-Chang Hsu, Niraj K. Jha, Yilin Shen, Hongxia Jin
NAACL-HLT5
2024 DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
abstract
Large Language Models (LLMs) have achieved remarkable success in various natural language processing tasks, including language modeling, understanding, and generation. However, the increased memory and computational costs associated with these models pose significant challenges for deployment on resource-limited devices. Structural pruning has emerged as a promising solution to reduce the costs of LLMs without requiring post-processing steps. Prior structural pruning methods either follow the dependence of structures at the cost of limiting flexibility, or introduce non-trivial additional parameters by incorporating different projection matrices. In this work, we propose a novel approach that relaxes the constraint imposed by regular structural pruning methods and eliminates the structural dependence along the embedding dimension. Our dimension-independent structural pruning method offers several benefits. Firstly, our method enables different blocks to utilize different subsets of the feature maps. Secondly, by removing structural dependence, we facilitate each block to possess varying widths along its input and output dimensions, thereby significantly enhancing the flexibility of structural pruning. We evaluate our method on various LLMs, including OPT, LLaMA, LLaMA-2, Phi-1.5, and Phi-2. Experimental results demonstrate that our approach outperforms other state-of-the-art methods, showing for the first time that structural pruning can achieve an accuracy similar to semi-structural pruning.
Shangqian Gao, Chi-Heng Lin, Ting Hua, Yilin Shen, Hongxia Jin, Yen-Chang Hsu
NeurIPS5
2024 Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View Benchmark
abstract
Thanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavorable visual conditions such as poor light or overexposure. Unfortunately, there are currently very few datasets available, including for example MVSeg dataset that focuses purely toward eye-level view; and it features the sparse annotation nature due to the intensive demands of labeling process. To address these key challenges of the MVSS task, this paper presents two major contributions: the introduction of MVUAV, a new MVSS benchmark dataset, and the development of a dedicated semi-supervised MVSS baseline - SemiMV. Our MVUAV dataset is captured via Unmanned Aerial Vehicles (UAV), which offers a unique oblique bird’s-eye view complementary to the existing MVSS datasets; it also encompasses a broad range of day/night lighting conditions and over 30 semantic categories. In the meantime, to better leverage the sparse annotations and extra unlabeled RGB-Thermal videos, a semi-supervised learning baseline, SemiMV, is proposed to enforce consistency regularization through a dedicated Cross-collaborative Consistency Learning (C3L) module and a denoised temporal aggregation strategy. Comprehensive empirical evaluations on both MVSeg and MVUAV benchmark datasets have showcased the efficacy of our SemiMV baseline.
Wei Ji 0011, Wenbo Li 0001, Yilin Shen, Li Cheng 0001, Hongxia Jin
NeurIPS4
2024 CIFD: Controlled Information Flow to Enhance Knowledge Distillation
abstract
Knowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediately sized models, Teacher Assistants (TAs) to help the transfer process. However, training TAs is expensive, as training these models is a knowledge transfer task in itself. Further, these TAs are larger than the student model and training them especially in large data settings can be computationally intensive. In this paper, we propose a novel framework called Controlled Information Flow for Knowledge Distillation (CIFD) consisting of two components. First, we propose a significantly smaller alternatives to TAs, the Rate-Distortion Module (RDM) which uses the teacher's penultimate layer embedding and a information rate-constrained bottleneck layer to replace the Teacher Assistant model. RDMs are smaller and easier to train than TAs, especially in large data regimes, since they operate on the teacher embeddings and do not need to relearn low level input feature extractors. Also, by varying the information rate across the bottleneck, RDMs can replace TAs of different sizes. Secondly, we propose the use of Information Bottleneck Module in the student model, which is crucial for regularization in the presence of a large number of RDMs. We show comprehensive state-of-the-art results of the proposed method over large datasets like Imagenet. Further, we show the significant improvement in distilling CLIP like models over a huge 12M image-text dataset. It outperforms CLIP specialized distillation methods across five zero-shot classification datasets and two zero-shot image-text retrieval datasets.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
NeurIPS6
2024 Token Fusion: Bridging the Gap between Token Pruning and Token Merging
abstract
Vision Transformers (ViTs) have emerged as powerful backbones in computer vision, outperforming many traditional CNNs. However, their computational overhead, largely attributed to the self-attention mechanism, makes deployment on resource-constrained edge devices challenging. Multiple solutions rely on token pruning or token merging. In this paper, we introduce "Token Fusion" (ToFu), a method that amalgamates the benefits of both token pruning and token merging. Token pruning proves advantageous when the model exhibits sensitivity to input interpolations, while token merging is effective when the model manifests close to linear responses to inputs. We combine this to propose a new scheme called Token Fusion. Moreover, we tackle the limitations of average merging, which doesn’t preserve the intrinsic feature norm, resulting in distributional shifts. To mitigate this, we introduce MLERP merging, a variant of the SLERP technique, tailored to merge multiple tokens while maintaining the norm distribution. ToFu is versatile, applicable to ViTs with or without additional training. Our empirical evaluations indicate that ToFu establishes new benchmarks in both classification and image generation tasks concerning computational efficiency and model accuracy.
Shangqian Gao, Yen-Chang Hsu, Yilin Shen, Hongxia Jin
WACV4
2024 Efficient Layout-Guided Image Inpainting for Mobile Use
abstract
The layout guidance, which specifies the pixel-wise object distribution, is beneficial to preserving the object boundaries in image inpainting while not hurting model’s generalization capability. We aim to design an efficient and robust layout-guided image inpainting method for mobile use, which can achieve the robustness in presence of the mixed scenes where objects with the delicate shape reside next to the hole. Our method is made up of two sub-models, which restore the pixel-information for the hole from coarse to fine, and support each other to overcome the practical challenges encountered when making the whole method lightweight. The layout mask guides the two sub-models, which thus enables the robustness of our method in mixed scenes. We demonstrate the efficiency and robustness of our method via both the experiments and a mobile demo.
Wenbo Li 0001, Yi Wei 0006, Yilin Shen, Hongxia Jin
WACV3
2023 GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision Transformer
abstract
The recently proposed Vision transformers (ViTs) have shown very impressive empirical performance in various computer vision tasks, and they are viewed as an important type of foundation model. However, ViTs are typically constructed with large-scale sizes, which then severely hinder their potential deployment in many practical resources constrained applications. To mitigate this challenging problem, structured pruning is a promising solution to compress model size and enable practical efficiency. However, unlike its current popularity for CNNs and RNNs, structured pruning for ViT models is little explored. In this paper, we propose GOHSP, a unified framework of Graph and Optimization-based Structured Pruning for ViT models. We first develop a graph-based ranking for measuring the importance of attention heads, and the extracted importance information is further integrated to an optimization-based procedure to impose the heterogeneous structured sparsity patterns on the ViT models. Experimental results show that our proposed GOHSP demonstrates excellent compression performance. On CIFAR-10 dataset, our approach can bring 40% parameters reduction with no accuracy loss for ViT-Small model. On ImageNet dataset, with 30% and 35% sparsity ratio for DeiT-Tiny and DeiT-Small models, our approach achieves 1.65% and 0.76% accuracy increase over the existing structured pruning methods, respectively.
Miao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin, Bo Yuan 0001
AAAI3
2023 One-stage Progressive Dichotomous Segmentation
Karim Ahmed, Wenbo Li 0001, Yilin Shen, Hongxia Jin
BMVC4
2023 Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching Masking
abstract
In multichannel speech enhancement (SE), time-frequency (T-F) mask-based neural beamforming algorithms take advantage of deep neural networks to predict T-F masks that represent speech and noise dominance. The predicted masks are subsequently leveraged to estimate the speech and noise power spectral density (PSD) matrices for computing the beamformer filter weights based on signal statistics. However, in the literature most networks are trained to estimate some pre-defined masks, e.g., the ideal binary mask (IBM) and ideal ratio mask (IRM) that lack direct connection to the PSD estimation. In this paper, we propose a new masking strategy to predict the Snapshot Matching Mask (SMM) that aims to minimize the distance between the predicted and the true signal snapshots, thereby estimating the PSD matrices in a more systematic way. Performance of SMM compared with existing IBM- and IRM-based PSD estimation for mask-based neural beamforming is presented on several datasets to demonstrate its effectiveness for the SE task.
Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP3
2023 To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive Refinement
abstract
Keyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP5
2023 Learning to Jointly Share and Prune Weights for Grounding Based Vision and Language Models
Shangqian Gao, Burak Uzkent, Yilin Shen, Heng Huang 0001, Hongxia Jin
ICLR3
2023 ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
abstract
The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, which generalizes poorly to novel objects in unknown environments. In this work, we present a novel zero-shot object navigation method, Exploration with Soft Commonsense constraints (ESC), that transfers commonsense knowledge in pre-trained models to open-world object navigation without any navigation experience nor any other training on the visual environments. First, ESC leverages a pre-trained vision and language model for open-world prompt-based grounding and a pre-trained commonsense language model for room and object reasoning. Then ESC converts commonsense knowledge into navigation actions by modeling it as soft logic predicates for efficient exploration. Extensive experiments on MP3D, HM3D, and RoboTHOR benchmarks show that our ESC method improves significantly over baselines, and achieves new state-of-the-art results for zero-shot object navigation (e.g., 288% relative Success Rate improvement than CoW on MP3D).
Kaiwen Zhou 0002, Kaizhi Zheng, Connor Pryor, Yilin Shen, Hongxia Jin, Lise Getoor, Xin Wang 0061
ICML4
2023 Compositional Generalization in Spoken Language Understanding
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH2
2023 Robust Keyword Spotting for Noisy Environments by Leveraging Speech Enhancement and Speech Presence Probability
Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Yilin Shen, Hongxia Jin
INTERSPEECH5
2023 CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive Loss
abstract
This paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar to Contrastive Language-Image Pre-training (CLIP) and Locked-image Tuning (LiT) that have recently gained considerable attention. Classical contrastive training employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a new contrastive loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to transfer the structure of the embedding space from one modality to another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. By using publicly available datasets, we achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification.
Rakshith Sharma Srinivasa, Chouchang Yang, Yashas Malur Saidutta, Ching Hua Lee, Yilin Shen, Hongxia Jin
NeurIPS6
2023 TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models
abstract
Large Language Models (LLMs) are progressively being utilized as machine learning services and interface tools for various applications. However, the security implications of LLMs, particularly in relation to adversarial and Trojan attacks, remain insufficiently examined. In this paper, we propose TrojLLM, an automatic and black-box framework to effectively generate universal and stealthy triggers. When these triggers are incorporated into the input data, the LLMs' outputs can be maliciously manipulated. Moreover, the framework also supports embedding Trojans within discrete prompts, enhancing the overall effectiveness and precision of the triggers' attacks. Specifically, we propose a trigger discovery algorithm for generating universal triggers for various inputs by querying victim LLM-based APIs using few-shot data samples. Furthermore, we introduce a novel progressive Trojan poisoning algorithm designed to generate poisoned prompts that retain efficacy and transferability across a diverse range of models. Our experiments and results demonstrate TrojLLM's capacity to effectively insert Trojans into text prompts in real-world black-box LLM APIs including GPT-3.5 and GPT-4, while maintaining exceptional performance on clean test sets. Our work sheds light on the potential security risks in current models and offers a potential defensive approach. The source code of TrojLLM is available at https://github.com/UCF-ML-Research/TrojLLM.
Mengxin Zheng, Ting Hua, Yilin Shen, Ladislau Bölöni, Qian Lou
NeurIPS4
2022 Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial Relations
abstract
Phrase grounding is a multi-modal problem that localizes a particular noun phrase in an image referred to by a text query. In the challenging zero-shot phrase grounding setting, the existing state-of-the-art grounding models have limited capacity in handling the unseen phrases. Humans, however, can ground novel types of objects in images with little effort, significantly benefiting from reasoning with commonsense. In this paper, we design a novel phrase grounding architecture that builds multi-modal knowledge graphs using external knowledge and then performs graph reasoning and spatial relation reasoning to localize the referred nouns phrases. We perform extensive experiments on different zero-shot grounding splits sub-sampled from the Flickr30K Entity and Visual Genome dataset, demonstrating that the proposed framework is orthogonal to backbone image encoders and outperforms the baselines by 2~3% in accuracy, resulting in a significant improvement under the standard evaluation metrics.
Yilin Shen, Hongxia Jin, Xiaodan Zhu 0001
AAAI2
2022 Text-Based Interactive Recommendation via Offline Reinforcement Learning
abstract
Interactive recommendation with natural-language feedback can provide richer user feedback and has demonstrated advantages over traditional recommender systems. However, the classical online paradigm involves iteratively collecting experience via interaction with users, which is expensive and risky. We consider an offline interactive recommendation to exploit arbitrary experience collected by multiple unknown policies. A direct application of policy learning with such fixed experience suffers from the distribution shift. To tackle this issue, we develop a behavior-agnostic off-policy correction framework to make offline interactive recommendation possible. Specifically, we leverage the conservative Q-function to perform off-policy evaluation, which enables learning effective policies from fixed datasets without further interactions. Empirical results on the simulator derived from real-world datasets demonstrate the effectiveness of our proposed offline training framework.
Ruiyi Zhang 0002, Tong Yu 0001, Yilin Shen, Hongxia Jin
AAAI3
2022 Lite-MDETR: A Lightweight Multi-Modal Detector
abstract
Recent multi-modal detectors based on transformers and modality encoders have successfully achieved impressive results on end-to-end visual object detection conditioned on a raw text query. However, they require a large model size and an enormous amount of computations to achieve high performance, which makes it difficult to deploy mobile applications that are limited by tight hardware resources. In this paper, we present a Lightweight modulated detector, Lite-MDETR, to facilitate efficient end-to-end multi-modal understanding on mobile devices. The key primitive is that Dictionary-Lookup-Transformormations (DLT) is proposed to replace Linear Transformation (LT) in multi-modal detectors where each weight in Linear Transformation (LT) is approximately factorized into a smaller dictionary, index, and coefficient. This way, the enormous linear projection with weights is converted into efficient linear projection with dictionaries, a few lookups and scalings with indices and coefficients. DLT can be applied to any pretrained multi-modal detectors, removing the need to perform expensive training from scratch. To tackle the challenging training of DLT due to non-differentiable index, we convert the index and coefficient into a sparse matrix, train this sparse matrix during the fine-tuning phase, and recover it back to index and coefficient during the inference phase. Our experiments on phrase grounding, referring expression comprehension and segmentation, and VQA show that our Lite-MDETR achieves similar accuracy as the prior multi-modal detectors with up to ~ 4.1 × model size reduction.
Qian Lou, Yen-Chang Hsu, Burak Uzkent, Ting Hua, Yilin Shen, Hongxia Jin
CVPR5
2022 Numerical Optimizations for Weighted Low-rank Estimation on Language Models
abstract
Singular value decomposition (SVD) is one of the most popular compression methods that approximate a target matrix with smaller matrices.However, standard SVD treats the parameters within the matrix with equal importance, which is a simple but unrealistic assumption.The parameters of a trained neural network model may affect the task performance unevenly, which suggests non-equal importance among the parameters.Compared to SVD, the decomposition method aware of parameter importance is the more practical choice in real cases.Unlike standard SVD, weighted value decomposition is a non-convex optimization problem that lacks a closed-form solution.We systematically investigated multiple optimization strategies to tackle the problem and examined our method by compressing Transformer-based language models.Further, we designed a metric to predict when the SVD may introduce a significant performance drop, for which our method can be a rescue strategy.The extensive evaluations demonstrate that our method can perform better than current SOTA methods in compressing Transformer-based language models.
Ting Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou, Yilin Shen, Hongxia Jin
EMNLP5
2022 Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, Hongxia Jin
ICLR5
2022 DictFormer: Tiny Transformer with Shared Dictionary
Qian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin
ICLR4
2021 Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU
abstract
Yilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin
ACL/IJCNLP (1)1
2021 Multi-Step Spoken Language Understanding System Based on Adversarial Learning
abstract
Most of the existing spoken language understanding systems can perform only semantic frame parsing based on a single-round user query. They cannot take users’ feedback to up-date/add/remove slot values through multiround interactions with users. In this paper, we introduce a novel multi-step spoken language understanding system based on adversarial learning that can leverage the multiround user’s feedback to update slot values. We perform two experiments on the benchmark ATIS dataset and demonstrate that the new system can improve parsing performance by at least 2.5% in terms of F1, with only one round of feedback. The improvement becomes even larger when the number of feedback rounds increases. Furthermore, we also compare the new system with state-of-the-art dialogue state tracking systems and demonstrate that the new interactive system can perform better on multiround spoken language understanding tasks in terms of slot- and sentence-level accuracy.
Yu Wang 0091, Yilin Shen, Hongxia Jin
ICASSP2
2021 An End-To-End Actor-Critic-Based Neural Coreference Resolution System
abstract
The target of a coreference resolution system is to cluster all mentions that refer to the same entity in a given context. All coreference resolution systems need to solve two subtasks; one task is to detect all of the potential mentions, and the other is to learn the linking of an antecedent for each possible mention. In this paper, we propose an actor-critic-based neural coreference resolution system, which can achieve both mention detection and mention clustering by leveraging an actor-critic deep reinforcement learning technique and a joint training algorithm. We experiment on the BERT model to generate different input span representations. Our model with the BERT span representation achieves the state-of-the-art performance among the models on the CoNLL-2012 Shared Task English Test Set.
Yu Wang 0091, Yilin Shen, Hongxia Jin
ICASSP2
2021 Always Be Dreaming: A New Approach for Data-Free Class-Incremental Learning
abstract
Modern computer vision applications suffer from catastrophic forgetting when incrementally learning new concepts over time. The most successful approaches to alleviate this forgetting require extensive replay of previously seen data, which is problematic when memory constraints or data legality concerns exist. In this work, we consider the high-impact problem of Data-Free Class-Incremental Learning (DFCIL), where an incremental learning agent must learn new concepts over time without storing generators or training data from past tasks. One approach for DFCIL is to replay synthetic images produced by inverting a frozen copy of the learner’s classification model, but we show this approach fails for common class-incremental benchmarks when using standard distillation strategies. We diagnose the cause of this failure and propose a novel incremental distillation strategy for DFCIL, contributing a modified cross-entropy training and importance-weighted feature distillation, and show that our method results in up to a 25.1% increase in final task accuracy (absolute difference) compared to SOTA DFCIL methods for common class-incremental benchmarks. Our method even outperforms several standard replay based methods which store a coreset of images. Our code is available at https://github.com/GT-RIPL/AlwaysBeDreaming-DFCIL
James Seale Smith, Yen-Chang Hsu, Jonathan C. Balloch, Yilin Shen, Hongxia Jin, Zsolt Kira
ICCV4
2021 SAFENet: A Secure, Accurate and Fast Neural Network Inference
Qian Lou, Yilin Shen, Hongxia Jin, Lei Jiang 0001
ICLR2
2021 Automatic Mixed-Precision Quantization Search of BERT
abstract
Pre-trained language models such as BERT have shown remarkable effectiveness in various natural language processing tasks. However, these models usually contain millions of parameters, which prevent them from the practical deployment on resource-constrained devices. Knowledge distillation, Weight pruning, and Quantization are known to be the main directions in model compression. However, compact models obtained through knowledge distillation may suffer from significant accuracy drop even for a relatively small compression ratio. On the other hand, there are only a few attempts based on quantization designed for natural language processing tasks, and they usually require manual setting on hyper-parameters. In this paper, we proposed an automatic mixed-precision quantization framework designed for BERT that can conduct quantization and pruning simultaneously. Specifically, our proposed method leverages Differentiable Neural Architecture Search to assign scale and precision for parameters in each sub-group automatically, and at the same pruning out redundant groups of parameters. Extensive evaluations on BERT downstream tasks reveal that our proposed method beats baselines by providing the same performance with much smaller model size. We also show the possibility of obtaining the extremely light-weight model by combining our solution with orthogonal methods such as DistilBERT.
Changsheng Zhao 0002, Ting Hua, Yilin Shen, Qian Lou, Hongxia Jin
IJCAI3
2021 Hyperparameter-free Continuous Learning for Domain Classification in Natural Language Understanding
abstract
Ting Hua, Yilin Shen, Changsheng Zhao, Yen-Chang Hsu, Hongxia Jin. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ting Hua, Yilin Shen, Changsheng Zhao 0002, Yen-Chang Hsu, Hongxia Jin
NAACL-HLT2
2020 Towards Hands-Free Visual Dialog Interactive Recommendation
abstract
With the recent advances of multimodal interactive recommendations, the users are able to express their preference by natural language feedback to the item images, to find the desired items. However, the existing systems either retrieve only one item or require the user to specify (e.g., by click or touch) the commented items from a list of recommendations in each user interaction. As a result, the users are not hands-free and the recommendations may be impractical. We propose a hands-free visual dialog recommender system to interactively recommend a list of items. At each time, the system shows a list of items with visual appearance. The user can comment on the list in natural language, to describe the desired features they further want. With these multimodal data, the system chooses another list of items to recommend. To understand the user preference from these multimodal data, we develop neural network models which identify the described items among the list and further predict the desired attributes. To achieve efficient interactive recommendations, we leverage the inferred user preference and further develop a novel bandit algorithm. Specifically, to avoid the system exploring more than needed, the desired attributes are utilized to reduce the exploration space. More importantly, to achieve sample efficient learning in this hands-free setting, we derive additional samples from the user's relative preference expressed in natural language and design a pairwise logistic loss in bandit learning. Our bandit model is jointly updated by the pairwise logistic loss on the additional samples derived from natural language feedback and the traditional logistic loss. The empirical results show that the probability of finding the desired items by our system is about 3 times as high as that by the traditional interactive recommenders, after a few user interactions.
Tong Yu 0001, Yilin Shen, Hongxia Jin
AAAI2
2020 Generalized ODIN: Detecting Out-of-Distribution Image Without Learning From Out-of-Distribution Data
abstract
Deep neural networks have attained remarkable performance when applied to data that comes from the same distribution as that of the training set, but can significantly degrade otherwise. Therefore, detecting whether an example is out-of-distribution (OoD) is crucial to enable a system that can reject such samples or alert users. Recent works have made significant progress on OoD benchmarks consisting of small image datasets. However, many recent methods based on neural networks rely on training or tuning with both in-distribution and out-of-distribution data. The latter is generally hard to define a-priori, and its selection can easily bias the learning. We base our work on a popular method ODIN, proposing two strategies for freeing it from the needs of tuning with OoD data, while improving its OoD detection performance. We specifically propose to decompose confidence scoring as well as a modified input pre-processing method. We show that both of these significantly help in detection performance. Our further analysis on a larger scale image dataset shows that the two types of distribution shifts, specifically semantic shift and non-semantic shift, present a significant difference in the difficulty of the problem, providing an analysis of when ODIN-like strategies do or do not work.
Yen-Chang Hsu, Yilin Shen, Hongxia Jin, Zsolt Kira
CVPR2
2020 Generating Dialogue Responses from a Semantic Latent Space
abstract
Existing open-domain dialogue generation models are usually trained to mimic the gold response in the training set using cross-entropy loss on the vocabulary.However, a good response does not need to resemble the gold response, since there are multiple possible responses to a given prompt.In this work, we hypothesize that the current models are unable to integrate information from multiple semantically similar valid responses of a prompt, resulting in the generation of generic and uninformative responses.To address this issue, we propose an alternative to the end-to-end classification on vocabulary.We learn the pair relationship between the prompts and responses as a regression task on a latent space instead.In our novel dialog generation model, the representations of semantically related sentences are close to each other on the latent space.Human evaluation showed that learning the task on a continuous space can generate responses that are both relevant and informative.
Wei-Jen Ko, Avik Ray, Yilin Shen, Hongxia Jin
EMNLP (1)3
2020 PGLP: Customizable and Rigorous Location Privacy Through Policy Graph
Yang Cao 0011, Yonghui Xiao, Li Xiong 0001, Masatoshi Yoshikawa, Yilin Shen, Jinfei Liu, Hongxia Jin
ESORICS (1)6
2020 A BI-Model Approach for Handling Unknown Slot Values in Dialogue State Tracking
abstract
In this paper, we present an end-to-end bi-model structure for dialogue state tracking, which can handle the scenarios when the spoken language understanding model with a predefined slot candidate list is absent. Furthermore, the model structure described in this paper can effectively extract unknown slot values and still maintain the state-of-the-art performance on DSTC2 benchmark. We also compare our model in detail with an existing end-to-end dialogue state tracking model using pointer network which can also handle the unknown slot values, and demonstrates that how the bi-model structure can benefit the task and hence gives better performance.
Yu Wang 0091, Yilin Shen, Hongxia Jin
ICASSP2
2020 An Interactive Adversarial Reward Learning-Based Spoken Language Understanding System
Yu Wang 0091, Yilin Shen, Hongxia Jin
INTERSPEECH2
2020 Activity Recommendation: Optimizing Life in the Long Term
abstract
College students every day decide and plan how to best spend their time to balance academic, physical, and social goals under uncertainty. This process is likely suboptimal where long-term life satisfaction and success is not guaranteed, and poor decision-making may lead to longer-term problems like depression. To support everyday planning, we introduce activity recommendation, a novel method that combines artificial intelligence, machine learning, and a psychology-informed approach to automatically generate activity-recommendations that optimize long-term life satisfaction. We tested our method with an existing dataset and derived activity recommendations for depressed and non-depressed students. We evaluated the recommendations through interviews with college students who rated the suggestions positively. Our model can be optimized for different goals and domains and is easy to interpret. Our results demonstrate the feasibility of our approach and lay the groundwork towards implementing a live system.
Julian Ramos 0001, Johana Rosas, Yilin Shen, Hongxia Jin, Anind K. Dey
PerCom3
2020 A New Concept of Multiple Neural Networks Structure Using Convex Combination
abstract
In this article, a new concept of convex-combined multiple neural networks (NNs) structure is proposed. This new approach uses the collective information from multiple NNs to train the model. Based on both theoretical and experimental analyses, the new approach is shown to achieve faster training convergence with a similar or even better test accuracy than a conventional NN structure. Two experiments are conducted to demonstrate the performance of our new structure: the first one is a semantic frame parsing task for spoken language understanding (SLU) on the Airline Travel Information System (ATIS) data set and the other is a handwritten digit recognition task on the Mixed National Institute of Standards and Technology (MNIST) data set. We test this new structure using both the recurrent NN and convolutional NNs through these two tasks. The results of both experiments demonstrate a 4× - 8× faster training speed with better or similar performance by using this new concept.
Yu Wang 0091, Yue Deng 0001, Yilin Shen, Hongxia Jin
IEEE Trans. Neural Networks Learn. Syst.3
2019 A Progressive Model to Enable Continual Learning for Semantic Slot Filling
abstract
Yilin Shen, Xiangyu Zeng, Hongxia Jin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yilin Shen, Xiangyu Zeng 0003, Hongxia Jin
EMNLP/IJCNLP (1)1
2019 Adversarial Multi-label Prediction for Spoken and Visual Signal Tagging
abstract
We introduce an adversarial multi-label classification (ADMLC) framework to improve the robustness and performance of existing algorithms on multi-domain signals. The core contribution of our ADMLC is the innovation of an `adversarial module' that serves as a critic to provide augmenting information to improve supervised learning in multi label classification (MLC) tasks. Our approach is not intended to be regarded as an emerging competitor for many well-established algorithms in the field. In fact, many existing deep and shallow architectures can all be adopted as building blocks integrated in the ADMLC framework. We show the performance and generalization ability of ADMLC on diverse tasks including audio and image tagging.
Yue Deng 0001, KaWai Chen, Yilin Shen, Hongxia Jin
ICASSP3
2019 SLiQA-I: Towards Cold-start Development of End-to-end Spoken Language Interface for Question Answering
abstract
Question answering (QA) has become a key capability for voice enabled personal assistants to automatically answer various user questions. However, the development of a spoken language interface for QA in a new domain is time consuming and requires a lot of human labors. Thus, it is crucially desirable to design an end-to-end system, referred to as SliQA, that can facilitate developers to easily and quickly build a QA interface from scratch and output a high quality plug-and-play QA engine. In this paper, we take the first step of SliQA system design, named SliQA-I, to support answering factoid questions regarding an entity over existing knowledge graphs. SliQA-I incorporates a novel iterative human-in-the-loop question generator and an enhanced deep coupled QA engine, thereby requiring light human workload. We implement the real system and evaluate it on three domains from different aspects. The results show that the QA performance of SliQA-I achieves up to 3.58% accuracy gain compared with baseline approaches which use existing QA engine on human generated data. More importantly, we show that SliQA-I only takes as low as 0.025 second to generate a question which has similar quality as human generated ones in terms of both naturalness and grammatical correctness.
Yilin Shen, Yu Wang 0091, Abhishek Patel, Hongxia Jin
ICASSP1
2019 Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded
abstract
Many vision and language models suffer from poor visual grounding -- often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called Human Importance-aware Network Tuning (HINT) that effectively leverages human demonstrations to improve visual grounding. HINT encourages deep networks to be sensitive to the same input regions as humans. Our approach optimizes the alignment between human attention maps and gradient-based network importances -- ensuring that models learn not just to look at but rather rely on visual concepts that humans found relevant for a task when making predictions. We apply HINT to Visual Question Answering and Image Captioning tasks, outperforming top approaches on splits that penalize over-reliance on language priors (VQA-CP and robust captioning) using human attention demonstrations for just 6% of the training data.
Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, Devi Parikh
ICCV3
2019 Learning Assistance from an Adversarial Critic for Multi-Outputs Prediction
abstract
We introduce an adversarial-critic-and-assistant (ACA) learning framework to improve the performance of existing supervised learning with multiple outputs. The core contribution of our ACA is the innovation of two novel modules, i.e. an `adversarial critic' and a `collaborative assistant', that are jointly designed to provide augmenting information for facilitating general learning tasks. Our approach is not intended to be regarded as an emerging competitor for tons of well-established algorithms in the field. In fact, most existing approaches, while implemented with different learning objectives, can all be adopted as building blocks seamlessly integrated in the ACA framework to accomplish various real-world tasks. We show the performance and generalization ability of ACA on diverse learning tasks including multi-label classification, attributes prediction and sequence-to-sequence generation.
Yue Deng 0001, Yilin Shen, Hongxia Jin
IJCAI2
2019 Iterative Delexicalization for Improved Spoken Language Understanding
abstract
Recurrent neural network (RNN) based joint intent classification and slot tagging models have achieved tremendous success in recent years for building spoken language understanding and dialog systems.However, these models suffer from poor performance for slots which often encounter large semantic variability in slot values after deployment (e.g.message texts, partial movie/artist names).While greedy delexicalization of slots in the input utterance via substring matching can partly improve performance, it often produces incorrect input.Moreover, such techniques cannot delexicalize slots with out-of-vocabulary slot values not seen at training.In this paper, we propose a novel iterative delexicalization algorithm, which can accurately delexicalize the input, even with out-of-vocabulary slot values.Based on model confidence of the current delexicalized input, our algorithm improves delexicalization in every iteration to converge to the best input having the highest confidence.We show on benchmark and in-house datasets that our algorithm can greatly improve parsing performance for RNN based models, especially for out-of-distribution slot values.
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH2
2019 Interpreting and Improving Deep Neural SLU Models via Vocabulary Importance
Yilin Shen, Wenhu Chen, Hongxia Jin
INTERSPEECH1
2019 A Visual Dialog Augmented Interactive Recommender System
abstract
Traditional recommender systems rely on user feedback such as ratings or clicks to the items, to analyze the user interest and provide personalized recommendations. However, rating or click feedback are limited in that they do not exactly tell why users like or dislike an item. If a user does not like the recommendations and can not effectively express the reasons via rating and clicking, the feedback from the user may be very sparse. These limitations lead to inefficient model learning of the recommender system. To address these limitations, more effective user feedback to the recommendations should be designed, so that the system can effectively understand a user's preference and improve the recommendations over time. In this paper, we propose a novel dialog-based recommender system to interactively recommend a list of items with visual appearance. At each time, the user receives a list of recommended items with visual appearance. The user can point to some items and describe their feedback, such as the desired features in the items they want in natural language. With this natural language based feedback, the recommender system updates and provides another list of items. To model the user behaviors of viewing, commenting and clicking on a list of items, we propose a visual dialog augmented cascade model. To efficiently understand the user preference and learn the model, exploration should be encouraged to provide more diverse recommendations to quickly collect user feedback on more attributes of the items. We propose a variant of the cascading bandits, where the neural representations of the item images and user feedback in natural language are utilized. In a task of recommending a list of footwear, we show that our visual dialog augmented interactive recommender needs around 41.03% rounds of recommendations, compared to the traditional interactive recommender only relying on the user click behavior.
Tong Yu 0001, Yilin Shen, Hongxia Jin
KDD2
2019 Vision-Language Recommendation via Attribute Augmented Multimodal Reinforcement Learning
abstract
Interactive recommenders have demonstrated the advantage over traditional recommenders with dynamic change of items. However, the traditional user feedback in the format of clicks or ratings, provides limited user preference information and limited history tracking capabilities. As a result, it takes a user many interactions to find a desired item. Data of other modalities, such as item visual appearance and user comments in natural language, may enable richer user feedback. However, there are several critical challenges to be addressed when utilizing these multimodal data: multimodal matching, user preference tracking, and adaptation to dynamic unseen items. Without properly handling these challenges, the recommendations can easily violate the users' preference from their past natural language feedback. In this paper, we introduce a novel approach, called vision-language recommendation, that enables users to provide natural language feedback on visual products to have more natural and effective interactions. To model more explicit and accurate multimodal matching, we propose a novel visual attribute augmented reinforcement learning approach that enhances the grounding of natural language to visual items. Furthermore, to effectively track the users' preference and overcome the performance deficiency on dynamic unseen items after deployment, we propose a novel history multimodal matching reward to continuously adapt the model on-the-fly. Empirical results show that, our system augmented by visual attribute and history multimodal matching can significantly increase the success rate, reduce the number of recommendations that violate the user's previous feedback, and need less number of user interactions to find the desired items.
Tong Yu 0001, Yilin Shen, Ruiyi Zhang 0002, Xiangyu Zeng 0003, Hongxia Jin
ACM Multimedia2
2019 Teach Once and Use Everywhere - Building AI Assistant Eco-Skills via User Instruction and Demonstration
abstract
Voice-enabled AI assistants rely on developers to build every single skill, although many skills share similar functions. We propose a concept and prototype system, \ksystem, to automatically build a set of similar skills in the ecosystem (eco-skills) with one-time teaching from end users. During teaching, a user only needs to demonstrate on the screen in one (native) mobile app and provides natural language (NL) instructions.
Yilin Shen, Sandeep Nama, Hongxia Jin
MobiSys1
2019 Text-Based Interactive Recommendation via Constraint-Augmented Reinforcement Learning
abstract
Text-based interactive recommendation provides richer user preferences and has demonstrated advantages over traditional interactive recommender systems. However, recommendations can easily violate preferences of users from their past natural-language feedback, since the recommender needs to explore new items for further improvement. To alleviate this issue, we propose a novel constraint-augmented reinforcement learning (RL) framework to efficiently incorporate user preferences over time. Specifically, we leverage a discriminator to detect recommendations violating user historical preference, which is incorporated into the standard RL objective of maximizing expected cumulative future rewards. Our proposed framework is general and is further extended to the task of constrained text generation. Empirical results show that the proposed method yields consistent improvement relative to standard RL methods.
Ruiyi Zhang 0002, Tong Yu 0001, Yilin Shen, Hongxia Jin, Changyou Chen
NeurIPS3
2018 Adversarial Active Learning for Sequences Labeling and Generation
abstract
We introduce an active learning framework for general sequence learning tasks including sequence labeling and generation. Most existing active learning algorithms mainly rely on an uncertainty measure derived from the probabilistic classifier for query sample selection. However, such approaches suffer from two shortcomings in the context of sequence learning including 1) cold start problem and 2) label sampling dilemma. To overcome these shortcomings, we propose a deep-learning-based active learning framework to directly identify query samples from the perspective of adversarial learning. Our approach intends to offer labeling priorities for sequences whose information content are least covered by existing labeled data. We verify our sequence-based active learning approach on two tasks including sequence labeling and sequence generation.
Yue Deng 0001, KaWai Chen, Yilin Shen, Hongxia Jin
IJCAI3
2018 Learning Out-of-Vocabulary Words in Intelligent Personal Agents
abstract
Semantic parsers play a vital role in intelligent agents to convert natural language instructions to an actionable logical form representation. However, after deployment, these parsers suffer from poor accuracy on encountering out-of-vocabulary (OOV) words, or significant accuracy drop on previously supported instructions after retraining. Achieving both goals simultaneously is non-trivial. In this paper, we propose novel neural networks based parsers to learn OOV words; one incorporating a new hybrid paraphrase generation model, and an enhanced sequence-to-sequence model. Extensive experiments on both benchmark and custom datasets show our new parsers achieve significant accuracy gain on OOV words and phrases, and in the meanwhile learn OOV words while maintaining accuracy on previously supported instructions.
Avik Ray, Yilin Shen, Hongxia Jin
IJCAI2
2018 Training Recurrent Neural Network through Moment Matching for NLP Applications
Yue Deng 0001, Yilin Shen, KaWai Chen, Hongxia Jin
INTERSPEECH2
2018 Robust Spoken Language Understanding via Paraphrasing
abstract
Learning intents and slot labels from user utterances is a fundamental step in all spoken language understanding (SLU) and dialog systems.State-of-the-art neural network based methods, after deployment, often suffer from performance degradation on encountering paraphrased utterances, and out-of-vocabulary words, rarely observed in their training set.We address this challenging problem by introducing a novel paraphrasing based SLU model which can be integrated with any existing SLU model in order to improve their overall performance.We propose two new paraphrase generators using RNN and sequence-to-sequence based neural networks, which are suitable for our application.Our experiments on existing benchmark and in house datasets demonstrate the robustness of our models to rare and complex paraphrased utterances, even under adversarial test distributions.
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH2
2018 User Information Augmented Semantic Frame Parsing Using Progressive Neural Networks
Yilin Shen, Xiangyu Zeng 0003, Yu Wang 0091, Hongxia Jin
INTERSPEECH1
2018 A Deep Reinforcement Learning Based Multimodal Coaching Model (DCM) for Slot Filling in Spoken Language Understanding(SLU)
Yu Wang 0091, Abhishek Patel, Yilin Shen, Hongxia Jin
INTERSPEECH3
2018 Interactive recommendation via deep neural memory augmented contextual bandits
abstract
Personalized recommendation with user interactions has become increasingly popular nowadays in many applications with dynamic change of contents (news, media, etc.). Existing approaches model user interactive recommendation as a contextual bandit problem to balance the trade-off between exploration and exploitation. However, these solutions require a large number of interactions with each user to provide high quality personalized recommendations. To mitigate this limitation, we design a novel deep neural memory augmented mechanism to model and track the history state for each user based on his previous interactions. As such, the user's preferences on new items can be quickly learned within a small number of interactions. Moreover, we develop new algorithms to leverage large amount of all users' history data for offline model training and online model fine tuning for each user with the focus of policy evaluation. Extensive experiments on different synthetic and real-world datasets validate that our proposed approach consistently outperforms a variety of state-of-the-art approaches.
Yilin Shen, Yue Deng 0001, Avik Ray, Hongxia Jin
RecSys1
2018 Accelerating Time Series Searching with Large Uniform Scaling
abstract
Similarity search is arguably the most important primitive in time series data mining. It is useful in its own right as an exploratory tool, and a subroutine in almost all higher level algorithms, such as motif discovery, anomaly detection, classification, clustering and summarization. Because of this, and the prevalence of time series data, the last decade has seen fast algorithms for time series similarity search under Dynamic Time Warping (DTW) and Uniform Scaling (US) distance measures. However, current state-of-the-art algorithms for US have only been demonstrated for the modest amounts of rescaling in datasets produced by human behaviors such as gestures, speech, music performance and physiological measurements such as heartbeats and respiration. As we shall show, in many industrial and commercial contexts we may encounter much greater amounts of rescaling, rendering current solutions little better than brute force search. To mitigate this problem we introduce novel lower bounds, LBnew, which, for the first time allows efficient search even in domains that exhibit more than a factor-of-two variability in scale. We demonstrate the utility of our ideas with both theoretical guarantees and comprehensive experiments on real data from commercial important domains, including power consumption monitoring and ECG monitoring. The results show the application of our lower bounds significantly outperforms state-of-the-art approaches for accelerating similarity searching of time series with more than a factor-of-two variability in scale as well as high-level time series mining tasks.
Yilin Shen, Yanping Chen 0005, Eamonn J. Keogh, Hongxia Jin
SDM1
2017 Searching Time Series with Invariance to Large Amounts of Uniform Scaling
abstract
Similarity search is arguably the most important primitive in time series data mining. Recent research has made significant progress on fast algorithms for time series similarity search under Dynamic Time Warping (DTW) and Uniform Scaling (US) distance measures. However, the current state-of-the-art algorithms cannot support greater amounts of rescaling in many practical applications. In this paper, we introduce a novel lower bound, LBnew, to allow efficient search even in domains that exhibit more than a factor-of-two variability in scale. The effectiveness of our idea is validated on various large-scale real datasets from commercial important domains.
Yilin Shen, Yanping Chen 0005, Eamonn J. Keogh, Hongxia Jin
ICDE1
2017 Disguise Adversarial Networks for Click-through Rate Prediction
abstract
We introduced an adversarial learning framework for improving CTR prediction in Ads recommendation. Our approach was motivated by observing the extremely low click-through rate and imbalanced label distribution in the historical Ads impressions. We hence proposed a Disguise-Adversarial-Networks (DAN) to improve the accuracy of supervised learning with limited positive-class information. In the context of CTR prediction, the rationality behind DAN could be intuitively understood as ``non-clicked Ads makeup''. DAN disguises the disliked Ads impressions (non-clicks) to be interesting ones and encourages a discriminator to classify these disguised Ads as positive recommendations. In an adversarial aspect, the discriminator should be sober-minded which is optimized to allocate these disguised Ads to their inherent classes according to an unsupervised information theoretic assignment strategy. We applied DAN to two Ads datasets including both mobile and display Ads for CTR prediction. The results showed that our DAN approach significantly outperformed other supervised learning and generative adversarial networks (GAN) in CTR prediction.
Yue Deng 0001, Yilin Shen, Hongxia Jin
IJCAI2
2017 Secure Pick Up: Implicit Authentication When You Start Using the Smartphone
abstract
We propose Secure Pick Up (SPU), a convenient, lightweight, in-device, non-intrusive and automatic-learning system for smartphone user authentication. Operating in the background, our system implicitly observes users' phone pick-up movements, the way they bend their arms when they pick up a smartphone to interact with the device, to authenticate the users.
Wei-Han Lee, Yilin Shen, Hongxia Jin, Ruby B. Lee
SACMAT3
2017 PCASA: Proximity Based Continuous and Secure Authentication of Personal Devices
abstract
User's personal portable devices such as smartphone, tablet and laptop require continuous authentication of the user to prevent against illegitimate access to the device and personal data. Current authentication techniques require users to enter password or scan fingerprint, making frequent access to the devices inconvenient. In this work, we propose to exploit user's on-body wearable devices to detect their proximity from her portable devices, and use the proximity for continuous authentication of the portable devices. We present PCASA which utilizes acoustic communication for secure proximity estimation with sub-meter level accuracy. PCASA uses Differential Pulse Position Modulation scheme that modulates data through varying the silence period between acoustic pulses to ensure energy efficiency even when authentication operation is being performed once every second. It yields an secure and accurate distance estimation even when user is mobile by utilizing Doppler effect for mobility speed estimation. We evaluate PCASA using smartphone and smartwatches, and show that it supports up to 34 hours of continuous authentication with a fully charged battery.
Pengfei Hu 0001, Parth H. Pathak, Yilin Shen, Hongxia Jin, Prasant Mohapatra
SECON3
2016 EpicRec: Towards Practical Differentially Private Framework for Personalized Recommendation
abstract
Recommender systems typically require users' history data to provide a list of recommendations and such recommendations usually reside on the cloud/server. However, the release of such private data to the cloud has been shown to put users at risk. It is highly desirable to provide users high-quality personalized services while respecting their privacy. In this paper, we develop the first Enhanced Privacy-built-In Client for Personalized Recommendation (EpicRec) system that performs the data perturbation on the client side to protect users' privacy. Our system needs no assumption of trusted server and no change on the recommendation algorithms on the server side; and needs minimum user interaction in their preferred manner, which makes our solution fit very well into real world practical use.
Yilin Shen, Hongxia Jin
CCS1
2016 Differentially Private User Data Perturbation with Multi-level Privacy Controls
abstract
Service providers typically collect user data for profiling users in order to provide high-quality services, yet this brings up user privacy concerns. One hand, service providers oftentimes need to analyze multiple user data attributes that usually have different privacy concern levels. On the other hand, users often pose different trusts towards different service providers based on their reputation. However, it is unrealistic to repeatedly ask users to specify privacy levels for each data attribute towards each service provider. To solve this problem, we develop the first lightweight and provably framework that not only guarantees differential privacy on both service provider and different data attributes but also allows configurable utility functions based on service needs. Using various large-scale real-world datasets, our solution helps to significantly improve the utility up to 5 times with negligible computational overhead, especially towards numerous low reputed service providers in practice.
Yilin Shen, Hongxia Jin
ECML/PKDD (2)1
2015 Private Analysis of Infinite Data Streams via Retroactive Grouping
abstract
With the rapid advances in hardware technology, data streams are being generated daily in large volumes, enabling a wide range of real-time analytical tasks. Yet data streams from many sources are inherently sensitive, and thus providing continuous privacy protection in data streams has been a growing demand. In this paper, we consider the problem of private analysis of infinite data streams under differential privacy. We propose a novel data stream sanitization framework that periodically releases histograms summarizing the event distributions over sliding windows to support diverse data analysis tasks. Our framework consists of two modules, a sampling-based change monitoring module and a continuous histogram publication module. The monitoring module features an adaptive Bernoulli sampling process to accurately track the evolution of a data stream. We for the first time conduct error analysis of sampling under differential privacy, which allows to select the best sampling rate. The publication module features three different publishing strategies, including a novel technique called retroactive grouping to enjoy reduced noise. We provide theoretical analysis of the utility, privacy and complexity of our framework. Extensive experiments over real datasets demonstrate that our solution substantially outperforms the state-of-the-art competitors.
Rui Chen 0012, Yilin Shen, Hongxia Jin
CIKM2
2014 Controllable Information Sharing for User Accounts Linkage across Multiple Online Social Networks
abstract
People have multiple accounts on Online Social Networks (OSNs) for various purposes. It is of great interest for third parties to collect more users' information by linking their accounts on different OSNs. Unfortunately, most users have not been aware of potential risks of such accounts linkage. Therefore, the design of a control methodology that allows users to share their information without the risk of being linked becomes an urgent need, yet still remains open.
Yilin Shen, Hongxia Jin
CIKM1
2014 Approximation Algorithms for Optimization Problems in Random Power-Law Graphs
Yilin Shen, Xiang Li 0016, My T. Thai
COCOA1
2014 Privacy-Preserving Personalized Recommendation: An Instance-Based Approach via Differential Privacy
abstract
Recommender systems become increasingly popular and widely applied nowadays. The release of users' private data is required to provide users accurate recommendations, yet this has been shown to put users at risk. Unfortunately, existing privacy-preserving methods are either developed under trusted server settings with impractical private recommender systems or lack of strong privacy guarantees. In this paper, we develop the first lightweight and provably private solution for personalized recommendation, under untrusted server settings. In this novel setting, users' private data is obfuscated before leaving their private devices, giving users greater control on their data and service providers less responsibility on privacy protections. More importantly, our approach enables the existing recommender systems (with no changes needed) to directly use perturbed data, rendering our solution very desirable in practice. We develop our data perturbation approach on differential privacy, the state-of-the-art privacy model with lightweight computation and strong but provable privacy guarantees. In order to achieve useful and feasible perturbations, we first design a novel relaxed admissible mechanism enabling the injection of flexible instance-based noises. Using this novel mechanism, our data perturbation approach, incorporating the noise calibration and learning techniques, obtains perturbed user data with both theoretical privacy and utility guarantees. Our empirical evaluation on large-scale real-world datasets not only shows its high recommendation accuracy but also illustrates the negligible computational overhead on both personal computers and smart phones. As such, we are able to meet two contradictory goals, privacy preservation and recommendation accuracy. This practical technology helps to gain user adoption with strong privacy protection and benefit companies with high-quality personalized services on perturbed user data.
Yilin Shen, Hongxia Jin
ICDM1
2013 Assessing network vulnerability in a community structure point of view
abstract
We introduce Community structure Vulnerability Assessment (CVA) problem to assess the network vulnerability under a community structure point of view. Given a positive number k, CVA aims to find out the k most vulnerable nodes whose removals maximally transform the current network community structure to a different one. As the first attempt, we suggest an approximation algorithm for the special case k = 1, and propose multiple greedy algorithms for CVA problem. To certify the effectiveness of suggested approaches, we test them on not only synthesized networks with known community structures but also on real-world social traces.
Nam P. Nguyen, Md Abdul Alim, Yilin Shen, My T. Thai
ASONAM3
2013 Network vulnerability assessment under cascading failures
abstract
The assessment of network vulnerability is of great importance in the presence of unexpected disruptive events or adversarial attacks, which will lead to a much more devastating consequence especially when failures can be cascaded. In this context, we study the Cascading Vulnerability Node Detection (CVND) optimization problem to identify the most vulnerable nodes in a network whose removals maximally destroy the network's functions after cascading failures, based on the recently proposed effective metric, total pairwise connectivity. Besides its NP-hardness on various graphs, we further show that the CVND problem is NP-hard to be approximated within Ω ((1+(d/n1-∈-1)2(n-k)/n∈)) equation with n vertices and k vulnerable nodes after d-hop failure cascades. Despite the intractability of this problem, we propose TRGA, a novel iterative two-phase algorithm, for efficiently solving the CVND problem in a timely manner. We also formulate the integer linear programming for obtaining the optimal solution. The effectiveness of our solutions is validated on various synthetic and real-world networks.
Yilin Shen, My T. Thai
GLOBECOM1
2013 Adaptive approximation algorithms for hole healing in hybrid wireless sensor networks
abstract
Region coverage and network connectivity are among the most important problems for the quality of service in wireless sensor networks. Unfortunately, due to the sensor failures and hostile environments, such as active volcanic regions or battle fields, the emergence of coverage holes and disconnections among sensors is unavoidable. One way to handle this problem is to deploy mobile sensors in the network, which is called hybrid sensor networks, so that these mobile sensors can be relocated to heal the holes or maintain the network connectivity. However, because of the low-power of mobile sensors, it is extremely challenging to design a fast and effective movement schedule for mobile sensors to (1) maintain both the region coverage and network connectivity at any time, and (2) minimize the moving energy consumption. In this paper, we develop an adaptive algorithm, AHCH algorithm, to adaptively heal the holes with the guarantee of network connectivity without recomputing from scratch. By comparing AHCH algorithm with the optimal solution at each time-slot, we show its expected adaptive approximation ratio as O(log |M|) with |M| mobile sensors in some special cases. In more general cases, we extend our AHCH algorithm to InAHCH and GenAHCH algorithms, handling insufficient mobile sensors as well as disconnected regions, along with the proof of their corresponding theoretical adaptive approximation ratios. The experimental evaluation shows the effectiveness of our proposed algorithms with respect to both low energy consumption and hole healing latency.
Yilin Shen, Dung T. Nguyen 0002, My T. Thai
INFOCOM1
2013 Efficient Multi-Link Failure Localization Schemes in All-Optical Networks
abstract
Link failure localization has been an important and challenging problem for all-optical networks. The most general monitoring structure, called m-trail, is a light-path into which optical signals are launched and monitored. How to minimize the number of required m-trails is critical to the expense of this technique. Existing solutions are limited to localizing single link failure or handling only small networks. Moreover, some practical constraints, like lacking of knowledge of the failure quantity, are ignored. To overcome these limitations is prospective but quite challenging. To this end, we provide novel theoretical solution frameworks toward the multi-failure localization problem. On one hand, for small dense networks, we provide a tree-decomposition based algorithm; on the other hand, a random walk based localized algorithm for large scale sparse networks is proposed. In addition, we further adapt these two algorithms to cope with three practical constraints. Theoretical analysis and simulation results are included to prove the correctness and efficiency of the proposed schemes.
Ying Xuan, Yilin Shen, Nam P. Nguyen, My T. Thai
IEEE Trans. Commun.2
2013 On the Discovery of Critical Links and Nodes for Assessing Network Vulnerability
abstract
The assessment of network vulnerability is of great importance in the presence of unexpected disruptive events or adversarial attacks targeting on critical network links and nodes. In this paper, we study Critical Link Disruptor (CLD) and Critical Node Disruptor (CND) optimization problems to identify critical links and nodes in a network whose removals maximally destroy the network's functions. We provide a comprehensive complexity analysis of CLD and CND on general graphs and show that they still remain NP-complete even on unit disk graphs and power-law graphs. Furthermore, the CND problem is shown NP-hard to be approximated within Ω([(n-k)/(nε)] ) on general graphs withnvertices andkcritical nodes. Despite the intractability of these problems, we propose HILPR, a novel LP-based rounding algorithm, for efficiently solving CLD and CND problems in a timely manner. The effectiveness of our solutions is validated on various synthetic and real-world networks.
Yilin Shen, Nam P. Nguyen, Ying Xuan, My T. Thai
IEEE/ACM Trans. Netw.1
2012 The walls have ears: optimize sharing for visibility and privacy in online social networks
abstract
With a rapid expansion of online social networks (OSNs), millions of users are tweeting and sharing their personal status daily without being aware of where that information eventually travels to. Likewise, with a huge magnitude of data available on OSNs, it poses a substantial challenge to track how a piece of information leaks to specific targets. In this paper, we study the problem of smartly sharing information to control the propagation of sensitive information in OSNs.
Thang N. Dinh, Yilin Shen, My T. Thai
CIKM2
2012 Interest-matching information propagation in multiple online social networks
abstract
Online social networks have become an imperative channel for extremely fast information propagation and influence. Thus, the problem of finding a minimum number of seed users who can eventually influence as many users in the network as possible has become one of the central research topics recently. Unfortunately, most of related works have only focused on the network topologies and largely ignored many other important factors such as the users' engagements and the negative or positive impacts between users. More challengingly, the behavior of information propagation across multiple networks simultaneously remains an untrodden area and becomes an urgent need. Our work is the first attempt to tackle the above problem in multiple networks, considering these lacking important factors. In order to capture the users' engagement, we propose to targeting the set of interest-matching users whose interests are similar to what we try to propagate. Then, we develop our Iterative Semi-Supervising Learning based approach to identify the minimum seed users. We validate the effectiveness of our solution by using real-world Twitter-Foursquare networks and academic collaboration multiple networks.
Yilin Shen, Thang N. Dinh, My T. Thai
CIKM1
2012 New techniques for approximating optimal substructure problems in power-law graphs
Yilin Shen, Dung T. Nguyen 0002, Ying Xuan, My T. Thai
Theor. Comput. Sci.1
2012 A Trigger Identification Service for Defending Reactive Jammers in WSN
abstract
During the last decade, Reactive Jamming Attack has emerged as a great security threat to wireless sensor networks, due to its mass destruction to legitimate sensor communications and difficulty to be disclosed and defended. Considering the specific characteristics of reactive jammer nodes, a new scheme to deactivate them by efficiently identifying all trigger nodes, whose transmissions invoke the jammer nodes, has been proposed and developed. Such a trigger-identification procedure can work as an application-layer service and benefit many existing reactive-jamming defending schemes. In this paper, on the one hand, we leverage several optimization problems to provide a complete trigger-identification service framework for unreliable wireless sensor networks. On the other hand, we provide an improved algorithm with regard to two sophisticated jamming models, in order to enhance its robustness for various network scenarios. Theoretical analysis and simulation results are included to validate the performance of this framework.
Ying Xuan, Yilin Shen, Nam P. Nguyen, My T. Thai
IEEE Trans. Mob. Comput.2
2011 Exploiting the Robustness on Power-Law Networks
Yilin Shen, Nam P. Nguyen, My T. Thai
COCOON1
2010 On the Hardness and Inapproximability of Optimization Problems on Power Law Graphs
Yilin Shen, Dung T. Nguyen 0002, My T. Thai
COCOA (1)1
2010 A Graph-Theoretic QoS-Aware Vulnerability Assessment for Network Topologies
abstract
How to assess the topology vulnerability of a network has attracted more and more attentions recently. Due to the rapid growing number of real- time internet applications developed since the last decade, the discovery of topology weakness related to its quality of service (QoS) is of more interest. In this paper, we provide a novel QoS-aware measurement for assessing the vulnerability of general network topologies. Specifically, we evaluate the vulnerability by detecting the minimum number of link failures that decrease the satisfactory level of the QoS-Optimal source- destination path to a given value, which means a topology with a smaller amount of such link failures is more vulnerable. We formulate this process as a graph optimization problem called QoSCE and provide several exact and heuristic algorithms for various QoS constraint amounts. To our best knowledge, this is the first graph-theoretical framework to evaluate QoS-aware topology vulnerability. Through extensive simulations, the performance of the proposed algorithms are validated in terms of assessment accuracy and time complexity.
Ying Xuan, Yilin Shen, Nam P. Nguyen, My T. Thai
GLOBECOM2
2009 On trigger detection against reactive Jamming Attacks: A clique-independent set based approach
abstract
Among existing countermeasures against Reactive Jamming Attacks in Wireless Sensor Networks, the detection of trigger nodes whose transmissions invoke the jammer nodes has been proposed and developed as an efficient jamming-resistent routing scheme in unreliable WSN environments. Sequential group testing techniques were adopted in our previous work to alleviate the overhead of trigger detection in terms of time and communication complexity. In this paper, we further improve the detection procedure by leveraging the classic models of randomized non-adaptive group testing and clique-independent set, which dramatically decreases the time complexity compared to our previous solution.
Ying Xuan, Yilin Shen, Incheol Shin, My T. Thai
IPCCC2
2005 A Middleware for Replicated Web Services
abstract
This paper presents a middleware that supports reliable Web services built on active replication. The middleware is responsible for maintaining the consistency of the replicas' states. A Java package for handling the interactions with the middleware and the failures of the Web services is provided for programmers to use when writing client applications. The package reduces the complexity in developing client applications.
Xinfeng Ye, Yilin Shen
ICWS2
2005 Replicating Multithreaded Web Services
Xinfeng Ye, Yilin Shen
ISPA2