VLDB 2026 Research / reviewers in the wild / expert
Yucheng Xie
dblp:254/5174
· DBLP profile ↗
30ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-5830-6625ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DivControl: Knowledge Diversion for Controllable Image GenerationabstractDiffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with entangled representations, resulting in poor generalization and high adaptation costs for novel conditions. To this end, we propose DivControl, a decomposable pretraining framework for unified controllable generation and efficient adaptation. DivControl factorizes ControlNet via SVD into basic components—pairs of singular vectors—which are disentangled into condition-agnostic learngenes and condition-specific tailors through knowledge diversion during multi-condition training. Knowledge diversion is implemented via a dynamic gate that performs soft routing over tailors based on the semantics of condition instructions, enabling zero-shot generalization and parameter-efficient adaptation to novel conditions. To further improve condition fidelity and training efficiency, we introduce a representation alignment loss that aligns condition embeddings with early diffusion features. Extensive experiments demonstrate that DivControl achieves state-of-the-art controllability with 36.4× less training cost, while simultaneously improving average performance on basic conditions. It also delivers strong zero-shot and few-shot performance on unseen conditions, demonstrating superior scalability, modularity, and transferability. Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001 |
AAAI | 1 |
| 2026 | Attacking mmWave-enabled Chest Vibration Sensing via Actuator-induced Mimicry
Xiaonan Guo 0003, Yucheng Xie, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
INFOCOM | 4 |
| 2026 | DietWatch: Fine-Grained and Robust Dietary Monitoring via Smartwatch in Real-World ScenariosabstractDietary behaviors play a pivotal role in promoting overall health and preventing chronic diseases (e.g., hypertension and diabetes). The widespread adoption of smartwatches offers a promising platform for continuous dietary monitoring. However, existing smartwatch-based dietary monitoring approaches struggle with challenges in real-world scenarios, including dynamic interference, gesture generalization, and user diversity. To address these limitations, we proposeDietWatch, a real-world dietary monitoring system that utilizes a commercial smartwatch to capture and analyze fine-grained dietary behaviors.DietWatchincorporates a dynamic interference mitigation module to suppress acoustic and inertial noise. It further employs a contrastive learning-based framework to distinguish eating gestures from diverse daily activities, without constraining users’ eating styles and activity types. To enhance generalizability across users,DietWatchadopts a cross-user adaptation mechanism to extract user-independent features. Furthermore, a clustering algorithm is designed to estimate dietary time, while an attention-based multimodal fusion method is employed to analyze biting and chewing frequencies and identify food categories. Experimental results demonstrate thatDietWatchachieves 79.95% temporal Intersection over Union for eating time detection, 85.68% accuracy in food classification, and mean absolute errors of 1.26 bites/min for biting frequency and 7.71 chews/min for chewing frequency estimation. Zhen Hou 0002, Yucheng Xie, Feng Li 0001, Honggang Wang 0001 |
IEEE Internet Things J. | 2 |
| 2025 | WAVE: Weight Templates for Adaptive Initialization of Variable-sized ModelsabstractThe growing complexity of model parameters underscores the significance of pre-trained models. However, deployment constraints often necessitate models of varying sizes, exposing limitations in the conventional pre-training and fine-tuning paradigm, particularly when target model sizes are incompatible with pre-trained ones. To address this challenge, we propose WAVE, a novel approach that reformulates variable-sized model initialization from a multitask perspective, where initializing each model size is treated as a distinct task. WAVE employs shared, size-agnostic weight templates alongside size-specific weight scalers to achieve consistent initialization across various model sizes. These weight templates, constructed within the Learngene framework, integrate knowledge from pre-trained models through a distillation process constrained by Kronecker-based rules. Target models are then initialized by concatenating and weighting these templates, with adaptive connection rules established by lightweight weight scalers, whose parameters are learned from minimal training data. Extensive experiments demonstrate the efficiency of WAVE, achieving state-of-the-art performance in initializing models of various depth and width. The knowledge encapsulated in weight templates is also task-agnostic, allowing for seamless transfer across diverse downstream datasets. Code will be made available at https://github.com/fu-feng/WAVE. Fu Feng, Yucheng Xie, Jing Wang 0113, Xin Geng 0001 |
CVPR | 2 |
| 2025 | Redefining in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generationabstract"Creative" remains an inherently abstract concept for both humans and diffusion models. While text-to-image (T2I) diffusion models can easily generate out-of-distribution concepts like "a blue banana", they struggle with generating combinatorial objects such as "a creative mixture that resembles a lettuce and a mantis", due to difficulties in understanding the semantic depth of "creative". Current methods rely heavily on synthesizing reference prompts or images to achieve a creative effect, typically requiring retraining for each unique creative output—a process that is computationally intensive and limits practical applications. To address this, we introduce CreTok, which brings meta-creativity to diffusion models by redefining "creative" as a new token,, thus enhancing models’ semantic understanding for combinatorial creativity. CreTok achieves such redefinition by iteratively sampling diverse text pairs from our proposed CangJie dataset to form adaptive prompts and restrictive prompts, and then optimizing the similarity between their respective text embeddings. Extensive experiments demonstrate thatenables the universal and direct generation of combinatorial creativity across diverse concepts without additional training, achieving state-of-the-art performance with improved text-image alignment and higher human preference ratings. Code will be made available at https://github.com/fu-feng/CreTok. Fu Feng, Yucheng Xie, Xu Yang 0021, Jing Wang 0113, Xin Geng 0001 |
CVPR | 2 |
| 2025 | Vigilante Defender: A Vaccination-based Defense Against Backdoor Attacks on 3D Point Clouds Using Particle Swarm OptimizationabstractBackdoor attacks on 3D Point Clouds (PCs) pose a serious threat by embedding hidden triggers into a subset of the training data. These triggers cause targeted misclassifications at inference time while leaving the model’s behavior unaffected in the absence of triggers, making them stealthy and difficult to detect. In distributed learning settings, where a central trainer aggregates data from multiple sources and offers only black-box access to the model, a single malicious contributor can compromise the model’s integrity if defenses are not in place. We propose a novel client-side defense that empowers individual contributors to act as vigilante defenders. By injecting benign ‘vaccination’ triggers—identified via Particle Swarm Optimization—into their local training data, defenders can proactively neutralize potential backdoors without prior knowledge of their location or structure. Experiments on standard benchmarks with PointNet and DGCNN show our method significantly reduces attack success while preserving classification accuracy, outperforming existing defenses. Agnideven Palanisamy Sundar, Feng Li 0001, Xukai Zou, Yucheng Xie, Ryan Hosler |
ICCCN | 4 |
| 2025 | KIND: Knowledge Integration and Diversion for Training Decomposable ModelsabstractPre-trained models have become the preferred backbone due to the increasing complexity of model parameters. However, traditional pre-trained models often face deployment challenges due to their fixed sizes, and are prone to negative transfer when discrepancies arise between training tasks and target tasks.
To address this, we propose **KIND**, a novel pre-training method designed to construct decomposable models.
KIND integrates knowledge by incorporating Singular Value Decomposition (SVD) as a structural constraint, with each basic component represented as a combination of a column vector, singular value, and row vector from $U$, $\Sigma$, and $V^\top$ matrices.
These components are categorized into **learngenes** for encapsulating class-agnostic knowledge and \textbf{tailors} for capturing class-specific knowledge, with knowledge diversion facilitated by a class gate mechanism during training.
Extensive experiments demonstrate that models pre-trained with KIND can be decomposed into learngenes and tailors, which can be adaptively recombined for diverse resource-constrained deployments.
Moreover, for tasks with large domain shifts, transferring only learngenes with task-agnostic knowledge, when combined with randomly initialized tailors, effectively mitigates domain shifts.
Code will be made available at https://github.com/Te4P0t/KIND. Yucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 0113, Yong Rui, Xin Geng 0001 |
ICML | 1 |
| 2025 | mmWave Testbed for Data Collection and Model Sharing in Contactless Concentration Monitoring SystemabstractMaintaining concentration is essential for productivity, learning and safety, yet it remains difficult to assess objectively in everyday settings. Traditional methods such as self-reporting and observational studies are subjective and labor-intensive. Wearable sensors can provide physiological data but require constant contact with the user, while camera-based systems raise privacy concerns and are sensitive to illumination and occlusion. Xiaonan Guo 0003, Yucheng Xie, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
SEC | 5 |
| 2025 | Exploring Cross-Environment modeling and Robustness in Palm-based User Authentication using mmWave TestbedabstractReliable and ubiquitous user authentication has become essential in smart cities, connected vehicles, and smart homes where users interact with multiple devices in their daily lives. However, existing biometric approaches, such as fingerprint, facial, or voice recognition, often require expensive hardware intrusive interaction, or raise privacy concerns, limiting their scalability in everyday settings [1–3]. To address these limitations, we explore a millimeter-wave (mmWave) testbed that enables palm-based user authentication through fine-grained sensing of palm geometry, skin thickness, and surface texture. By leveraging the widespread integration of mmWave technology in WiGig and 5G, this approach provides a low-cost, contactless, and privacy-preserving alternative to conventional biometrics. This work presents how the mmWave testbed is utilized to investigate cross-environment modeling and robustness in palm-based user authentication. Our system, named mmPalm, captures the reflections of Frequency-Modulated Continuous Wave (FMCW) signals from a user's palm to construct a distinctive palm profile that represents both structural and material characteristics of the hand. These reflections contain rich information about the three-dimensional geometry of the palm, sub-surface tissue variations, and fine surface textures, allowing unique identification without visual or physical contact. The mmWave testbed allows us to systematically collect palm data under varied distances, angles, and environments, providing a consistent platform for model development and evaluation. Yucheng Xie, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Tianfang Zhang, Yingying Chen 0001 |
SEC | 1 |
| 2025 | MCD-CLIP: Multi-view Chest Disease Diagnosis with Disentangled CLIPabstractPre-trained methods for multi-view chest X-ray images have demonstrated impressive performance in chest disease diagnosis, but there are still some limitations that need to be addressed. Firstly, many pre-trained methods require full fine-tuning pre-trained models to induce significant computational resource usage and the prior knowledge destruction. Secondly, many pre-trained methods cannot efficiently balance consistency and complementarity among views, leading to information loss and performance degradation. To tackle these issues, we propose MCD-CLIP, a CLIP-based multi-view chest disease diagnosis method. It uses visual prompts and a Prompt-Aligner to align prompts across views, along with the additional text representation for efficient transfer. Moreover, we employ Adapters to disentangle the image representation, maintaining consistency and complementarity from different views. Experimental results on the chest X-ray dataset demonstrate that MCD-CLIP achieves comparable or better performance on a variety of tasks with 94.31% fewer tunable parameters compared to state-of-the-art methods. The source codes are released at https://github.com/YuzunoKawori/MCD-CLIP. Songyue Cai, Yujie Mo, Yucheng Xie, Tao Tong, Xiaofeng Zhu 0001 |
IJCAI | 4 |
| 2025 | Meta Label Correction with Generalization RegularizerabstractDeep neural networks can easily lead to the over-fitting issue due to the influence of noisy labels. However, previous label correction methods for dealing with noisy labels often need expensive computation cost to achieve effectiveness and ignore the generalization ability of the model. To address these issues, in this paper, we propose a new meta-based self-correction method to achieve accurate filtering of noisy labels and to enhance the generalization ability of the label correction model. Specifically, we first investigate a new gradient score method to filter noisy labels with less computation cost, and then theoretically design a new generalization regularizer into the meta-learner and the base learner, for correcting noisy labels as well as achieving the generalization ability. Experimental results on real datasets verify the effectiveness of our proposed method in terms of different classification tasks. Tao Tong, Yujie Mo, Yucheng Xie, Songyue Cai, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
IJCAI | 3 |
| 2025 | Seeking Proxy Point via Stable Feature Space for Noisy Correspondence LearningabstractTo meet the growing demand for cross-modal training data, directly collecting multimodal data from the Internet has become prevalent. However, such data inevitably suffer from Noisy Correspondence. Previous works focused on recasting soft labels to mitigate noise's negative impact. We explore a novel perspective to solve this problem: pursuing proxy representation for noisy data to enable reliable feature learning. To this end, we propose a novel framework: Seeking Proxy Point via Stable Feature Space (SPS). This framework employs a fine-grained partitioning strategy to obtain a high-confidence reliable set. By imposing intermodal cross-transformation consistency constraints and intramodal metric consistency constraints, a stable feature space is constructed. Building on this foundation, SPS seeks proxy points for noisy data, enabling even noisy data to be accurately embedded into appropriate positions within the feature space. Combined with partial alignment for partially matched data pairs, SPS ultimately achieves robust learning under Noisy Correspondence. Experiments on three widely used cross-modal datasets demonstrate that SPS significantly outperforms previous methods. Our code is available at https://github.com/C-TeaRanger/SPS. Yucheng Xie, Songyue Cai, Tao Tong, Ping Hu 0001, Xiaofeng Zhu 0001 |
IJCAI | 1 |
| 2025 | RansomNet: Ransomware Classification Using Sub-Graph Mining of Function Call GraphsabstractThe classification of ransomware remains a critical yet challenging task in cybersecurity. Motivated by the increasing sophistication and overlap in behaviors between ransomware and general malware, this work addresses the need for more precise differentiation methods to facilitate targeted mitigation efforts. Our study proposes an innovative approach for ransomware classification using sub-graph mining of Function Call Graphs (FCGs). We employ Cuckoo Sandbox™ to extract dynamic API calls and construct detailed FCGs. Through focused subgraph mining, we isolate critical API call patterns specifically relevant to ransomware behavior. These extracted patterns are then vectorized and classified using a Convolutional Neural Network (RansomNet-CNN), achieving high precision in distinguishing ransomware from general malware. Unlike full-graph or flat-sequence models, our subgraph-level approach precisely captures ransomware-relevant behaviors. The RansomNet-CNN model demonstrates superior performance, achieving a precision of 99% and a recall of 100%, thus underscoring its practical effectiveness in ransomware identification. The dataset and code are publicly available at our Zenodo Repository1. Garvit Agarwal, Yucheng Xie, Yousef Mohammed Y. Alomayri, Feng Li 0001 |
MASS | 2 |
| 2025 | ECO: Evolving Core Knowledge for Efficient TransferabstractKnowledge in modern neural networks is often entangled and structurally opaque, making current transfer methods—typically based on reusing entire parameter sets—inefficient and inflexible. Efforts to improve flexibility by reusing partial parameters frequently depend on handcrafted heuristics or rigid structural assumptions, which constrain generalization. In contrast, biological evolution enables efficient knowledge transfer by encoding only essential information into genes through iterative refinement under environmental pressure. Inspired by this principle, we propose **ECO**, a framework that **E**volves **CO**re knowledge into modular, reusable neural components—termed *learngenes*—through similar evolutionary dynamics. To this end, we redefine learngenes as neural circuits and introduce Genetic Transfer Learning (GTL), a biologically inspired paradigm that establishes a genetic mechanism within neural networks in the context of supervised learning. GTL simulates evolutionary processes by generating diverse network populations, selecting high-performing individuals, and transferring their learngenes to subsequent generations. Through iterative refinement, GTL enables learngenes to accumulate transferable common knowledge. Extensive experiments show that ECO achieves efficient initialization and strong generalization across diverse models and tasks, while significantly reducing computational and memory costs compared to conventional methods. Fu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen, Jing Wang 0113, Xin Geng 0001 |
NeurIPS | 2 |
| 2025 | RankMatch: A Novel Approach to Semi-Supervised Label Distribution Learning Leveraging Rank Correlation between LabelsabstractPseudo label based semi-supervised learning (SSL) for single-label and multi-label classification tasks has been extensively studied; however, semi-supervised label distribution learning (SSLDL) remains a largely unexplored area. Existing SSL methods fail in SSLDL because the pseudo-labels they generate only ensure overall similarity to the ground truth but do not preserve the ranking relationships between true labels, as they rely solely on KL divergence as the loss function during training. These skewed pseudo-labels lead the model to learn incorrect semantic relationships, resulting in reduced performance accuracy. To address these issues, we propose a novel SSLDL method called \textit{RankMatch}. \textit{RankMatch} fully considers the ranking relationships between different labels during the training phase with labeled data to generate higher-quality pseudo-labels. Furthermore, our key observation is that a flexible utilization of pseudo-labels can enhance SSLDL performance. Specifically, focusing solely on the ranking relationships between labels while disregarding their margins helps prevent model overfitting. Theoretically, we prove that incorporating ranking correlations enhances SSLDL performance and establish generalization error bounds for \textit{RankMatch}. Finally, extensive real-world experiments validate its effectiveness. Zhiqiang Kou, Yucheng Xie, Hailin Wang 0001, Jing Wang 0113, Ming-Kun Xie, Shuo Chen 0003, Yuheng Jia, Tongliang Liu, Xin Geng 0001 |
NeurIPS | 2 |
| 2024 | Palm-Based User Authentication Through mmWaveabstractBiometric authentication systems are increasingly needed across a broad range of applications including in smart city environments (e.g., entering hotels, high-rise buildings, train stations, hospitals, and personalizing vehicles settings), and in smart home environments (e.g., controlling smart devices, en-hancing VR/AR experience). Traditional methods, such as face-based and fingerprint-based authentication, usually incur high cost to be installed in all this kind of environments, making them hard to become a ubiquitous authentication approach. In this paper, we develop a ubiquitous low-effort user authentication approach based on palm recognition using millimeter wave (mmWave) signals. Extensive experiments demonstrate that our system achieves 99% authentication accuracy. Yucheng Xie, Tianfang Zhang, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
ICDCS | 1 |
| 2023 | Secure and Efficient Mobile DNN Using Trusted Execution EnvironmentsabstractMany mobile applications have resorted to deep neural networks (DNNs) because of their strong inference capabilities. Since both input data and DNN architectures could be sensitive, there is an increasing demand for secure DNN execution on mobile devices. Towards this end, hardware-based trusted execution environments on mobile devices (mobile TEEs), such as ARM TrustZone, have recently been exploited to execute CNN securely. However, running entire DNNs on mobile TEEs is challenging as TEEs have stringent resource and performance constraints. In this work, we develop a novel mobile TEE-based security framework that can efficiently execute the entire DNN in a resource-constrained mobile TEE with minimal inference time overhead. Specifically, we propose a progressive pruning to gradually identify and remove the redundant neurons from a DNN while maintaining a high inference accuracy. Next, we develop a memory optimization method to deallocate the memory storage of the pruned neurons utilizing the low-level programming technique. Finally, we devise a novel adaptive partitioning method that divides the pruned model into multiple partitions according to the available memory in the mobile TEE and loads the partitions into the mobile TEE separately with a minimal loading time overhead. Our experiments with various DNNs and open-source datasets demonstrate that we can achieve 2-30 times less inference time with comparable accuracy compared to existing approaches securing entire DNNs with mobile TEE. Bin Hu 0016, Yan Wang 0003, Jerry Q. Cheng, Tianming Zhao 0001, Yucheng Xie, Xiaonan Guo 0003, Yingying Chen 0001 |
AsiaCCS | 5 |
| 2023 | Universal Targeted Adversarial Attacks Against mmWave-based Human Activity Recognition
Yucheng Xie, Ruizhe Jiang, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
INFOCOM | 1 |
| 2022 | mmFit: Low-Effort Personalized Fitness Monitoring Using Millimeter WaveabstractThere is a growing trend for people to perform work-outs at home due to the global pandemic of COVID-19 and the stay-at-home policy of many countries. Since a self-designed fitness plan often lacks professional guidance to achieve ideal outcomes, it is important to have an in-home fitness monitoring system that can track the exercise process of users. Traditional camera-based fitness monitoring may raise serious privacy concerns, while sensor-based methods require users to wear dedicated devices. Recently, researchers propose to utilize RF signals to enable non-intrusive fitness monitoring, but these approaches all require huge training efforts from users to achieve a satisfactory performance, especially when the system is used by multiple users (e.g., family members). In this work, we design and implement a fitness monitoring system using a single COTS mm Wave device. The proposed system integrates workout recognition, user identification, multi-user monitoring, and training effort reduction modules and makes them work together in a single system. In particular, we develop a domain adaptation framework to reduce the amount of training data collected from different domains via mitigating impacts caused by domain characteristics embedded in mm Wave signals. We also develop a GAN-assisted method to achieve better user identification and workout recognition when only limited training data from the same domain is available. We propose a unique spatialtemporal heatmap feature to achieve personalized workout recognition and develop a clustering-based method for concurrent workout monitoring. Extensive experiments with 14 typical workouts involving 11 participants demonstrate that our system can achieve 97% average workout recognition accuracy and 91% user identification accuracy. Yucheng Xie, Ruizhe Jiang, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
ICCCN | 1 |
| 2022 | Universal targeted attacks against mmWave-based human activity recognition systemabstractMillimeter wave (mmWave)-based human activity recognition (HAR) systems have emerged in recent years due to their better privacy preservation and higher-resolution sensing. However, these systems are vulnerable to adversarial attacks. In this work, we propose a universal targeted attack method for mmWave-based HAR system. In particular, a universal perturbation is generated in advance which can be added to new-coming mmWave data to deceive the HAR system, causing it to output our desired label. We validate our proposed attack using a public mmWave dataset. We demonstrate the effectiveness of our proposed universal attack with a high attack success rate of over 95%. Yucheng Xie, Ruizhe Jiang, Xiaonan Guo 0003, Yan Wang 0003, Jerry Q. Cheng, Yingying Chen 0001 |
MobiSys | 1 |
| 2022 | A Survey of Deep Learning on Mobile Devices: Applications, Optimizations, Challenges, and Research OpportunitiesabstractDeep learning (DL) has demonstrated great performance in various applications on powerful computers and servers. Recently, with the advancement of more powerful mobile devices (e.g., smartphones and touch pads), researchers are seeking DL solutions that could be deployed on mobile devices. Compared to traditional DL solutions using cloud servers, deploying DL on mobile devices have unique advantages in data privacy, communication overhead, and system cost. This article provides a comprehensive survey for the current studies of adopting and deploying DL on mobile devices. Specifically, we summarize and compare the state-of-the-art DL techniques on mobile devices in various application domains involving vision, speech/speaker recognition, human activity recognition, transportation mode detection, and security. We generalize an optimization pipeline for bringing DL to mobile devices, including model-oriented optimization mechanisms (e.g., pruning and quantization) and nonmodel-oriented optimization mechanisms (e.g., software accelerator and hardware design). Moreover, we summarize popular DL libraries regarding their support to state-of-the-art models (software) and processors (hardware). Based on our summarization, we further provide insights into potential research opportunities for developing DL for mobile devices. Tianming Zhao 0001, Yucheng Xie, Yan Wang 0003, Jerry Q. Cheng, Xiaonan Guo 0003, Bin Hu 0016, Yingying Chen 0001 |
Proc. IEEE | 2 |
| 2022 | Learning semantic alignment from image for text-guided image inpainting
Yucheng Xie, Zehang Lin, Zhenguo Yang, Xingcai Wu, Xudong Mao, Qing Li 0001, Wenyin Liu |
Vis. Comput. | 1 |
| 2021 | MIXP: Efficient Deep Neural Networks Pruning for Further FLOPs Compression via Neuron BondabstractNeuron networks pruning is effective in compressing pre-trained CNNs for their deployment on low-end edge devices. However, few works have focused on reducing the computational cost of pruning and inference. We find that existing pruning methods usually remove parameters without fine-grained impact analysis, making it hard to achieve an optimal solution. This work develops a novel mixture pruning mechanism, MIXP, which can effectively reduce the computational cost of CNNs while maintaining a high weight compression ratio and model accuracy. We propose to remove neuron bond that can effectively reduce convolution computations and weight size in CNNs. We also design an influence factor to analyze the importance of neuron bonds and weights in a fine-grained way so that MIXP could achieve precise pruning with few retraining iterations. Experiments with MNIST, CIFAR-10, and ImageNet datasets demonstrate that MIXP could achieve significantly fewer FLOPs and retraining iterations on four widely-used CNNs than existing pruning methods. Bin Hu 0016, Tianming Zhao 0001, Yucheng Xie, Yan Wang 0003, Xiaonan Guo 0003, Jerry Q. Cheng, Yingying Chen 0001 |
IJCNN | 3 |
| 2021 | Environment-independent In-baggage Object Identification Using WiFi SignalsabstractLow-cost in-baggage object identification is highly demanded in enhancing public safety and smart manufacturing. Existing approaches usually require specialized equipment and heavy deployment overhead, making them hard to scale for wide deployment. The recent WiFi-based approach is unsuitable for practical deployment as it did not address dynamic environmental impacts. In this work, we propose an environment-independent in-baggage object identification system by leveraging low-cost WiFi. We exploit the channel state information (CSI) to capture material and shape characteristics to facilitate fine-grained inbaggage object identification. A major challenge of building such a system is that CSI measurements are sensitive to real-world dynamics, such as different types of baggage, time-varying ambient noises and interferences, and different deployment environments. To tackle these problems, we develop WiFi features based on polarized directional antennas that can capture objects’ material and shape characteristics. A convolutional neural network-based model is developed to constructively integrate the WiFi features and perform accurate in-baggage object identification. We also develop a material-based domain adaptation using adversarial learning to facilitate fast deployments in different environments. We conduct extensive experiments involving 14 representation objects, 4 types of bags in 3 different room environments. The results show that our system can achieve over 97% in the same environment, and our domain adaptation method can improve the object identification accuracy by 42% when the system is deployed in a new environment with little training. Cong Shi 0004, Tianming Zhao 0001, Yucheng Xie, Tianfang Zhang, Yan Wang 0003, Xiaonan Guo 0003, Yingying Chen 0001 |
MASS | 3 |
| 2021 | Adversarial Learning with Mask Reconstruction for Text-Guided Image InpaintingabstractText-guided image inpainting aims to complete the corrupted patches coherent with both visual and textual context. On one hand, existing works focus on surrounding pixels of the corrupted patches without considering the objects in the image, resulting in the characteristics of objects described in text being painted on non-object regions. On the other hand, the redundant information in text may distract the generation of objects of interest in the restored image. In this paper, we propose an adversarial learning framework with mask reconstruction (ALMR) for image inpainting with textual guidance, which consists of a two-stage generator and dual discriminators. The two-stage generator aims to restore coarse-grained and fine-grained images, respectively. In particular, we devise a dual-attention module (DAM) to incorporate the word-level and sentence-level textual features as guidance on generating the coarse-grained and fine-grained details in the two stages. Furthermore, we design a mask reconstruction module (MRM) to penalize the restoration of the objects of interest with the given textual descriptions about the objects. For adversarial training, we exploit global and local discriminators for the whole image and corrupted patches, respectively. Extensive experiments conducted on CUB-200-2011, Oxford-102 and CelebA-HQ show the outperformance of the proposed ALMR (e.g., FID value is reduced from 29.69 to 14.69 compared with the state-of-the-art approach on CUB-200-2011). Codes are available at https://github.com/GaranWu/ALMR Xingcai Wu, Yucheng Xie, Jiaqi Zeng, Zhenguo Yang, Yi Yu 0001, Qing Li 0001, Wenyin Liu |
ACM Multimedia | 2 |
| 2020 | Mobile Device Usage Recommendation based on User Context Inference Using Embedded SensorsabstractThe proliferation of mobile devices along with their rich functionalities/applications have made people form addictive and potentially harmful usage behaviors. Though this problem has drawn considerable attention, existing solutions (e.g., text notification or setting usage limits) are insufficient and cannot provide timely recommendations or control of inappropriate usage of mobile devices. This paper proposes a generalized context inference framework, which supports timely usage recommendations using low-power sensors in mobile devices Comparing to existing schemes that rely on detection of single type user contexts (e.g., merely on location or activity), our framework derives a much larger-scale of user contexts that characterize the phone usages, especially those causing distraction or leading to dangerous situations. We propose to uniformly describe the general user context with context fundamentals, i.e., physical environments, social situations, and human motions, which are the underlying constituent units of diverse general user contexts. To mitigate the profiling efforts across different environments, devices, and individuals, we develop a deep learning-based architecture to learn transferable representations derived from sensor readings associated with the context fundamentals. Based on the derived context fundamentals, our framework quantifies how likely an inferred user context would lead to distractions/dangerous situations, and provides timely recommendations for mobile device access/usage. Extensive experiments during a period of 7 months demonstrate that the system can achieve 95% accuracy on user context inference while offering the transferability among different environments, devices, and users. Cong Shi 0004, Xiaonan Guo 0003, Ting Yu 0001, Yingying Chen 0001, Yucheng Xie, Jian Liu 0001 |
ICCCN | 5 |
| 2020 | WiEat: Fine-grained Device-free Eating Monitoring Leveraging Wi-Fi SignalsabstractEating well plays a key role in people's overall health and wellbeing. Studies have shown that many health-related problems such as obesity, diabetes and anemia are closely associated with people's unhealthy eating habits (e.g., skipping meals, eating irregularly and overeating). Thus, keeping track of diet is becoming more important. Traditional eating monitoring solutions relying on self-report remain an onerous task, while the recent trends requiring users to wear dedicated yet expensive hardware are cumbersome. To overcome these limitations, in this paper, we develop a device-free eating monitoring system using WiFi-enabled devices (e.g., smartphone or laptop). Our system aims to automatically monitor users' eating activities by identifying the fine-grained eating motions and detecting the minute movements during chewing and swallowing. In particular, our system distinguishes eating from non-eating activities by using K-means clustering with principal component analysis on the extracted Channel State Information (CSI) from WiFi signals. It further adopts a soft decision-based eating motion classification through identifying the utensils (e.g., using a folk, knife, spoon or bare hands) in use. Moreover, we propose a minute motion reconstruction method to identify chewing and swallowing through detecting users' minute facial muscle movements. The derived fine-grained eating monitoring results are beneficial to the understanding of users' eating behaviors and estimation of food intake types and amounts. Extensive experiments with 20 users over 1600-minute eating show that the proposed system can recognize the user's eating motions with up to 95% accuracy and estimate the chewing and swallowing amount within 10% percentage error. Zhenzhe Lin, Yucheng Xie, Xiaonan Guo 0003, Yanzhi Ren, Yingying Chen 0001, Chen Wang 0009 |
ICCCN | 2 |
| 2020 | LiveScreen: Video Chat Liveness Detection Leveraging Skin ReflectionabstractThe rapid advancement of social media and communication technology enables video chat to become an important and convenient way of daily communication. However, such convenience also makes personal video clips easily obtained and exploited by malicious users who launch scam attacks. Existing studies only deal with the attacks that use fabricated facial masks, while the liveness detection that targets the playback attacks using a virtual camera is still elusive. In this work, we develop a novel video chat liveness detection system, LiveScreen, which can track the weak light changes reflected off the skin of a human face leveraging chromatic eigenspace differences. We design an inconspicuous challenge frame with minimal intervention to the video chat and develop a robust anomaly frame detector to verify the liveness of the remote user in the video chat using the response to the challenge frame. Furthermore, we propose resilient defense strategies to defeat both naive and intelligent playback attacks leveraging spatial and temporal verification. We implemented a prototype over both laptop and smartphone platforms and conducted extensive experiments in various realistic scenarios. We show that our system can achieve robust liveness detection with accuracy and false detection rates 97.7% (94.8%) and 1% (1.6%) on smartphones (laptops), respectively. Hongbo Liu 0002, Yucheng Xie, Ruizhe Jiang, Yan Wang 0003, Xiaonan Guo 0003, Yingying Chen 0001 |
INFOCOM | 3 |
| 2020 | MU-ID: Multi-user Identification Through Gaits Using Millimeter Wave RadiosabstractMulti-user identification could facilitate various large-scale identity-based services such as access control, automatic surveillance system, and personalized services, etc. Although existing solutions can identify multiple users using cameras, such vision-based approaches usually raise serious privacy concerns and require the presence of line-of-sight. Differently, in this paper, we propose MU-ID, a gait-based multi-user identification system leveraging a single commercial off-the-shelf (COTS) millimeter-wave (mmWave) radar. Particularly, MU-ID takes as input frequency-modulated continuous-wave (FMCW) signals from the radar sensor. Through analyzing the mmWave signals in the range-Doppler domain, MU-ID examines the users' lower limb movements and captures their distinct gait patterns varying in terms of step length, duration, instantaneous lower limb velocity, and inter-lower limb distance, etc. Additionally, an effective spatial-temporal silhouette analysis is proposed to segment each user's walking steps. Then, the system identifies steps using a Convolutional Neural Network (CNN) classifier and further identifies the users in the area of interest. We implement MU-ID with the TI AWR1642BOOST mmWave sensor and conduct extensive experiments involving 10 people. The results show that MU-ID achieves up to 97% single-person identification accuracy, and over 92% identification accuracy for up to four people, while maintaining a low false positive rate. Jian Liu 0001, Yingying Chen 0001, Xiaonan Guo 0003, Yucheng Xie |
INFOCOM | 5 |
| 2019 | Poster: Video Chat Scam Detection Leveraging Screen Light ReflectionabstractThe rapid advancement of social media and communication technology enables video chat to become an important and convenient way of daily communication. However, such convenience also makes personal video clips easily obtained and exploited by malicious users who launch scam attacks. Existing studies only deal with the attacks that use fabricated facial masks, while the liveness detection that targets the playback attacks using a virtual camera is still elusive. In this work, we develop a novel video chat liveness detection system, which can track the weak light changes reflected off the skin of a human face leveraging chromatic eigenspace differences. We design an inconspicuous challenge frame with minimal intervention to the video chat and develop a robust anomaly frame detector to verify the liveness of remote user in a video chat session. Furthermore, we propose a resilient defense strategy to defeat both naive and intelligent playback attacks leveraging spatial and temporal verification. The evaluation results show that our system can achieve accurate and robust liveness detection with the accuracy and false detection rate as high as 97.7% (94.8%) and 1% (1.6%) on smartphones (laptops), respectively. Hongbo Liu 0002, Yucheng Xie, Ruizhe Jiang, Yan Wang 0003, Xiaonan Guo 0003, Yingying Chen 0001 |
MobiCom | 3 |