VLDB 2026 Research / reviewers in the wild / expert
Zongxing Xie
dblp:79/6475
· DBLP profile ↗
13ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-3464-1610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ADL-CLIP: Text-Aligned RF Representation for Continuous ADL Detection in Home EnvironmentsabstractPatterns of activities of daily living (ADLs) can provide early indicators of changes for many diseases and health conditions. Radio frequency (RF) sensing is promising for its contactless and privacy-preserving properties. However, most RF-based human activity recognition (HAR) studies focus on short, well-segmented actions, and their applicability to continuous, longer-duration ADLs remains largely uncertain. Existing continuous activity detection methods are typically trained in a closed set with labeled boundaries, limiting their ability to handle unseen, unsegmented ADLs in real life. In this work, we address both challenges by introducing ADL-CLIP, a framework for RF-based continuous ADL detection in home environments without pre-known boundaries and generalizable to unseen activities. We observe that intra-ADL radio frames cluster closely in the representation space, while those across ADLs are more separated. Building on this insight, we formulate ADL segmentation as a label-free optimization problem that identifies the boundaries by minimizing intra-segment while maximizing across-segment distances. To reconcile the discrepancy between data-driven segmentation and human annotated boundaries, we alternate between refining segmentation boundaries in the representation space and fine-tuning the representation model combining both boundaries using CLIP-style contrastive learning, progressively aligning the two while preserving prior knowledge from language models. We evaluate ADL-CLIP on a dataset collected from 46 non-researcher participants performing 25 ADLs continuously with 16 UWB sensors in an instrumented one-bedroom, one-bathroom apartment. In an open-set setting where 5 ADLs are held out during training, ADL-CLIP achieves 76.6% accuracy on all 25 ADLs, outperforming temporal activity detection baselines by 26% and NLS-based HAR by 34%. To our knowledge, this is the first RF-based open-set, boundary-free ADL detection study in an instrumented home environment. Mengjing Liu, Zongxing Xie, Fan Ye 0003 |
MobiSys | 2 |
| 2025 | Language-Agnostic Speech Biomarker Exploration for Early Dementia ScreeningabstractEarly dementia detection is a global healthcare priority in diverse populations. In this study, we propose a language-agnostic screening pipeline for dementia detection in the early stage. First, we use speaker diarization to isolate the speech of the target subject from a conversational recording. From the extracted speech segments, we derive a set of acoustic features (e.g., spectral centroid, pitch mean, mel-frequency cepstral coefficients) and linguistic features (e.g., normalized tone contrast, articulation clarity coefficient, articulatory effort coefficient). These features are used to train a ResNet-based binary classifier to distinguish between Healthy Controls (HC) and individuals with Mild Cognitive Impairment (MCI). We evaluated the trained model on a held-out test set comprising speakers of previously unseen languages, achieving an accuracy of 70%. This cross-lingual transfer performance highlights the potential of our approach for scalable, language-independent dementia screening. Josh Ashik, Zongxing Xie, Chenhan Xu, Huining Li |
BSN | 2 |
| 2025 | Proteus: An Easily Managed Home-Based Health Monitoring InfrastructureabstractA data collection infrastructure is vital for generating sufficient amounts and diversity of data necessary for developing algorithms in home-based health monitoring. However, the manageability—deployment and operation efforts—of such an infrastructure has long been overlooked. Even a small size of a dozen homes may incur enormous manual efforts on the research team. In this article, we present Proteus, an easily managed infrastructure designed to automate much of the work in deploying and operating such systems. We develop new components and combine with mature technologies to minimize the human efforts required. Proteus includes: 1) scalable, continuous deployment, operation, and update of devices with automatic bootstrapping; 2) automatic fault and error monitoring and recovery with watchdogs and LED feedback, and complementary edge and cloud storage backups; and 3) an easy-to-use data-agnostic pipeline for integrating new modalities. We demonstrate our system’s robustness through different sets of experiments: three sensor nodes (SNs) running for 24 days sending data (17.4 Mb/s aggregate rate), 10 SNs for 14 days (58 Mb/s aggregate rate), and 32 emulated sensors (419.2 Mb/s aggregate rate). All such experiments have data loss rates less than 1%. Further we reduce human efforts by 25-fold and code required for adding new data modality by 25-fold. We also share our experience and lessons learned during the design, development, and pilot deployment of Proteus. Our results show that Proteus is a promising solution for enabling research teams to effectively manage home-based health monitoring at small to medium sizes. Mengjing Liu, Mohammed Elbadry, Yindong Hua, Zongxing Xie, Suvab Baral, Isac Park, Fan Ye 0003 |
IEEE Internet Things J. | 4 |
| 2023 | HeartInsightify: Interpreting Longitudinal Heart Rate Data for Health Insights through Conformal ClusteringabstractHeart rate, a commonly accessible health data from most wearables, carries rich information of a person’s well-being, yet remains of limited deep health applications, due to the lack of groundtruth of health events and their impact on heart rate patterns. Specifically, standard health analytics usually are designed based on well-modeled health conditions thus known data patterns and rich training data. To bridge the gap, we propose HeartInsightify, an exploratory framework that facilitates the process of deriving health-relevant measurable indicators from longitudinal heart rate data, without any of the above knowledge. HeartInsightify focuses on comparative and qualitative study, using model-free statistical methods such as conformal prediction, to study similarities, perform clustering and detect outliers, and build multi-resolutional data summaries, allowing human experts to efficiently examine and verify their health relevance. We conduct extensive experiments to evaluate HeartInsightify using individuals’ free-living heart rate data collected through Fitbit over 6 years. We illustrate the process of analyzing heart rate data for its health relevance and demonstrate the effectiveness of HeartInsightify. We envision that HeartInsightify lays the groundwork for personalized health analytics with continuous monitoring data from wearables. Prathamesh Dharangutte, Zongxing Xie, Jie Gao 0001, Elinor Schoenfeld, Yindong Hua, Fan Ye 0003 |
BIBM | 2 |
| 2022 | Scaling Device-Free Indoor Tracking Based on Self CalibrationabstractThe democratization of indoor tracking systems lays the groundwork for a wide spectrum of smart home applications. While prior work on RF-based device-free localization/tracking have shown preferable features and promising results, they heavily relied on well-calibrated sensor placements, which require hours of intensive manual setup and respective expertise, making it prohibitively expensive to scale deployments to wide range (e.g., tens or hundreds of real homes). We propose SCALING, a plug-and-play indoor tracking system, of which the key enabler is a self calibrating algorithm that estimates the distributed sensor locations through their distance measurements to a person walking a trajectory, a trivial effort without taxing layman users physically or cognitively. We have experimentally evaluated SCALING via real world testbeds and shown an 80-percentile tracking accuracy of 40.5 cm, only 1% degradation compared to the classical multilateration with known sensor locations (anchors), which costs hours of intensive calibrating effort. Zongxing Xie, Fan Ye 0003 |
SenSys | 1 |
| 2022 | Passive and Context-Aware In-Home Vital Signs Monitoring Using Co-Located UWB-Depth Sensor FusionabstractBasic vital signs such as heart and respiratory rates (HR and RR) are essential bio-indicators. Their longitudinal in-home collection enables prediction and detection of disease onset and change, providing for earlier health intervention. In this article, we propose a robust, non-touch vital signs monitoring system using a pair of co-located Ultra-Wide Band (UWB) and depth sensors. By extensive manual examination, we identify four typical temporal and spectral signal patterns and their suitable vital sign estimators. We devise a probabilistic weighted framework (PWF) that quantifies evidence of these patterns to update the weighted combination of estimator output to track the vital signs robustly. We also design a “heatmap”-based signal quality detector to exclude the disturbed signal from inadvertent motions. To monitor multiple co-habiting subjects in-home, we build a two-branch long short-term memory (LSTM) neural network to distinguish between individuals and their activities, providing activity context crucial to disambiguating critical from normal vital sign variability. To achieve reliable context annotation, we carefully devise the feature set of the consecutive skeletal poses from the depth data, and develop a probabilistic tracking model to tackle non-line-of-sight (NLOS) cases. Our experimental results demonstrate the robustness and superior performance of the individual modules as well as the end-to-end system for passive and context-aware vital sign monitoring. Zongxing Xie, Bing Zhou 0001, Elinor Schoenfeld, Fan Ye 0003 |
ACM Trans. Comput. Heal. | 1 |
| 2022 | Robust Human Face Authentication Leveraging Acoustic Sensing on SmartphonesabstractUser authentication on smartphones is the key to many applications, which must satisfy both security and convenience. We propose a novel user authentication systemEchoPrint, which leverages acoustics and vision for secure and convenient user authentication, without requiring any special hardware.EchoPrintactively emits almost inaudible acoustic signals from the earpiece speaker to “illuminate” the user's face and authenticates the user by the unique features extracted from the echoes bouncing off the 3D facial contour. To combat changes in phone-holding poses thus echoes, a convolutional neural network (CNN) is trained to extract reliable acoustic features, which are further combined with visual facial features extracted from state-of-the-art face recognition deep models to feed a binary support vector machine (SVM) classifier for final authentication. Because the echo features depend on 3D facial geometries,EchoPrintis not easily spoofed by images or videos like 2D visual face recognition systems. It needs only commodity hardware, thus avoiding the extra costs of special sensors in solutions like FaceID. Experiments with 62 volunteers and non-human objects such as images, photos, and sculptures show thatEchoPrintachieves 93.75 percent balanced accuracy and 93.50 percent F-score, while the average precision is 98.05 percent using acoustic features and basic facial landmarks. The precision is further improved to 99.96 percent with sophisticated visual features. Bing Zhou 0001, Zongxing Xie, Jay Lohokare, Ruipeng Gao, Fan Ye 0003 |
IEEE Trans. Mob. Comput. | 2 |
| 2019 | Multi-Modal Face Authentication using Deep Visual and Acoustic FeaturesabstractUser authentication on smartphones is the key to many applications, which must satisfy both security and convenience. We propose a multi-modal face authentication system, which pushes the limit of state-of-the-art image based face recognition solutions by incorporating a new dimension of sensing modality - acoustics. It actively emits almost inaudible acoustic signals from the earpiece speaker to "illuminate" the user's face and extracts features from the echoes using a customized convolutional neural network, which are fused with sophisticated visual features extracted from state-of-the-art face recognition models, for secure face authentication. Because the echo features depend on 3D facial geometries and material, our multi-modal design is not easily spoofed by images or videos like image based face recognition systems. It does not require any special sensors thus eliminating the extra costs in solutions like FaceID. Experiments show that our design achieves comparable face recognition performance to the state-of-the-art image based face authentication, while able to block image/video spoofing. Bing Zhou 0001, Zongxing Xie, Fan Ye 0003 |
ICC | 2 |
| 2008 | Video Semantic Event/Concept Detection Using a Subspace-Based Multimedia Data Mining FrameworkabstractIn this paper, a subspace-based multimedia data mining framework is proposed for video semantic analysis, specifically video event/concept detection, by addressing two basic issues, i.e.,semantic gapandrare event/concept detection. The proposed framework achieves full automation via multimodal content analysis and intelligent integration of distance-based and rule-based data mining techniques. The content analysis process facilitates the comprehensive video analysis by extracting low-level and middle-level features from audio/visual channels. The integrated data mining techniques effectively address these two basic issues by alleviating the class imbalance issue along the process and by reconstructing and refining the feature dimension automatically. The promising experimental performance on goal/corner event detection and sports/commercials/building concepts extraction from soccer videos and TRECVID news collections demonstrates the effectiveness of the proposed framework. Furthermore, its unique domain-free characteristic indicates the great potential of extending the proposed multimedia data mining framework to a wide range of different application domains. Mei-Ling Shyu, Zongxing Xie, Min Chen 0009, Shu-Ching Chen |
IEEE Trans. Multim. | 2 |
| 2007 | Video Event Detection with Combined Distance-Based and Rule-Based Data Mining TechniquesabstractIn this paper, the rare event detection issue in video event detection is addressed through the proposed data mining framework which can be generalized to be domain independent. The fully automatic process via the combination of distance-based and rule-based data mining techniques can greatly reduce the number of negative (non-event) instances and the feature dimension to facilitate the final event detection, without pruning away any positive (event) testing instance along the process. The effectiveness and efficiency of the proposed framework are demonstrated over the goal event detection application based on a large collection of soccer videos with different styles. Zongxing Xie, Mei-Ling Shyu, Shu-Ching Chen |
ICME | 1 |
| 2007 | Network intrusion detection through Adaptive Sub-Eigenspace Modeling in multiagent systemsabstractRecently, network security has become an extremely vital issue that beckons the development of accurate and efficient solutions capable of effectively defending our network systems and the valuable information journeying through them. In this article, a distributed multiagent intrusion detection system (IDS) architecture is proposed, which attempts to provide an accurate and lightweight solution to network intrusion detection by tackling issues associated with the design of a distributed multiagent system, such as poor system scalability and the requirements of excessive processing power and memory storage. The proposed IDS architecture consists of (i) the Host layer with lightweight host agents that perform anomaly detection in network connections to their respective hosts, and (ii) the Classification layer whose main functions are to perform misuse detection for the host agents, detect distributed attacks, and disseminate network security status information to the whole network. The intrusion detection task is achieved through the employment of the lightweight Adaptive Sub-Eigenspace Modeling (ASEM)-based anomaly and misuse detection schemes. Promising experimental results indicate that ASEM-based schemes outperform the KNN and LOF algorithms, with high detection rates and low false alarm rates in the anomaly detection task, and outperform several well-known supervised classification methods such as C4.5 Decision Tree, SVM, NN, KNN, Logistic, and Decision Table (DT) in the misuse detection task. To assess the performance in a real-world scenario, the Relative Assumption Model, feature extraction techniques, and common network attack generation tools are employed to generate normal and anomalous traffic in a private LAN testbed. Furthermore, the scalability performance of the proposed IDS architecture is investigated through the simulation of the proposed agent communication scheme, and satisfactory linear relationships for both degradation of system response time and agent communication generated network traffic overhead are achieved. Mei-Ling Shyu, Thiago Quirino, Zongxing Xie, Shu-Ching Chen, LiWu Chang |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2006 | Collateral Representative Subspace Projection Modeling for Supervised ClassificationabstractIn this paper, a novel supervised classification approach called collateral representative subspace projection modeling (C-RSPM) is presented. C-RSPM facilitates schemes for collateral class modeling, class-ambiguity solving, and classification, resulting a multi-class supervised classifier with high detection rate and various operational benefits including low training and classification times and low processing power and memory requirements. In addition, C-RSPM is capable of adaptively selecting nonconsecutive principal dimensions from the statistical information of the training data set to achieve an accurate modeling of a representative subspace. Experimental results have shown that the proposed C-RSPM approach outperforms other supervised classification methods such as SIMCA, C4.5 decision tree, decision table (DT), nearest neighbor (NN), KNN, support vector machine (SVM), I-NN best warping window DTW, I-NN DTW with no warping window, and the well-known classifier boosting method AdaBoost with SVM Thiago Quirino, Zongxing Xie, Mei-Ling Shyu, Shu-Ching Chen, LiWu Chang |
ICTAI | 2 |
| 2006 | UNPCC: A Novel Unsupervised Classification Scheme for Network Intrusion DetectionabstractThe development of effective classification techniques, particularly unsupervised classification, is important for real-world applications since information about the training data before classification is relatively unknown. In this paper, a novel unsupervised classification algorithm is proposed to meet the increasing demand in the domain of network intrusion detection. Our proposed UNPCC (unsupervised principal component classifier) algorithm is a multiclass unsupervised classifier with absolutely no requirements for any a priori class related data information (e.g., the number of classes and the maximum number of instances belonging to each class), and an inherently natural supervised classification scheme, both which present high detection rates and several operational advantages (e.g., lower training time, lower classification time, lower processing power requirement, and lower memory requirement). Experiments have been conducted with the KDD Cup 99 data and network traffic data simulated from our private network testbed, and the promising results demonstrate that our UNPCC algorithm outperforms several well-known supervised and unsupervised classification algorithms Zongxing Xie, Thiago Quirino, Mei-Ling Shyu, Shu-Ching Chen, LiWu Chang |
ICTAI | 1 |