VLDB 2026 Research / reviewers in the wild / expert
Tiantian Feng
dblp:121/1188
· DBLP profile ↗
57ranked-venue papers
20as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 26 since 2021Artificial intelligence and machine learning · 24 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease IdentificationabstractRespiratory diseases remain a leading cause of global mortality, where timely and accurate diagnosis is critical to improving patient outcomes and reducing healthcare burdens.While prior work has explored audio-based models for respiratory disease detection, such unimodal approaches often suffer from limited generalizability and diagnostic precision.In this paper, we propose RespiraMFM, a Multimodal Foundation Model that integrates respiratory sounds with patient medical history and symptoms to enhance diagnostic accuracy and disease detection capabilities.We introduce an effective contrastive alignment strategy for audio-text multimodal integration, allowing the model to learn better cross-modal representations between respiratory sounds and corresponding textual clinical information.We evaluate RespiraMFM across five major respiratory diseases using seven real-world datasets in both supervised fine-tuning and zero-shot settings, achieving a 9.15% improvement in AU-ROC on supervised tasks and a 20.98% gain on zero-shot tasks over existing baselines.These findings underscore the potential of our framework to advance early diagnosis and improve clinical decision-making in respiratory disease management. Shakhrul Iman Siam, Tiantian Feng, Shri Narayanan, Mi Zhang 0002 |
ACL (1) | 2 |
| 2026 | Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the GlobeabstractWe present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the RAIL license at https://github.com/tiantiaf0627/voxlect. Tiantian Feng, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun, Yoonjeong Lee, Dani Byrd, Shri Narayanan |
KDD (1) | 1 |
| 2026 | Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model
Xuan Shi, Tiantian Feng, Jay Park, Christina Hagedorn, Louis Goldstein, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2026 | LiteTrack: Towards efficient vision-language tracking with parameter freezing and feature selection
Liqiang Liu, Lingling Yang, Yanfang Fu, Tiantian Feng, Anyuan Xie, Zijian Cao 0001 |
Pattern Recognit. | 5 |
| 2025 | Directional Source Separation for Robust Speech Recognition on Smart GlassesabstractModern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance. Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide |
ICASSP | 1 |
| 2025 | Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence PredictionabstractBrain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We propose a novel approach to enhance listened speech decoding from electroencephalography (EEG) signals by utilizing an auxiliary phoneme predictor that simultaneously decodes textual phoneme sequences. The proposed model architecture consists of three main parts: EEG module, speech module, and phoneme predictor. The EEG module learns to properly represent EEG signals into EEG embeddings. The speech module generates speech waveforms from the EEG embeddings. The phoneme predictor outputs the decoded phoneme sequences in text modality. Our proposed approach allows users to obtain decoded listened speech from EEG signals in both modalities (speech waveforms and textual phoneme sequences) simultaneously, eliminating the need for a concatenated sequential pipeline for each modality. The proposed approach also outperforms previous methods in both modalities. The source code and speech samples are publicly available1. Tiantian Feng, Aditya Kommineni, Sudarsana Reddy Kadiri, Shri Narayanan |
ICASSP | 2 |
| 2025 | Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during SpeechabstractUnderstanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using an speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate. The source code is available for the public at: https://github.com/Hong7Cong/SPAN-rtmri.git Hong Nguyen, Sean Foley, Xuan Shi, Tiantian Feng, Shri Narayanan |
ICASSP | 5 |
| 2025 | Data Efficient Child-Adult Speaker Diarization with Simulated ConversationsabstractAutomating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies "who spoke when", is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available. Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Catherine Lord, Shri Narayanan |
ICASSP | 2 |
| 2025 | Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
Tiantian Feng, Thanathai Lertpetchpun, Dani Byrd, Shri Narayanan |
INTERSPEECH | 1 |
| 2025 | Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
Tiantian Feng, Anfeng Xu, Xuan Shi, Somer Bishop, Shri Narayanan |
INTERSPEECH | 1 |
| 2025 | Can Multimodal Foundation Models Help Analyze Child-Inclusive Autism Diagnostic Videos?
Aditya Kommineni, Digbalay Bose, Tiantian Feng, So Hyun Kim, Helen Tager-Flusberg, Somer Bishop, Catherine Lord, Sudarsana Reddy Kadiri, Shri Narayanan |
INTERSPEECH | 3 |
| 2025 | Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shri Narayanan |
INTERSPEECH | 2 |
| 2025 | Examining Test-Time Adaptation for Personalized Child Speech Recognition
Zhonghao Shi, Xuan Shi, Anfeng Xu, Tiantian Feng, Harshvardhan Srivastava, Shri Narayanan, Maja J. Mataric |
INTERSPEECH | 4 |
| 2025 | 75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment
Xuan Shi, Yubin Zhang, Yijing Lu, Marcus Ma, Tiantian Feng, Asterios Toutios, Haley Hsu, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 5 |
| 2025 | Large Language Models based ASR Error Correction for Child Conversations
Anfeng Xu, Tiantian Feng, So Hyun Kim, Somer Bishop, Catherine Lord, Shri Narayanan |
INTERSPEECH | 2 |
| 2025 | Convex Hull-based Algebraic Constraint for Visual Quadric SLAMabstractUsing Quadrics as the object representation has the benefits of both generality and closed-form projection derivation between image and world spaces. Although numerous constraints have been proposed for dual quadric reconstruction, we found that many of them are imprecise and provide minimal improvements to localization. After scrutinizing the existing constraints, we introduce a concise yet more precise convex hull-based algebraic constraint for object landmarks, which is applied to object reconstruction, frontend pose estimation, and backend bundle adjustment. This constraint is designed to fully leverage precise semantic segmentation, effectively mitigating mismatches between complex-shaped object contours and dual quadrics. Experiments on public datasets demonstrate that our approach is applicable to both monocular and RGB-D SLAM and achieves improved object mapping and localization than existing quadric SLAM methods. The implementation of our method is available at https://github.com/tiev-tongji/convexhull-based-algebraic-constraint. Junqiao Zhao, Shuangfu Song, Zhongyang Zhu, Zihan Yuan, Chen Ye 0002, Tiantian Feng, Qiankun Yu |
IROS | 7 |
| 2025 | MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive ClusteringabstractVisual Place Recognition (VPR) enables robust localization through image retrieval based on learned descriptors.
However, drastic appearance variations of images at the same place caused by viewpoint changes can lead to inconsistent supervision signals, thereby degrading descriptor learning.
Existing methods either rely on manually defined cropping rules or labeled data for view differentiation, but they suffer from two major limitations:
(1) reliance on labels or handcrafted rules restricts generalization capability;
(2) even within the same view direction, occlusions can introduce feature ambiguity.
To address these issues, we propose MutualVPR, a mutual learning framework that integrates unsupervised view self-classification and descriptor learning.
We first group images by geographic coordinates, then iteratively refine the clusters using K-means to dynamically assign place categories without manual labeling.
Specifically, we adopt a DINOv2-based encoder to initialize the clustering.
During training, the encoder and clustering co-evolve, progressively separating drastic appearance variations of the same place and enabling consistent supervision.
Furthermore, we find that capturing fine-grained image differences at a place enhances robustness.
Experiments demonstrate that MutualVPR achieves state-of-the-art (SOTA) performance across multiple datasets, validating the effectiveness of our framework in improving view direction generalization, occlusion robustness. Qiwen Gu, Xufei Wang, Junqiao Zhao, Siyue Tao, Tiantian Feng, Guang Chen 0001 |
NeurIPS | 5 |
| 2025 | ModalityMirror: Enhancing Audio Classification in Modality Heterogeneity Federated Learning via Multimodal DistillationabstractMultimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with audio is often assumed to be the weaker modality in recognition tasks. To address this challenge, we introduce ModalityMirror to improve audio model performance by leveraging knowledge distillation from an audiovisual federated learning model. ModalityMirror involves two phases: a modality-wise FL stage to aggregate unimodal encoders; and a federated knowledge distillation stage on multimodality clients to train a unimodal student model. Our results demonstrate that ModalityMirror significantly improves the audio classification compared to the state-of-the-art FL methods such as Harmony, particularly in audiovisual FL facing video missing. Our approach unlocks the potential for exploiting the diverse modality spectrum inherent in multimodal FL. Tiantian Feng, Amir Salman Avestimehr, Shri Narayanan |
NOSSDAV | 1 |
| 2025 | A Dual-Branch Architecture for Adaptive Loss Multitask Mapping Based on AI4Arctic Sea Ice Challenge DatasetabstractAutomated sea ice mapping is increasingly critical for global climate change research and Arctic shipping route planning. In this article, a dual-branch deep learning architecture with channel attention is proposed for multisource sea ice mapping, incorporating an adaptive multitask loss function weighting mechanism. The Ready-To-Train (RTT) AI4Arctic Sea Ice Challenge dataset is used to evaluate the performance of the proposed model. Experimental results demonstrate that compared with the current state-of-the-art model, the proposed model achieves a 1.03% improvement in the combined score. Specifically, the stage of development (SOD)$F1$score increases by 1.33%, the floe size (FLOE)$F1$score improves by 4.39%, whereas$R^{2}$for sea ice concentration (SIC) decreases by 1.35%. Finally, ablation experiments are conducted to validate the effectiveness of the proposed model and the adaptive multitask loss function. Tiantian Feng, Peng Jiang 0033 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Automated Prediction of Gamburtsev Subglacial Lakes in East Antarctica With Optimized Stacking Ensemble LearningabstractThe development in machine learning (ML) technology has brought new horizons for the prediction of subglacial lakes (SLs) using radio-echo sounding (RES) data, offering fresh perspectives toward the automated identification of SLs. Nonetheless, the inherent data imbalance across various classes within the dataset presents significant analytical challenges. To address this limitation, the artificial bee colony (ABC) optimization algorithm is introduced to automatically predict SLs in Gamburtsev Province in East Antarctica, using an optimized stacking ensemble learning approach. The proposed method predicts SLs by using five representative features selected through importance and correlation analyses of eight features derived from RES data. The experimental outcomes demonstrate the superiority of this method in overcoming the significant imbalance of RES data, successfully identifying known lakes in the validation dataset. Furthermore, this study summarizes an inventory of SLs across the Gamburtsev subglacial mountains in East Antarctica, and a total of 55 new candidate SLs with lengths ranging from 108 to 38130 m have been predicted using our novel method. The source code is publicly available athttps://github.com/vivian-ma97/ABC-Stacking-for-Subglacial-Lakes Tiantian Feng, Gang Qiao, Asoke K. Nandi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion RecognitionabstractRecent studies have explored using pre-trained embeddings for speech emotion recognition, achieving comparable performance to conventional methods that rely on low-level knowledge-inspired acoustic features. These embeddings are often generated from models trained on large-scale speech datasets using self-supervised or weakly-supervised learning objectives. Despite the significant advancements made in SER through pre-trained embeddings, there is a limited understanding of the trustworthiness of these methods, including privacy breaches, unfair performance, vulnerability to adversarial attacks, and computational cost, all of which may hinder the real-world deployment of these systems. In response, we introduce TrustSER, a general framework designed to evaluate the trustworthiness of SER systems using deep learning methods, focusing on privacy, safety, fairness, and sustainability, offering unique insights into future research in the field of SER. Our code is publicly available under: https://github.com/usc-sail/trust-ser. Tiantian Feng, Rajat Hebbar, Shri Narayanan |
ICASSP | 1 |
| 2024 | Foundation Model Assisted Automatic Speech Emotion Recognition: Transcribing, Annotating, and AugmentingabstractSignificant advances are being made in speech emotion recognition (SER) using deep learning models. Nonetheless, training SER systems remains challenging, requiring both time and costly resources. Like many other machine learning tasks, acquiring datasets for SER requires substantial data annotation efforts, including transcription and labeling. These annotation processes present challenges when attempting to scale up conventional SER systems. Recent developments in foundational models have had a tremendous impact, giving rise to applications such as ChatGPT. These models have enhanced human-computer interactions including bringing unique possibilities for streamlining data collection in fields like SER. In this research, we explore the use of foundational models to assist in automating SER from transcription and annotation to augmentation. Our study demonstrates that these models can generate transcriptions to enhance the performance of SER systems that rely solely on speech data. Furthermore, we note that annotating emotions from transcribed speech remains a challenging task. However, combining outputs from multiple LLMs enhances the quality of annotations. Lastly, our findings suggest the feasibility of augmenting existing speech emotion datasets by annotating unlabeled speech samples. Tiantian Feng, Shri Narayanan |
ICASSP | 1 |
| 2024 | Emotion-Aligned Contrastive Learning Between Images and MusicabstractTraditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and speech. While most approaches aim to match general music semantics to the input queries, only a few focus on affective qualities. In this work, we address the task of retrieving emotionally-relevant music from image queries by learning an affective alignment between images and music audio. Our approach focuses on learning an emotion-aligned joint embedding space between images and music. This embedding space is learned via emotion-supervised contrastive learning, using an adapted cross-modal version of the SupCon loss. We evaluate the joint embeddings through cross-modal retrieval tasks (image-to-music and music-to-image) based on emotion labels. Furthermore, we investigate the generalizability of the learned music embeddings via automatic music tagging. Our experiments show that the proposed approach successfully aligns images and music, and that the learned embedding space is effective for cross-modal retrieval applications. Shanti Stewart, Kleanthis Avramidis, Tiantian Feng, Shri Narayanan |
ICASSP | 3 |
| 2024 | Audio-Visual Child-Adult Speaker Classification in Dyadic InteractionsabstractInteractions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer scale and robustness across diverse and wide-ranging conditions. Identifying the speech segments belonging to the child is a critical step in such modeling. Conventional child-adult speaker classification typically relies on audio modeling approaches, overlooking visual signals that convey speech articulation information, such as lip motion. Building on the foundation of an audio-only child-adult speaker classification pipeline, we propose incorporating visual cues through active speaker detection and visual processing models. Our framework involves video preprocessing, utterance-level child-adult speaker detection, and late fusion of modality-specific predictions. We demonstrate from extensive experiments that a visually aided classification pipeline enhances the accuracy and robustness of the classification. We show relative improvements of 2.38% and 3.97% in F1 macro score when one face and two faces are visible, respectively. Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Shri Narayanan |
ICASSP | 3 |
| 2024 | Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?abstractMulti-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains. Despite its success, the robustness of multi-modal learning for visual recognition is often challenged by the unavailability of a subset of modalities, especially the visual modality. Conventional approaches to mitigate missing modalities in multi-modal learning rely heavily on modality fusion schemes. In contrast, this paper explores the use of text-to-image models to assist multi-modal learning. Specifically, we propose and explore a simple but effective multi-modal learning framework GTI-MM to enhance the data efficiency and model robustness against missing visual modality by imputing the missing data with generative models. Using multiple multi-modal datasets with visual recognition tasks, we present a comprehensive analysis of diverse conditions involving missing visual data. Our findings show that synthetic images benefit training data efficiency with missing visual data during training and improve model robustness with visual data missing during both training and testing. Moreover, we demonstrate GTI-MM is effective with lower generation quantity and simple prompt techniques. Our code base and synthetic images are at https://github.com/usc-sail/GTI-MM. Tiantian Feng, Daniel Yang, Digbalay Bose, Shri Narayanan |
ICMI | 1 |
| 2024 | Mangrove Extraction in Shenzhen Area Based on Multi-Source Temporal Remote Sensing DataabstractIn this study, Sentinel-1 image data with all-weather observation characteristics and Sentinel-2 image data with medium and high spatial and temporal resolution were used to explore the optimal feature scheme suitable for coastal wetland classification. The optimal feature combination scheme will be put into the UNet network and applied to the mangrove classification in the coastal area of Shenzhen. In addition, we added DCN structure to the model, which can better extract mangrove features compared with traditional methods. The experimental results showed that the mangrove area in Shenzhen showed an overall increasing trend in the past four years, which proved that the local mangrove restoration process had a certain effect. Yarong Zou, Tiantian Feng |
IGARSS | 2 |
| 2024 | Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
Tiantian Feng, Dimitrios Dimitriadis, Shri Narayanan |
INTERSPEECH | 1 |
| 2024 | Toward Fully-End-to-End Listened Speech Decoding from EEG SignalsabstractSpeech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli.We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals.Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required.The proposed method consists of an EEG module and a speech module along with a connector.The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations.The connector learns to bridge the distributions of the latent spaces of EEG and speech.The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics.A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding.The source code is available here: github.com/lee-jhwn/fesde. Aditya Kommineni, Tiantian Feng, Kleanthis Avramidis, Xuan Shi, Sudarsana Reddy Kadiri, Shri Narayanan |
INTERSPEECH | 3 |
| 2024 | Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
Anfeng Xu, Tiantian Feng, Lue Shen, Helen Tager-Flusberg, Shri Narayanan |
INTERSPEECH | 3 |
| 2024 | A Multiscale Segment-Based Subglacial Water Body Identification Method According to Dry-Wet Transition CharacteristicsabstractMapping the distribution of subglacial water bodies (SWBs) in Antarctica using radio echo sounding (RES) data is of great significance for understanding subglacial hydrology, the mass balance of Antarctic ice sheet, biology, and so on. However, existing methods for identifying SWBs mostly focus on utilizing the characteristics of radar echoes at the individual trace level to pinpoint areas that consist of subglacial water, leading to discontinuities in the identification results. In this article, a novel multiscale segment-based SWB identification method according to basal dry-wet transition characteristics is proposed and applied to the RES data collected in the Antarctic’s Gamburtsev Province (AGAP) in East Antarctica. The identification results of SWBs with high, medium, and low confidence levels are provided and compared to the existing inventories of SWBs. It is found that 96.06% of the SWBs in the inventories can be identified, with 87.40% in quantity and 84.73% in length identified with high confidence, and 8.66% in quantity and 8.45% in length identified with medium confidence. In addition, some SWBs with high confidence that are not recorded in the inventories are identified by using the proposed method. The distribution of identified SWBs with high confidence is clustered and shows high consistency with inventorial SWBs. Three characteristics, including subglacial depth, hydraulic gradient, and abruptness index, of identified SWBs with different confidence levels are statistically analyzed, demonstrating rationality and consistency with the characteristics of the subglacial ice-water interface. The SWB identification results contribute valuable insights for a better understanding of subglacial environment in the AGAP region. Dailiang Wang, Tiantian Feng, Jinyu Jia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Scaling Representation Learning From Ubiquitous ECG With State-Space ModelsabstractUbiquitous sensing from wearable devices in the wild holds promise for enhancing human well-being, from diagnosing clinical conditions and measuring stress to building adaptive health promoting scaffolds. But the large volumes of data therein across heterogeneous contexts pose challenges for conventional supervised learning approaches. Representation Learning from biological signals is an emerging realm catalyzed by the recent advances in computational modeling and the abundance of publicly shared databases. The electrocardiogram (ECG) is the primary researched modality in this context, with applications in health monitoring, stress and affect estimation. Yet, most studies are limited by small-scale controlled data collection and over-parameterized architecture choices. We introduce WildECG, a pre-trained state-space model for representation learning from ECG signals. We train this model in a self-supervised manner with 275 000 10 s ECG recordings collected in the wild and evaluate it on a range of downstream tasks. The proposed model is a robust backbone for ECG analysis, providing competitive performance on most of the tasks considered, while demonstrating efficacy in low-resource regimes. Kleanthis Avramidis, Dominika Kunc, Bartosz Perz, Kranti Adsul, Tiantian Feng, Przemyslaw Kazienko, Stanislaw Saganowski, Shri Narayanan |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech ModelsabstractMany recent studies have focused on fine-tuning pretrained models for speech emotion recognition (SER), resulting in promising performance compared to traditional methods that rely largely on low-level, knowledge-inspired acoustic features. These pre-trained speech models learn general-purpose speech representations using self-supervised or weakly-supervised learning objectives from large-scale datasets. Despite the significant advances made in SER through the use of pre-trained architecture, fine-tuning these large pre-trained models for different datasets requires saving copies of entire weight parameters, rendering them impractical to deploy in real-world settings. As an alternative, this work explores parameter-efficient fine-tuning (PEFT) approaches for adapting pre-trained speech models for emotion recognition. Specifically, we evaluate the efficacy of adapter tuning, embedding prompt tuning, and LoRa (Low-rank approximation) on four popular SER testbeds. Our results reveal that LoRa achieves the best fine-tuning performance in emotion recognition while enhancing fairness and requiring only a minimal extra amount of weight parameters. Furthermore, our findings offer novel insights into future research directions in SER, distinct from existing approaches focusing on directly fine-tuning the model architecture. Our code is publicly available under: https://github.com/usc-sail/peft-ser. Tiantian Feng, Shri Narayanan |
ACII | 1 |
| 2023 | Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal DataabstractThe sensitivity of data collected via ambulatory monitoring, which regularly involve the recording of speech signals and sensor information, can cause strong privacy concerns. We investigate user re-identification risk in a corpus of such data collected to observe the interplay between behavior, physiology, and well-being of healthcare workers in their daily life. We then develop a user anonymization approach that preserves well-being information (i.e., anxiety), but eliminates user identify (ID) information. We formulate this via an auto-encoder that learns a transformed version of the original feature set in an adversarial manner so that it minimizes the anxiety estimation loss and maximizes the user classification loss. Results indicate that the original features bear a large user re-identification risk, while also having a good ability to classify a user’s anxiety. After removing the most prone features to user re-identification from the original feature set, the user classification accuracy decreases, while the anxiety classification performance is preserved. The final features transformed via the auto-encoder further reduce evidence of user ID and preserve anxiety classification ability. Findings from this study can contribute to the design privacy-aware bio-behavioral models that can be used for responsible ambulatory monitoring in healthcare and beyond. Ravi Pranjal, Ranjana Seshadri, Rakesh Kumar Sanath Kumar Kadaba, Tiantian Feng, Shri Narayanan, Theodora Chaspari |
ICASSP | 4 |
| 2023 | FedAudio: A Federated Learning Benchmark for Audio TasksabstractFederated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio data and audio-related tasks. In this paper, we fill this critical gap by introducing a new FL benchmark for audio tasks which we refer to as FedAudio. FedAudio includes four representative and commonly used audio datasets from three important audio tasks that are well aligned with FL use cases. In particular, a unique contribution of FedAudio is the introduction of data noises and label errors to the datasets to emulate challenges when deploying FL systems in real-world settings. FedAudio also includes the benchmark results of the datasets and a PyTorch library with the objective of facilitating researchers to fairly compare their algorithms. We hope FedAudio could act as a catalyst to inspire new FL research for audio tasks and thus benefit the acoustic and speech research community. The datasets and benchmark results can be accessed at https://github.com/zhang-tuo-pdf/FedAudio. Tiantian Feng, Samiul Alam, Sunwoo Lee 0001, Mi Zhang 0002, Shri Narayanan, Amir Salman Avestimehr |
ICASSP | 2 |
| 2023 | Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism
Rimita Lahiri, Tiantian Feng, Rajat Hebbar, Catherine Lord, So Hyun Kim, Shri Narayanan |
INTERSPEECH | 2 |
| 2023 | Understanding Spoken Language Development of Children with ASD Using Pre-trained Speech Embeddings
Anfeng Xu, Rajat Hebbar, Rimita Lahiri, Tiantian Feng, Lindsay Butler, Lue Shen, Helen Tager-Flusberg, Shri Narayanan |
INTERSPEECH | 4 |
| 2023 | FedMultimodal: A Benchmark for Multimodal Federated LearningabstractOver the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal. Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan |
KDD | 1 |
| 2023 | MM-AU: Towards Multimodal Understanding of Advertisement VideosabstractAdvertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches. Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan |
ACM Multimedia | 3 |
| 2022 | Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) techniques have gained considerable interest in many applications including smart virtual assistants and health state tracking. SER systems often acquire and transmit speech data collected at the client-side to remote cloud platforms for inference and decision making. However, speech data carries rich information not only about emotions conveyed in vocal expressions, but also other sensitive demographic traits, such as gender, age, and language background. It is desirable to select only features that are necessary for the emotion classification while protecting sensitive features. However, there are some features that are necessary for emotion classification. These features may also reveal other demographic traits. In this work, we propose a method to improve inference privacy for sensitive features by injecting noise into the input speech data, but without degrading the SER system performance. The approach combines a noise representation learning architecture, called Cloak [1], with adversarial training to keep relevant information inside the data for emotion classification while removing information that would enable inferring sensitive demographic attributes. Experimental results show that our method can effectively prevent inference of sensitive demographic information, and that the improved privacy comes at a cost of only a minor utility loss for the emotion classification. Tiantian Feng, Hanieh Hashemi, Murali Annavaram, Shri Narayanan |
ICASSP | 1 |
| 2022 | Semi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On Federated Learning using Multiview Pseudo-LabelingabstractSpeech Emotion Recognition (SER) application is frequently associated with privacy concerns as it often acquires and transmits speech data at the client-side to remote cloud platforms for further processing. These speech data can reveal not only speech content and affective information but the speaker's identity, demographic traits, and health status. Federated learning (FL) is a distributed machine learning algorithm that coordinates clients to train a model collaboratively without sharing local data. This algorithm shows enormous potential for SER applications as sharing raw speech or speech features from a user's device is vulnerable to privacy attacks. However, a major challenge in FL is limited availability of high-quality labeled data samples. In this work, we propose a semi-supervised federated learning framework, Semi-FedSER, that utilizes both labeled and unlabeled data samples to address the challenge of limited labeled data samples in FL. We show that our Semi-FedSER can generate desired SER performance even when the local label rate l=20 using two SER benchmark datasets: IEMOCAP and MSP-Improv. Tiantian Feng, Shri Narayanan |
INTERSPEECH | 1 |
| 2022 | User-Level Differential Privacy against Attribute Inference Attack of Speech Emotion Recognition on Federated LearningabstractMany existing privacy-enhanced speech emotion recognition (SER) frameworks focus on perturbing the original speech data through adversarial training within a centralized machine learning setup. However, this privacy protection scheme can fail since the adversary can still access the perturbed data. In recent years, distributed learning algorithms, especially federated learning (FL), have gained popularity to protect privacy in machine learning applications. While FL provides good intuition to safeguard privacy by keeping the data on local devices, prior work has shown that privacy attacks, such as attribute inference attacks, are achievable for SER systems trained using FL. In this work, we propose to evaluate the user-level differential privacy (UDP) in mitigating the privacy leaks of the SER system in FL. UDP provides theoretical privacy guarantees with privacy parameters $\epsilon$ and $\delta$. Our results show that the UDP can effectively decrease attribute information leakage while keeping the utility of the SER system with the adversary accessing one model update. However, the efficacy of the UDP suffers when the FL system leaks more model updates to the adversary. We make the code publicly available to reproduce the results in https://github.com/usc-sail/fed-ser-leakage. Tiantian Feng, Raghuveer Peri, Shri Narayanan |
INTERSPEECH | 1 |
| 2022 | Scale Estimation with Dual Quadrics for Monocular Object SLAMabstractThe scale ambiguity problem is inherently unsolvable to monocular SLAM without the metric baseline between moving cameras. In this paper, we present a novel scale estimation approach based on an object-level SLAM system. To obtain the absolute scale of the reconstructed map, we formulate an optimization problem to make the scaled dimensions of objects conform to the distribution of their sizes in the physical world, without relying on any prior information about gravity direction. The dual quadric is adopted to represent objects for its ability to describe objects compactly and accurately, thus providing reliable dimensions for scale estimation. In the proposed monocular object-level SLAM system, semantic objects are initialized first from fitted 3-D oriented bounding boxes and then further optimized under constraints of 2-D detections and 3-D map points. Experiments on indoor and outdoor public datasets show that our approach outperforms existing methods in terms of accuracy and robustness. Shuangfu Song, Junqiao Zhao, Tiantian Feng, Chen Ye 0002, Lu Xiong 0001 |
IROS | 3 |
| 2022 | PMDRnet: A Progressive Multiscale Deformable Residual Network for Multi-Image Super-Resolution of AMSR2 Arctic Sea Ice ImagesabstractThe extent of the area covered by polar sea ice is an important indicator of global climate change. Continuous monitoring of Arctic sea ice concentration (SIC) primarily relies on passive microwave images. However, passive microwave images have coarse spatial resolution, resulting in SIC production with significant blurring at the ice–water divides. In this article, a novel multi-image super-resolution (MISR) network called progressive multiscale deformable residual network (PMDRnet) is proposed to improve the spatial resolution of sea ice passive microwave images according to the characteristics of both passive microwave images and sea ice motions. To achieve image alignment with complex and large Arctic sea ice motions, we design a novel alignment module that includes a progressive alignment strategy and a multiscale deformable convolution alignment unit. In addition, the temporal attention mechanism is used to adaptively fuse the effective spatiotemporal information across image sequence. The sea ice-related loss function is designed to provide more detailed sea ice information of the network to improve super-resolution performance and further benefit finer Arctic SIC results. Experimental results demonstrate that PMDRnet significantly outperforms the current state-of-the-art MISR methods and can generate super-resolved SIC products with finer texture features and much sharper sea ice edges. The code and datasets of PMDRnet are available athttps://doi.org/10.5061/dryad.k3j9kd590. Tiantian Feng, Xiaofan Shen, Rongxing Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | G-VIDO: A Vehicle Dynamics and Intermittent GNSS-Aided Visual-Inertial State Estimator for Autonomous DrivingabstractThis paper proposes G-VIDO, a vehicle dynamics, and intermittent Global Navigation Satellite System (GNSS)-aided visual-inertial state estimator, to address the state estimation problem of autonomous vehicle localization (i.e., position and orientation estimation in the global coordinate system) under various GNSS states. A dynamics pre-integration theory is proposed on the basis of a two-degree-of-freedom (DOF) vehicle dynamics model, and dynamics constraints are built in the optimization back-end, considering the unobservable problem of the monocular visual-inertial system under degenerate motions. The proposed highly nonlinear system can be robustly initialized by loosely aligning the monocular structure from motion (SfM) results, pre-integrated IMU measurements, and vehicle motion information. GNSS is used for reference frame transformation and constraint construction in the sliding window. The cumulative error can be corrected with the aid of GNSS, and the vehicle’s position in the global coordinate system can be determined. A GNSS anomaly detection algorithm is proposed to improve the system robustness under intermittent GNSS. Experiments have shown that G-VIDO can provide real-time, robust, and seamless localization in multiple GNSS states, with an RMSE of less than 30 cm (with GNSS). Moreover, we proved that the initialization and local odometry modules in G-VIDO outperform several state-of-the-art VIO systems and our preliminary work VINS-Vehicle. Lu Xiong 0001, Rong Kang, Junqiao Zhao, Peizhi Zhang, Ran Ju, Chen Ye 0002, Tiantian Feng |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2021 | Privacy and Utility Preserving Data Transformation for Speech Emotion RecognitionabstractSpeech carries rich information not only about an individual’s intent but about demographic traits, physical and psychological state among other things. Notably, continuously worn wearable sensors enable researchers to collect egocentric speech data to study and assess real-life expressed emotions, offering unprecedented opportunities for applications in the field of assistive agents, medical diagnoses, and personalized education. Many existing systems collect and transmit these speech data, either processed or unprocessed, from users’ devices to a central server for post analysis. However, egocentric audio sensing for speech emotion recognition has created concerns and risks to privacy, where unintended/improper inferences of sensitive information and demographic information may occur without user consent. Toward addressing these concerns, in this work, we propose a privacy-preserving data transformation technique to mitigate potential threats associated with sensitive information and demographic inferences. The proposed mechanism combines an autoencoder architecture, called replacement autoencoder, with gradient reversal layer to remove sensitive information inside the data, such as sensitive labels and demographics. We empirically validate our approach for predicting emotions using three commonly used datasets for speech emotion recognition. We show that our method can effectively prevent inferences of sensitive emotions and demographic information. We further show that the improved privacy comes at a cost of a minor utility loss for the target application. Tiantian Feng, Shri Narayanan |
ACII | 1 |
| 2021 | Distributed Energy Transaction Model Based on the Alliance Blockchain in Case of ChinaabstractDistributed energy, mainly composed of new energy, plays an important role in promoting the development of new energy. At present, the development of distributed energy is greatly hindered by imperfect trading platform and unstable output of new energy. Blockchain is decentralized, autonomous and requires collaborative management. Its own technical characteristics have the inherent advantages of reconstructing the energy system. The alliance chain in the blockchain is more suitable for building a distributed energy trading platform. The paper constructs a distributed energy transaction model based on alliance blockchain, studies the integration mode of blockchain and distributed energy transaction, and explores the application of blockchain in distributed transaction. The paper provides a new idea for optimizing and reconstructing the traditional distributed energy trading platform, and providing decision support for promoting distributed energy trading. Yongxiu He, Binyou Yang, Ming-li Cui, Tiantian Feng, Yi-er Sun |
J. Web Eng. | 6 |
| 2021 | Temporal Dynamics of Workplace Acoustic Scenes: Egocentric Analysis and PredictionabstractIdentification of the acoustic environment from an audio recording, also known as acoustic scene classification, is an active area of research. In this paper, we study dynamically-changing background acoustic scenes from the egocentric perspective of an individual in a workplace. In a novel data collection setup, wearable sensors were deployed on individuals to collect audio signals within a built environment, while Bluetooth-based hubs continuously tracked the individual's location which represents the acoustic scene at a certain time. The data of this paper come from 170 hospital workers gathered continuously during work shifts for a 10 week period. In the first part of our study, we investigate temporal patterns in the egocentric sequence of acoustic scenes encountered by an employee, and the association of those patterns with factors such as job-role and daily routine of the individual. Motivated by evidence of multifaceted effects of ambient sounds on human psychology, we also analyze the association of the temporal dynamics of the perceived acoustic scenes with particular behavioral traits of the individual. Experiments reveal rich temporal patterns in the acoustic scenes experienced by the individuals during their work shifts, and a strong association of those patterns with various constructs related to job-roles and behavior of the employees. In the second part of our study, we employ deep learning models to predict the temporal sequence of acoustic scenes from the egocentric audio signal. We propose a two-stage framework where a recurrent neural network is trained on top of the latent acoustic representations learned by a segment-level neural network. The experimental results show the efficacy of the proposed system in predicting sequence of acoustic scenes, highlighting the existence of underlying temporal patterns in the acoustic scenes experienced in workplace. Arindam Jati, Amrutha Nadarajan, Raghuveer Peri, Karel Mundnich, Tiantian Feng, Benjamin Girault, Shri Narayanan |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor StudyabstractWearable sensors today can unobtrusively collect rich time-series of physiological states and human movement patterns over a prolonged period. Gaining a better understanding of how an individual's physiological responses vary in different workplace environments can be valuable in understanding human behavior related to wellness and performance. In this work, we describe our exploration in discovering the correlation between one's physiological responses and movement patterns within different indoor locations using data collected from nurses in a hospital workplace for a ten week period. In this work, we use simple heuristics to empirically validate the idea that such a relationship may exist and then quantify it using mutual information analysis. We propose and demonstrate a data analysis approach that can also detect variations in the level of mutual dependency between different locations and physiological responses. The mutual dependency measures derived from our method are empirically shown to provide valuable information for improving modeling of self-reported work behavior patterns compared to using features derived from a single data stream. Tiantian Feng, Brandon M. Booth, Shri Narayanan |
ICASSP | 1 |
| 2020 | Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral SignalsabstractContinuously-worn wearable sensors provide an unprecedented opportunity to unobtrusively measure rich bio-behavioral time-series recordings in natural settings such as the workplace. These time-series data can be helpful in inferring broad patterns of behavior such as common routines and daily stress. Many existing approaches either rely on rigid pre-defined notions of activities or use sensitive contextual measurements, such as GPS location or localization within the home, that present privacy concerns and measurement challenges. In this work, we introduce a novel data processing pipeline to model behavioral consistency in a large real-world wearable recording data-set collected in a hospital workplace setting from nurses and direct clinical providers for a period of ten weeks. We use a non-parametric clustering method to generate time series clusters and capture behavioral consistency via the activity curve model. We evaluate the behavioral consistency model under different work roles and conditions such as between different groups of nursing professions and day versus night shift individuals. We also demonstrate that the learned behavioral consistency feature can assist in predicting self-reported work behaviors and anxiety levels. Tiantian Feng, Shri Narayanan |
ICASSP | 1 |
| 2020 | Occlusion aware unsupervised learning of optical flow from videoabstractIn this paper, we proposed an unsupervised learning method for estimating the optical flow between video frames, especially to solve the occlusion problem. Occlusion is caused by the movement of an object or the movement of the camera, defined as when certain pixels are visible in one video frame but not in adjacent frames. Due to the lack of pixel correspondence between frames in the occluded area, incorrect photometric loss calculation can mislead the optical flow training process. In the video sequence, we found that the occlusion in the forward (t→t+1) and backward (t→t-1) frame pairs are usually complementary. That is, pixels that are occluded in subsequent frames are often not occluded in the previous frame and vice versa. Therefore, by using this complementarity, a new weighted loss is proposed to solve the occlusion problem. Our method achieves competitive optical flow accuracy compared to the baseline and some supervised methods on KITTI and Sintel benchmarks. Junqiao Zhao, Tiantian Feng |
ICMV | 3 |
| 2019 | Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace SettingabstractGaining a better understanding of how people move about and interact with their environment is an important piece of understanding human behavior. Careful analysis of individuals' deviations or variations in movement over time can provide an awareness about changes to their physical or mental state and may be helpful in tracking performance and well-being especially in workplace settings. We propose a technique for clustering and discovering patterns in human movement data by extracting motifs from the time series of durations where participants linger at different locations. Using a data set of over 200 participants moving around a hospital for ten weeks, we show this technique intuitively captures local temporal relationships between hospital rooms and also clusters them in a fashion consistent with the room type labels (e.g. lounge, break room, etc.) without using prior knowledge. Machine learning features derived from these clusters are empirically shown to provide information similar to features attained using domain knowledge of the room type labels directly when predicting mental wellness from self-reports. Brandon M. Booth, Tiantian Feng, Abhishek Jangalwa, Shri Narayanan |
ICASSP | 2 |
| 2019 | Discovering Optimal Variable-length Time Series Motifs in Large-scale Wearable Recordings of Human Bio-behavioral SignalsabstractContinuously-worn wearable sensors produce copious amounts of rich bio-behavioral time series recordings. Exploring recurring patterns, often known as motifs, in wearable time series offers critical insights into understanding the nature of human behavior. Challenges in discovering motifs from wearable recordings include noise removal, pattern generalization, and accounting for subtle variations between subsequences in one motif set. In this work, we introduce a time series processing pipeline to summarize an optimal set of variable-length motifs in a real-world wearable recording data-set collected in a hospital workplace setting. We propose the use of the Savitzky-Golay filter for noise removal without significant data distortion. We then combine the previously developed HierarchIcal based Motif Enumeration (HIME) algorithm with a principled optimization approach to obtain the most repetitive patterns in long-term wearable time-series. We also describe challenges in using just a single method to detect motifs in wearable time series in our experiments. We demonstrate our pipeline can effectively identify meaningful variable-length motifs in large-scale heart rate signals collected continuously from over 100 individuals both at and outside their workplace over 10 weeks through two machine learning experiments. Tiantian Feng, Shri Narayanan |
ICASSP | 1 |
| 2019 | Super Resolution Reconstruction Technique in Passive Microwave Images of Arctic Sea IceabstractPolar sea ice is one of the key parameters of cryosphere and polar environmental change, which plays an important role in the study of global climate change. High-resolution monitoring of polar sea ice relies mainly on optical satellite imagery and synthetic aperture radar (SAR) data, with limited spatial and temporal coverages for many applications. Passive microwave data is an important data source for continuous observations of polar sea ice, thanks to its working ability in all-sky conditions and its wide coverage. However, it is difficult to achieve high-resolution monitoring of polar sea ice using passive microwave data due to its coarse resolution. In order to solve this problem, super resolution (SR) reconstruction technique is adopted in this paper to improve the spatial resolution of passive microwave images. SR reconstruction technique based on both single-image and multi-image are attempted. AMSR2 level 3 (L3) Products of Brightness Temperatures (BTs) for Arctic sea ice are used as experimental data, and the reconstruction results obtained from different SR methods are compared and discussed. Tiantian Feng, Junqiao Zhao, Rongxing Li |
IGARSS | 2 |
| 2018 | Vision-based Semantic Mapping and Localization for Autonomous Indoor ParkingabstractIn this paper, we proposed a novel and practical solution for the real-time indoor localization of autonomous driving in parking lots. High-level landmarks, the parking slots, are extracted and enriched with labels to avoid the aliasing of low-level visual features. We then proposed a robust method for detecting incorrect data associations between parking slots and further extended the optimization framework by dynamically eliminating suboptimal data associations. Visual fiducial markers are introduced to improve the overall precision. As a result, a semantic map of the parking lot can be established fully automatically and robustly. We experimented the performance of real-time localization based on the map using our autonomous driving platform TiEV, and the average accuracy of 0.3m track tracing can be achieved at a speed of 10kph. Yewei Huang 0001, Junqiao Zhao, Shaoming Zhang, Tiantian Feng |
Intelligent Vehicles Symposium | 5 |
| 2016 | Inverse Synthetic Aperture Radar imaging of maneuvering targets based on joint time-frequency analysisabstractIn this paper, the joint time-frequency(JTF) analysis techniques were applied to Inverse Synthetic Aperture Radar(ISAR) imaging and motion compensation (MOCOM P), which compensated for the blurring effect in imaging maneuvering targets using traditional RD algorithm. During the implementation of the JTF algorithm, a Gabor-wavelet transform was employed as the JTF tool to obtain a corresponding 3D time-range-Doppler cube. The simulation results showed different time snapshots of moving target, namely target pose in particular instants. Then, ISAR image formed separately by traditional methodology without any compensation and JTF algorithm were compared. The former is highly distorted and its spectrogram demonstrates the severe frequency shifts between the pulses while the latter is very well focused due to the JTF tool used to compensate motion effect and the frequencies of time pulses are well aligned. Tiantian Feng |
IGARSS | 1 |
| 2016 | Quality assessment of existing antarctic remote sensing productsabstractThere are a variety of remote sensing products of Antarctica that cover different time periods and published by different research groups. The resolution and accuracy for these products are quite different due to original data sources, extraction methods, etc. The aim of this research is to assess the quality and investigate changes that can be detected in three main categories of Antarctic remote sensing products including groundling lines, coastlines and surface elevation models (DEM). All the products are available at National Snow and Ice Data Center (NSIDC). In order to identify changes and uncertainties of these products, cross-validation strategies are developed for the quality assessment when there is a lack of ground truth data. The quality assessment results for each remote sensing product are presented in the paper. Rongxing Li, Yixiang Tian, Tiantian Feng, Huan Xie 0001, Haifeng Xiao, Hexia Weng, Da Lv, Xiaohua Tong |
IGARSS | 3 |
| 2014 | Novel Image Registration Method Based on Local Structure ConstraintsabstractThis letter presents an effective approach to reduce the ambiguity of matching results for image registration based on a coarse-to-fine strategy. In the coarse registration stage, we compute initial transformation parameters via the descriptors. In the fine registration stage, we propose a new matching strategy for an iterative closest point framework, in which the matching pairs are determined by a bidirectional matching criterion in terms of feature similarity and spatial consistency. In this letter, the spatial consistency includes not only spatial distance but also local structure constraints on reference and sensed images. Comparative experiments on multispectral and viewpoint-altered images show that the proposed algorithm achieves higher performance in accuracy and robustness. Aixia Li, Xiaojun Cheng, Haiyan Guan, Tiantian Feng, Zequn Guan |
IEEE Geosci. Remote. Sens. Lett. | 4 |