Xinlei Chen

dblp:41/9922 · DBLP profile ↗
← Back
159ranked-venue papers
24as first author
106since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 75 · 14 first-author · 47 since 2021Computer networks · 71 · 9 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 12 first-author · 20 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DIMM: Decoupled Multi-hierarchy Kalman Filter via Reinforcement Learning
abstract
State estimation is challenging for target tracking with high maneuverability, as the target's state transition function changes rapidly, irregularly, and is unknown to the estimator. Existing work based on interacting multiple model (IMM) achieves more accurate estimation than single-filter approaches through model combination, aligning appropriate models for different motion modes of the target over time. However, two limitations of conventional IMM remain unsolved. First, the solution space of the model combination is constrained as the target's diverse kinematic properties in different directions are ignored. Second, the model combination weights calculated by the observation likelihood are not accurate enough due to the measurement uncertainty. In this paper, we propose a novel framework, DIMM, to effectively combine estimates from different motion models in each direction, thus increasing the target tracking accuracy. First, DIMM extends the model combination solution space of conventional IMM from a hyperplane to a hypercube by designing a 3D-decoupled multi-hierarchy filter bank, which describes the target's motion with various-order linear models. Second, DIMM generates more reliable combination weight matrices through a differentiable adaptive fusion network for importance allocation rather than solely relying on the observation likelihood; it contains an attention-based twin delayed deep deterministic policy gradient (TD3) method with a hierarchical reward. Experiments demonstrate that DIMM significantly improves the tracking accuracy of existing state estimation methods by 31.61%~99.23%.
Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen
AAAI6
2026 AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
abstract
Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions. To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception.
Jirong Zha, Yuxuan Fan, Chen Gao 0001, Xinlei Chen
AAAI7
2026 FireSentry: A Multi-Modal Spatio-temporal Benchmark Dataset for Fine-Grained Wildfire Spread Forecasting
abstract
Fine-grained wildfire spread prediction is crucial for enhancing emergency response efficacy and decision-making precision. However, existing research predominantly focuses on coarse spatiotemporal scales and relies on low-resolution satellite data, capturing only macroscopic fire states while fundamentally constraining high-precision localized fire dynamics modeling capabilities. To bridge this gap, we present FireSentry, a provincial-scale multi-modal wildfire dataset characterized by sub-meter spatial and sub-second temporal resolution. Collected using synchronized UAV platforms, FireSentry provides visible and infrared video streams, in-situ environmental measurements, and manually validated fire masks. Building on FireSentry, we establish a comprehensive benchmark encompassing physics-based, data-driven, and generative models, revealing the limitations of existing mask-only approaches. Our analysis proposes FiReDiff, a novel dual-modality paradigm that first predicts future video sequences in the infrared modality, and then precisely segments fire masks in the mask modality based on the generated dynamics. FiReDiff achieves state-of-the-art performance, with video quality gains of 39.2% in PSNR, 36.1% in SSIM, 50.0% in LPIPS, 29.4% in FVD, and mask accuracy gains of 3.3% in AUPRC, 59.1% in F1 score, 42.9% in IoU, and 62.5% in MSE when applied to generative models. The FireSentry benchmark dataset and FiReDiff paradigm collectively advance fine-grained wildfire forecasting and dynamic disaster simulation. The processed benchmark dataset is publicly available at: https://github.com/Munan222/FireSentry-Benchmark-Dataset.
Huandong Wang, Yali Song, Qiuhua Wang, Yong Li 0008, Xinlei Chen
KDD (1)8
2026 Log Anomaly Detection in Kubernetes Using Graph Neural Networks and Retrieval-Augmented Generation
Dongyi Fan, Suqiong Zhang, Yifan Huo, Xinlei Chen
KSEM (7)5
2026 VocabLog: A Vocabulary-Driven and LLM-Augmented Framework for High-Performance Log Parsing
Dongyi Fan, Suqiong Zhang, Yifan Huo, Xinlei Chen
KSEM (2)5
2026 Count Every Rotation and Every Rotation Counts: Exploring Drone Dynamics via Propeller Sensing
abstract
As drone-based applications proliferate, paramount contactless sensing of airborne drones from the ground becomes indispensable. This work demonstrates concentrating on propeller rotational speed will substantially improve drone sensing performance and proposes an event-camera-based solution, EventPro. EventPro features two components: Count Every Rotation achieves accurate, real-time propeller speed estimation by mitigating ultra-high sensitivity of event cameras to environmental noise. Every Rotation Counts leverages these speeds to infer both internal and external drone dynamics. Extensive evaluations in real-world drone delivery scenarios show that EventPro achieves a sensing latency of 3 ms and a rotational speed estimation error of merely 0.23%. Additionally, EventPro infers drone flight commands with 96.5% precision and improves drone tracking accuracy by over 22% when combined with other sensing modalities. Demo: https://eventpro25.github.io/EventPro/.
Xuecheng Chen, Jingao Xu, Wenhua Ding, Haoyang Wang 0012, Xinyu Luo, Ruiyang Duan, Xueqian Wang 0001, Yunhao Liu 0001, Xinlei Chen
SenSys10
2026 QUIDS: Quality-Informed Incentive-Driven Multiagent Dispatching System for Mobile Crowdsensing
abstract
This paper addresses the challenges of achieving optimal quality of information (QoI) in non-dedicated vehicular mobile crowdsensing (NVMCS) system, where vehicles not originally designed for sensing are leveraged to collect real-time data as they traverse urban environments. These challenges are exacerbated by the interrelated issues of sensing coverage, sensing reliability, and the inherently dynamic nature of participating vehicles. To tackle these challenges, we propose QUIDS, a QUality-informed Incentive-driven multi-agent Dispatching System, which ensures high sensing coverage and sensing reliability under budget constraints in NVMCS systems. QUIDS improves QoI by introducing a novel metric, Aggregated Sensing Quality (ASQ), designed to quantitatively capture the concept of QoI by integrating both sensing coverage and sensing reliability. Moreover, we develop a Mutually Assisted Belief-aware Vehicle Dispatching algorithm that estimates sensing reliability and allocates monetary incentives under uncertain vehicle conditions, thereby further improving ASQ. Evaluation using real-world data collected from a deployed NVMCS system in a metropolitan area demonstrates the effectiveness of QUIDS. The ASQ metric shows a 38% improvement over non-dispatching scenarios and a 10% enhancement over state-of-the-art methods. Additionally, QUIDS reduces reconstruction map errors by 39–74% across various reconstruction algorithms, validating its efficacy in improving QoI within NVMCS systems. Addressing the often-overlooked issue of sensing reliability in existing studies, the QUIDS system leverages non-dedicated vehicles and incorporates a quality-informed incentive-driven dispatching system to jointly optimize sensing coverage and sensing reliability. This enables low-cost, high-quality, and scalable urban environmental monitoring without the need for dedicated sensing infrastructure, and makes the system applicable to diverse smart-city scenarios such as traffic monitoring and environmental sensing.
Zuxin Li, Fanhang Man, Xuecheng Chen, Susu Xu, Fan Dang 0001, Chaopeng Hong, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen
IEEE Internet Things J.10
2026 STeP-Diff: Spatio-Temporal Physics-Informed Diffusion Models for Mobile Fine-Grained Pollution Forecasting
abstract
Fine-grained air pollution forecasting is crucial for urban management and the development of healthy buildings. Deploying portable sensors on mobile platforms such as cars and buses offers a low-cost, easy-to-maintain, and wide-coverage data collection solution. However, due to the random and uncontrollable movement patterns of these non-dedicated mobile platforms, the resulting sensor data are often incomplete and temporally inconsistent. By exploring potential training patterns in the reverse process of diffusion models, we proposeSpatio-TemporalPhysics-InformedDiffusion Models (STeP-Diff). STeP-Diff leverages DeepONet to model the spatial sequence of measurements along with a PDE-informed diffusion model to forecast the spatio-temporal field from incomplete and time-varying data. Through a PDE-constrained regularization framework, the denoising process asymptotically converges to the convection-diffusion dynamics, ensuring that predictions are both grounded in real-world measurements and aligned with the fundamental physics governing pollution dispersion. To assess the performance of the system, we deployed 59 self-designed portable sensing devices in two cities, operating for 14 days to collect air pollution data. Compared to the second-best performing algorithm, our model achieved improvements of up to 89.12% in MAE, 82.30% in RMSE, and 25.00% in MAPE, with extensive evaluations demonstrating that STeP-Diff effectively captures the spatio-temporal dependencies in air pollution fields.
Weijie Hong, Huandong Wang, Qiuhua Wang, Yali Song, Xiao-Ping Zhang 0002, Yong Li 0008, Xinlei Chen
IEEE Trans. Knowl. Data Eng.9
2026 mmE-Loc: Facilitating Accurate Drone Landing With Ultra-High-Frequency Localization
abstract
For precise, efficient, and safe drone landings, ground platforms should real-time, accurately locate descending drones and guide them to designated spots. While mmWave sensing combined with cameras improves localization accuracy, lower sampling frequency of traditional frame cameras compared to mmWave radar creates bottlenecks in system throughput. In this work, we upgrade traditional frame camera with event camera, a novel sensor that harmonizes in sampling frequency with mmWave radar within ground platform setup, and introduce mmE-Loc, a high-precision, low-latency ground localization system designed for precise drone landings. To fully exploit thetemporal consistencyandspatial complementaritybetween these two modalities, we propose two innovative modules:(i)the Consistency-instructed Collaborative Tracking module, which further leverages the drone's physical knowledge of periodic micro-motions and structure for accurate measurements extraction, and(ii)the Graph-informed Adaptive Joint Optimization module, which integrates drone motion information for efficient sensor fusion and drone localization. Extensive experiments (30+ hours) demonstrate that mmE-Loc attains 0.083$m$localization accuracy and 5.12$ms$end-to-end latency, outperforming four state-of-the-art methods by over 48% and 62%, respectively.
Haoyang Wang 0012, Jingao Xu, Xinyu Luo, Xuecheng Chen, Ruiyang Duan, Yunhao Liu 0001, Weijie Hong, Xiaoqiang Ji 0001, Xinlei Chen
IEEE Trans. Mob. Comput.12
2026 Aerial Shepherds: Enabling Hierarchical Localization in Heterogeneous MAV Swarms
abstract
A heterogeneous micro aerial vehicles (MAV) swarm consists of resource-intensive but expensive advanced MAVs (AMAVs) and resource-limited but cost-effective basic MAVs (BMAVs), offering opportunities in diverse fields. Accurate and real-time localization is crucial for MAV swarms, but current practices lack a low-cost, high-precision, and real-time solution, especially for lightweight BMAVs. We find an opportunity to accomplish the task by transforming AMAVs into mobile localization infrastructures for BMAVs. However, translating this insight into a practical system is challenging due to issues in estimating locations with diverse and unknown localization errors of BMAVs, and allocating resources of AMAVs considering interconnected influential factors. This work introduces TransformLoc, a new framework that transforms AMAVs into mobile localization infrastructures, specifically designed for low-cost and resource-constrained BMAVs. We design an error-aware joint location estimation model to perform intermittent joint estimation for BMAVs and introduce a similarity-instructed adaptive grouping-scheduling strategy to allocate resources of AMAVs dynamically. TransformLoc achieves a collaborative, adaptive, and cost-effective localization system suitable for large-scale heterogeneous MAV swarms. We implement and validate TransformLoc on industrial drones. Results show it outperforms all baselines by up to 68% in localization performance, improving navigation success rates by 60%. Extensive robustness and ablation experiments further highlight superiority of its design.
Haoyang Wang 0012, Jingao Xu, Chenyu Zhao 0002, Yuhan Cheng, Xuecheng Chen, Chaopeng Hong, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen
IEEE Trans. Mob. Comput.9
2026 A Novel Integrated Sensing and Communication Scheme in UAVs-Enabled Vehicular Networks With MARL-Driven Adaptive Control
abstract
In this paper, we propose a novel integrated sensing and communication (ISAC) scheme tailored for UAVs-enabled vehicular networks, which leverages the information coverage capabilities of multiple UAVs and addresses critical challenges posed by multiple moving users. Unlike many traditional scheme, our scheme efficiently leverages ISAC signal echoes and real-time data uploads to provide communication services while achieving accurate sensing, thereby overcoming issues of resource waste and low operational efficiency. In the scheme, we aim to optimize both communication and sensing indicators, taking into account practical issues such as energy saving and collision avoidance for UAVs. However, the inherent complexity of multi-objective stochastic optimization in dynamic environments and limited communication resources render centralized UAV control inconvenient. To address the above challenges, we propose a novel multi-agent reinforcement learning (MARL) algorithm based on local information to realize the distributed adaptive control of motion decision, power selection, and channel allocation for UAVs. The algorithm combines random network distillation (RND) and dynamic data augmentation with multi-agent deep deterministic policy gradient (MADDPG) to encourage agents to explore effectively under sparse rewards and improve MADDPG's policy learning ability in finite data, thus approaching the global optimal solution. Experimental results demonstrate that the proposed algorithm can improve communication and sensing performance by more than 16.71% and 68.26% compared with other baselines and satisfy the set constraints. Furthermore, by adjusting hyperparameters, we can optimize the ISAC performance while achieving different energy savings levels for UAVs, proving that the designed scheme can reduce the waste of resources and improve the ISAC operation efficiency.
Ziyuan Wang 0002, Xiao-Ping Zhang 0002, Wenbo Ding 0001, Yuhan Dong, Xinlei Chen
IEEE Trans. Mob. Comput.5
2026 Scalable UAV Multi-Hop Networking via Multi-Agent Reinforcement Learning With Large Language Models
abstract
In disaster scenarios, establishing robust emergency communication networks is critical, and unmanned aerial vehicles (UAVs) offer a promising solution to rapidly restore connectivity. However, organizing UAVs to form multi-hop networks in large-scale dynamic environments presents significant challenges, including limitations in algorithmic scalability and the vast exploration space required for coordinated decision-making. To address these issues, we propose MRLMN, a novel framework that integrates multi-agent reinforcement learning (MARL) and large language models (LLMs) to jointly optimize UAV agents toward achieving optimal networking performance. The framework incorporates a grouping strategy with reward decomposition to enhance algorithmic scalability and balance decision-making across UAVs. In addition, behavioral constraints are applied to selected key UAVs to improve the robustness of the network. Furthermore, the framework integrates LLM agents, leveraging knowledge distillation to transfer their high-level decision-making capabilities to MARL agents. This enhances both the efficiency of exploration and the overall training process. In the distillation module, a Hungarian algorithm-based matching scheme is applied to align the decision outputs of the LLM and MARL agents and define the distillation loss. Extensive simulation results validate the effectiveness of our approach, demonstrating significant improvements in network performance over the MAPPO baseline and other comparison methods, including enhanced coverage and communication quality.
Yanggang Xu, Jirong Zha, Weijie Hong, Xiangmin Yi, Chen-Chun Hsia, Xinlei Chen
IEEE Trans. Mob. Comput.8
2026 Breaking the Communication-Accuracy Trade-Off: A Sparsified Information Diffusion Framework for Multi-Agent Collaborative Perception
abstract
The growing relevance of multi-agent systems has drawn increasing focus on communication-efficient filters for collaborative perception to alleviate the system's communication burden. While the event-triggered (ET) mechanism can improve communication efficiency in collaborative state estimation, an inevitable trade-off exists between estimation accuracy and communication cost in ET filters. This paper proposes a fast and accurate ET diffusion-based filter for real-time multi-agent collaborative target tracking, aiming to reduce the system's data transmission without compromise in tracking performance. The proposed filter achieves improved tracking accuracy, reduced data transmission, and accelerated convergence using an error-minimized ET cubature information filter (CIF) for local estimation, and a correlation-aware diffusion strategy for global fusion. The experimental results confirm the scalability of the proposed EDC-CIF algorithm and demonstrate its efficacy in simultaneously reducing estimation error and computation time while significantly enhancing communication efficiency.
Jirong Zha, Chenyu Zhao 0002, Zhenyu Liu 0003, Tao Sun 0013, Xinlei Chen
IEEE Trans. Mob. Comput.8
2026 BlueKey: Exploiting Bluetooth Low Energy for Enhanced Physical-Layer Key Generation
abstract
Bluetooth Low Energy (BLE) is a prevalent technology in various applications due to its low power consumption and wide device compatibility. Despite its numerous advantages, the encryption methods of BLE often expose devices to potential attacks. To fortify security, we investigate the application of Physical-layer Key Generation (PKG), a promising technology that enables devices to generate a shared secret key from their shared physical environment. Although extensively investigated, PKG is generally discussed in the context of Wi-Fi, and existing solutions for BLE demonstrate significantly lower performance. To bridge this gap, we propose a distinctive approach that capitalizes on the inherent characteristics of BLE to facilitate efficient PKG. We utilize the constant tone extension within BLE protocols to extract comprehensive physical layer information and introduce an innovative method that employs Legendre polynomial quantization for PKG. This method facilitates the exchange of secret keys with a high key matching rate and a high key generation rate. The efficacy of our approach is validated through extensive experiments on a software-defined radio platform, underscoring its potential to enhance security in the rapidly expanding field of BLE applications. A pilot study on commercial off-the-shelf BLE devices further validates the system's practicality, revealing important trade-offs between performance and hardware constraints in real-world deployments.
Fan Dang 0001, Jinyan Jiang, Xu Wang 0018, Lin Wang 0023, Kebin Liu 0001, Xinlei Chen, Yunhao Liu 0001
IEEE Trans. Mob. Comput.8
2026 SniffySquad: Patchiness-Aware Gas Source Localization with Multi-Robot Collaboration
abstract
Gas source localization is pivotal for the rapid mitigation of gas leakage disasters, where mobile robots emerge as a promising solution. However, existing methods predominantly schedule robots’ movements based on reactive stimuli or simplified gas plume models. These approaches typically excel in idealized, simulated environments but fall short in real-world gas environments characterized by their patchy distribution. In this work, we introduce SniffySquad , a multi-robot olfaction-based system designed to address the inherent patchiness in gas source localization. SniffySquad incorporates a patchiness-aware active sensing approach that enhances the quality of data collection and estimation. Moreover, it features an innovative collaborative role adaptation strategy to boost the efficiency of source-seeking endeavors. Extensive evaluations demonstrate that our system achieves an increase in the success rate by \(20\%+\) and an improvement in path efficiency by \(30\%+\) , outperforming state-of-the-art gas source localization solutions.
Yuhan Cheng, Xuecheng Chen, Haoyang Wang 0012, Jingao Xu, Chaopeng Hong, Susu Xu, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen
ACM Trans. Sens. Networks10
2025 Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents
abstract
Fanhang Man, Huandong Wang, Jianjie Fang, Zhaoyi Deng, Baining Zhao, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fanhang Man, Huandong Wang, Jianjie Fang, Zhaoyi Deng, Baining Zhao, Xinlei Chen, Yong Li 0008
ACL (1)6
2025 CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
abstract
Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng, Baining Zhao, Qian Zhang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chen Gao 0001, Shiquan Yu, Ruiying Peng, Baining Zhao, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)8
2025 UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
abstract
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Chen Gao 0001, Yue Wang 0007, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)10
2025 Transformers without Normalization
abstract
Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation DyT(x) = tanh(αx), as a dropin replacement for normalization layers in Transformers. DyT is inspired by the observation that layer normalization in Transformers often produces tanh-like, S-shaped input-output mappings. By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning. We validate the effectiveness of Transformers with DyT across diverse settings, from recognition to generation, supervised to self-supervised learning, and computer vision to language models. These findings challenge the conventional understanding that normalization layers are indispensable in modern neural networks, and offer new insights into their role in deep networks.
Jiachen Zhu 0002, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu 0003
CVPR2
2025 Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions
abstract
Accurate modeling and control of response length is essential for optimizing large language model (LLM) deployment, impacting computational efficiency, user experience, and system reliability.We develop a statistical framework based on extreme value theory, analyzing 14,301 GPT-4o responses across temperature settings and prompting strategies, with cross-validation on Qwen and DeepSeek architectures.Our analysis reveals that response lengths follow Weibull-type generalized extreme value (GEV) distributions, exhibiting heavier tails under stochastic generation conditions.The key contributions include:(1) a novel GEV-generalized Pareto (GPD) hybrid model that achieves superior tail fit (R 2 CDF = 0.9993 vs standalone GEV's 0.998) while preserving architectural generalizability;(2) quantitative characterization of prompt anchoring effects, showing reduced dispersion but increased outlier propensity under randomization; and (3) identification of temperaturedependent response patterns that remain consistent across architectures, where higher temperatures amplify length variability while maintaining the underlying extreme-value mechanisms.The proposed hybrid model's adaptive threshold selection enables precise verbosity control in production systems, regardless of the specific LLM architecture employed.These findings provide both theoretical insights into LLM generation patterns and practical tools for response length optimization.
Liuxuan Jiao, Chen Gao 0001, Yiqian Yang, Chenliang Zhou, YiXian Huang, Xinlei Chen, Yong Li 0008
EMNLP6
2025 Scaling Language-Free Visual Representation Learning
abstract
Visual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP models are often trained on different data. In this work, we ask the question: "Do visual self-supervised approaches lag behind CLIP due to the lack of language supervision, or differences in the training data?" We study this question by training both visual SSL and CLIP models on the same MetaCLIP data, and leveraging VQA as a diverse testbed for vision encoders. In this controlled setup, visual SSL models scale better than CLIP models in terms of data and model capacity, and visual SSL performance does not saturate even after scaling up to 7B parameters. Consequently, we observe visual SSL methods achieve CLIP-level performance on a wide range of VQA and classic vision benchmarks. These findings demonstrate that pure visual SSL can match language-supervised visual pretraining at scale, opening new opportunities for vision-centric representation learning.
David Fan 0001, Shengbang Tong, Jiachen Zhu 0002, Koustuv Sinha, Zhuang Liu 0003, Xinlei Chen, Michael G. Rabbat, Nicolas Ballas, Yann LeCun, Amir Bar, Saining Xie
ICCV6
2025 PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera Deraining
abstract
Event cameras excel in high temporal resolution and dynamic range but suffer from dense noise in rainy conditions. Existing event deraining methods face trade-offs between temporal precision, deraining effectiveness, and computational efficiency. In this paper, we propose PRE-Mamba, a novel point-based event camera deraining framework that fully exploits the spatiotemporal characteristics of raw event and rain. Our framework introduces a 4D event cloud representation that integrates dual temporal scales to preserve high temporal precision, a Spatio-Temporal Decoupling and Fusion module (STDF) that enhances deraining capability by enabling shallow decoupling and interaction of temporal and spatial information, and a Multi-Scale State Space Model (MS3M) that captures deeper rain dynamics across dual-temporal and multi-spatial scales with linear computational complexity. Enhanced by frequency-domain regularization, PRE-Mamba achieves superior performance (0.95 SR, 0.91 NR, and 0.4s/M events) with only 0.26M parameters on EventRain-27K, a comprehensive dataset with labeled synthetic and real-world sequences. Moreover, our method generalizes well across varying rain intensities, viewpoints, and even snowy conditions.
Ciyu Ruan, Ruishan Guo, Zihang Gong, Jingao Xu, Wenhan Yang, Xinlei Chen
ICCV6
2025 MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
abstract
In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph into an unified autoregressive model capable of generating both text and visual tokens. VPiT teaches an LLM to predict discrete text tokens and continuous visual tokens from any input sequence of image and text data curated in an instruction-following format. Our empirical investigation reveals several intriguing properties of VPiT: (1) visual generation ability emerges as a natural byproduct of improved visual understanding, and can be unlocked efficiently with a small amount of generation data; (2) while we find understanding and generation to be mutually beneficial, understanding data contributes to both capabilities more effectively than generation data. Building upon these findings, we train our MetaMorph model and achieve competitive performance on both visual understanding and generation. In visual generation, MetaMorph can leverage the world knowledge and reasoning abilities gained from LLM pretraining, and overcome common failure modes exhibited by other generation models. Our results suggest that LLMs may have strong "prior" vision capabilities that can be efficiently adapted to both visual understanding and generation with a relatively simple instruction tuning process.
Shengbang Tong, David Fan 0001, Jiachen Zhu 0002, Yunyang Xiong, Xinlei Chen, Koustuv Sinha, Michael G. Rabbat, Yann LeCun, Saining Xie, Zhuang Liu 0003
ICCV5
2025 Deconstructing Denoising Diffusion Models for Self-Supervised Learning
abstract
In this study, we examine the representation learning abilities of Denoising Diffusion Models (DDM) that were originally purposed for image generation. Our philosophy is to deconstruct a DDM, gradually transforming it into a classical Denoising Autoencoder (DAE). This deconstructive process allows us to explore how various components of modern DDMs influence self-supervised representation learning. We observe that only a very few modern components are critical for learning good representations, while many others are nonessential. Our study ultimately arrives at an approach that is highly simplified and to a large extent resembles a classical DAE. We hope our study will rekindle interest in a family of classical methods within the realm of modern self-supervised learning.
Xinlei Chen, Zhuang Liu 0003, Saining Xie, Kaiming He
ICLR1
2025 An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
abstract
This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a token and achieve highly performant results. This is substantially different from the popular design in Vision Transformer, which maintains the inductive bias from ConvNets towards local neighborhoods (e.g., by treating each 16x16 patch as a token). We showcase the effectiveness of pixels-as-tokens across three well-studied computer vision tasks: supervised learning for classification and regression, self-supervised learning via masked autoencoding, and image generation with diffusion models. Although it's computationally less practical to directly operate on individual pixels, we believe the community must be made aware of this surprising piece of knowledge when devising the next generation of neural network architectures for computer vision.
Duy-Kien Nguyen, Mido Assran, Unnat Jain, Martin R. Oswald, Cees Snoek, Xinlei Chen
ICLR6
2025 Learning to (Learn at Test Time): RNNs with Expressive Hidden States
abstract
Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with linear complexity and expressive hidden states. The key idea is to make the hidden state a machine learning model itself, and the update rule a step of self-supervised learning. Since the hidden state is updated by training even on test sequences, our layers are called Test-Time Training (TTT) layers. We consider two instantiations: TTT-Linear and TTT-MLP, whose hidden state is a linear model and a two-layer MLP respectively. We evaluate our instantiations at the scale of 125M to 1.3B parameters, comparing with a strong Transformer and Mamba, a modern RNN. Similar to Transformer, TTT-Linear and TTT-MLP can keep reducing perplexity by conditioning on more tokens, while Mamba cannot after 16k context. TTT-MLP still faces challenges in memory I/O, but shows larger potential in long context, pointing to a promising direction for future research.
Yu Sun 0020, Karan Dalal, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang 0004, Oluwasanmi Koyejo, Tatsunori B. Hashimoto, Carlos Guestrin
ICML8
2025 LLMs can see and hear without any training
abstract
We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ability to perform multi-step reasoning, MILS prompts the LLM to generate candidate outputs, each of which are scored and fed back iteratively, eventually generating a solution to the task. This enables various applications that typically require training specialized models on task-specific data. In particular, we establish a new state-of-the-art on emergent zero-shot image, video and audio captioning. MILS seamlessly applies to media generation as well, discovering prompt rewrites to improve text-to-image generation, and even edit prompts for style transfer! Finally, being a gradient-free optimization approach, MILS can invert multimodal embeddings into text, enabling applications like cross-modal arithmetic.
Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen, Ishan Misra, Rohit Girdhar
ICML3
2025 Highly Compressed Tokenizer Can Generate Without Training
abstract
Commonly used image tokenizers produce a 2D grid of spatially arranged tokens. In contrast, so-called 1D image tokenizers represent images as highly compressed one-dimensional sequences of as few as 32 discrete tokens. We find that the high degree of compression achieved by a 1D tokenizer with vector quantization enables image editing and generative capabilities through heuristic manipulation of tokens, demonstrating that even very crude manipulations – such as copying and replacing tokens between latent representations of images – enable fine-grained image editing by transferring appearance and semantic attributes. Motivated by the expressivity of the 1D tokenizer’s latent space, we construct an image generation pipeline leveraging gradient-based test-time optimization of tokens with plug-and-play loss functions such as reconstruction or CLIP similarity. Our approach is demonstrated for inpainting and text-guided image editing use cases, and can generate diverse and realistic samples without requiring training of any generative model.
L. Lao Beyer, Tianhong Li, Xinlei Chen, Sertac Karaman, Kaiming He
ICML3
2025 Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
abstract
Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions remain about how auto-encoder design impacts reconstruction and downstream generative performance. This work explores scaling in auto-encoders for reconstruction and generation by replacing the convolutional backbone with an enhanced Vision Transformer for Tokenization (ViTok). We find scaling the auto-encoder bottleneck correlates with reconstruction but exhibits a nuanced relationship with generation. Separately, encoder scaling yields no gains, while decoder scaling improves reconstruction with minimal impact on generation. As a result, we determine that scaling the current paradigm of auto-encoders is not effective for improving generation performance. Coupled with Diffusion Transformers, ViTok achieves competitive image reconstruction and generation performance on 256p and 512p ImageNet-1K. In videos, ViTok achieves SOTA reconstruction and generation performance on 16-frame 128p UCF-101.
Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang 0001, Tingbo Hou, Sriram Vishwanath, Peter Vajda, Xinlei Chen
ICML10
2025 Underwater Motions Analysis and Control of a Coupling-Tiltable Unmanned Aerial-Aquatic Vehicle
abstract
Coupling-Tiltable Unmanned Aerial-Aquatic Vehicles (UAAVs) have gained increasing importance, yet lack comprehensive analysis and suitable controllers. This paper analyzes the underwater motion characteristics of a self-designed UAAV, Mirs-Alioth, and designs a controller for it. The effectiveness of the controller is validated through experiments. The singularities of Mirs-Alioth are derived as Singular Thrust Tilt Angle (STTA), which serve as an essential tool for an analysis of its underwater motion characteristics. The analysis reveals several key factors for designing the controller. These include the need for logic switching, using a Nussbaum function to compensate control direction uncertainty in the auxiliary channel, and employing an auxiliary controller to mitigate coupling effects. Based on these key points, a control scheme is designed. It consists of a controller that regulates the thrust tilt angle to the singular value, an auxiliary controller incorporating a Saturated Nussbaum function, and a logic switch. Eventually, two sets of experiments are conducted to validate the effectiveness of the controller and demonstrate the necessity of the Nussbaum function.
Dongyue Huang, Minghao Dou, Xuchen Liu 0001, Xinlei Chen, Ben M. Chen
ICRA7
2025 How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
abstract
3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks, showing potential to surpass traditional computer vision methods. In this survey, we present a comprehensive review of methods integrating LLMs with 3D spatial understanding. We propose a taxonomy that categorizes existing methods into three branches: image-based methods deriving 3D understanding from 2D visual data, point cloud-based methods working directly with 3D representations, and hybrid modality-based methods combining multiple data streams. We systematically review representative methods along these categories, covering data representations, architectural modifications, and training strategies that bridge textual and 3D modalities. Finally, we discuss current limitations, including dataset scarcity and computational challenges, while highlighting promising research directions in spatial perception, multi-modal fusion, and real-world applications.
Jirong Zha, Yuxuan Fan, Xinlei Chen
IJCAI5
2025 Open3D-VQA: A Benchmark for Embodied Spatial Concept Reasoning with Multimodal Large Language Model in Open Space
abstract
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs across seven general spatial reasoning tasks, offered in multiple-choice, true/false, and short-answer formats, and supports both visual and point cloud modalities. The questions are automatically generated from spatial relations extracted from both real-world and simulated aerial scenes. Evaluation on 13 popular MLLMs reveals that: 1) Models are generally better at answering questions about relative spatial relations than absolute distances, 2) 3D LLMs fail to demonstrate significant advantages over 2D LLMs, and 3) Fine-tuning solely on the simulated dataset can significantly improve the model's spatial reasoning performance in real-world scenarios. The benchmark, generation pipeline, and evaluation toolkit are released on this page.
Zile Zhou, Xuchen Liu 0001, Jianjie Fang, Chen Gao 0001, Jinqiang Cui, Yong Li 0008, Xinlei Chen, Xiao-Ping Zhang 0002
ACM Multimedia9
2025 AirScape: An Aerial Generative World Model with Motion Controllability
abstract
How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the first world model designed for six-degree-of-freedom aerial agents. AirScape predicts future observation sequences based on current visual inputs and motion intentions. Specifically, we construct a dataset for aerial world model training and testing, which consists of 11k video-intention pairs. This dataset includes first-person-view videos capturing diverse drone actions across a wide range of scenarios, with over 1,000 hours spent annotating the corresponding motion intentions. Then we develop a two-phase schedule to train a foundation model-initially devoid of embodied spatial knowledge-into a world model that is controllable by motion intentions and adheres to physical spatio-temporal constraints. Experimental results demonstrate that AirScape significantly outperforms existing foundation models in 3D spatial imagination capabilities, especially with over a 50% improvement in metrics reflecting motion alignment. The project is available at: https://embodiedcity.github.io/AirScape/.
Baining Zhao, Rongze Tang, Mingyuan Jia, Ziyou Wang, Fanhang Man, Xin Zhang 0123, Wei Wu 0021, Chen Gao 0001, Xinlei Chen, Yong Li 0008
ACM Multimedia11
2025 Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
abstract
Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilities, especially high-level reasoning, remains unclear. This paper introduces Embodied-R, a collaborative framework combining large-scale Vision-Language Models (VLMs) for perception and small-scale Language Models (LMs) for reasoning. Using Reinforcement Learning (RL) with a novel reward system considering think-answer logical consistency, the model achieves slow-thinking capabilities with limited computational resources. After training on only 5k embodied video samples, Embodied-R with a 3B LM matches state-of-the-art multimodal reasoning models (OpenAI-o1, Gemini-2.5-pro) on both in-distribution and out-of-distribution embodied spatial reasoning tasks. Embodied-R also exhibits emergent thinking patterns such as systematic analysis and contextual integration. We further explore research questions including response length, training on VLM, strategies for reward design, and differences in model generalization after SFT (Supervised Fine-Tuning) and RL training. The project page is available at: https://embodiedcity.github.io/Embodied-R/.
Baining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao 0001, Fanhang Man, Jinqiang Cui, Xin Wang 0019, Xinlei Chen, Yong Li 0008, Wenwu Zhu 0001
ACM Multimedia8
2025 Palantir: Towards Efficient Super Resolution for Ultra-high-definition Live Streaming
abstract
Neural enhancement through super-resolution (SR) deep neural networks (DNNs) opens up new possibilities for ultra-high-definition (UHD) live streaming. Yet, the heavy SR DNN inference overhead leads to severe deployment challenges. To reduce the overhead, existing systems propose to apply DNN-based SR only on carefully selected anchor frames while upscaling non-anchor frames via the lightweight reusing-based SR approach. However, frame-level scheduling is coarse-grained and fails to deliver optimal efficiency. In this work, we propose Palantír, the first neural-enhanced UHD live streaming system with fine-grained patch-level scheduling.
Xinqi Jin, Zhui Zhu, Xikai Sun, Fan Dang 0001, Jiangchuan Liu, Jingao Xu, Kebin Liu 0001, Xinlei Chen, Yunhao Liu 0001
MMSys8
2025 Demo: HawkEye: Practical In-Flight Obstacle Avoidance with Event Camera and LiDAR Fusion
abstract
Drones are increasingly used in applications such as last-mile delivery and infrastructure inspection, but their safe operation, especially in high-speed scenarios, remains a critical challenge. Existing vision- and LiDAR-based obstacle localization methods suffer from motion blur, latency, and low spatio-temporal resolution, making them inadequate for detecting and tracking fast-moving objects. In this work, we present HawkEye, a drone obstacle avoidance system that fuses event cameras and LiDAR to achieve high-frequency, accurate 3D tracking of dynamic objects. By leveraging the complementary strengths of both sensors, Hawkeye enables robust real-time sensing and safe evasive maneuvers, addressing a key requirement for the large-scale deployment of autonomous drones. Demo: https://wenhua00.github.io/HawkEye/.
Wenhua Ding, Zhengli Zhang, Haoyang Wang 0012, Yinan Zhu, Shilong Ji, Xin Zhou 0015, Jingao Xu, Dongyue Huang, Xinlei Chen
MobiCom10
2025 Poster: Skyshield: Event-Driven Submillimeter Thin Obstacle Detection for Drone Flight Safety
abstract
Drones operating in complex environments face a significant threat from thin obstacles, such as steel wires and kite strings at the submillimeter level, which are notoriously difficult for conventional sensors like RGB cameras, LiDAR, and depth cameras to detect. This paper introduces SkyShield, an event-driven, end-to-end framework designed for the perception of submillimeter scale obstacles. Drawing upon the unique features that thin obstacles present in the event stream, our method employs a lightweight U-Net architecture and an innovative Dice-Contour Regularization Loss to ensure precise detection. Experimental results demonstrate that our event-based approach achieves mean F1 Score of 0.7088 with a low latency of 21.2 ms, making it ideal for deployment on edge and mobile platforms.
Zhengli Zhang, Xinyu Luo, Wenhua Ding, Dongyue Huang, Xinlei Chen
MobiCom6
2025 Meta CLIP 2: A Worldwide Scaling Recipe
abstract
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP.
Yung-Sung Chuang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James R. Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu 0003, Saining Xie, Scott Yih, Shang-Wen Li 0001, Hu Xu 0001
NeurIPS11
2025 Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
abstract
Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image tokens results in significant computational overhead, and the use of dynamic high-resolution inputs further increases this burden. Previous approaches have attempted to reduce the number of image tokens through token pruning, typically by selecting tokens based on attention scores or image token diversity. Through empirical studies, we observe that existing methods often overlook the joint impact of pruning on both the current layer's output (local) and the outputs of subsequent layers (global), leading to suboptimal pruning decisions. To address this challenge, we propose Balanced Token Pruning (BTP), a plug-and-play method for pruning vision tokens. Specifically, our method utilizes a small calibration set to divide the pruning process into multiple stages. In the early stages, our method emphasizes the impact of pruning on subsequent layers, whereas in the deeper stages, the focus shifts toward preserving the consistency of local outputs. Extensive experiments across various LVLMs demonstrate the broad effectiveness of our approach on multiple benchmarks. Our method achieves a 78\% compression rate while preserving 96.7\% of the original models' performance on average. Our code is available at https://github.com/EmbodiedCity/NeurIPS2025-Balanced-Token-Pruning.
Chen Gao 0001, Yong Li 0008, Xinlei Chen
NeurIPS5
2025 What Can RL Bring to VLA Generalization? An Empirical Study
abstract
Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under distribution shifts. Reinforcement learning (RL) offers a path to overcome these limitations by optimizing for task objectives via trial-and-error, yet a systematic understanding of its specific generalization benefits for VLAs compared to SFT is lacking. To address this, our study introduces a comprehensive benchmark for evaluating VLA generalization and systematically investigates the impact of RL fine-tuning across diverse visual, semantic, and execution dimensions. Our extensive experiments reveal that RL fine-tuning, particularly with PPO, significantly enhances generalization in semantic understanding and execution robustness over SFT, while maintaining comparable visual robustness. We identify PPO as a more effective RL algorithm for VLAs than LLM-derived methods like DPO and GRPO. We also develop a simple recipe for efficient PPO training on VLAs, and demonstrate its practical utility for improving VLA generalization. The project page is at https://rlvla.github.io
Jijia Liu, Bingwen Wei, Xinlei Chen, Qingmin Liao, Yi Wu 0013, Chao Yu 0005, Yu Wang 0002
NeurIPS4
2025 VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
abstract
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering.These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations.We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play.We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy.To highlight VolleyBots’ sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones.
Zelai Xu, Ruize Zhang 0001, Chao Yu 0005, Huining Yuan 0002, Xiangmin Yi, Shilong Ji, Chuqi Wang, Wenbo Ding 0001, Xinlei Chen, Yu Wang 0002
NeurIPS11
2025 Ultra-High-Frequency Harmony: mmWave Radar and Event Camera Orchestrate Accurate Drone Landing
abstract
For precise, efficient, and safe drone landings, ground platforms should real-time, accurately locate descending drones and guide them to designated spots. While mmWave sensing combined with cameras improves localization accuracy, lower sampling frequency of traditional frame cameras compared to mmWave radar creates bottlenecks in system throughput. In this work, we replace traditional frame camera with event camera, a novel sensor that harmonizes in sampling frequency with mmWave radar within ground platform setup, and introduce mmE-Loc, a high-precision, low-latency ground localization system designed for drone landings. To fully leverage the temporal consistency and spatial complementarity between these modalities, we propose two innovative modules, consistency-instructed collaborative tracking and graph-informed adaptive joint optimization, for accurate drone measurement extraction and efficient sensor fusion. Real-world experiments in landing scenarios from a drone delivery company demonstrate that mmE-Loc outperforms SOTA methods in both accuracy and latency.
Haoyang Wang 0012, Jingao Xu, Xinyu Luo, Xuecheng Chen, Ruiyang Duan, Yunhao Liu 0001, Xinlei Chen
SenSys8
2025 Test-Time Training on Video Streams
abstract
Prior work has established Test-Time Training (TTT) as a general framework to further improve a trained model at test time. Before making a prediction on each test instance, the model is first trained on the same instance using a self-supervised task such as reconstruction. We extend TTT to the streaming setting, where multiple test instances - video frames in our case - arrive in temporal order. Our extension is online TTT: The current model is initialized from the previous model, then trained on the current frame and a small window of frames immediately before. Online TTT significantly outperforms the fixed-model baseline for four tasks, on three real-world datasets. The improvements are more than 2.2x and 1.5x for instance and panoptic segmentation. Surprisingly, online TTT also outperforms its offline variant that accesses strictly more information, training on all frames from the entire test video regardless of temporal order. This finding challenges those in prior work using synthetic videos. We formalize a notion of locality as the advantage of online over offline TTT, and analyze its role with ablations and a theory based on bias-variance trade-off.
Renhao Wang, Yu Sun 0020, Arnuv Tandon, Yossi Gandelsman, Xinlei Chen, Alexei A. Efros, Xiaolong Wang 0004
J. Mach. Learn. Res.5
2025 A fast and accurate 3D lung tumor segmentation algorithm
Ziwei Han, Xinlei Chen
Pattern Anal. Appl.3
2025 SmartSpr: A Physics-Informed Mobile Sprinkler Scheduling System for Reducing Urban Particulate Matter Pollution
abstract
Urban particulate pollution presents considerable public health hazards, underscoring the need for effective control measures in various cities. This paper proposes SmartSpr, a physics-informed urban mobile sprinkler scheduling system designed for enhanced efficiency in reducing particulate pollution. SmartSpr incorporates a Physics-Informed Neural Network (PINN)-based model, enriched with Bayesian optimization, to accurately simulate the impact of mobile sprinklers on particulate matter (PM) dispersion. Building on this sprinkling effect model, a selective sprinkling strategy considering the replenish process is proposed. This strategy employs a sparsity-driven decoupling simulated annealing algorithm to refine sprinkler routes, prioritizing areas with substantial environmental benefits. Extensive field experiments and simulations have validated SmartSpr, demonstrating a 64.8% reduction in prediction error of SmartSpr's sprinkling model compared to the leading baseline and an 18% enhancement in pollutant reduction efficiency of the proposed scheduling algorithm.
Zijian Xiao, Zuxin Li, Xuecheng Chen, Chaopeng Hong, Xiao-Ping Zhang 0002, Xinlei Chen
IEEE Trans. Mob. Comput.7
2025 CatUA: Catalyzing Urban Air Quality Intelligence Through Mobile Crowd-Sensing
abstract
Mobile air pollution sensing methods have emerged to collect air quality data with improved spatial and temporal resolutions. However, existing methodologies struggle to effectively process spatially mixed gas samples due to the highly dynamic fluctuations experienced by sensors, resulting in significant measurement deviations. We identify an opportunity to address this issue by exploring potential patterns within sensor measurements. To this end, we propose CatUA, a novel city-scale fine-grained air quality estimation system designed to deliver accurate mobile air quality data. First, we design AirBERT, a representation learning model specifically aimed at discerning mixed gas concentrations from sensor data. Second, we implement a Prompt-informed Training Strategy that leverages extensive unlabeled and minimal labeled city-scale data to enhance the performance of CatUA. Notably, the Auto-Prompt mechanism allows CatUA to conveniently acquire new knowledge tailored to specific downstream tasks. To ensure the practicality of CatUA, we have invested considerable effort in developing the software stack on our meticulously crafted Sensing Front-end, which has successfully gathered city-scale air quality data for over 1,200 hours. Experiments conducted on the collected data demonstrate that CatUA reduces sensing errors by 96.9% with a latency of only 44.9ms, outperforming the state-of-the-art baseline by 42.6%.
Yuxuan Liu 0010, Haoyang Wang 0012, Fanhang Man, Jingao Xu, Fan Dang 0001, Chaopeng Hong, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Yali Song, Qiuhua Wang, Xinlei Chen
IEEE Trans. Mob. Comput.12
2024 R-MAE: Regions Meet Masked Autoencoders
abstract
In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.
Duy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald, Alexander Kirillov, Cees Snoek, Xinlei Chen
ICLR7
2024 QUEST: Quality-informed Multi-agent Dispatching System for Optimal Mobile Crowdsensing
abstract
We address the challenges in achieving optimal Quality of Information (QoI) for non-dedicated vehicular Mobile Crowdsensing (MCS) systems, by utilizing vehicles not originally designed for sensing purposes to provide real-time data while moving around the city. These challenges include the coupled sensing coverage and sensing reliability, as well as the uncertainty and time-varying vehicle status. To tackle these issues, we propose QUEST, a QUality-informed multi-agEnt diSpaTching system, that ensures high sensing coverage and sensing reliability in non-dedicated vehicular MCS. QUEST optimizes QoI by introducing a novel metric called ASQ (aggregated sensing quality), which considers both sensing coverage and sensing reliability jointly. Additionally, we design a mutual-aided truth discovery dispatching method to estimate sensing reliability and improve ASQ under uncertain vehicle statuses. Real-world data from our deployed MCS system in a metropolis is used for evaluation, demonstrating that QUEST achieves up to 26% higher ASQ improvement, leading to a reduction of reconstruction map errors by 32-65% for different reconstruction algorithms.
Zuxin Li, Fanhang Man, Xuecheng Chen, Susu Xu, Fan Dang 0002, Xiao-Ping Zhang 0002, Xinlei Chen
INFOCOM7
2024 TransformLoc: Transforming MAVs into Mobile Localization Infrastructures in Heterogeneous Swarms
abstract
A heterogeneous micro aerial vehicles (MAV) swarm consists of resource-intensive but expensive advanced MAVs (AMAVs) and resource-limited but cost-effective basic MAVs (BMAVs), offering opportunities in diverse fields. Accurate and real-time localization is crucial for MAV swarms, but current practices lack a low-cost, high-precision, and real-time solution, especially for lightweight BMAVs. We find an opportunity to accomplish the task by transforming AMAVs into mobile localization infrastructures for BMAVs. However, turning this insight into a practical system is non-trivial due to challenges in location estimation with BMAVs’ unknown and diverse localization errors and resource allocation of AMAVs given coupled influential factors. This study proposes TransformLoc, a new framework that transforms AMAVs into mobile localization infrastructures, specifically designed for low-cost and resource- constrained BMAVs. We first design an error-aware joint location estimation model to perform intermittent joint location estimation for BMAVs and then design a proximity-driven adaptive grouping-scheduling strategy to allocate resources of AMAVs dynamically. TransformLoc achieves a collaborative, adaptive, and cost-effective localization system suitable for large-scale heterogeneous MAV swarms. We implement TransformLoc on industrial drones and validate its performance. Results show that TransformLoc outperforms baselines including SOTA up to 68% in localization performance, motivating up to 60% navigation success rate improvement.
Haoyang Wang 0012, Jingao Xu, Chenyu Zhao 0002, Zihong Lu, Yuhan Cheng, Xuecheng Chen, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen
INFOCOM9
2024 BlueKey: Exploiting Bluetooth Low Energy for Enhanced Physical-Layer Key Generation
abstract
Bluetooth Low Energy (BLE) is a prevalent technology in various applications due to its low power consumption and wide device compatibility. Despite its numerous advantages, the encryption methods of BLE often expose devices to potential attacks. To fortify security, we investigate the application of Physical-layer Key Generation (PKG), a promising technology that enables devices to generate a shared secret key from their shared physical environment. We propose a distinctive approach that capitalizes on the inherent characteristics of BLE to facilitate efficient PKG. We harness the constant tone extension within BLE protocols to extract comprehensive physical layer information and introduce an innovative method that employs Legendre polynomial quantization for PKG. This method facilitates the exchange of secret keys with a high key matching rate and a high key generation rate. The efficacy of our approach is validated through extensive experiments on a software-defined radio platform, underscoring its potential to enhance security in the rapidly expanding field of BLE applications.
Fan Dang 0001, Jinyan Jiang, Xu Wang 0018, Lin Wang 0023, Kebin Liu 0001, Xinlei Chen, Yunhao Liu 0001
INFOCOM8
2024 Demo Abstract: CARL: Collaborative Altitude-Adaptive Reinforcement Learning for Active Search with UAV Swarms
abstract
Sensing noise and complex decision-making pose critical challenges to active search for lost persons amid disasters, impeding efficient rescue efforts. We introduce CARL, a collaborative altitude-adaptive reinforcement learning framework for UAV swarms. CARL integrates confidence-informed assessment with Sparse Bayesian Learning to diminish the noise impact on sensor performance, and an altitude-adaptive planner for collaborative active search strategy. Simulation experiments with up to 50 targets and 10 UAVs demonstrate CARL’s superior performance compared to baseline methods in lost person active search scenarios.
Chen-Chun Hsia, Yanggang Xu, Jiyuan Ren, Xinlei Chen
IPSN4
2024 Poster Abstract: Generative Modeling of Post-Disaster POI Visits Recovery
abstract
The development of Internet of Things (IoT) systems has enabled disaster perception and prediction to be highly accurate. On this basis, high-quality post-disaster Point of Interest (POI) visit data can help city decision-makers develop more sophisticated recovery plans to minimize the cost of recovery. This work focuses on the problem of POI visits generation in post-disaster recovery scenarios, utilizing diffusion model to generate visit recovery curves base on the data from sensor networks. We take the disaster severity as condition and propose a disaster mapping method to map the sensor data to each POI.
Yan Zhuo, Huandong Wang, Xinlei Chen
IPSN5
2024 Demo Abstract: Range-SLAM: UWB based Realtime Indoor Location and Mapping
abstract
Simultaneous localization and mapping (SLAM) systems frequently employ LiDAR and cameras as essential sensing components. However, these sensors are proved to be unreliable in environments with poor visibility or reflective surfaces. And UWB (Ultra Wide Band) sensor with a longer wavelength shows better potential to achieve perception tasks. However, since UWB sensors can only obtain distance information from the anchors, it is difficult to densely construct the geometric structure of the environment. In this paper, We propose Range-SLAM, a method based on received signal strength indicator (RSSI) recognition and binary filtering to complete the mapping task and enhance positioning based on the map, and only require UWB as external perception sensor. Real-world experiments are conducted and prove the effectiveness, real-time performance and robustness of the Range-SLAM algorithm.
Zhuozhu Jian, Junbo Tan, Lunfei Liang, Houde Liu, Xinlei Chen
IPSN6
2024 Demo Abstract: A Spatio-Temporal System for Public Transit-Guided Volunteer Task Matching
abstract
Volunteer activity often undergoes unique transformations with the constant changes in society. The information behind volunteer data was created to enhance public welfare efficiently and boost governmental organization productivity. This research aims to utilize public transit systems for volunteer services, reducing inequality in volunteer service provision across different regions and improving overall service efficiency. We collected and processed large-scale data related to public transit and volunteer services, conducting in-depth analysis using data mining techniques and deep learning methods. Through LDA, we annotated a large amount of volunteer data, and via data analysis, discovered patterns related to population distribution, spatial distribution, and temporal distribution. Combining public transit data and the mined features, we propose a novel spatio-temporal embedding model based on the transformer architecture, which can effectively classify and predict the matching between volunteer service demands and public transit systems. Studying the coupling between volunteer services and transportation systems helps establish a new data-driven mindset, better utilize urban resources, and provide high-quality volunteer services to the public.
Xuzhe Wang, Chengzhao Yu, Chenyu Zhao 0002, Xinlei Chen
IPSN6
2024 Poster Abstract: TCT: Zero-training two staged Contrastive Transformer network for SSVEP classification
abstract
Steady-State Evoked Potential (SSVEP) is a brain response to specific frequency visual stimuli, used in brain-computer interfaces due to its robust and easily detectable signals. Researchers have long applied methods like Canonical Correlation Analysis and deep learning for SSVEP signal decomposition and classification. However, those methods struggle to classify SSVEP signals without new subject’s data, and calibration is time-consuming. In this paper, we propose a two-stage, two-Transformer streams network to address the challenge of classifying SSVEP signals from new subjects. We utilize hierarchical contrastive learning to project features into a more discriminable feature space before classification. The comparative experiment demonstrates that our approach exhibits superior performance relative to alternative methods in processing SSVEP signals from new subjects.
Yan Zhuo, Xinlei Chen
IPSN4
2024 Poster Abstract: Sprinkler-UAV Cooperative Active Scheduling System
abstract
Urban particulate pollution presents considerable public health hazards, underscoring the need for effective control measures in various cities. A prevalent approach involves employing mobile sprinkling trucks. This paper proposes a Sprinkler-UAV Cooperative Active Scheduling System for enhanced efficiency in reducing particulate pollution. The system employs ground-based sprinkler trucks and airborne air pollution detection drones to actively explore and reduce PM2.5 in environments with dynamic and unknown pollution distributions. Preliminary experiments have demonstrated the effectiveness of using sprinklers for urban particulate matter control.
Zijian Xiao, Xuecheng Chen, Yuhan Cheng, Haoyang Wang 0012, Xinlei Chen
IPSN6
2024 Poster Abstract: Emergency Networking Using UAVs: A Reinforcement Learning Approach with Large Language Model
abstract
Utilizing unmanned aerial vehicles (UAVs) as mobile access points can assist urban communication systems in establishing emergency networks in disaster scenarios. In this paper, to organize UAVs in large-scale environments for networking purposes, we propose a multi-agent reinforcement learning (MARL) model, in which the design of a selective parameter sharing mechanism and a grouping strategy enhances the model’s scalability. Furthermore, the model adopts a reward mechanism based on intrinsic motivation, using the Large Language Model (LLM), to accelerate the optimization process. Numerical results demonstrate that this algorithm outperforms existing alternatives.
Yanggang Xu, Zhuozhu Jian, Jirong Zha, Xinlei Chen
IPSN4
2024 Demo Abstract: Embodied Aerial Agent for City-level Visual Language Navigation Using Large Language Model
abstract
As unmanned aerial vehicles (UAVs) become more prevalent in smart cities, their capacity for visual language navigation (VLN) is garnering increasing interest. VLN in cities has significant applications in delivery, rescue, and security patrol, among other fields. One of the most representative tasks is to navigate to specific locations following the language instructions. While some current methods have achieved notable results in indoor settings, challenges persist outdoors, including agents’ inaccurate spatial understanding and ambiguous language instructions. In this work, we explore an embodied navigation agent design, in which a fine-grained spatial verbalizer and a history path memory are proposed to guarantee accurate VLN in open 3D urban environments.
Yuxuan Liu 0010, Xuzhe Wang, Xuecheng Chen, Chen Gao 0001, Xinlei Chen
IPSN6
2024 Demo Abstract: Bio-inspired Tactile Sensing for MAV Landing with Extreme Low-cost Sensors
abstract
MAV (Micro Aerial Vehicle) requires landing on a docking platform for recharging during or after missions due to their limited energy capacity. Inspired by biological tactile sensing, we propose a proprioceptive sensing system that allows MAV to "touch", recognize, and locate the landing platform even when visual or other positioning systems are not functioning properly. We leverage a physical phenomenon: as the MAV approaches a beneath obstacle, it experiences attitude disturbances caused by the airflow generated by the rotor’s reflections from the ground. By employing traditional signal processing and learning-based techniques to analyze signals from the IMU (Inertial Measurement Unit) and motors, the MAV can sense the edges of the platform and further calculate the precise landing coordinates. With a power consumption of less than 40 mW, our system achieves an edge detection error of less than 2 cm and a landing success rate exceeding 90%.CCS CONCEPTS• Applied computing → Aerospace; • Computing methodologies → Machine learning approaches; • Computer systems organization → Sensors and actuators.
Chenyu Zhao 0002, Ciyu Ruan, Jirong Zha, Haoyang Wang 0012, Jiaqi Li 0028, Yuxuan Liu 0010, Xuzhe Wang, Xinlei Chen
IPSN9
2024 Poster Abstract: Adaptive Chirps Domain Window Order of MM-Wave Radar for UAV Motion Capture
abstract
Accurate motion capture of aerial robots in 3D is a key enabler for autonomous operation. Recently, some research considers using MM-Wave radar sensors for drone motion capture. However, due to the high noise and difficulty in capturing the center of an object in MM-Wave radar, the existing traditional methods have achieved unsatisfactory results. We develop a novel adaptive chirps domain window order method for MM-Wave radar data and customize a neural network architecture.
Yan Zhuo, Xinlei Chen
IPSN4
2024 Physics-informed Neural ODE for Post-disaster Mobility Recovery
abstract
Urban mobility undergoes a profound decline in the aftermath of a disaster, subsequently exhibiting a complex recovery trajectory. Effectively capturing and predicting this dynamic recovery process holds paramount importance for devising more efficient post-disaster recovery strategies, such as resource allocation to areas with protracted recovery periods. Existing models for post-disaster mobility recovery predominantly employ basic mathematical methods, which are strongly based on simplifying assumptions, and their limited parameters restrict their capacity to fully capture the mobility recovery patterns. In response to this gap, we introduce the Coupled Dynamic Graph ODE Network (CDGON) to model the intricate dynamics of post-disaster mobility recovery. Our model seamlessly integrates existing physical knowledge pertaining to post-disaster mobility recovery and incorporates the nuanced interactions between intra-regional and inter-regional population flows. Extensive experimental results demonstrate the efficiency of our model in capturing the dynamic recovery patterns of urban population mobility in post-disaster scenarios, surpassing the capabilities of current dynamic graph prediction models.
Huandong Wang, Xinlei Chen
KDD3
2024 Multi-Agent Target Pursuit Using Perception Uncertainty-Aware Reinforcement Learning
abstract
Existing target pursuit systems are able to coordinate a team of mobile agents to capture or intercept unauthorized targets. Multi-agent reinforcement learning (MARL) further empowers pursuit strategies with the potential to emerge complex behaviors. However, existing solutions lack the ability to handle the perception uncertainty caused by relative position measurement noises, which blurs the understanding of the target's state and complicates the pursuit strategy learning process. This study proposes PUARL, which enhances the learning under the perception uncertainty process by guiding exploration with probabilistic estimation and adapting the policy based on awareness of perception uncertainty. We validate its performance in terms of both accuracy and efficiency. PUARL achieves a success rate increase of 12.3%+ and a reduction in total steps by 58.3%+, outperforming both state-of-the-art heuristic and learning-based solutions.
Yuhan Cheng, Jirong Zha, Renjue Yang, Susu Xu, Xinlei Chen
MobiCom6
2024 FormerReckoning: Physics Inspired Transformer for Accurate Inertial Navigation
abstract
Although modern localization methods have achieved remarkable accuracy with various sensors, there are still some circumstances where only proprioceptive sensing works (Inertial Navigation). However, localization and navigation using only IMU sensors (costing less than $1000) still face significant challenges such as low accuracy and large cumulative errors when using traditional filter methods. Furthermore, AI-based approaches, while promising, often yield unpredictable and unreliable outputs. This paper proposes FormerReckoning, an inertial localization estimation framework for wheeled robotics that incorporates physical prompts into a Transformer framework to enhance translation estimation accuracy. Our tests show that FormerReckoning not only reduces mean translation errors to 0.72% but also surpasses all baseline models in performance, demonstrating its potential to provide reliable and precise localization in a cost-effective manner.
Jiaqi Li 0028, Chenyu Zhao 0002, Yuzhu Mao, Xinlei Chen, Wenbo Ding 0001, Xiaoyang Qu, Jianzong Wang
MobiCom4
2024 EventTracker: 3D Localization and Tracking of High-Speed Object with Event and Depth Fusion
abstract
Accurately localizing high-speed dynamic objects in 3D space with low latency is crucial for various robotic applications. Current methods face challenges due to extended exposure times and limited sensor resolution, hindering precise object detection and localization. Event cameras, known for their high temporal resolution and asynchronous nature, offer a promising solution. To leverage the potential of the event camera, we propose EventTracker, a novel framework that integrates event and depth measurements for precise and low-latency 3D localization and tracking of the high-speed dynamic object. EventTracker incorporates a collaborative object detection and tracking algorithm optimized for both event and depth data, overcoming detection and registration challenges. Additionally, a graph-instructed optimization algorithm enhances accuracy by fusing heterogeneous sensor data effectively. Experimental evaluation in dynamic environments demonstrates significant improvements in localization performance compared to baseline methods.
Xinyu Luo, Haoyang Wang 0012, Ciyu Ruan, Chenxin Liang, Jingao Xu, Xinlei Chen
MobiCom6
2024 Distill Drops into Data: Event-based Rain-Background Decomposition Network
abstract
Event cameras excel in high-speed and high-dynamic-range scenarios but are highly sensitive to rain, which introduces significant noise while also revealing detailed rain features. This paper introduces a novel Event-based Rain-Background Decomposition Network that integrates Spiking Neural Networks (SNNs) and Convolutional Neural Networks (CNNs). By "Distilling Rain," we reconstruct a rain-free background for downstream tasks, and by "Collecting Rain," we extract the physical characteristics of rain. Experimental evaluations demonstrate the network's effectiveness in both background reconstruction and rain modeling. This work extends the capabilities of event cameras by mitigating the adverse effects of rain while also leveraging rain-induced noise to extract valuable environmental data, enhancing their utility in both challenging weather conditions and detailed environmental analysis.
Ciyu Ruan, Chenyu Zhao 0002, Chenxin Liang, Xinyu Luo, Jingao Xu, Xinlei Chen
MobiCom6
2024 Scalable Multi-Agent Reinforcement Learning for Effective UAV Scheduling in Multi-Hop Emergency Networks
abstract
Utilizing unmanned aerial vehicles (UAVs) as mobile access points can assist urban communication systems in establishing emergency networks in disaster scenarios. However, in large-scale dynamic environments, the extensive exploration space makes effective collaboration among a large number of UAVs challenging. In this paper, to schedule the deployment of UAVs for networking purposes, we propose a novel approach, MAEN, using multi-agent reinforcement learning. The grouping and information sharing mechanisms in MAEN enable the algorithm to easily scale up the number of UAVs to dozens and address the issue of strategy equilibrium. Additionally, a reward decomposition module is designed to handle coordination and task allocation among UAVs. Experimental results demonstrate that the algorithm outperforms existing algorithms in terms of ground device coverage and communication quality.
Yanggang Xu, Jirong Zha, Jiyuan Ren, Xintao Jiang, Xinlei Chen
MobiCom6
2024 Foes or Friends: Embracing Ground Effect for Edge Detection on Lightweight Drones
abstract
Drone-based rapid and accurate environmental edge detection is highly advantageous for tasks such as disaster relief and autonomous navigation. Current methods, using radar or cameras, raise deployment costs and burden lightweight drones with high computational demands. In this paper, we propose AirTouch, a system that transforms the ground effect from a stability "foe" in traditional flight control views, into a "friend" for accurate and efficient edge detection. Our key insight is that analyzing drone sensor readings and flight commands allows us to detect ground effect changes. Such changes typically indicate the drone flying over an edge, making this information valuable for edge detection. We approach this insight through theoretical analysis, algorithm design, and implementation, fully leveraging the ground effect as a new sensing modality without compromising drone flight stability, thereby achieving accurate and efficient scene edge detection. Extensive evaluations demonstrate that our system achieves a high detection accuracy with mean detection distance errors of 0.051m, outperforming the baseline performance by 86%.
Chenyu Zhao 0002, Ciyu Ruan, Jingao Xu, Haoyang Wang 0012, Jiaqi Li 0028, Jirong Zha, Zheng Yang 0002, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen
MobiCom11
2024 MobiAir: Unleashing Sensor Mobility for City-scale and Fine-grained Air-Quality Monitoring with AirBERT
abstract
Mobile air pollution sensing methods are developed to collect air quality data with higher spatial-temporal resolutions. However, existing methods cannot process the spatially mixed gas samples effectively due to the highly dynamic temporal and spatial fluctuations experienced by the sensor, leading to significant measurement deviations. We find an opportunity to tackle the problem by exploring the potential patterns from sensor measurements. In light of this, we propose MobiAir, a novel city-scale fine-grained air quality estimation system to deliver accurate mobile air quality data. First, we design AirBERT, a representation learning model to discern mixed gas concentrations. Second, we design a knowledge-informed training strategy leveraging massive unlabeled city-scale data to enhance the AirBERT performance. To ensure the practicality of MobiAir, we have invested significant efforts in implementing the software stack on our meticulously crafted Sensing Front-end, which has successfully gathered air quality data at a city-scale for more than 1200 hours. Experiments conducted on collected data show that MobiAir reduces sensing errors by 96.7% with only 44.9ms latency, outperforming the SOTA baseline by 39.5%.
Yuxuan Liu 0010, Haoyang Wang 0012, Fanhang Man, Jingao Xu, Fan Dang 0001, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen
MobiSys8
2024 On the Surprising Effectiveness of Attention Transfer for Vision Transformers
abstract
Conventional wisdom suggests that pre-training Vision Transformers (ViT) improves downstream performance by learning useful representations. Is this actually true? We investigate this question and find that the features and representations learned during pre-training are not essential. Surprisingly, using only the attention patterns from pre-training (i.e., guiding how information flows between tokens) is sufficient for models to learn high quality features from scratch and achieve comparable downstream performance. We show this by introducing a simple method called attention transfer, where only the attention patterns from a pre-trained teacher ViT are transferred to a student, either by copying or distilling the attention maps. Since attention transfer lets the student learn its own features, ensembling it with a fine-tuned teacher also further improves accuracy on ImageNet. We systematically study various aspects of our findings on the sufficiency of attention maps, including distribution shift settings where they underperform fine-tuning. We hope our exploration provides a better understanding of what pre-training accomplishes and leads to a useful alternative to the standard practice of fine-tuning.
Alexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak, Xinlei Chen
NeurIPS5
2024 Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
abstract
One of the roadblocks for training generalist robotic models today is heterogeneity. Previous robot learning methods often collect data to train with one specific embodiment for one task, which is expensive and prone to overfitting. This work studies the problem of learning policy representations through heterogeneous pre-training on robot data across different embodiments and tasks at scale. We propose Heterogeneous Pre-trained Transformers (HPT), which pre-train a large, shareable trunk of a policy neural network to learn a task and embodiment agnostic shared representation. This general architecture aligns the specific proprioception and vision inputs from distinct embodiments to a short sequence of tokens and then processes such tokens to map to control robots for different tasks. Leveraging the recent large-scale multi-embodiment real-world robotic datasets as well as simulation, deployed robots, and human video datasets, we investigate pre-training policies across heterogeneity. We conduct experiments to investigate the scaling behaviors of training objectives, to the extent of 52 datasets. HPTs outperform several baselines and enhance the fine-tuned policy performance by over 20% on unseen tasks in multiple simulator benchmarks and real-world settings. See the project website (liruiw.github.io/hpt) for code and videos.
Lirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming He
NeurIPS2
2024 A novel image inpainting method based on a modified Lengyel-Epstein model
Mengyu Luo, Xinlei Chen, Heming Xu
Comput. Vis. Image Underst.3
2024 SOScheduler: Toward Proactive and Adaptive Wildfire Suppression via Multi-UAV Collaborative Scheduling
abstract
Multi-UAV systems have shown immense potential in handling complex tasks in large-scale, dynamic, and cold-start (i.e., limited prior knowledge) scenarios, such as wildfire suppression. Due to the dynamic and stochastic environmental conditions, the scheduling for sensing tasks (i.e., fire monitoring) and operation tasks (i.e., fire suppression) should be executed concurrently to enable real-time information collection and timely intervention of the environment. However, the planning inclinations of sensing and operation tasks are typically inconsistent and evolve over time, complicating the task of identifying the optimal strategy for each UAV. To solve this problem, this paper proposes SOScheduler, a collaborative multi-UAV scheduling framework for integrated sensing and operation in large-scale and dynamic wildfire environments. We introduce a spatio-temporal confidence-aware assessment model to dynamically and directly pinpoint locations that can optimally enhance the understanding of environmental dynamics and operational effectiveness, as well as a priority graph-instructed scalable scheduler to coordinate multi-UAV in an efficient manner. Experiments on real multi-UAV testbeds and large-scale physical feature-based simulations show that our SOScheduler reduces the fire expansion ratio by 59% and enhances the fire coverage ratio by 190% compared to state-of-the-art (SOTA) solutions.
Xuecheng Chen, Zijian Xiao, Yuhan Cheng, Chen-Chun Hsia, Haoyang Wang 0012, Jingao Xu, Susu Xu, Fan Dang 0001, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen
IEEE Internet Things J.11
2024 StreamingTag: A Scalable Piracy Tracking Solution for Mobile Streaming Services
abstract
Streaming services have billions of mobile subscribers, yet video piracy has cost service providers billions. Digital Rights Management (DRM), however, is still far from satisfactory. Unlike DRM, which attempts to prohibit the creation of pirated copies, fingerprinting may be used to track out the source of piracy. Nevertheless, existing fingerprinting-based streaming systems are not widely used since they fail to serve numerous users. In this paper, we present the design and evaluation of StreamingTag, a scalable piracy tracing system for mobile streaming services. StreamingTag adopts a segment-level fingerprint embedding scheme to remove the need of re-embedding the fingerprint into the video for each new viewer. The key innovations of StreamingTag include a scalable and CDN-friendly delivery framework, an accurate and lightweight temporal synchronization scheme, a polarized and randomized SVD watermarking scheme, and a collusion-resistant fingerprinting scheme. Experiment results show the good QoS of StreamingTag in terms of preparation latency, bandwidth consumption, and video fidelity. Compared with existing methods, the proposed three schemes improve the re-identification accuracy by 4-49x, the watermark extraction accuracy by 2.25x at most and 1.5x on average, and the recall rate of catching colluders by 26%.
Fan Dang 0001, Xinqi Jin, Qi-An Fu, Lingkun Li, Guanyan Peng, Xinlei Chen, Kebin Liu 0001, Yunhao Liu 0001
IEEE Trans. Mob. Comput.6
2024 LSync: A Universal Timeline-Synchronizing Solution for Live Streaming
abstract
The widespread use of intelligent devices and the development of mobile networks have led to the increasing popularity of live-streaming services worldwide. In addition to video and audio transmissions, a wide range of media content is also sent to audiences, such as player statistics for sports streams and subtitles for live news. However, due to the diverse transmission process between live streams and other media content, synchronizing them has become a significant challenge. Unfortunately, existing commercial solutions are not universal, requiring specific server cloud services or CDNs and limiting users’ free choices of web infrastructures. To address this issue, we propose a lightweight and universal solution called LSync, which inserts a series of audio signals containing metadata into the original audio stream. Based on the embedded metadata, a well-designed timeline-synchronizing solution helps to synchronize the information stream to the live stream. It brings no modifications to the original live broadcast process and thus fits prevalent live broadcast infrastructures. Evaluations show that the proposed solution reduces the signal processing delay to around 5% of an audio buffer length in mobile phones and ensures real-time signal processing. It achieves a channel utilization of more than 150 bps/kHz in a specific configuration, greatly outperforming recent works. Furthermore, the proposed synchronization mechanism reaches a precision of 24.84 ms on average, which matches people’s viewing habits.
Fan Dang 0001, Yifan Xu 0023, Rongwu Xu, Xinlei Chen, Yunhao Liu 0001
IEEE/ACM Trans. Netw.4
2024 BEANet: An Energy-efficient BLE Solution for High-capacity Equipment Area Network
abstract
The digital transformation of factories has greatly increased the number of peripherals that need to connect to a network for sensing or control, resulting in a growing demand for a new network category known as the Equipment Area Network (EAN). The EAN is characterized by its cable-free, high-capacity, low-latency, and low-power features. To meet these expectations, we presentBEANet, a novel solution designed specifically for EAN that combines a two-stage synchronization mechanism with a time division protocol. We implemented the system using commercially available Bluetooth Low Energy (BLE) modules and evaluated its performance. Our results show that the network can support up to 150 peripherals with a packet reception rate of 95.4%, which is only 0.9% lower than collision-free BLE transmission. When the cycle time is set to 2 s, the average transmission latency for all peripherals is 0.1 s, while the power consumption is 18.9 μW, which is only half that of systems using LLDN or TSCH. Simulation results also demonstrate that BEANet has the potential to accommodate over 30,000 peripherals under certain configurations.
Yifan Xu 0023, Fan Dang 0001, Kebin Liu 0001, Zhui Zhu, Xinlei Chen, Xu Wang 0018, Haitian Zhao
ACM Trans. Sens. Networks5
2023 Improving Selective Visual Question Answering by Learning from Your Peers
abstract
Despite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%.
Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach
CVPR6
2023 ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders
abstract
Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt [33], have demonstrated strong performance in various scenarios. While these models were originally designed for supervised learning with ImageNet labels, they can also potentially benefit from self-supervised learning techniques such as masked autoencoders (MAE) [14]. However, we found that simply combining these two approaches leads to subpar performance. In this paper, we propose a fully convolutional masked autoencoder framework and a new Global Response Normalization (GRN) layer that can be added to the ConvNeXt architecture to enhance inter-channel feature competition. This co-design of self-supervised learning techniques and architectural improvement results in a new model family called ConvNeXt V2, which significantly improves the performance of pure ConvNets on various recognition benchmarks, including ImageNet classification, COCO detection, and ADE20K segmentation. We also provide pre-trained ConvNeXt V2 models of various sizes, ranging from an efficient 3.7M-parameter Atto model with 76.7% top-1 accuracy on ImageNet, to a 650M Huge model that achieves a state-of-the-art 88.9% accuracy using only public training data.
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu 0003, In-So Kweon, Saining Xie
CVPR4
2023 UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
abstract
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts on connecting these two related tasks with highly task-specific neural modules, it remains understudied how to explicitly depict their shared nature to learn them simultaneously. In this work, we propose UniT3D, a simple yet effective fully unified transformer-based architecture for jointly solving 3D visual grounding and dense captioning. UniT3D enables learning a strong multimodal representation across the two tasks through a supervised joint pre-training scheme with bidirectional and seq-to-seq objectives. With a generic architecture design, UniT3D allows expanding the pre-training scope to more various training sources such as the synthesized data from 2D prior knowledge to benefit 3D vision-language tasks. Extensive experiments and analysis demonstrate that UniT3D obtains significant gains for 3D dense captioning and visual grounding.
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, Angel X. Chang
ICCV3
2023 Poster Abstract: TENG-enabled Self-powered Human-machine Interfaces for the Metaverse
abstract
Human-machine interface (HMI) of high degrees of freedom (DoF) is one of the most critical bases of the metaverse. The ideal HMI for the metaverse should be cheap, robust, customizable, and ergonomically friendly. In light of this, we propose a triboelectric nanogenerator (TENG)-based sensing system. We developed a low-cost, soft, light, and customizable TENG sensor to collect data from the human body. We then used an artificial neural network (ANN) to obtain the corresponding human motion from collected sensory data. The effectiveness of the proposed system is demonstrated with experiments of a working prototype.
Haoyang Wang 0012, Fanhang Man, Yuxuan Liu 0010, Xinlei Chen, Wenbo Ding 0001
IPSN5
2023 Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement Learning
abstract
The learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictability, and non-linearity. Despite considerable progress made in swarm robotics, addressing these challenges remains a significant issue. In this work, we model the stochastic dynamics of a swarm robot system and then propose a novel control framework based on a mean-field control (MFC) embedding multi-agent reinforcement learning (MARL) approach named MF-MARL to deal with these challenges. While MARL is able to deal with stochasticity statistically, we integrate MFC, allowing MF-MARL to cope with large-scale robots. Moreover, we apply statistical moments of robots' state and control action to discretize continuous input and enable MF-MARL to be applied in continuous scenarios. To demonstrate the effectiveness of MF-MARL, we evaluate the performance of the robots on a specific swarm simulation platform. The experimental results show that our algorithm outperforms the traditional algorithms both in navigation and manipulation tasks. Finally, we demonstrate the adaptability of the proposed algorithm through the component failure test.
Huaze Tang, Hengxi Zhang, Zhenpeng Shi, Xinlei Chen, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IROS4
2023 Image manipulation detection by multiple tampering traces and edge artifact enhancement
Xun Lin, Shuai Wang 0049, Jiahao Deng, Ying Fu 0001, Xiao Bai 0001, Xinlei Chen, Xiaolei Qu, Wenzhong Tang
Pattern Recognit.6
2022 Point-Level Region Contrast for Object Detection Pre-Training
abstract
In this work we present point-level region contrast, a self-supervised pre-training approach for the task of object detection. This approach is motivated by the two key factors in detection: localization and recognition. While accurate localization favors models that operate at the pixel- or point-level, correct recognition typically relies on a more holistic, region-level view of objects. Incorporating this perspective in pre-training, our approach performs contrastive learning by directly sampling individual point pairs from different regions. Compared to an aggregated representation per region, our approach is more robust to the change in input region quality, and further enables us to implicitly improve initial region assignments via online knowledge distillation during training. Both advantages are important when dealing with imperfect regions encountered in the unsupervised setting. Experiments show point-level region contrast improves on state-of-the-art pre-training methods for object detection and segmentation across multiple tasks and datasets, and we provide extensive ablation studies and visualizations to aid understanding. Code will be made available.
Yutong Bai, Xinlei Chen, Alexander Kirillov, Alan L. Yuille, Alexander C. Berg
CVPR2
2022 Masked Autoencoders Are Scalable Vision Learners
abstract
This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3× or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pretraining and shows promising scaling behavior.
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross B. Girshick
CVPR2
2022 On the Importance of Asymmetry for Siamese Representation Learning
abstract
Many recent self-supervised frameworks for visual representation learning are based on certain forms of Siamese networks. Such networks are conceptually symmetric with two parallel encoders, but often practically asymmetric as numerous mechanisms are devised to break the symmetry. In this work, we conduct a formal study on the importance of asymmetry by explicitly distinguishing the two encoders within the network - one produces source encodings and the other targets. Our key insight is keeping a relatively lower variance in target than source generally benefits learning. This is empirically justified by our results from five case studies covering different variance-oriented designs, and is aligned with our preliminary theoretical analysis on the baseline. Moreover, we find the improvements from asymmetric designs generalize well to longer training schedules, multiple other frameworks and newer backbones. Finally, the combined effect of several asymmetric designs achieves a state-of-the-art accuracy on ImageNet linear probing and competitive results on downstream transfer. We hope our exploration will inspire more research in exploiting asymmetry for Siamese representation learning.
Haoqi Fan 0001, Yuandong Tian, Daisuke Kihara, Xinlei Chen
CVPR5
2022 NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training
Chengyue Gong, Dilin Wang, Meng Li 0004, Xinlei Chen, Zhicheng Yan 0001, Yuandong Tian, Qiang Liu 0001, Vikas Chandra
ICLR4
2022 LSync: A Universal Event-synchronizing Solution for Live Streaming
abstract
The widespread of smart devices and the development of mobile networks brings the growing popularity of live streaming services worldwide. In addition to the video and audio transmission, a lot more media content is sent to the audiences as well, including player statistics for a sports stream, subtitles for living news, etc. However, due to the diverse transmission process between live streams and other media content, the synchronization of them has grown to be a great challenge. Unfortunately, the existing commercial solutions are not universal, which require specific server cloud services or CDN and limit the users’ free choices of web infrastructures. To address the issue, we propose a lightweight universal event-synchronizing solution for live streaming, called LSync, which inserts a series of audio signals containing metadata into the original audio stream. It brings no modification to the original live broadcast process and thus fits prevalent live broadcast infrastructure. Evaluations on real system show that the proposed solution reduces the signal processing delay by at most 5.62% of an audio buffer length in mobile phones and ensures real-time signal processing. It also achieves a data rate of 156.25 bps in a specific configuration and greatly outperforms recent works.
Yifan Xu 0023, Fan Dang 0001, Rongwu Xu, Xinlei Chen, Yunhao Liu 0001
INFOCOM4
2022 StreamingTag: a scalable piracy tracking solution for mobile streaming services
abstract
Streaming services have billions of mobile subscribers, yet video piracy has cost service providers billions. Digital Rights Management (DRM), however, is still far from satisfactory. Unlike DRM, which attempts to prohibit the creation of pirated copies, fingerprinting may be used to track out the source of piracy. Nevertheless, the idea of piracy tracing is not widely used at the moment, since existing fingerprinting-based streaming systems fail to serve numerous users. In this paper, we present the design and evaluation of StreamingTag, a scalable piracy tracing system for mobile streaming services. StreamingTag adopts a segment-level fingerprint embedding scheme to remove the need of re-embedding the fingerprint into the video for each new viewer. The key innovations of StreamingTag include a scalable and CDN-friendly delivery framework, a polarized and randomized SVD watermarking scheme suitable for short segments, and a collusion-resistant fingerprinting scheme optimized for large-scale streaming services. Experiment results show the good QoS of StreamingTag in terms of preparation latency, bandwidth consumption, and video fidelity. Compared with existing SVD watermarking schemes, the proposed watermarking scheme improves the watermark extraction accuracy by 2.25x at most and 1.5x on average. Compared with existing collusion-resistant fingerprinting schemes, the proposed scheme catches more colluders and improves the recall rate by 26%.
Xinqi Jin, Fan Dang 0001, Qi-An Fu, Lingkun Li, Guanyan Peng, Xinlei Chen, Kebin Liu 0001, Yunhao Liu 0001
MobiCom6
2022 ST-ICM: spatial-temporal inference calibration model for low cost fine-grained mobile sensing
abstract
In order to reduce the measurement error of low cost sensor in the real-time mobile sensing network, rendezvous calibration mechanism is widely used. To tackle the sparsity of reference data and the lack of calibration opportunities, we propose ST-ICM: a Spatial-Temporal Inference Calibration Model based on Gaussian Process Regression, assisting the calibration task by creating more calibration grids in both spatial and temporal dimensions. By using the GPR, the inferred grids generated by ST-ICM are associated with various confidence levels. Based on this property, we propose to make use of a hyperparameter, i.e., variance threshold, to balance the tradeoff between the quantity and quality of the inferred grids. Specifically, only the grids with variances below the threshold will be employed. We conducted experiments using a real-world dataset collected in Nanjing, China, to evaluate the performance of the proposed ST-ICM. The experimenal results show that our model achieves 24% improvement on error calibration compared to the baseline.
Chengzhao Yu, Rongye Shi, Xinyu Liu 0003, Fan Dang 0001, Xinlei Chen
MobiCom6
2022 Test-Time Training with Masked Autoencoders
abstract
Test-time training adapts to a new test distribution on the fly by optimizing a model for each test input using self-supervision.In this paper, we use masked autoencoders for this one-sample learning problem.Empirically, our simple method improves generalization on many visual benchmarks for distribution shifts.Theoretically, we characterize this improvement in terms of the bias-variance trade-off.
Yossi Gandelsman, Yu Sun 0020, Xinlei Chen, Alexei A. Efros
NeurIPS3
2022 Riemannian Geometric Instance Filtering for Transfer Learning in Brain-Computer Interfaces
abstract
Due to the inter-subject variability of Electroencephalogram(EEG) signals, a long calibration time is required to collect a large number of labeled trials to calibrate classifier parameters before using the Brain-computer Interface(BCI). This challenge greatly limits the practical roll-out of BCIs. To address this problem, we propose a novel instance-based transfer learning framework named Riemannian Geometric Instance Filtering (RGIF) to reduce calibration time without sacrificing accuracy. A new inter-subject similarity metric based on Riemannian geometry is proposed to measure the similarity between a few trials from the target subject and adequate trials from source subjects. The classification model for the target subject is then trained with the help of abundant trials from similar source subjects with high similarity to the target subject. We evaluate our method on two open-source EEG datasets. The results show that our approach improves significantly compared with other baselines. Furthermore, compared with using all source subjects data, our method reduces the training time by at least half and achieves slightly better accuracy.
Qianxin Hui, Yang Li 0104, Susu Xu, Shuailei Zhang, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng
SenSys8
2022 Fine-Grained Air Pollution Data Enables Smart Living and Efficient Management
abstract
Fine-grained air pollution data is essential for smart living and efficient city management. However, it is arduous to obtain accurate air pollution data with high spatial and temporal resolutions via mobile crowdsensing (MCS) under limited budgets. Thus, we propose FAD, a system fully using fine-grained air pollution data to provide diverse services. Moreover, a low-cost yet highly accurate portable sensing device is designed for MCS applications to enhance data resolutions. Finally, we demonstrate various FAD-based services for citizens and governments in the real world.
Yuxuan Liu 0010, Xinyu Liu 0003, Fanhang Man, Chenye Wu, Xinlei Chen
SenSys5
2022 C-RIDGE: Indoor CO2 Data Collection System for Large Venues Based on prior Knowledge
abstract
CO2 concentration data with high resolution in large venues is highly required during indoor sport events for in-time environment adjustment to guarantee the athlete performances and audience experience. However, the limited battery energy of the wireless sensors cannot support high data resolution and long time coverage simultaneously. Besides, there also lacks effective embedded methods to clean anomaly data caused by the human and environmental factors probably occurring in large venues. Thus, in this paper, we propose C-RIDGE, a low-power sensing system for high resolution CO2 data collection in large venues. Based on prior knowledge, firstly, an adaptive sampling rate adjustment policy is developed for lower energy consumption to extend the time coverage of data. Secondly, CO2 physical property (CPP) aided data cleaning algorithm is designed to improve data quality as well, using Pearson Correlation Coefficient (PCC) and standard deviation with sliding windows. C-RIDGE has been deployed in one venue during a world-class event. The experiments and collected data have shown the system power consumption can be reduced by 36.1%, with measurement error less than 10.2%. The outliers and anomaly trends can also be detected and calibrated effectively via CPP algorithm. The dataset is available at https://doi.org/10.5281/zenodo.7160830.
Yuxuan Liu 0010, Xiaolei Qu, Dezhi Zheng, Xinlei Chen
SenSys6
2022 H-SwarmLoc: Efficient Scheduling for Localization of Heterogeneous MAV Swarm with Deep Reinforcement Learning
abstract
Emergency rescue scenarios are considered to be high-risk scenarios. Using a micro air vehicle (MAV) swarm to explore the environment can provide valuable environmental information. However, due to the absence of localization infrastructure and the limited on-board capabilities, it's challenging for the low-cost MAV swarm to maintain precise localization. In this paper, a collaborative localization system for the low-cost heterogeneous MAV swarm is proposed. This system takes full advantage of advanced MAV to effectively achieve accurate localization of the heterogeneous MAV swarm through collaboration. Subsequently, H-SwarmLoc, a reinforcement learning-based planning method is proposed to plan the advanced MAV with a non-myopic objective in real-time. The experimental results show that the localization performance of our method improves 40% on average compared with baselines.
Haoyang Wang 0012, Xuecheng Chen, Yuhan Cheng, Chenye Wu, Fan Dang 0001, Xinlei Chen
SenSys6
2022 Non-Acoustic Speech Sensing System Based on Flexible Piezoelectric
abstract
Speech is one of the most important biological signals to complement human-human and human-computer interaction. Traditional speech datasets were collected by air microphones, but using these datasets in noisy environments such as factories is practically challenging. Therefore, speech recognition in noisy environments poses higher requirements. The non-acoustic speech dataset plays a significant role in robust speech recognition under high background noise. Existing datasets suffered from dull sound, low intelligibility and poor recognition accuracy due to hardware and computer technology limitations. This paper presents a non-acoustic speech sensing system based on flexible piezoelectric. The system collected vibration signals from the jaws of six males and five females, and the corpus contained ten different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and capable of achieving 93.7% recognition accuracy by calculation. With the aforementioned benefits, this dataset is an essential tool for studying human-computer interaction in high-noise environments, analyzing human acoustic properties, and aiding medical rehabilitation.
Shiji Yuan, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng, Shangchun Fan
SenSys4
2022 A Wearable Low-Power Collaborative Sensing System for High-Quality SSVEP-BCI Signal Acquisition
abstract
The brain–computer interface (BCI) technology improves the communication efficiency between people and Internet of Things (IoT) devices. BCI based on the steady-state visual evoked potential (SSVEP-BCI) is the preferred scheme for controlling devices because of its convenient operation, low training requirement, and high information transmission rate (ITR). Most signal acquisition devices for BCIs are used for medical diagnosis and scientific research and utilize multiple channels and wet electrodes to obtain high-quality signals. However, the practicability, wearability, and cost of the signal acquisition devices for real-life applications need to be considered, resulting in new requirements for the acquisition mode, the number of electrodes, power consumption, and signal processing methods. This article presents a wearable low-power collaborative sensing system based on a time mask window canonical correlation analysis method (TMW-CCA). An 8-array spring dry electrode signal acquisition device based on a flexible circuit board is designed to address the shortcomings of traditional wet electrode acquisition devices, such as high-power consumption, discomfort, and being unsuitable for long-time use. The proposed TMW-CCA method, which uses a dry electrode sensor to evaluate the time domain’s signal quality dynamically, exhibits 12.5% higher steady-state visual evoked potential recognition accuracy and 40% lower average power consumption (only 740 mW) than the benchmark.
Rui Na, Dezhi Zheng, Ying Sun 0012, Mingzhe Han, Shuai Wang 0049, Shuailei Zhang, Qianxin Hui, Xinlei Chen, Jun Zhang 0007, Chun Hu
IEEE Internet Things J.8
2022 Ultralow-Power Sensing Framework for Internet of Things: A Smart Gas Meter as a Case
abstract
Gas serves as one of the most indispensable energy sources for industrial production and household life. In order to improve the gas utilization efficiency, smart gas meters for the Internet of Things (IoT) has been designed to achieve two-way communication and remote control functions. However, the existing smart gas meter does not consider further reduction of power consumption. Thus, we designed an ultralow-power sensing framework for IoT and applied it to the smart gas meter. Based on a thorough analysis of the metering system, we propose an ultralow power system framework and a low-power peripheral management solution to reach ultralow power consumption. Also, we propose a cooperative sensing scheme to achieve stability and accuracy of gas volume detection. Finally, we implement a real smart gas manage system to evaluate our low-power smart gas meter solution. Verified by a comparative test, the proposed gas meter successfully reduces the power consumption by at least 37% compared to the baseline. The experimental results verify the innovative design and confirm that the proposed gas meter features ultralow power consumption, high precision, and high reliability.
Ziteng Wang 0007, Chun Hu, Dezhi Zheng, Xinlei Chen
IEEE Internet Things J.4
2022 TCACNet: Temporal and channel attention convolutional network for motor imagery classification of EEG-based BCI
abstract
Brain–computer interface (BCI) is a promising intelligent healthcare technology to improve human living quality across the lifespan, which enables assistance of movement and communication, rehabilitation of exercise and nerves, monitoring sleep quality, fatigue and emotion. Most BCI systems are based on motor imagery electroencephalogram (MI-EEG) due to its advantages of sensory organs affection, operation at free will and etc. However, MI-EEG classification, a core problem in BCI systems, suffers from two critical challenges: the EEG signal’s temporal non-stationarity and the nonuniform information distribution over different electrode channels. To address these two challenges, this paper proposes TCACNet, a temporal and channel attention convolutional network for MI-EEG classification. TCACNet leverages a novel attention mechanism module and a well-designed network architecture to process the EEG signals. The former enables the TCACNet to pay more attention to signals of task-related time slices and electrode channels, supporting the latter to make accurate classification decisions. We compare the proposed TCACNet with other state-of-the-art deep learning baselines on two open source EEG datasets. Experimental results show that TCACNet achieves 11.4% and 7.9% classification accuracy improvement on two datasets respectively. Additionally, TCACNet achieves the same accuracy as other baselines with about 50% less training data. In terms of classification accuracy and data efficiency, the superiority of the TCACNet over advanced baselines demonstrates its practical value for BCI systems.
Rongye Shi, Qianxin Hui, Susu Xu, Shuai Wang 0049, Rui Na, Ying Sun 0012, Wenbo Ding 0001, Dezhi Zheng, Xinlei Chen
Inf. Process. Manag.10
2022 Cyber-Resilient Multi-Energy Management for Complex Systems
abstract
Resilience problems from cyber-attacks on information communication technologies exist under their wide usage. False data injection (FDI) judiciously designed by attackers may cause severe consequences such as uneconomic operation and blackouts, particularly multivector energy distribution systems (MEDS), which are closely linked and interdependent. This article addresses the cyber resilient issues of an MEDS caused by FDI, considering the uncertainty from renewable resources. A novel two-stage distributionally robust optimization (DRO) is proposed to realize the day-ahead and real-time resilience improvement. The ambiguity set is based on both the Wasserstein distance and moment information. Compared to robust optimization which considers the worst case, DRO yields less-conservative solutions and thus provides more economic operation schemes. The Wasserstein metric-based ambiguity set enables to provide additional flexibility hedging against renewable uncertainty. Case studies are demonstrated on two representative MEDS networked with energy hubs, illustrating the effectiveness of the proposed cybersecured model. The produced adaptive robust economic operation for MEDS can reduce load shedding and enhance system resilience against severe cyberattacks.
Alexis Pengfei Zhao, Zhidong Cao, Daniel Dajun Zeng, Chenghong Gu, Zhaoyu Wang 0001, Yue Xiang, Meysam Qadrdan, Xinlei Chen, Xiaohe Yan, Shuangqi Li
IEEE Trans. Ind. Informatics8
2022 Adaptive Hybrid Model-Enabled Sensing System (HMSS) for Mobile Fine-Grained Air Pollution Estimation
abstract
Fine-grained city-scale outdoor air pollution maps provide important environmental information for both city managers and residents. Installing portable sensors on vehicles (e.g., taxis, Ubers) provides a low-cost, easy-maintenance, and high-coverage approach to collecting data for air pollution estimation. However, as non-dedicated platforms, vehicles like taxis usually prefer gathering at busy areas of a city where it is more likely to pick up riders. This leaves many parts of the city unsensed or less-sensed. In addition, due to the natural changes in a city and the movements of the vehicles, the sensed and unsensed areas change over time. Consequently, challenges of air pollution estimation with data collected by non-dedicated mobile platforms are twofold:i.data coverage is sparse;ii.data coverage changes over time. Therefore, the major research question is: how can we derive accurate and robust fine-grained field (e.g., air pollution) estimation given dynamic and sparse data collected from uncontrollable mobile sensing platforms? This paper presents adaptiveHMSS, an adaptivehybridmodel-enabledsensingsystem for fine-grained air pollution estimation with dynamic and sparse data collected from uncontrollable mobile sensing platforms, which is achieved by combining the advantages of aphysics guided modeland adata driven model. To address the challenge of sparse coverage, the physical understanding of the spatiotemporal correlation for air pollution distribution in thephysics guided modelis utilized to infer values at unsensed sparse areas. Meanwhile, thedata driven modelis adopted to estimate the air pollution influential factors (e.g., buildings) not included in thephysics guided model. To address the challenge of time-varying coverage, an adaptive model combination algorithm is designed to enable the system bias to either of the two models according to the amount of data collection and uncertainty of the model. To evaluate the system performance, we deployed 47 air pollution sensing devices on taxis and fixed locations in 2 cities for both controlled and uncontrolled experiments for over two weeks. The results show that with a resolution of$500 \;\mathrm m$by$500\;\mathrm m$by$1\;\mathrm {hour}$, our system achieves up to$3.2\times$error reduction when compared to the baseline approaches.
Xinlei Chen, Susu Xu, Xinyu Liu 0003, Xiangxiang Xu 0001, Hae Young Noh, Lin Zhang 0001, Pei Zhang 0001
IEEE Trans. Mob. Comput.1
2021 Exploring Simple Siamese Representation Learning
abstract
Siamese networks have become a common structure in various recent models for unsupervised visual representation learning. These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing. We provide a hypothesis on the implication of stop-gradient, and further show proof-of-concept experiments verifying it. Our "SimSiam" method achieves competitive results on ImageNet and downstream tasks. We hope this simple baseline will motivate people to rethink the roles of Siamese architectures for unsupervised representation learning. Code is made available.1
Xinlei Chen, Kaiming He
CVPR1
2021 KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA
abstract
One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image. In this work we study open-domain knowledge, the setting when the knowledge required to answer a question is not given/annotated, neither at training nor test time. We tap into two types of knowledge representations and reasoning. First, implicit knowledge which can be learned effectively from unsupervised language pretraining and supervised training data with transformer-based models. Second, explicit, symbolic knowledge encoded in knowledge bases. Our approach combines both—exploiting the powerful implicit reasoning of transformer models for answer prediction, and integrating symbolic representations from a knowledge graph, while never losing their explicit semantics to an implicit embedding. We combine diverse sources of knowledge to cover the wide variety of knowledge needed to solve knowledge-based questions. We show our approach, KRISP (Knowledge Reasoning with Implicit and Symbolic rePresentations), significantly out-performs state-of-the-art on OK-VQA, the largest available dataset for open-domain knowledge-based VQA. We show with extensive ablations that while our model successfully exploits implicit knowledge reasoning, the symbolic answer module which explicitly connects the knowledge graph to the answer vocabulary is critical to the performance of our method and generalizes to rare answers.1
Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta 0001, Marcus Rohrbach
CVPR2
2021 An Empirical Study of Training Self-Supervised Vision Transformers
abstract
This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know baseline given the recent progress in computer vision: self-supervised learning for Vision Transformers (ViT). While the training recipes for standard convolutional networks have been highly mature and robust, the recipes for ViT are yet to be built, especially in the self-supervised scenarios where training becomes more challenging. In this work, we go back to basics and investigate the effects of several fundamental components for training self-supervised ViT. We observe that instability is a major issue that degrades accuracy, and it can be hidden by apparently good results. We reveal that these results are indeed partial failure, and they can be improved when training is made more stable. We benchmark ViT results in MoCo v3 and several other self-supervised frameworks, with ablations in various aspects. We discuss the currently positive evidence as well as challenges and open questions. We hope that this work will provide useful data points and experience for future research.
Xinlei Chen, Saining Xie, Kaiming He
ICCV1
2021 MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen
ICLR3
2021 Understanding self-supervised learning dynamics without contrastive pairs
abstract
While contrastive approaches of self-supervised learning (SSL) learn representations by minimizing the distance between two augmented views of the same data point (positive pairs) and maximizing views from different data points (negative pairs), recent \emph{non-contrastive} SSL (e.g., BYOL and SimSiam) show remarkable performance {\it without} negative pairs, with an extra learnable predictor and a stop-gradient operation. A fundamental question rises: why they do not collapse into trivial representation? In this paper, we answer this question via a simple theoretical study and propose a novel approach, \ourmethod{}, that \emph{directly} sets the linear predictor based on the statistics of its inputs, rather than trained with gradient update. On ImageNet, it performs comparably with more complex two-layer non-linear predictors that employ BatchNorm and outperforms linear predictor by $2.5%$ in 300-epoch training (and $5%$ in 60-epoch). \ourmethod{} is motivated by our theoretical study of the nonlinear learning dynamics of non-contrastive SSL in simple linear networks. Our study yields conceptual insights into how non-contrastive SSL methods learn, how they avoid representational collapse, and how multiple factors, like predictor networks, stop-gradients, exponential moving averages, and weight decay all come into play. Our simple theory recapitulates the results of real-world ablation studies in both STL-10 and ImageNet. Code is released\footnote{\url{https://github.com/facebookresearch/luckmatters/tree/master/ssl}}.
Yuandong Tian, Xinlei Chen, Surya Ganguli
ICML2
2021 Accelerometer-Based Alcohol Consumption Detection from Physical Activity
abstract
Smartphones have become a common tool for researchers to collect, process, and analyze large quantities of data. This will lead to the creation of solutions that will mostly come in the form of smartphone apps, which will help solve real-life problems. One such real-life problem is the over-consumption of alcohol, since it can lead to many problems including fatality. Currently, there are very expensive or tedious alternative procedures for testing blood alcohol consumption in the market. This paper offers a cheaper alternative to address this problem by detecting if the user has consumed alcohol or not by using a smartphone. We describe an experiment and propose two features derived from accelerometer data that can help us distinguish between sober and intoxicated individuals.
Deeptaanshu Kumar, Ajmal Thanikkal, Prithvi Krishnamurthy, Xinlei Chen, Pei Zhang 0001
WiMob4
2021 Data-Driven Multi-Energy Investment and Management Under Earthquakes
abstract
Seismic events can severely damage both electricity and natural gas systems, causing devastating consequences. Ensuring the secure and reliable operation of the integrated energy system (IES) is of high importance to avoid potential damage to the infrastructure and reduce economic losses. This article proposes a new optimal two-stage optimization to enhance the reliability of IES planning and operation against seismic attacks. In the first stage, hardening investment on the IES is conducted, featuring preventive measures for seismic attacks. The second stage minimizes the expected operation cost of emergency response. The random seismic attack is modeled as uncertainty, which is realized after the first stage. An explicit damage assessment model is developed to define the budget set of the uncertain seismic activity. Based on the survivability of transmission lines and gas pipelines of IES, an optimal system investment plan is developed. The problem is formulated as a two-stage distributionally robust optimization (DRO) model, which is tested on an integrated IEEE 30-bus system and 20-node gas network. Case studies demonstrate that the two-stage DRO outperforms robust optimization and a single-stage optimization model in terms of minimizing the investment cost and expected economic loss. This article can help system operators to make economical hardening and operation strategies to improve the reliability of IES under seismic attacks, thus managing a more robust and secure energy system.
Alexis Pengfei Zhao, Chenghong Gu, Zhidong Cao, Yichen Shen 0002, Fei Teng 0005, Xinlei Chen, Chenye Wu, Da Huo 0001, Shuangqi Li
IEEE Trans. Ind. Informatics6
2020 In Defense of Grid Features for Visual Question Answering
abstract
Popularized as `bottom-up' attention, bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well -- running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available.
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, Xinlei Chen
CVPR5
2020 ImVoteNet: Boosting 3D Object Detection in Point Clouds With Image Votes
abstract
3D object detection has seen quick progress thanks to advances in deep learning on point clouds. A few recent works have even shown state-of-the-art performance with just point clouds input (e.g. VoteNet). However, point cloud data have inherent limitations. They are sparse, lack color information and often suffer from sensor noise. Images, on the other hand, have high resolution and rich texture. Thus they can complement the 3D geometry provided by point clouds. Yet how to effectively use image information to assist point cloud based detection is still an open question. In this work, we build on top of VoteNet and propose a 3D detection architecture called ImVoteNet specialized for RGB-D scenes. ImVoteNet is based on fusing 2D votes in images and 3D votes in point clouds. Compared to prior work on multi-modal detection, we explicitly extract both geometric and semantic features from the 2D images. We leverage camera parameters to lift these features to 3D. To improve the synergy of 2D-3D feature fusion, we also propose a multi-tower training scheme. We validate our model on the challenging SUN RGB-D dataset, advancing state-of-the-art results by 5.7 mAP. We also provide rich ablation studies to analyze the contribution of each design choice.
Charles R. Qi, Xinlei Chen, Or Litany, Leonidas J. Guibas
CVPR2
2020 Seeing the Un-Scene: Learning Amodal Semantic Maps for Room Navigation
Medhini Narasimhan, Erik Wijmans, Xinlei Chen, Trevor Darrell, Dhruv Batra, Devi Parikh, Amanpreet Singh
ECCV (18)3
2020 PAS: Prediction-Based Actuation System for City-Scale Ridesharing Vehicular Mobile Crowdsensing
abstract
Vehicular mobile crowdsensing (MCS) enables many smart city applications. Ridesharing vehicle fleets provide promising solutions to MCS due to the advantages of low cost, easy maintenance, high mobility, and long operational time. However, as nondedicated mobile sensing platforms, the first priorities of these vehicles are delivering passengers, which may lead to poor sensing coverage quality. Therefore, to help MCS derive good (large and balanced) sensing coverage quality, an actuation system is required to dispatch vehicles with a limited amount of monetary budget. This article presents PAS, a prediction-based actuation system for city-wide ridesharing vehicular MCS to achieve optimal sensing coverage quality with a limited budget. In PAS, two prediction models forecast probabilities of potential near-future vehicle routes and ride requests across the city. Based on prediction results, a prediction-based actuation planning algorithm is proposed to decide which vehicles to actuate and the corresponding routes. Experiments on city-scale deployments and physical feature-based simulations show that our PAS achieves up to 40% more improvement in sensing coverage quality and up to 20% higher ride request matching rate than baselines. In addition, to achieve a similar level of sensing coverage quality as the baseline, our PAS only needs 10% budget.
Xinlei Chen, Susu Xu, Jun Han 0001, Haohao Fu, Xidong Pi, Carlee Joe-Wong, Yong Li 0008, Lin Zhang 0001, Hae Young Noh, Pei Zhang 0001
IEEE Internet Things J.1
2020 iLOCuS: Incentivizing Vehicle Mobility to Optimize Sensing Distribution in Crowd Sensing
abstract
Vehicular crowd sensing systems are designed to achieve large spatio-temporal sensing coverage with low-cost in deployment and maintenance. For example, taxi platforms can be utilized for sensing city-wide air quality. However, the goals of vehicle agents are often inconsistent with the goal of the crowdsourcer. Vehicle agents like taxis prioritize searching for passenger ride requests (defined as task requests), which leads them to gather in busy regions. In contrast, sensing systems often need to sample data over the entire city with a desired distribution (e.g., Uniform distribution, Gaussian Mixture distribution, etc.) to ensure sufficient spatio-temporal information for further analysis. This inconsistency decreases the sensing coverage quality and thus impairs the quality of the collected information. A simple approach to reduce the inconsistency is to greedily incentivize the vehicle agents to different regions. However, incentivization brings challenges, including the heterogeneity of desired target distributions, limited budget to incentivize more vehicle agents, and the high computational complexity of optimizing incentivizing strategies. To this end, we present a vehicular crowd sensing system to efficiently incentivize the vehicle agents to match the sensing distribution of the sampled data to the desired target distribution with a limited budget. To make the system flexible to various desired target distributions, we formulate the incentivizing problem as a new type of non-linear multiple-choice knapsack problem, with the dissimilarity between the collected data distribution and the desired distribution as the objective function. To utilize the budget efficiently, we design a customized incentive by combining monetary incentives and potential task (ride) requests at the destination. Meanwhile, an efficient optimization algorithm, iLOCuS, is presented to plan the incentivizing policy for vehicle agents to decompose the sensing distribution into two distinct levels: time-location level and vehicle level, to approximate the optimal solution iteratively and reduce the dissimilarity objective. Our experimental results based on real-world data show that our system can reduce up to 26.99 percent of the dissimilarity between the sensed and target distributions compared to benchmark methods.
Susu Xu, Xinlei Chen, Xidong Pi, Carlee Joe-Wong, Pei Zhang 0001, Hae Young Noh
IEEE Trans. Mob. Comput.2
2020 H-DrunkWalk: Collaborative and Adaptive Navigation for Heterogeneous MAV Swarm
abstract
Large-scale micro-aerial vehicle (MAV) swarms provide promising solutions for situational awareness in applications such as environmental monitoring, urban surveillance, search and rescue, and so on. However, these scenarios do not provide localization infrastructure and limit cost and size of on-board capabilities of individual nodes, which makes it challenging for nodes to autonomously navigate to suitable preassigned locations. In this article, we present H-DrunkWalk , a collaborative and adaptive technique for heterogeneous MAV swarm navigation in environments not formerly preconditioned for operation. Working with heterogeneous MAV swarm, the H-DrunkWalk achieves high accuracy through collaboration but still maintains a low cost of the entire swarm. The heterogeneous MAV swarm consists of two types of nodes: (1) basic MAVs with limited sensing, communication, computing capabilities and (2) advanced MAVs with premium sensing, communication, computing capabilities. The key focus behind this networked MAV swarm research is to (1) rely on collaboration to overcome limitations of individual nodes and efficiently achieve system-wide sensing objectives and (2) fully take advantage of advanced MAVs to help basic MAVs improve their performance. The evaluations based on real MAV testbed experiments and large-scale physical-feature-based simulations show that compared to the traditional non-collaborative and non-adaptive method (dead reckoning with map bias), our system achieves up to 6× reductions in location estimation errors, and as much as 3× improvements in navigation success rate under the given time and accuracy constraints. In addition, by comprehensively considering the environment, heterogeneous structure, and quality of location estimation, our H-DrunkWalk brings 2× performance improvement (on average) as that of a hardware upgrade.
Xinlei Chen, Carlos Ruiz Dominguez, Sihan Zeng, Liyao Gao, Aveek Purohit, Stefano Carpin, Pei Zhang 0001
ACM Trans. Sens. Networks1
2019 CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
abstract
Jin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach, Byoung-Tak Zhang, Yuandong Tian, Dhruv Batra, Devi Parikh. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach, Byoung-Tak Zhang, Yuandong Tian, Dhruv Batra, Devi Parikh
ACL (1)3
2019 Cycle-Consistency for Robust Visual Question Answering
abstract
Despite significant progress in Visual Question Answer-ing over the years, robustness of today’s VQA models leave much to be desired. We introduce a new evaluation protocol and associated dataset (VQA-Rephrasings) and show that state-of-the-art VQA models are notoriously brittle to linguistic variations in questions. VQA-Rephrasings contains 3 human-provided rephrasings for 40k questions-image pairs from the VQA v2.0 validation dataset. As a step towards improving robustness of VQA models, we propose a model-agnostic framework that exploits cycle consistency. Specifically, we train a model to not only answer a question, but also generate a question conditioned on the answer, such that the answer predicted for the generated question is the same as the ground truth answer to the original question. Without the use of additional supervision, we show that our approach is significantly more robust to linguistic variations than state-of-the-art VQA models, when evaluated on the VQA-Rephrasings dataset. In addition, our approach also outperforms state-of-the-art approaches on the standard VQA and Visual Question Generation tasks on the challenging VQA v2.0 dataset. Code and models will be made publicly available.
Meet Shah 0001, Xinlei Chen, Marcus Rohrbach, Devi Parikh
CVPR2
2019 Towards VQA Models That Can Read
abstract
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today’s VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new “TextVQA” dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.
Amanpreet Singh, Vivek Natarajan, Meet Shah 0001, Xinlei Chen, Dhruv Batra, Devi Parikh, Marcus Rohrbach
CVPR5
2019 Multi-Target Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA as introduced in [8] makes the fundamental assumption that every question, e.g., ``what color is the car?", has exactly one target (``car") being inquired about. This assumption puts a direct limitation on the abilities of the agent. We present a generalization of EQA -- Multi-Target EQA (MT-EQA). Specifically, we study questions that have multiple targets in them, such as ``Is the dresser in the bedroom bigger than the oven in the kitchen?", where the agent has to navigate to multiple locations (``dresser in bedroom", ``oven in kitchen") and perform comparative reasoning (``dresser" bigger than ``oven") before it can answer a question. Such questions require the development of entirely new modules or components in the agent. To address this, we propose a modular architecture composed of a program generator, a controller, a navigator, and a VQA module. The program generator converts the given question into sequential executable sub-programs; the navigator guides the agent to multiple locations pertinent to the navigation-related sub-programs; and the controller learns to select relevant observations along its path. These observations are then fed to the VQA module to predict the answer. We perform detailed analysis for each of the model components and show that our joint model can outperform previous methods and strong baselines by a significant margin.
Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L. Berg, Dhruv Batra
CVPR2
2019 Grounded Video Description
abstract
Video description is one of the most challenging problems in vision and language understanding due to the large variability both on the video and language side. Models, hence, typically shortcut the difficulty in recognition and generate plausible sentences that are based on priors but are not necessarily grounded in the video. In this work, we explicitly link the sentence to the evidence in the video by annotating each noun phrase in a sentence with the corresponding bounding box in one of the frames of a video. Our dataset, ActivityNet-Entities, augments the challenging ActivityNet Captions dataset with 158k bounding box annotations, each grounding a noun phrase. This allows training video description models with this data, and importantly, evaluate how grounded or "true" such model are to the video they describe. To generate grounded captions, we propose a novel video description model which is able to exploit these bounding box annotations. We demonstrate the effectiveness of our model on our dataset, but also show how it can be applied to image description on the Flickr30k Entities dataset. We achieve state-of-the-art performance on video description, video paragraph description, and image description and demonstrate our generated sentences are better grounded in the video.
Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J. Corso, Marcus Rohrbach
CVPR3
2019 nocaps: novel object captioning at scale
abstract
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supervision. To encourage the development of image captioning models that can learn visual concepts from alternative data sources, such as object detection datasets, we present the first large-scale benchmark for this task. Dubbed `nocaps', for novel object captioning at scale, our benchmark consists of 166,100 human-generated captions describing 15,100 images from the Open Images validation and test sets. The associated training data consists of COCO image-caption pairs, plus Open Images image-level labels and object bounding boxes. Since Open Images contains many more classes than COCO, nearly 400 object classes seen in test images have no or very few associated training captions (hence, nocaps). We extend existing novel object captioning models to establish strong baselines for this benchmark and provide analysis to guide future work.
Harsh Agrawal, Peter Anderson 0001, Karan Desai, Yufei Wang 0003, Xinlei Chen, Mark Johnson 0001, Dhruv Batra, Devi Parikh, Stefan Lee
ICCV5
2019 TensorMask: A Foundation for Dense Object Segmentation
abstract
Sliding-window object detectors that generate bounding-box object predictions over a dense, regular grid have advanced rapidly and proven popular. In contrast, modern instance segmentation approaches are dominated by methods that first detect object bounding boxes, and then crop and segment these regions, as popularized by Mask R-CNN. In this work, we investigate the paradigm of dense sliding-window instance segmentation, which is surprisingly under-explored. Our core observation is that this task is fundamentally different than other dense prediction tasks such as semantic segmentation or bounding-box object detection, as the output at every spatial location is itself a geometric structure with its own spatial dimensions. To formalize this, we treat dense instance segmentation as a prediction task over 4D tensors and present a general framework called TensorMask that explicitly captures this geometry and enables novel operators on 4D tensors. We demonstrate that the tensor view leads to large gains over baselines that ignore this structure, and leads to results comparable to Mask R-CNN. These promising results suggest that TensorMask can serve as a foundation for novel advances in dense mask prediction and a more complete understanding of the task. Code will be made available.
Xinlei Chen, Ross B. Girshick, Kaiming He, Piotr Dollár
ICCV1
2019 Order-Aware Generative Modeling Using the 3D-Craft Dataset
abstract
Research on 2D and 3D generative models typically focuses on the final artifact being created, e.g., an image or a 3D structure. Unlike 2D image generation, the generation of 3D objects in the real world is commonly constrained by the process and order in which the object is constructed. For instance, gravity needs to be taken into account when building a block tower. In this paper, we explore the prediction of ordered actions to construct 3D objects. Instead of predicting actions based on physical constraints, we propose learning through observing human actions. To enable large-scale data collection, we use the Minecraft1 environment. We introduce 3D-Craft, a new dataset of 2,500 Minecraft houses each built by human players sequentially from scratch. To learn from these human action sequences, we propose an order-aware 3D generative model called VoxelCNN. In contrast to other 3D generative models which either have no explicit order (e.g. holistic generation with 3DGAN [35]), or follow a simple heuristic order (e.g. raster-scan), VoxelCNN is trained to imitate human building order with spatial awareness. We also transferred the order to other dataset such as ShapeNet[10]. The 3D-Craft dataset, models, and benchmark system will be made publicly available, which may inspire new directions for future research exploration. https://github.com/facebookresearch/VoxelCNN.
Zhuoyuan Chen, Kavya Srinet, Charles R. Qi, Haoqi Fan 0001, Jerry Ma, C. Lawrence Zitnick, Demi Guo, Tong Xiao 0003, Saining Xie, Xinlei Chen, Arthur Szlam, Shubham Tulsiani, Haonan Yu, Jonathan Gray
ICCV10
2019 Embodied Amodal Recognition: Learning to Move to Perceive Objects
abstract
Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this work, we introduce the task of Embodied Amodel Recognition (EAR): an agent is instantiated in a 3D environment close to an occluded target object, and is free to move in the environment to perform object classification, amodal object localization, and amodal object segmentation. To address this problem, we develop a new model called Embodied Mask R-CNN for agents to learn to move strategically to improve their visual recognition abilities. We conduct experiments using a simulator for indoor environments. Experimental results show that: 1) agents with embodiment (movement) achieve better visual recognition performance than passive ones and 2) in order to improve visual recognition abilities, agents can learn strategic paths that are different from shortest paths.
Zhile Ren, Xinlei Chen, David Crandall, Devi Parikh, Dhruv Batra
ICCV4
2019 Prior-Aware Neural Network for Partially-Supervised Multi-Organ Segmentation
abstract
Accurate multi-organ abdominal CT segmentation is essential to many clinical applications such as computer-aided intervention. As data annotation requires massive human labor from experienced radiologists, it is common that training data is usually partially-labeled. However, these background labels can be misleading in multi-organ segmentation since the ``background'' usually contains some other organs of interest. To address the background ambiguity in these partially-labeled datasets, we propose Prior-aware Neural Network (PaNN) via explicitly incorporating anatomical priors on abdominal organ sizes, guiding the training process with domain-specific knowledge. More specifically, PaNN assumes that the average organ size distributions in the abdomen should approximate their empirical distributions, a prior statistics obtained from the fully-labeled dataset. As our objective is difficult to be directly optimized using stochastic gradient descent, it is reformulated as a min-max form and optimized via the stochastic primal-dual gradient algorithm. PaNN achieves state-of-the-art performance on the MICCAI2015 challenge ``Multi-Atlas Labeling Beyond the Cranial Vault'', a competition on organ segmentation in the abdomen. We report an average Dice score of 84.97%, surpassing the prior art by a large margin of 3.27%. Code and models will be made publicly available.
Yuyin Zhou, Song Bai 0001, Xinlei Chen, Elliot K. Fishman, Alan L. Yuille
ICCV4
2019 Vehicle dispatching for sensing coverage optimization in mobile crowdsensing systems: poster abstract
abstract
Mobile crowd sensing (MCS) collects city-scale sensing data with low cost and high efficiency. One important goal of MCS is to ensure high quality of sensing coverage to provide sufficient information to data analysis end. However, the goal of the MCS may be inconsistent with the goal of vehicles. This inconsistency between goals results in a bad sensing coverage and decreases the quality of the collected information. Key challenges to resolve this inconsistency include the heterogeneous target desired spatio-temporal distributions, limited budget constraining the ability to incentivize more taxis, and high computational complexity.
Susu Xu, Xinlei Chen, Xidong Pi, Carlee Joe-Wong, Pei Zhang 0001, Hae Young Noh
IPSN2
2019 Poster: FlexDP-Flexible Data Plane for ENFV
abstract
We design a flexible data plane to integrate SDN and ENFV to manage complex network services for future mobile services. We build our prototype and show its feasibility in terms of latency and throughput by using mobile edge computing as an example use case in a realistic mobile networking testbed.
Amit Samanta 0001, Xinlei Chen, Yong Li 0008
MobiCom2
2019 Quantitative analysis for capabilities of vehicular fog computing
Xuefeng Xiao 0002, Xueshi Hou, Xinlei Chen, Chenhao Liu, Yong Li 0008
Inf. Sci.3
2018 Iterative Visual Reasoning Beyond Convolutions
abstract
We present a novel framework for iterative visual reasoning. Our framework goes beyond current recognition systems that lack the capability to reason beyond stack of convolutions. The framework consists of two core modules: a local module that uses spatial memory [4] to store previous beliefs with parallel updates; and a global graph-reasoning module. Our graph module has three components: a) a knowledge graph where we represent classes as nodes and build edges to encode different types of semantic relationships between them; b) a region graph of the current image where regions in the image are nodes and spatial relationships between these regions are edges; c) an assignment graph that assigns regions to classes. Both the local module and the global module roll-out iteratively and cross-feed predictions to each other to refine estimates. The final predictions are made by combining the best of both modules with an attention mechanism. We show strong performance over plain ConvNets, e.g. achieving an 8.4% absolute improvement on ADE [55] measured by per-class average precision. Analysis also shows that the framework is resilient to missing regions for reasoning.
Xinlei Chen, Li-Jia Li 0001, Li Fei-Fei 0001, Abhinav Gupta 0001
CVPR1
2018 Locally Differentially Private Participant Recruitment for Mobile Crowdsourcing
abstract
Location-aware mobile crowdsourcing tasks like urban sensing always require exposing users' location, which lead to serious privacy breaches. In this poster, we propose a locally differentially private participants recruitment system to maximize spatial coverage of the mobile crowdsourcing task while preserving location privacy. Based on the mechanism of randomized response, our system preserves the privacy in a local way, which eliminates the need for a trusted server. With guaranteed location privacy protection, a heuristic algorithm is proposed to solve the maximum spatial coverage problem efficiently given the obfuscated reports. Extensive experiments on real-world user trajectories demonstrate the feasibility of our proposed system, which improves the spatial coverage by more than 10% on average compared with the state-of-the-art solutions.
Fengli Xu, Yong Li 0008, Xinlei Chen, Pei Zhang 0001
SenSys4
2018 Social Trust Aided D2D Communications: Performance Bound and Implementation Mechanism
abstract
In a device-to-device (D2D) communications underlaying cellular network, any user is a potential eavesdropper for the transmissions of others that occupy the same spectrum. The physical-layer security mechanism of theoretical secure capacity, which maximizes the rate of reliable communication from the source user to the legitimate receiver and ensure unauthorized users learn as little as information as possible, is typically employed to guarantee secure communications. As hand-held devices are carried by human beings, we may leverage their social trust to decrease the number of potential eavesdroppers. Aiming to establish a new paradigm for solving the challenging problem of security and efficiency tradeoff, we propose a social trust-aware D2D communication architecture that exploits the social-domain trust for securing the physical-domain communication. In order to understand the impact of social trust on the security of transmissions, we analyze the system ergodic rate of social trust aided communications via stochastic geometry, and our result based on a real data set shows that the proposed social trust aided D2D communication increases the system secrecy rate by about 63% compared with the scheme without considering social trust relation. Furthermore, in order to provide implementation mechanism, we utilize matching theory to implement efficient resource allocation among multiple users. Numerical results show that our proposed mechanism increases the system secrecy rate by 28% with fast convergence over the social oblivious approach.
Xinlei Chen, Yulei Zhao, Yong Li 0008, Xu Chen 0004, Ning Ge 0001, Sheng Chen 0001
IEEE J. Sel. Areas Commun.1
2018 Device-to-Device Communications Enabled Energy Efficient Multicast Scheduling in mmWave Small Cells
abstract
To keep pace with the rapid growth of mobile traffic demands, dense deployment of small cells in millimeter wave (mmWave) bands has become a promising candidate for next-generation wireless communication systems. With a greatly increased data rate from huge bandwidth of mmWave communications, energy consumption should be mitigated for higher energy efficiency. Due to content popularity, many content-based mobile applications can be supported by the multicast service. mmWave communications exploit directional antennas to overcome high path loss, and concurrent transmissions can be enabled for better multicast service. On the other hand, device-to-device (D2D) communications in physical proximity should be exploited to improve multicast performance. In this paper, we propose an energy-efficient multicast scheduling scheme, referred to as EMS, which utilizes both D2D communications and concurrent transmissions to achieve high energy efficiency. In EMS, a D2D path planning algorithm establishes multi-hop D2D transmission paths, and a concurrent scheduling algorithm allocates the links on the D2D paths into different pairings. Then, the transmission power of links is adjusted by the power control algorithm. Furthermore, we theoretically analyze the roles of D2D communications and concurrent transmissions in reducing energy consumption. Extensive simulations under various system parameters demonstrate the superior performance of EMS in terms of energy consumption compared with the state-of-the-art schemes. Furthermore, we also investigate the choice of the interference threshold to optimize network performance.
Yong Niu, Yu Liu 0016, Yong Li 0008, Xinlei Chen, Zhangdui Zhong, Zhu Han 0001
IEEE Trans. Commun.4
2018 Socially Aware Secrecy-Ensured Resource Allocation in D2D Underlay Communication: An Overlapping Coalitional Game Scheme
abstract
With the popularity of proximity-based services, device-to-device (D2D) communication underlaying cellular networks is a promising technology to cope with the growing demands by improving network resource utilization. However, the wireless communication's broadcast nature is vulnerable to eavesdropping, and thus, ensuring a secrecy communication for both cellular user equipments (CUEs) and D2D pairs in an underlay network is a challenging issue. We investigate the problem of physical-layer secure transmission jointly with resource allocation in D2D communications. Different from existing works, we framed overlapping (partial) coalitional game where each D2D pair can access multiple CUEs' spectral resources. Moreover, the multiple D2D pairs can share single CUE subchannel in multiple eavesdroppers scenario to ensure information security for both CUEs and D2D pairs and to maximize system sum rate in a socially aware D2D network. We incorporate the mutual interference and propose different transmission modes for a secrecy-ensured resource allocation-based overlapping coalition formation scheme with transferable utility to obtain a final stable partition. We further prove the proposed algorithm stability, convergence, and computational complexity. Both analytical and numerical results demonstrate the effectiveness of our proposed scheme, which ensures a system-wide security and at the same time improves the performance by maximizing the system sum rate.
Manzoor Ahmed, Xinlei Chen, Yong Li 0008, Muhammad Waqas 0001, Depeng Jin
IEEE Trans. Wirel. Commun.3
2018 Mobility-Aware Transmission Scheduling Scheme for Millimeter-Wave Cells
abstract
With the explosive growth of mobile traffic, diverse mobile applications with high throughput demands have gained considerable attention from both academia and industry. However, dynamics due to human mobility impose serious issues on high throughput communications. Although millimeter-wave (mm-wave), relays, and concurrent transmission have been applied for throughput improvement, how to achieve high throughput transmission in mobility-aware scenarios is still challenging. In this paper, we propose a throughput-efficient service scheduling (TESS) scheme, which exploits multi-hop relay and concurrent transmissions with the consideration of human mobility. In TESS, we first propose a mobility-aware transmission scheduling scheme in single mm-wave cell. The proposed scheduling scheme consists of a relay path planning algorithm and a global time scheduling algorithm. In the relay path planning algorithm, we establish multi-hop relay transmission paths from base station to service points. In the global time scheduling algorithm, we schedule concurrent transmission in relay paths. Furthermore, the proposed TESS scheme is extended for multi-cell scenarios. Through extensive performance evaluations under realistic human mobility trajectories, we demonstrate the superior performance of TESS in terms of system throughput compared with the state-of-the-art schemes.
Yu Liu 0016, Xinlei Chen, Yong Niu, Bo Ai 0001, Yong Li 0008, Depeng Jin
IEEE Trans. Wirel. Commun.2
2017 Spatial Memory for Context Reasoning in Object Detection
Xinlei Chen, Abhinav Gupta 0001
ICCV1
2017 Hybrid and adaptive drone identification through motion actuation and vision feature matching: poster abstract
abstract
Unmanned aerial vehicle (UAV) swarms provide situation awareness in tasks such as emergency response, search and rescue, etc. However, most of these scenarios take place in GPS-denied environments, where accurately localizing each UAV is challenging. Heterogeneous UAV swarms, in which only a subset of the drones carry cameras, face the additional challenge of identifying each individual UAV to avoid sending position updates to the wrong drone, thus crashing. This work presents an identification mechanism based on the correlation between motion observed from external camera, and acceleration measured on each UAV's accelerometer.
Carlos Ruiz Dominguez, Xinlei Chen, Pei Zhang 0001
IPSN2
2017 E-loc: indoor localization through building electric wiring: poster abstract
abstract
E-Loc is an indoor localization system, which, through using existing indoor electric wiring, detects occupants' location. While many indoor localization technologies require intensive infrastructural supports, E-Loc obtain locations by injecting a signal into the protected earth line of existing residential power network. Caused by human body inside a room, the electromagnetic character changes can be detected to deduce a resident's location. We evaluate our system through experiments inside multiple rooms and our system is able to reach meter-level accuracy.
Yue Zhang 0044, Xinlei Chen, Pei Zhang 0001, Lin Zhang 0001
IPSN3
2017 Delay Effect in Mobile Sensing System for Urban Air Pollution Monitoring
abstract
In this paper, given the scenario of a mobile sensing system for air pollution monitoring, we aim at the cause and influence of delay effect on measurement and present a filter-based solution to calibrate the sensing data. We also validate the idea and solution by a real-data experiment. It indicates that the solution decreases deviation on spatial measurement and can be applied in mobile sensing systems to improve the sensing data quality.
Xinyu Liu 0003, Xinlei Chen, Xiangxiang Xu 0001, Enhan Mai, Hae Young Noh, Pei Zhang 0001, Lin Zhang 0001
SenSys2
2017 Individualized Calibration of Industrial-Grade Gas Sensors in Air Quality Sensing System
abstract
Low-cost sensors are widely used to realize large-scale deployment for sensing systems. In this paper, we discuss challenges in using industrial-grade gas sensors for air quality monitoring. To overcome variation due to system errors, we present a framework for individualized calibration. Within the framework, multiple regression and interpolation methods are prepared for alternative optimization on fitting sensors' response to gas concentration.
Xinyu Liu 0003, Xiangxiang Xu 0001, Xinlei Chen, Enhan Mai, Hae Young Noh, Pei Zhang 0001, Lin Zhang 0001
SenSys3
2017 Design Experiences in Minimalistic Flying Sensor Node Platform through SensorFly
abstract
Indoor emergency response situations, such as urban fire, are characterized by dangerous constantly changing operating environments with little access to situational information for first responders. In situ information about the conditions, such as the extent and evolution of an indoor fire, can augment rescue efforts and reduce risk to emergency personnel. Static sensor networks that are pre-deployed or manually deployed have been proposed but are less practical due to need for large infrastructure, lack of adaptivity, and limited coverage. Controlled-mobility in sensor networks, that is, the capability of nodes to move as per network needs can provide the desired autonomy to overcome these limitations. In this article, we present SensorFly, a controlled-mobile aerial sensor network platform for indoor emergency response application. The miniature, low-cost sensor platform has capabilities to self deploy, achieve three-dimensional sensing, and adapt to node and network disruptions in harsh environments. We describe hardware design trade-offs, the software architecture, and the implementation that enables limited-capability nodes to collectively achieve application goals. Through the indoor fire monitoring application scenario, we validate that the platform can achieve coverage and sensing accuracy that matches or exceeds static sensor networks and provide higher adaptability and autonomy.
Xinlei Chen, Aveek Purohit, Shijia Pan, Carlos Ruiz Dominguez, Jun Han 0001, Zheng Sun 0003, Frank Mokaya, Patrick Tague, Pei Zhang 0001
ACM Trans. Sens. Networks1
2016 Learning Visual Storylines with Skipping Recurrent Neural Networks
Gunnar A. Sigurdsson, Xinlei Chen, Abhinav Gupta 0001
ECCV (5)2
2016 Visualizing and Understanding Neural Models in NLP
abstract
While neural networks have been successfully applied to many NLP tasks the resulting vectorbased models are very difficult to interpret.For example it's not clear how they achieve compositionality, building sentence meaning from the meanings of words and phrases.In this paper we describe strategies for visualizing compositionality in neural models for NLP, inspired by similar work in computer vision.We first plot unit values to visualize compositionality of negation, intensification, and concessive clauses, allowing us to see wellknown markedness asymmetries in negation.We then introduce methods for visualizing a unit's salience, the amount that it contributes to the final composed meaning from first-order derivatives.Our general-purpose methods may have wide applications for understanding compositionality and other semantic properties of deep networks.
Jiwei Li 0001, Xinlei Chen, Eduard H. Hovy, Daniel Jurafsky
HLT-NAACL2
2016 HAP: Fine-Grained Dynamic Air Pollution Map Reconstruction by Hybrid Adaptive Particle Filter: Poster Abstract
abstract
This paper presents a hybrid adaptive particle filter (HAP) with online feedback to dynamically reconstruct high spatial-temporal resolution air pollution information from sparse vehicular based sensors. To deal with data sparsity, we apply both spatial and temporal correlation of air dispersion to reduce data dimension requirement. HAP adaptively predicts when the accumulated prediction error is low and then uses data compensation for correction whenever the prediction error becomes high. The preliminary results based on the city scale deployments with 10 taxis show that our system achieves up to 50% reduction on system errors.
Xinlei Chen, Xiangxiang Xu 0001, Xinyu Liu 0003, Hae Young Noh, Lin Zhang 0001, Pei Zhang 0001
SenSys1
2016 Collaborative Localization and Navigation in Heterogeneous UAV swarms: Demo Abstract
abstract
Resilient localization and navigation for autonomous Unmanned Aerial Vehicles (UAVs) still remains a challenge in certain scenarios, like GPS-deprived environments such as indoors or urban canyons. In this work, we explore a heterogeneous UAV swarm design, in which a small number of sensor and computationally powerful UAVs collaborate with the remaining resource-constrained UAVs to guarantee optimal localization accuracy.
Carlos Ruiz Dominguez, Xinlei Chen, Lin Zhang 0001, Pei Zhang 0001
SenSys2
2016 Gotcha II: Deployment of a Vehicle-based Environmental Sensing System: Poster Abstract
abstract
According to the World Health Organization (WHO), outdoor air pollution led to an estimated 3.7 million premature deaths worldwide in 2012. To address this problem, it is necessary for both residents and city administrations to understand air quality in their immediate environment with fine-grained temporal-spatial resolution. Currently both fixed and mobile systems are used to attempt to sense the pollution field. However, they generally are expensive, cover small areas and thus result in lower accuracy.
Xiangxiang Xu 0001, Xinlei Chen, Xinyu Liu 0003, Hae Young Noh, Pei Zhang 0001, Lin Zhang 0001
SenSys2
2015 Never-Ending Learning
abstract
Whereas people learn many different types of knowledge from diverse experiences over many years, most current machine learning systems acquire just a single function or data model from just a single data set. We propose a never-ending learning paradigm for machine learning, to better reflect the more ambitious and encompassing type of learning performed by humans. As a case study, we describe the Never-Ending Language Learner (NELL), which achieves some of the desired properties of a never-ending learner, and we discuss lessons learned. NELL has been learning to read the web 24 hours/day since January 2010, and so far has acquired a knowledge base with over 80 million confidence-weighted beliefs (e.g., servedWith(tea, biscuits)). NELL has also learned millions of features and parameters that enable it to read these beliefs from the web. Additionally, it has learned to reason over these beliefs to infer new beliefs, and is able to extend its ontology by synthesizing new relational predicates. NELL can be tracked online at http://rtw.ml.cmu.edu, and followed on Twitter at @CMUNELL.
Tom M. Mitchell, William W. Cohen, Estevam Hruschka, Partha P. Talukdar, Justin Betteridge, Andrew Carlson, Bhavana Dalvi, Matt Gardner 0001, Bryan Kisiel, Jayant Krishnamurthy, Ni Lao, Kathryn Mazaitis, Thahir Mohamed, Ndapandula Nakashole, Emmanouil A. Platanios, Alan Ritter, Mehdi Samadi, Burr Settles, Richard C. Wang, Derry Wijaya, Abhinav Gupta 0001, Xinlei Chen, Abulhair Saparov, Malcolm Greaves, Joel Welling
AAAI22
2015 Sense discovery via co-clustering on images and text
abstract
We present a co-clustering framework that can be used to discover multiple semantic and visual senses of a given Noun Phrase (NP). Unlike traditional clustering approaches which assume a one-to-one mapping between the clusters in the text-based feature space and the visual space, we adopt a one-to-many mapping between the two spaces. This is primarily because each semantic sense (concept) can correspond to different visual senses due to viewpoint and appearance variations. Our structure-EM style optimization not only extracts the multiple senses in both semantic and visual feature space, but also discovers the mapping between the senses. We introduce a challenging dataset (CMU Polysemy-30) for this problem consisting of 30 NPs (∼5600 labeled instances out of ∼22K total instances). We have also conducted a large-scale experiment that performs sense disambiguation for ∼2000 NPs.
Xinlei Chen, Alan Ritter, Abhinav Gupta 0001, Tom M. Mitchell
CVPR1
2015 Mind's eye: A recurrent visual representation for image caption generation
abstract
In this paper we explore the bi-directional mapping between images and their sentence-based descriptions. Critical to our approach is a recurrent neural network that attempts to dynamically build a visual representation of the scene as a caption is being generated or read. The representation automatically learns to remember long-term visual concepts. Our model is capable of both generating novel captions given an image, and reconstructing visual features given an image description. We evaluate our approach on several tasks. These include sentence generation, sentence retrieval and image retrieval. State-of-the-art results are shown for the task of generating novel image descriptions. When compared to human generated captions, our automatically generated captions are equal to or preferred by humans 21.0% of the time. Results are better than or comparable to state-of-the-art results on the image and sentence retrieval tasks for methods using similar visual features.
Xinlei Chen, C. Lawrence Zitnick
CVPR1
2015 Webly Supervised Learning of Convolutional Networks
abstract
We present an approach to utilize large amounts of web data for learning CNNs. Specifically inspired by curriculum learning, we present a two-step approach for CNN training. First, we use easy images to train an initial visual representation. We then use this initial CNN and adapt it to harder, more realistic images by leveraging the structure of data and categories. We demonstrate that our two-stage CNN outperforms a fine-tuned CNN trained on ImageNet on Pascal VOC 2012. We also demonstrate the strength of webly supervised learning by localizing objects in web images and training a R-CNN style [19] detector. It achieves the best performance on VOC 2007 where no VOC training data is used. Finally, we show our approach is quite robust to noise and performs comparably even when we use image search results from March 2013 (pre-CNN image search era).
Xinlei Chen, Abhinav Gupta 0001
ICCV1
2015 DrunkWalk: Collaborative and Adaptive Planning for Navigation of Micro-Aerial Sensor Swarms
abstract
Micro-aerial vehicle (MAV) swarms are a new class of mobile sensor networks with many applications, including search and rescue, urban surveillance, radiation monitoring, etc. These sensing applications require autonomously navigating a high number of low-cost, low-complexity MAV sensor nodes in hazardous environments. The lack of preexisting localization infrastructure and the limited sensing, computing, and communication abilities of individual nodes makes it challenging for nodes to autonomously navigate to suitable preassigned locations. In this paper, we present a collaborative and adaptive algorithm for resource-constrained MAV nodes to quickly and efficiently navigate to preassigned locations. Using radio fingerprints between flying and landed MAVs acting as radio beacons, the algorithm detects intersections in trajectories of mobile nodes. The algorithm combines noisy dead-reckoning measurements from multiple MAVs at detected intersections to improve the accuracy of the MAVs' location estimations. In addition, the algorithm plans intersecting trajectories of MAV nodes to aid the location estimation and provide desired performance in terms of timeliness and accuracy of navigation. We evaluate the performance of our algorithm through a real testbed implementation and large-scale physical feature based simulations. Our results show that, compared to existing autonomous navigation strategies, our algorithm achieves up to 6X reduction in location estimation errors, and as much as 3X improvement in navigation success rate under the given time and accuracy constraints.
Xinlei Chen, Aveek Purohit, Carlos Ruiz Dominguez, Stefano Carpin, Pei Zhang 0001
SenSys1
2015 Large Scale Spectral Clustering Via Landmark-Based Sparse Representation
abstract
Spectral clustering is one of the most popular clustering approaches. However, it is not a trivial task to apply spectral clustering to large-scale problems due to its computational complexity of O(n(3)), where n is the number of samples. Recently, many approaches have been proposed to accelerate the spectral clustering. Unfortunately, these methods usually sacrifice quite a lot information of the original data, thus result in a degradation of performance. In this paper, we propose a novel approach, called landmark-based spectral clustering, for large-scale clustering problems. Specifically, we select p ( << n) representative data points as the landmarks and represent the original data points as sparse linear combinations of these landmarks. The spectral embedding of the data can then be efficiently computed with the landmark-based representation. The proposed algorithm scales linearly with the problem size. Extensive experiments show the effectiveness and efficiency of our approach comparing to the state-of-the-art methods.
Deng Cai 0001, Xinlei Chen
IEEE Trans. Cybern.2
2014 Enriching Visual Knowledge Bases via Object Discovery and Segmentation
abstract
There have been some recent efforts to build visual knowledge bases from Internet images. But most of these approaches have focused on bounding box representation of objects. In this paper, we propose to enrich these knowledge bases by automatically discovering objects and their segmentations from noisy Internet images. Specifically, our approach combines the power of generative modeling for segmentation with the effectiveness of discriminative models for detection. The key idea behind our approach is to learn and exploit top-down segmentation priors based on visual subcategories. The strong priors learned from these visual subcategories are then combined with discriminatively trained detectors and bottom up cues to produce clean object segmentations. Our experimental results indicate state-of-the-art performance on the difficult dataset introduced by [29] Rubinstein et al. We have integrated our algorithm in NEIL for enriching its knowledge base [5]. As of 14th April 2014, NEIL has automatically generated approximately 500K segmentations using web data.
Xinlei Chen, Abhinav Shrivastava, Abhinav Gupta 0001
CVPR1
2014 Spatially correlated nonnegative matrix factorization
Xinlei Chen, Haifeng Liu 0001, Deng Cai 0001
Neurocomputing1
2013 NEIL: Extracting Visual Knowledge from Web Data
abstract
We propose NEIL (Never Ending Image Learner), a computer program that runs 24 hours per day and 7 days per week to automatically extract visual knowledge from Internet data. NEIL uses a semi-supervised learning algorithm that jointly discovers common sense relationships (e.g., "Corolla is a kind of/looks similar to Car", "Wheel is a part of Car") and labels instances of the given visual categories. It is an attempt to develop the world's largest visual structured knowledge base with minimum human labeling effort. As of 10th October 2013, NEIL has been continuously running for 2.5 months on 200 core cluster (more than 350K CPU hours) and has an ontology of 1152 object categories, 1034 scene categories and 87 attributes. During this period, NEIL has discovered more than 1700 relationships and has labeled more than 400K visual instances.
Xinlei Chen, Abhinav Shrivastava, Abhinav Gupta 0001
ICCV1
2012 Metric learning with two-dimensional smoothness for visual analysis
abstract
In recent years, metric learning methods based on pairwise side information have attracted considerable interests, and lots of efforts have been devoted to utilize these methods for visual analysis like content based image retrieval and face identification. When applied to image analysis, these methods merely look on an n1× n2image as a vector in Rn1×n2space and the pixels of the image are considered as independent. They fail to consider the fact that an image represented in the plane is intrinsically a matrix, and pixels spatially close to each other may probably be correlated. Even though we have n1× n2pixels per image, this spatial correlation suggests the real number of freedom is far less. In this paper, we introduce a regularized metric learning framework, Two-Dimensional Smooth Metric Learning (2DSML), which uses a discretized Laplacian penalty to restrict the coefficients to be two-dimensional smooth. Many existing metric learning algorithms can fit into this framework and learn a spatially smooth metric which is better for image applications than their original version. Recognition, clustering and retrieval can be then performed based on the learned metric. Experimental results on benchmark image datasets demonstrate the effectiveness of our method.
Xinlei Chen, Zifei Tong, Haifeng Liu 0001, Deng Cai 0001
CVPR1
2012 Semi-supervised Mesh Segmentation and Labeling
abstract
Abstract Recently, approaches have been put forward that focus on the recognition of mesh semantic meanings. These methods usually need prior knowledge learned from training dataset, but when the size of the training dataset is small, or the meshes are too complex, the segmentation performance will be greatly effected. This paper introduces an approach to the semantic mesh segmentation and labeling which incorporates knowledge imparted by both segmented, labeled meshes, and unsegmented, unlabeled meshes. A Conditional Random Fields (CRF) based objective function measuring the consistency of labels and faces, labels of neighbouring faces is proposed. To implant the information from the unlabeled meshes, we add an unlabeled conditional entropy into the objective function. With the entropy, the objective function is not convex and hard to optimize, so we modify the Virtual Evidence Boosting (VEB) to solve the semi‐supervised problem efficiently. Our approach yields better results than those methods which only use limited labeled meshes, especially when many unlabeled meshes exist. The approach reduces the overall system cost as well as the human labelling cost required during training. We also show that combining knowledge from labeled and unlabeled meshes outperforms using either type of meshes alone.
Jiajun Lv, Xinlei Chen, Jin Huang 0001, Hujun Bao
Comput. Graph. Forum2
2011 Large Scale Spectral Clustering with Landmark-Based Representation
abstract
Spectral clustering is one of the most popular clustering approaches. Despite its good performance, it is limited in its applicability to large-scale problems due to its high computational complexity. Recently, many approaches have been proposed to accelerate the spectral clustering. Unfortunately, these methods usually sacrifice quite a lot information of the original data, thus result in a degradation of performance. In this paper, we propose a novel approach, called Landmark-based Spectral Clustering (LSC), for large scale clustering problems. Specifically, we select $p\ (\ll n)$ representative data points as the landmarks and represent the original data points as the linear combinations of these landmarks. The spectral embedding of the data can then be efficiently computed with the landmark-based representation. The proposed algorithm scales linearly with the problem size. Extensive experiments show the effectiveness and efficiency of our approach comparing to the state-of-the-art methods.
Xinlei Chen, Deng Cai 0001
AAAI1
2011 Sparse structured probabilistic projections for factorized latent spaces
abstract
Building a common representation for several related data sets is an important problem in multi-view learning. CCA and its extensions have shown that they are effective in finding the shared variation among all data sets. However, these models generally fail to exploit the common structure of the data when the views are with private information. Recently, methods explicitly modeling the information into shared part and private parts have been proposed, but they presume to know the prior knowledge about the latent space, which is usually impossible to obtain. In this paper, we propose a probabilistic model, which could simultaneously learn the structure of the latent space whilst factorize the information correctly, therefore the prior knowledge of the latent space is unnecessary. Furthermore, as a probabilistic model, our method is able to deal with missing data problem in a natural way. We show that our approach attains the performance of state-of-art methods on the task of human pose estimation when the motion capture view is completely missing, and significantly improves the inference accuracy with only a few observed data.
Xinquan Qu, Xinlei Chen
CIKM2
2011 A Heterogeneous High Speed Wireless Body Sensor Network Based on SC-UWB and ZIGBEE
abstract
Some new medical applications, such as wireless endoscopy system for the diagnoses of whole human digestive tract and real-time endoscopic image monitoring, demand high speed transmission in wireless body sensor network (WBSN). This paper proposes a heterogeneous high speed wireless body sensor network based on ZIGBEE and single carrier ultra wideband (SC-UWB). Our system can choose high speed mode or low speed mode automatically according to the type of data source. Several measures are taken to miniaturize the nodes, lower the system complexity and reduce power consumption, offering system users great mobility and flexibility. We develop and validate a prototype WBSN, supporting four real-time high speed video streams and six low speed sensor data. The performance of the developed system is experimentally evaluated using prototype WBSN, and the ASIC has been designed and fabricated in the standard 0.18im CMOS technology, occupying a die area of 3.0mm * 3.0mm and consuming a dynamic power of 43.4280mw.
Xinlei Chen, Xiyu Lu, Zhongjin Liu, Shaoxia Fang, Depeng Jin, Lieguang Zeng
GLOBECOM1
2011 Channel Modeling of UWB-Based Wireless Body Area Networks
abstract
This paper presents channel measurements for wireless body area network (WBAN) and provides performance evaluation using the model derived from measurement and the model in the final document of the IEEE802.15.6 channel modeling subcommittee. We measure the radio propagation from 3.0GHz to 5.0GHz which falls in ultra-wideband (UWB) frequency range in an anechoic chamber and give out a static model. 10 positions for medical entertainment and sports application are chosen to conduct measurement on a fat body, a normal body and a thin body in order to compare different path loss of different types of body for Chinese people. Moreover, we compare the bit error ratio (BER) of a single carrier UWB (SC-UWB) system which employs different receiving policies such as rake, equalizer and channel coding. The results show that slimmer people cause less path loss in the valid distance from 100mm to 1000mm. The measured model we derived has less multi-path effect than that of the model in the final report given by IEEE802.15.6 channel modeling subcommittee. Even though, rake receiver, equalizer and channel coding are needed to get a satisfying system performance.
Xinlei Chen, Xiyu Lu, Depeng Jin, Li Su 0001, Lieguang Zeng
ICC1
2011 UWB-based Wireless Body Area Networks channel modeling and performance evaluation
abstract
In the Wireless Body Area Network (WBAN), the wireless channel is complex and distinctive because of the irregular shape of human body. The channel modeling methods and results are different from those in traditional narrow-band communication environments. In this paper, we present channel models for WBAN in Ultra wide band (UWB) frequency range 3-9 GHz. The channels are modeled statistically and the channel model parameters are derived from actual measured data in an office environment. An interesting common result is that different shape people have significantly differences in multipath parameters. Taking into account the characteristics of body shape, we provide a new classified small-scale channel model which has three kind of channel models: sparse, medium and dense multipath channel models. These models can be applied to different scenarios and improve the system design. We evaluated these models by delays and average number of mUltipaths with the measured data. Results prove the effectiveness of our models.
Xiyu Lu, Xinlei Chen, Guang Sun, Depeng Jin, Ning Ge 0001, Lieguang Zeng
IWCMC2
2002 Two-handed drawing on augmented desk system
abstract
This paper describes a two-handed drawing tool developed on our augmented desk system. Using our real-time finger tracking method, a user can draw and manipulate objects interactively by his/her own finger/hand. Based on the former work on two-handed interaction, different roles are assigned to each hand. The right hand is used to draw and to manipulate objects. Using gesture recognition, primitive objects can be drawn by users' handwriting. On the other hand, the left hand is used to manipulate menus and to assist the right hand. By closing all left hand fingers, users can initiate the appearance of structural radial menus around their left hands, and can select appropriate items by using a left hand finger. The left hand is also used to assist in the performance of drawing tasks, e.g., specifying the center of a circle or top-left corner of a rectangle, or specifying the object to be copied.
Xinlei Chen, Hideki Koike, Yasuto Nakanishi, Kenji Oka, Yoichi Sato 0001
AVI1