EDBT 2026 Demo / reviewers in the wild / expert
Xiaoran Fan
dblp:197/0141
· DBLP profile ↗
37ranked-venue papers
9as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 17 since 2021Computer networks · 15 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyabstractSpeech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui |
AAAI | 1 |
| 2026 | MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention Across Vision-Language ModelsabstractAs vision-language models (VLMs) tackle increasingly complex and multimodal tasks, the rapid growth of Key-Value (KV) cache imposes significant memory and computational bottlenecks during inference. While Multi-Head Latent Attention (MLA) offers an effective means to compress the KV cache and accelerate inference, adapting existing VLMs to the MLA architecture without costly pretraining remains largely unexplored. In this work, we present \textbf{MHA2MLA-VLM}, a parameter-efficient and multimodal-aware framework for converting off-the-shelf VLMs to MLA. Our approach features two core techniques: (1) a modality-adaptive partial-RoPE strategy that supports both traditional and multimodal settings by selectively masking nonessential dimensions, and (2) a modality-decoupled low-rank approximation method that independently compresses the visual and textual KV spaces. Furthermore, we introduce parameter-efficient fine-tuning to minimize adaptation cost and demonstrate that minimizing output activation error, rather than parameter distance, substantially reduces performance loss. Extensive experiments on three representative VLMs show that MHA2MLA-VLM restores original model performance with minimal supervised data, significantly reduces KV cache footprint, and integrates seamlessly with KV quantization. Xiaoran Fan, Lixing Shen, Tao Gui |
AAAI | 1 |
| 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data ContaminationabstractReasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods. Mingqi Wu, Zhihao Zhang 0002, Qiaole Dong, Zhiheng Xi, Jun Zhao 0019, Senjie Jin, Xiaoran Fan, Yuhao Zhou 0005, Huijie Lv, Ming Zhang 0030, Yanwei Fu 0001, Qin Liu 0010, Songyang Zhang 0001, Qi Zhang 0001 |
AAAI | 7 |
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2026 | Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentabstractYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2026 | FlowGait: Enabling Robust Long-Term Gait Recognition Across Real-World Covariates with mmWave RadarabstractGait recognition enables proactive and personalized smart home interactions, but its long-term reliability is challenged by the non-static nature of gait. Covariates like carrying items and clothing induce a persistent domain shift that degrades traditional, static models. To solve this, we introduce FlowGait, a mmWave-based framework designed for robust, long-term adaptation. It combines self-training with continual learning, allowing the model to daily align with a user’s evolving gait by learning from readily available unlabeled data. It features a specialized transformer network for radar spectrogram analysis and a novel two-stage labeling algorithm that leverages the gait’s hierarchical nature to assign pseudo-labels to the unlabeled data accurately. Evaluated on three challenging datasets from 47 volunteers (covering 12 gait-covariates, 11 routes, and two weeks), FlowGait achieves high accuracies of 94.8 (cross-covariate), 98.6% (cross-route), and 95.5% (cross-day). Notably, for the long-term dataset, it reduced performance decay from 13.6% to just 1.4%, demonstrating its real-world robustness. Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 5 |
| 2026 | FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion SensingabstractVoice has been a primary interaction mode with LLM-powered assistants. Beyond semantics, voice carries emotional cues with potential to guide empathetic system responses. Yet, robust vocal emotion sensing in noise and its use in optimizing interactions remain underexplored. In response, we present FeelWave, which achieves empathetic voice interaction through noise-robust mmWave emotion sensing and structured LLM prompts. It extracts robust vocal information from mmWave signals, applies audio-to-mmWave transfer learning for efficient emotion recognition, and employs chain-of-thought-based query optimization to enable emotion-adaptive responses. Evaluations show that FeelWave achieves 92.3% emotion recognition accuracy and remains robust in noisy environments, yielding a 62.9 percentage-point gain over audio-based models. In voice interaction studies, 74.3% of users prefer FeelWave, reporting significantly higher satisfaction than a baseline without emotion sensing (4.37 vs. 3.22). A SUS score of 88.3 confirms FeelWave’s high usability in real-world deployment. We hope this work will inspire more empathetic, user-centered AI-driven assistants. You Zuo, Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 7 |
| 2026 | Needle in a Haystack: Tracking UAVs from Massive Noise in Real-World 5G-A Base Station Data
Chengzhen Meng, Chenming He, Yidong Jiang, Xiaoran Fan, Dequan Wang, Jianmin Ji, Yanyong Zhang |
MobiSys | 4 |
| 2026 | ZA-SLAM: Leveraging Vision-Language Model for Zero-Shot Acoustic SLAMabstractExisting acoustic indoor location sensing systems are limited by the need for extensive data collection and model retraining in unseen environments. This paper introduces ZA-SLAM, a novel zero-shot acoustic Simultaneous Localization and Mapping (SLAM) system that can be deployed in unseen environments without model retraining. Our core idea is to train an acoustic encoder that inherits the generalization capabilities of pre-trained Vision-Language Models (VLMs), which show superiority in tasks like zero-shot visual SLAM. To achieve this goal, we perform Acoustic-Visual Feature Alignment to enable the acoustic encoder to generate features aligned with visual features from VLMs. To select high-quality images for effective alignment, we design a Semantic-Guided Image Selection that filters out low-quality collected images caused by factors like abrupt view changes, occlusions, and uninformative views. Furthermore, we address the challenge of false positive loop closures in structurally similar locations with the Learning-Based Trajectory Reachability Matching that validates loop closures leveraging IMU trajectory features. Extensive real-world experiments demonstrate that our system achieves comparable SLAM performance to retraining-based acoustic SLAM, and much improved performance compared to existing zero-shot Wi-Fi and geomagnetic SLAM systems. Our system achieves a mean mapping error of 0.56 m and a localization error of 0.78 m across multiple unseen environments. Zhuochen Yu, David K. Y. Yau, Yijie Shen, Xiaoran Fan, Tao Chen 0033, Qun Song 0001 |
MobiSys | 4 |
| 2025 | Heart Rate Monitoring Through ANC Headphones in Unconstrained EnvironmentsabstractThis paper introduces CLEAR-APG, a novel acoustic sensing approach that enables reliable heart rate monitoring in unconstrained environments using off-the-shelf active noise cancellation (ANC) headphones. By emitting ultrasonic signals into the user's ear canal via the headphone speaker and analyzing their echoes, which can detect the frequency of a pulsating vein along the canal wall. However, everyday activities such as exercising, speaking, or eating cause jaw movements that deform the ear canal, overwhelming the subtle deformation caused by blood flowing. To overcome this challenge, we employ the ANC headphone's built-in gyroscope to capture body motion and identify how various motion patterns influence the heartbeat waveform. Building on this insight, we propose a multi-modal method that effectively denoises the heartbeat waveform measurements and further accurately extracts heart rate. We implement CLEARAPG on ANC earbuds and conduct comprehensive field studies on 14 users. The results show that CLEAR-APG achieves an average heart rate error of 4.01% across seven different activities, satisfying industry-required margin of 10% heart rate error. Maanya Shanker, Tao Chen 0033, Xiaoran Fan, Longfei Shangguan |
BSN | 4 |
| 2025 | ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world ScenariosabstractExisting evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs’ tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes. Junjie Ye 0005, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
COLING | 7 |
| 2025 | Have the VLMs Lost Confidence? A Study of Sycophancy in VLMsabstractIn the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue. Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang 0001, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 3 |
| 2025 | Ghost Points Matter: Far-Range Vehicle Detection with a Single mmWave Radar in TunnelabstractVehicle detection in tunnels is crucial for traffic monitoring and accident response, yet remains underexplored. In this paper, we develop mmTunnel, a millimeter-wave radar system that achieves far-range vehicle detection in tunnels. The main challenge here is coping with ghost points caused by multi-path reflections, which lead to severe localization errors and false alarms. Instead of merely removing ghost points, we propose correcting them to true vehicle positions by recovering their signal reflection paths, thus reserving more data points and improving detection performance, even in occlusion scenarios. However, recovering complex 3D reflection paths from limited 2D radar points is highly challenging. To address this problem, we develop a multi-path ray tracing algorithm that leverages the ground plane constraint and identifies the most probable reflection path based on signal path loss and spatial distance. We also introduce a curve-to-plane segmentation method to simplify tunnel surface modeling such that we can significantly reduce the computational delay and achieve real-time processing. Chenming He, Chengzhen Meng, Xiaoran Fan, Dequan Wang, Haojie Ren, Jianmin Ji, Yanyong Zhang |
MobiCom | 4 |
| 2025 | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning DatasetabstractIn this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community. Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 6 |
| 2025 | LeakyFeeder: In-Air Gesture Control Through Leaky Acoustic WavesabstractWe present LeakyFeeder, a mobile application that explores the acoustic signals leaked from headphones to reconstruct gesture motions around the ear for fine-grained gesture control. To achieve this goal, LeakyFeeder repurposes the speaker and a single feedforward microphone on active noise cancellation (ANC) headphones as a SONAR system, using inaudible frequency-modulated continuous-wave (FMCW) signals to track gesture reflections for accurate sensing. Since this single-receiver SONAR system is unable to differentiate reflection angles and further disentangle signal reflections from different gesture parts, we draw on principles of multi-modal learning to frame gesture motion reconstruction as a multi-modal translation task and propose a deep learning-based approach to fill the information gap between low-dimensional FMCW ranging readings and high-dimensional 3D hand movements. We implement LeakyFeeder on a pair of Google Pixel Buds and conduct experiments to examine the efficacy and robustness of LeakyFeeder in various conditions. Experiments based on six gesture types inspired by Apple Vision Pro demonstrate that LeakyFeeder achieves a PCK performance of 89% at 3cm across ten users, with an average MPJPE and MPJRPE error of 2.71cm and 1.88cm, respectively. Yongjie Yang 0008, Tao Chen 0033, Zhenlin An, Shirui Cao, Xiaoran Fan, Longfei Shangguan |
SenSys | 5 |
| 2025 | Toward Sensor-In-the-Loop LLM Agent: Benchmarks and ImplicationsabstractThis paper explores sensor-informed personal agents that can take advantage of sensor hints on wearables to enhance the personal agent's response. We demonstrate that such a sensor-in-the-loop AI agent design can be easily integrated into existing LLM agents by building a prototype named WellMax based on existing well-developed techniques such as structured prompt templates and few-shot prompting. The head-to-head comparison with a non-sensor-informed agent across five use scenarios demonstrates that this sensor-in-the-loop design can effectively improve users' needs and their overall experience. The deep-dive into agents' replies and participants' feedback further reveals that sensor-in-the-loop agents not only provide more contextually relevant responses but also exhibit a better understanding of user priorities and situational nuances. In addition, we conduct two case studies to examine the potential pitfalls and distill key insights from this sensor-in-the-loop agent. We hope this work can spawn new ideas for building more intelligent, empathetic, and effective AI-driven personal assistants. Zhiwei Ren, Minjia Zhang, Di Wang 0003, Xiaoran Fan, Longfei Shangguan |
SenSys | 5 |
| 2025 | The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui |
Sci. China Inf. Sci. | 12 |
| 2024 | StepCoder: Improving Code Generation with Reinforcement Learning from Compiler FeedbackabstractShihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2024 | LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style PluginabstractShihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Enyu Zhou, Yan Liu 0002, Songyang Gao, Limao Xiong, Yuhao Zhou 0005, Xiao Wang 0001, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2024 | Leveraging Foundation Models for Zero-Shot IoT SensingabstractDeep learning models are increasingly deployed on edge Internet of Things (IoT) devices. However, these models typically operate under supervised conditions and fail to recognize unseen classes different from training. To address this, zero-shot learning (ZSL) aims to classify data of unseen classes with the help of semantic information. Foundation models (FMs) trained on web-scale data have shown impressive ZSL capability in natural language processing and visual understanding. However, leveraging FMs’ generalized knowledge for zero-shot IoT sensing using signals such as mmWave, IMU, and Wi-Fi has not been fully investigated. In this work, we align the IoT data embeddings with the semantic embeddings generated by an FM’s text encoder for zero-shot IoT sensing. To utilize the physics principles governing the generation of IoT sensor signals to derive more effective prompts for semantic embedding extraction, we propose to use cross-attention to combine a learnable soft prompt that is optimized automatically on training data and an auxiliary hard prompt that encodes domain knowledge of the IoT sensing task. To address the problem of IoT embeddings biasing to seen classes due to the lack of unseen class data during training, we propose using data augmentation to synthesize unseen class IoT data for fine-tuning the IoT feature extractor and embedding projector. We evaluate our approach on multiple IoT sensing tasks. Results show that our approach achieves superior open-set detection and generalized zero-shot learning performance compared with various baselines. Our code is available at https://github.com/schrodingho/FM_ZSL_IoT. Dinghao Xue, Xiaoran Fan, Tao Chen 0033, Guohao Lan, Qun Song 0001 |
ECAI | 2 |
| 2024 | RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool LearningabstractJunjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Junjie Ye 0005, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Xiaoran Fan, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 7 |
| 2024 | Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningabstractIn this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$. Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICML | 13 |
| 2024 | Exploring the Feasibility of Remote Cardiac Auscultation Using EarphonesabstractThe elderly over 65 accounts for 80% of COVID deaths in the United States. In response to the pandemic, the federal, state governments, and commercial insurers are promoting video visits, through which the elderly can access specialists at home over the Internet, without the risk of COVID exposure. However, the current video visit practice barely relies on video observation and talking. The specialist could not assess the patient's health conditions by performing auscultations. Tao Chen 0033, Yongjie Yang 0008, Xiaoran Fan, Xiuzhen Guo, Jie Xiong 0001, Longfei Shangguan |
MobiCom | 3 |
| 2024 | See Through Vehicles: Fully Occluded Vehicle Detection with Millimeter Wave RadarabstractA crucial task in autonomous driving is to continuously detect nearby vehicles. Problems thus arise when a vehicle is occluded and becomes "unseeable", which may lead to accidents. In this study, we develop mmOVD, a system that can detect fully occluded vehicles by involving millimeter-wave radars to capture the ground-reflected signals passing beneath the blocking vehicle's chassis. The foremost challenge here is coping with ghost points caused by frequent multi-path reflections, which highly resemble the true points. We devise a set of features that can efficiently distinguish the ghost points by exploiting the neighbor points' spatial and velocity distributions. We also design a cumulative clustering algorithm to effectively aggregate the unstable ground-reflected radar points over consecutive frames to derive the bounding boxes of the vehicles. Chenming He, Chengzhen Meng, Chunwang He, Xiaoran Fan, Yubo Yan, Yanyong Zhang |
MobiCom | 4 |
| 2024 | Enabling Hands-Free Voice Assistant Activation on EarphonesabstractWe present the design and implementation of EarVoice, a lightweight mobile service that enables hands-free voice assistant activation on commodity earphones. EarVoice comprises two design modules: one for joint speech detection and primary user identification that explores the attributes of the air channel and in-body audio pathway to differentiate between the primary user and others nearby; and another for accurate wakeup word enhancement, which employs a "copy, paste, and adapt" approach to reconstruct the missing high-frequency component in speech recordings. To minimize false positives, enhance agility, and preserve privacy, we deploy EarVoice on a dongle where the proposed signal processing algorithms are streamlined with a gating mechanism to permit only the primary user's speech to enter the pairing device (e.g., a smartphone) for wakeup word recognition, preventing unintended disclosure of ambient conversations. We implemented the dongle on a 4-layer PCB board and conducted extensive experiments with 23 participants in both controlled and uncontrolled scenarios. The experiment results show that EarVoice achieves around 90% wakeup word recognition accuracy in stationary scenarios, which is on par with the high-end, multi-sensor fusion-based Airpods Pro earbud. EarVoice's performance drops to 84% on mobile cases, slightly worse than Airpods (around 90%). Tao Chen 0033, Yongjie Yang 0008, Chonghao Qiu, Xiaoran Fan, Xiuzhen Guo, Longfei Shangguan |
MobiSys | 4 |
| 2023 | AmbiSense: Acoustic Field Based Blindspot-Free Proximity Detection and Bearing EstimationabstractIn this paper, we present AmbiSense, an acoustic field based sensing system that performs proximity detection and bearing estimation for safer physical human-robot interactions. A single low cost piezoelectric transducer is used to setup this novel acoustic sensing modality to create a blindspot-free sound field engulfing a robot arm. Two detection algorithms leveraging spectral information from reflected audio waves of objects entering the acoustic field are proposed to infer object presence and bearing. We also present a new receiver structure which improves signal to noise ratio (SNR). AmbiSense is paired with a collision avoidance inverse kinematic solver for real world deployment on a Kinova Gen3 robot. Validation is performed using ten test objects generating 2000 proximity and bearing estimation events in real world settings, we show that AmbiSense detects proximity with 93.8% sensitivity and 96.6 % specificity. It estimates bearing and maps it to three zones on a robot link with 100% sensitivity and specificity, while using fewer sensors than state of the art methods for similar coverage. Siddharth Rupavatharam, Xiaoran Fan, Caleb Escobedo, Dae-Won Lee, Lawrence D. Jackel, Richard E. Howard, Colin Prepscius, Daniel D. Lee, Volkan Isler |
IROS | 2 |
| 2023 | APG: Audioplethysmography for Cardiac Monitoring in HearablesabstractThis paper presents Audioplethysmography (APG), a novel cardiac monitoring modality for active noise cancellation (ANC) headphones. APG sends a low intensity ultrasound probing signal using an ANC headphone's speakers and receives the echoes via the on-board feedback microphones. We observed that, as the volume of ear canals slightly changes with blood vessel deformations, the heartbeats will modulate these ultrasound echoes. We built mathematical models to analyze the underlying physics and propose a multi-tone APG signal processing pipeline to derive the heart rate and heart rate variability in both constrained and unconstrained settings. APG enables robust monitoring of cardiac activities using mass-market ANC headphones in the presence of music playback and body motion such as running. Xiaoran Fan, David Pearl, Richard E. Howard, Longfei Shangguan, Trausti Thormundsson |
MobiCom | 1 |
| 2022 | PoseKernelLifter: Metric Lifting of 3D Human Pose using SoundabstractReconstructing the 3D pose of a person in metric scale from a single view image is a geometrically ill-posed problem. For example, we can not measure the exact distance of a person to the camera from a single view image without additional scene assumptions (e.g., known height). Existing learning based approaches circumvent this issue by reconstructing the 3D pose up to scale. However, there are many applications such as virtual telepresence, robotics, and augmented reality that require metric scale reconstruction. In this paper, we show that audio signals recorded along with an image, provide complementary information to reconstruct the metric 3D pose of the person. The key insight is that as the audio signals traverse across the 3D space, their interactions with the body provide metric information about the body's pose. Based on this insight, we introduce a time-invariant transfer function called pose kernel-the impulse response of audio signals induced by the body pose. The main properties of the pose kernel are that (1) its envelope highly correlates with 3D pose, (2) the time response corresponds to arrival time, indicating the metric distance to the microphone, and (3) it is invariant to changes in the scene geometry configurations. Therefore, it is readily generalizable to unseen scenes. We design a multistage 3D CNN that fuses audio and visual signals and learns to reconstruct 3D pose in a metric scale. We show that our multi-modal method produces accurate metric reconstruction in realworld scenes, which is not possible with state-of-the-art lifting approaches including parametric mesh regression and depth regression. Zhijian Yang, Xiaoran Fan, Volkan Isler, Hyunsoo Park |
CVPR | 2 |
| 2022 | NSNet: Non-saliency Suppression Sampler for Efficient Video Recognition
Boyang Xia, Dongliang He, Haosen Yang 0004, Xiaoran Fan, Wanli Ouyang |
ECCV (34) | 7 |
| 2022 | Towards Remote Auscultation with Commodity EarphonesabstractVirtual visits (a.k.a., telehealth) have been promoted in response to the COVID pandemic since early 2020. Despite its convenience, the current virtual visit practice barely relies on video observation and talking. The specialist, however, cannot accurately assess the patient's health condition by listening to acoustic cardiopulmonary signals emanating from the patient's heart with a stethoscope. In this poster, we explore the feasibility of remote auscultation in virtual visits settings by reusing the patient's earphones as a stethoscope. The proposed hardware-software system captures the minute heartbeats from the patient's ear canal. It then offloads these noisy cardiac signals to the pairing device (e.g., a smartphone or a laptop) to reconstruct fine-grained Phonocardiogram (PCG) signals. By listening to the reconstructed PCG signals, the specialist can easily assess the patient's health condition and make the most informed diagnosis. We describe the design challenges and explain our technical roadmap. Tao Chen 0033, Xiaoran Fan, Yongjie Yang 0008, Longfei Shangguan |
SenSys | 2 |
| 2022 | HeadFi II: Toward More Resilient Earable Computing PlatformabstractEarables are embedded devices that can be placed in, on, or around the ear to sense human motions and physiological activities over an extended period of time. However, today's earable design principle heavily relies on dedicated sensors (e.g., accelerometer, gyroscope, proximity sensor), which inevitably adds cost, weight, and power consumption to earable devices, constituting a critical bottleneck in their wide adoption. Moreover, the tight coupling of sensors with onboard microcontrollers makes existing earables difficult to program, raising the barrier of entry to earable computing. Xueteng Qian, Xiuzhen Guo, Yongjie Yang 0008, Xiaoran Fan, Longfei Shangguan |
SenSys | 4 |
| 2021 | AuraSense: Robot Collision Avoidance by Full Surface Proximity DetectionabstractPerceiving obstacles and avoiding collisions is fundamental to the safe operation of a robot system, particularly when the robot must operate in highly dynamic human environments. Proximity detection using on-robot sensors can be used to avoid or mitigate impending collisions. However, existing proximity sensing methods are orientation and placement dependent, resulting in blind spots even with large numbers of sensors. In this paper, we introduce the phenomenon of the Leaky Surface Wave (LSW), a novel sensing modality, and present AuraSense, a proximity detection system using the LSW. AuraSense is the first system to realize no-dead-spot proximity sensing for robot arms. It requires only a single pair of piezoelectric transducers, and can easily be applied to off-the-shelf robots with minimal modifications. We further introduce a set of signal processing techniques and a lightweight neural network to address the unique challenges in using the LSW for proximity sensing. Finally, we demonstrate a prototype system consisting of a single piezoelectric element pair on a robot manipulator, which validates our design. We conducted several micro benchmark experiments and performed more than 2000 on-robot proximity detection trials with various potential robot arm materials, colliding objects, approach patterns, and robot movement patterns. AuraSense achieves 100% and 95.3% true positive proximity detection rates when the arm approaches static and mobile obstacles respectively, with a true negative rate over 99%, showing the real-world viability of this system. Xiaoran Fan, Riley Simmons-Edler, Dae-Won Lee, Lawrence D. Jackel, Richard E. Howard, Daniel D. Lee |
IROS | 1 |
| 2021 | HeadFi: bringing intelligence to all headphonesabstractHeadphones continue to become more intelligent as new functions (e.g., touch-based gesture control) appear. These functions usually rely on auxiliary sensors (e.g., accelerometer and gyroscope) that are available in smart headphones. However, for those headphones that do not have such sensors, supporting these functions becomes a daunting task. This paper presents HeadFi, a new design paradigm for bringing intelligence to headphones. Instead of adding auxiliary sensors into headphones, HeadFi turns the pair of drivers that are readily available inside all headphones into a versatile sensor to enable new applications spanning across mobile health, user-interface, and context-awareness. HeadFi works as a plug-in peripheral connecting the headphones and the pairing device (e.g., a smartphone). The simplicity (can be as simple as only two resistors) and small form factor of this design lend itself to be embedded into the pairing device as an integrated circuit. We envision HeadFi can serve as a vital supplementary solution to existing smart headphone design by directly transforming large amounts of existing "dumb" headphones into intelligent ones. We prototype HeadFi on PCB and conduct extensive experiments with 53 volunteers using 54 pairs of non-smart headphones under the institutional review board (IRB) protocols. The results show that HeadFi can achieve 97.2%--99.5% accuracy on user identification, 96.8%--99.2% accuracy on heart rate monitoring, and 97.7%--99.3% accuracy on gesture recognition. Xiaoran Fan, Longfei Shangguan, Siddharth Rupavatharam, Yanyong Zhang, Jie Xiong 0001, Richard E. Howard |
MobiCom | 1 |
| 2020 | Acoustic Collision Detection and Localization for Robot ManipulatorsabstractCollision detection is critical for safe robot operation in the presence of humans. Acoustic information originating from collisions between robots and objects provides opportunities for fast collision detection and localization; however, audio information from microphones on robot manipulators needs to be robustly differentiated from motors and external noise sources. In this paper, we present Panotti, the first system to efficiently detect and localize on-robot collisions using low-cost microphones. We present a novel algorithm that can localize the source of a collision with centimeter level accuracy and is also able to reject false detections using a robust spectral filtering scheme. Our method is scalable, easy to deploy, and enables safe and efficient control for robot manipulator applications. We implement and demonstrate a prototype that consists of 8 miniature microphones on a 7 degree of freedom (DOF) manipulator to validate our design. Extensive experiments show that Panotti realizes near perfect on-robot true positive collision detection rate with almost zero false detections even in high noise environments. In terms of accuracy, it achieves an average localization error of less than 3.8 cm under various experimental settings. Xiaoran Fan, Dae-Won Lee, Yuan Chen 0006, Colin Prepscius, Volkan Isler, Lawrence D. Jackel, H. Sebastian Seung, Daniel D. Lee |
IROS | 1 |
| 2020 | Towards flexible wireless charging for medical implants using distributed antenna systemabstractThis paper presents the design, implementation and evaluation of In-N-Out, a software-hardware solution for far-field wireless power transfer. In-N-Out can continuously charge a medical implant residing in deep tissues at near-optimal beamforming power, even when the implant moves around inside the human body. To accomplish this, we exploit the unique energy ball pattern of distributed antenna array and devise a backscatter-assisted beamforming algorithm that can concentrate RF energy on a tiny spot surrounding the medical implant. Meanwhile, the power levels on other body parts stay in low level, reducing the risk of overheating. We proto-type In-N-Out on 21 software-defined radios and a printed circuit board (PCB). Extensive experiments demonstrate that In-N-Out achieves 0.37 mW average charging power inside a 10 cm-thick pork belly, which is sufficient to wirelessly power a range of commercial medical devices. Our head-to-head comparison with the state-of-the-art approach shows that In-N-Out achieves 5.4X-18.1X power gain when the implant is stationary, and 5.3X-7.4X power gain when the implant is in motion. Xiaoran Fan, Longfei Shangguan, Richard E. Howard, Yanyong Zhang, Yao Peng 0002, Jie Xiong 0001, Xiang-Yang Li 0001 |
MobiCom | 1 |
| 2018 | Secret-Focus: A Practical Physical Layer Secret Communication System by Perturbing Focused Phases in Distributed BeamformingabstractEnsuring confidentiality of communication is fundamental to securing the operation of a wireless system, where eavesdropping is easily facilitated by the broadcast nature of the wireless medium. By applying distributed beamforming among a coalition, we show that a new approach for assuring physical layer secrecy, without requiring any knowledge about the eavesdropper or injecting any additional cover noise, is possible if the transmitters frequently perturb their phases around the proper alignment phase while transmitting messages. This approach is readily applied to amplitude-based modulation schemes, such as PAM or QAM. We present our secrecy mechanisms, prove several important secrecy properties, and develop a practical secret communication system design. We further implement and deploy a prototype that consists of 16 distributed transmitters using USRP N210s in a 20×20×3m3area. By sending more than 160M bits over our system to the receiver, depending on system parameter settings, we measure that the eavesdroppers failed to decode 30%-60% of the bits cross multiple locations while the intended receiver has an estimated bit error ratio of 3×10-6. Xiaoran Fan, Wade Trappe, Yanyong Zhang, Richard E. Howard, Zhu Han 0001 |
INFOCOM | 1 |
| 2018 | Enabling Concurrent IoT Transmissions in Distributed C-RANabstractAs rapid expansion of the low-cost next billion devices, wireless sensor networks (WSN) undertake much denser low-end internet of things (IoT) nodes nowadays. In the meantime, the future next 5 generation (5G) radio base stations (BS) are granted more capabilities. Distributed cloud radio access network (C-RAN) is becoming available for the future massive WSN. However, real-world distributed C-RAN is less explored for low-end IoT based WSN due to its difficulties in implementation. In this paper, we built a distributed C-RAN which has tens of distributed radio frontends using USRP N210s in a 20 × 20 × 3 m3 area. By exploiting the inherent hardware properties of low-end IoT devices and the spatial diversity of distributed C-RAN system, we show the distributed C-RAN can potentially decode collided signals from low-end IoT devices with all signal processing been done on the cloud. Xiaoran Fan, Zhenzhou Qi, Zhenhua Jia, Yanyong Zhang |
SenSys | 1 |