VLDB 2026 Research / reviewers in the wild / expert
Siheng Li
dblp:312/9450
· DBLP profile ↗
21ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement Learning on Pre-Training DataabstractSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang, Kun Li, Haoyuan Wu, Wujiajia, Zihao Zheng, Chenchen Zhang, Kun Shi, Xue Gong, Qi Yi, Ruibin Xiong, Tingqiang Xu, Yuhao Jiang, Jianfeng Yan, Yuyuan Zeng, Guanghui Xu, Jinbao Xue, Zhijiang xu, Zheng Fang, Shuai LI, Qibin Liu, Xiaoxue Li, Zhuoyu Li, Yangyu Tao, Fei Gao, Cheng Jiang, Bochao Wang, Kai Liu, Jianchen Zhu, Wai Lam, Bo Zhou, Di Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Siheng Li, Kejiao Li 0001, Zenan Xu, Guanhua Huang, Haoyuan Wu, Qi Yi, Ruibin Xiong, Tingqiang Xu, Jianfeng Yan, Yuyuan Zeng, Jinbao Xue, Zhijiang Xu, Qibin Liu, Zhuoyu Li, Yangyu Tao, Bochao Wang, Kai Liu 0052, Jianchen Zhu, Wai Lam, Di Wang 0052 |
ACL (1) | 1 |
| 2026 | Empowering Contactless Sleep Health Monitoring with Multi-task Learning
Zeyu Long, Beihong Jin, Siheng Li, Zhi Wang 0016, Xiaoyong Ren, Haiqin Liu |
PAKDD (1) | 3 |
| 2025 | ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationabstractWe introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs, requiring LMMs to generate the corresponding code for chart rendering.
ChartMimic includes $4,800$ human-curated (figure, instruction, code) triplets, which represent the authentic chart use cases found in scientific papers across various domains (e.g., Physics, Computer Science, Economics, etc). These charts span $18$ regular types and $4$ advanced types, diversifying into $201$ subcategories.
Furthermore, we propose multi-level evaluation metrics to provide an automatic and thorough assessment of the output code and the rendered charts.
Unlike existing code generation benchmarks, ChartMimic places emphasis on evaluating LMMs' capacity to harmonize a blend of cognitive capabilities, encompassing visual understanding, code generation, and cross-modal reasoning. The evaluation of $3$ proprietary models and $14$ open-weight models highlights the substantial challenges posed by ChartMimic. Even the advanced GPT-4o, InternVL2-Llama3-76B only achieved an average score across Direct Mimic and Customized Mimic tasks of $82.2$ and $61.6$, respectively, indicating significant room for improvement.
We anticipate that ChartMimic will inspire the development of LMMs, advancing the pursuit of artificial general intelligence. Cheng Yang 0002, Chufan Shi, Bo Shui, Junjie Wang 0011, Mohan Jing, Linran Xu, Siheng Li, Gongye Liu, Xiaomei Nie, Deng Cai 0002, Yujiu Yang 0001 |
ICLR | 9 |
| 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityabstractMathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens – elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Zicheng Lin, Qiuzhi Liu, Xing Wang 0007, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang 0001, Zhaopeng Tu |
ICML | 8 |
| 2025 | MULoc: Towards Millimeter-Accurate Localization for Unlimited UWB Tags via Anchor OverhearingabstractRecent years have seen rapid advancements in ultra-wideband (UWB)-based localization systems. However, most existing solutions offer only centimeter-level accuracy and support a limited number of UWB tags, which fails to meet the growing demands of emerging sensing applications (e.g., virtual reality). This paper presents MULoc, the first system that can localize an unlimited number of UWB tags with millimeter-level accuracy. At the core of MULoc is the innovative use of UWB phase, which can provide finer-grained distance measurement than traditional time-of-f1ight (ToF) estimates. To accurately obtain phase estimates from unsynchronized devices, we introduce a novel localization scheme called anchor overhearing (AO) and eliminate raw signal errors through a signal-difference-based technique. For precise tag localization, we resolve phase ambiguity by combining a fusion-based filtering method and frequency hopping. We implement MULoc on commercial UWB modules. Extensive experiments demonstrate that our system achieves a median localization error of 0.47 mm and 90-th percentile error of 1.02 cm, reducing the error of traditional method by 91.12%. Junqi Ma 0002, Fusang Zhang, Beihong Jin, Siheng Li, Zhi Wang 0016 |
INFOCOM | 4 |
| 2025 | Self-Supervised Human Mesh Recovery from Partial Point Cloud via a Self-Improving LoopabstractAccurate 3D human mesh recovery from point clouds remains challenging. Most existing methods depend on full 3D supervision or complete input data, both of which are difficult to obtain in practice. %cr update This calls for robust solutions capable of handling partial point clouds in a self-supervised manner. However, the incompleteness of point clouds and the absence of supervision signals pose dual challenges. To tackle these challenges, This calls for robust solutions to handle partial point clouds in a self-supervised manner. To tackle the dual challenges of point cloud incompleteness and the absence of supervision signals, we propose a novel method named SS-HMR, which offers three key insights. First, we estimate point-wise semantics in a self-supervised manner to match partial inputs with a canonical template. The resulting correspondences serve as supervision signals for the regression network in human mesh recovery. Second, we incorporate regression-based and optimization-based paradigms into a self-improving loop: the regression network provides strong initialization for optimization, while the optimization routine generates pseudo-labels that, in turn, enhance the regression network. This mutual feedback enables more accurate and stable mesh recovery over time. Third, generating multiple initializations and selecting the best result mitigates the optimization routine's sensitivity to initialization, improving robustness to sparse and noisy data. %cr update Third, to mitigate sensitivity to initialization in the optimization routine, we generate diverse initialization candidates and transform the challenge of escaping local optima into a controllable selection task, improving robustness against sparse and noisy data. Extensive experiments are conducted on three public datasets and results demonstrate that SS-HMR outperforms existing methods. Notably, SS-HMR performs excellently on different test data, whether from original point clouds captured by depth cameras or LiDAR devices, or from noise-added ones. This shows that SS-HMR has strong generalization ability and robustness across different data sources. Codes are available at https://github.com/suchang-99/SS-HMR. Beihong Jin, Fusang Zhang, Siheng Li, Zhi Wang 0016 |
ACM Multimedia | 4 |
| 2024 | Parallel Vertex Diffusion for Unified Visual GroundingabstractUnified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling object box and contour prediction and provide a text-powered interface to vast related multi-modal tasks, e.g., visual question answering and captioning. However, these methods typically generate vertexes sequentially through autoregression, which is prone to be trapped in error accumulation and heavy computation, especially for high-dimension sequence generation in complex scenarios. In this paper, we develop Parallel Vertex Diffusion (PVD) based on the parallelizability of diffusion models to accurately and efficiently generate vertexes in a parallel and scalable manner. Since the coordinates fluctuate greatly, it typically encounters slow convergence when training diffusion models without geometry constraints. Therefore, we consummate our PVD by two critical components, i.e., center anchor mechanism and angle summation loss, which serve to normalize coordinates and adopt a differentiable geometry descriptor from the point-in-polygon problem of computational geometry to constrain the overall difference of prediction and label vertexes. These innovative designs empower our PVD to demonstrate its superiority with state-of-the-art performance across various grounding tasks. Zesen Cheng, Kehan Li 0002, Peng Jin 0001, Siheng Li, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
AAAI | 4 |
| 2024 | Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-ContrastabstractMixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activates a different subset of experts determined by a routing mechanism. However, the unchosen experts in MoE models do not contribute to the output, potentially leading to underutilization of the model's capacity.
In this work, we first conduct exploratory studies to demonstrate that increasing the number of activated experts does not necessarily improve and can even degrade the output quality. Then, we show that output distributions from an MoE model using different routing strategies substantially differ, indicating that different experts do not always act synergistically.
Motivated by these findings, we propose **S**elf-**C**ontrast **M**ixture-**o**f-**E**xperts (SCMoE), a training-free strategy that utilizes unchosen experts in a self-contrast manner during inference.
In SCMoE, the next-token probabilities are determined by contrasting the outputs from strong and weak activation using the same MoE model.
Our method is conceptually simple and computationally lightweight, as it incurs minimal latency compared to greedy decoding.
Experiments on several benchmarks (GSM8K, StrategyQA, MBPP and HumanEval) demonstrate that SCMoE can consistently enhance Mixtral 8x7B’s reasoning capability across various domains. For example, it improves the accuracy on GSM8K from 61.79 to 66.94.
Moreover, combining SCMoE with self-consistency yields additional gains, increasing major@20 accuracy from 75.59 to 78.31. Chufan Shi, Cheng Yang 0002, Jiahao Wang 0005, Taiqiang Wu, Siheng Li, Deng Cai 0002, Yujiu Yang 0001, Yu Meng 0001 |
NeurIPS | 6 |
| 2024 | An Energy-based Model for Word-level AutoCompletion in Computer-aided TranslationabstractAbstract Word-level AutoCompletion (WLAC) is a rewarding yet challenging task in Computer-aided Translation. Existing work addresses this task through a classification model based on a neural network that maps the hidden vector of the input context into its corresponding label (i.e., the candidate target word is treated as a label). Since the context hidden vector itself does not take the label into account and it is projected to the label through a linear classifier, the model cannot sufficiently leverage valuable information from the source sentence as verified in our experiments, which eventually hinders its overall performance. To alleviate this issue, this work proposes an energy-based model for WLAC, which enables the context hidden vector to capture crucial information from the source sentence. Unfortunately, training and inference suffer from efficiency and effectiveness challenges, therefore we employ three simple yet effective strategies to put our model into practice. Experiments on four standard benchmarks demonstrate that our reranking-based approach achieves substantial improvements (about 6.07%) over the previous state-of-the-art model. Further analyses show that each strategy of our approach contributes to the final performance.1 Cheng Yang 0007, Guoping Huang, Mo Yu, Zhirui Zhang, Siheng Li, Shuming Shi 0001, Yujiu Yang 0001, Lemao Liu |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | Leveraging Attention-reinforced UWB Signals to Monitor Respiration during SleepabstractThe respiration state during overnight sleep is an important indicator of human health. However, existing contactless solutions for sleep respiration monitoring either perform in controlled environments and have low usability in practical scenarios or only provide coarse-grained respiration rates, being unable to accurately detect abnormal events in patients. In this article, we propose Respnea, a non-intrusive sleep respiration monitoring system using an ultra-wideband device. Particularly, we propose a profiling algorithm, which can locate the sleep positions in non-controlled environments and identify different subject states. Further, we construct a deep learning model that adopts a multi-head self-attention mechanism and learns the patterns implicit in the respiration signals to distinguish sleep respiration events at a granularity of seconds. To improve the generalization of the model, we propose a contrastive learning strategy to learn a robust representation of the respiration signals. We deploy our system in hospital and home scenarios and conduct experiments on data from healthy subjects and patients with sleep disorders. The experimental results show that Respnea achieves high temporal coverage and low errors (a median error of 0.27 bpm) in respiration rate estimation and reaches an accuracy of 94.44% on diagnosing the severity of sleep apnea-hypopnea syndrome. Siheng Li, Beihong Jin, Zhi Wang 0016, Fusang Zhang, Xiaoyong Ren, Haiqin Liu |
ACM Trans. Sens. Networks | 1 |
| 2023 | Out-of-Candidate Rectification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffer from the unsolicited Out-of-Candidate (OC) error predictions that do not belong to the label candidates, which could be avoidable since the contradiction with image-level class tags is easy to be detected. In this paper, we develop a group ranking-based Out-of-f;Candidate Rectification (OCR) mechanism in a plug-and-play fashion. Firstly, we adaptively split the semantic categories into In-Candidate (IC) and OC groups for each OC pixel according to their prior annotation correlation and posterior prediction correlation. Then, we derive a differentiable rectification loss to force OC pixels to shift to the IC group. Incorporating OCR with seminal baselines (e.g., AffinityNet, SEAM, MCTformer), we can achieve remarkable performance gains on both Pascal VOC (+3.2%, +3.3%, +0.8% mIoU) and MS COCO (+1.0%, +1.3%, +0.5% mIoU) datasets with negligible extra training overhead, which jus-tifies the effectiveness and generality of OCR.††Ŋ github.com/sennnnn/Out-of-Candidate-Rectification Zesen Cheng, Pengchong Qiao, Kehan Li 0002, Siheng Li, Pengxu Wei, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
CVPR | 4 |
| 2023 | Question Answering as Programming for Solving Time-Sensitive QuestionsabstractQuestion answering plays a pivotal role in human daily life because it involves our acquisition of knowledge about the world.However, due to the dynamic and ever-changing nature of real-world facts, the answer can be completely different when the time constraint in the question changes.Recently, Large Language Models (LLMs) have shown remarkable intelligence in question answering, while our experiments reveal that the aforementioned problems still pose a significant challenge to existing LLMs.This can be attributed to the LLMs' inability to perform rigorous reasoning based on surfacelevel text semantics.To overcome this limitation, rather than requiring LLMs to directly answer the question, we propose a novel approach where we reframe the Question Answering task as Programming (QAaP).Concretely, by leveraging modern LLMs' superior capability in understanding both natural language and programming language, we endeavor to harness LLMs to represent diversely expressed text as wellstructured code and select the best matching answer from multiple candidates through programming.We evaluate our QAaP framework on several time-sensitive question answering datasets and achieve decent improvement, up to 14.5% over strong baselines.1 Cheng Yang 0007, Bei Chen 0008, Siheng Li, Jian-Guang Lou, Yujiu Yang 0001 |
EMNLP | 4 |
| 2023 | WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image SegmentationabstractThe top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturbed by Inferior Positive (IP) errors due to the lack of prior object information. Nevertheless, we discover that two types of methods are highly complementary for restraining respective weaknesses but the direct average combination leads to harmful interference. In this context, we build Win-win Cooperation (WiCo) to exploit complementary nature of two types of methods on both interaction and integration aspects for achieving a win-win improvement. For the interaction aspect, Complementary Feature Interaction (CFI) introduces prior object information to bottom-up branch and provides fine-grained information to top-down branch for complementary feature enhancement. For the integration aspect, Gaussian Scoring Integration (GSI) models the gaussian performance distributions of two branches and weighted integrates results by sampling confident scores from the distributions. With our WiCo, several prominent bottom-up and top-down combinations achieve remarkable improvements on three common datasets with reasonable extra costs, which justifies effectiveness and generality of our method. Zesen Cheng, Peng Jin 0001, Hao Li 0073, Kehan Li 0002, Siheng Li, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 5 |
| 2023 | Modeling Fine-grained Information via Knowledge-aware Hierarchical Graph for Zero-shot Entity RetrievalabstractZero-shot entity retrieval, aiming to link mentions to candidate entities under the zero-shot setting, is vital for many tasks in Natural Language Processing. Most existing methods represent mentions/entities via the sentence embeddings of corresponding context from the Pre-trained Language Model. However, we argue that such coarse-grained sentence embeddings can not fully model the mentions/entities, especially when the attention scores towards mentions/entities are relatively low. In this work, we propose GER, a Graph enhanced Entity Retrieval framework, to capture more fine-grained information as complementary to sentence embeddings. We extract the knowledge units from the corresponding context and then construct a mention/entity centralized graph. Hence, we can learn the fine-grained information about mention/entity by aggregating information from these knowledge units. To avoid the graph bottleneck for the central mention/entity node, we construct a hierarchical graph and design a novel Hierarchical Graph Attention Network~(HGAN). Experimental results on popular benchmarks demonstrate that our proposed GER framework performs better than previous state-of-the-art models. Taiqiang Wu, Xingyu Bai, Weigang Guo, Weijie Liu 0002, Siheng Li, Yujiu Yang 0001 |
WSDM | 5 |
| 2022 | Sleep Respiration Monitoring Using Attention-reinforced Radar SignalsabstractExisting contactless solutions on sleep respiration monitoring are either performed in controlled environments, having poor usability in practical scenarios, or only provide coarse-grained respiration rates, being unable to accurately detect abnormal events of patients. In this paper, we propose Respnea, a non-invasive sleep respiration monitoring system using an impulse-radio ultra-wideband (IR-UWB) radar. Particularly, we propose a profiling algorithm, which can locate the sleep positions in non-controlled environments and identify different states of subjects. Further, we construct a deep learning model which adopts a multi-headed self-attention and learn the patterns implicit in the respiration signal so as to distinguish sleep respiration events at a granularity of seconds. We conduct experiments on data collected from patients with sleep disorders and healthy subjects. The experimental results show that Respnea achieves a low error (less than 0.27 bpm) in respiration rate estimation and reaches the accuracy of 88.89% diagnosing the severity of Sleep Apnea-Hypopnea Syndrome. Siheng Li, Zhi Wang 0016, Beihong Jin, Fusang Zhang, Xiaoyong Ren |
BIBM | 1 |
| 2022 | MCSCSet: A Specialist-annotated Dataset for Medical-domain Chinese Spelling CorrectionabstractChinese Spelling Correction (CSC) is gaining increasing attention in recent years. Despite its extensive use in many applications, such as search engine and optical character recognition system, little has been explored in medical scenarios in which complex and uncommon medical entities are easily misspelled. Correcting the misspellings of medical entities is arguably more difficult than those in the open domain due to its requirements of specific domain knowledge. In this work, we define the task of Medical-domain Chinese Spelling Correction (MCSC) and propose MCSCSet, a large-scale specialist-annotated dataset that contains about 200k samples. In contrast to existing open-domain CSC datasets, MCSCSet involves: i) extensive real-world medical queries collected from Tencent Yidian, ii) corresponding misspelled sentences manually annotated by medical specialists. Our work further offers a medical-domain confusion set consisting of the common error-prone characters in medicine and their corresponding misspellings. Extensive empirical studies have shown significant gaps between the open-domain and medical-domain spelling correction, highlighting the need to develop high-quality datasets that allow for CSC in specific domains. Moreover, our work benchmarks several representative methods, establishing baselines for future work. Wangjie Jiang, Zhihao Ye, Zijing Ou, Ruihui Zhao, Jianguang Zheng, Yi Liu 0057, Bang Liu 0003, Siheng Li, Yujiu Yang 0001, Yefeng Zheng 0001 |
CIKM | 8 |
| 2022 | Multi-Turn Incomplete Utterance Restoration As Object DetectionabstractIn this paper, we investigate the task of multi-turn incomplete utterance restoration to tackle the issue of frequent coreference and information omission in multi-turn dialogues. Recent works mainly focus on edit-based approaches which have been proven to outperform traditional generation-based models in terms of accuracy and efficiency. However, they only model token-level edit relationships while ignoring span-level edit relationships. Our experiments find this breaks the semantic integrity of edit span, which causes inaccurate edit span prediction and disfluent utterance restoration. To address the problem, we propose a novel approach to directly model span-level edit relationships between the incomplete utterance and context. Specifically, we build an edit matrix in which each rectangular region represents a span-level edit operation. Then, we detect the region with a well-designed dual-branch detection module inspired by object detection. Empirical results demonstrate that our method outperforms state-of-the-art methods significantly on two public datasets. In addition, further studies verify that our method is capable of preserving the semantic integrity of edit span. Wangjie Jiang, Siheng Li, Jiayi Li 0002, Yujiu Yang 0001 |
ICASSP | 2 |
| 2022 | A New Approach to Training Multiple Cooperative Agents for Autonomous DrivingabstractTraining multiple agents to perform safe and coop-erative control in the complex scenarios of autonomous driving has been a challenge. For a small fleet of cars moving together, this paper proposes Lepus, a new approach to training multiple agents. Lepus adopts a pure cooperative manner for training multiple agents, featured with the shared parameters of policy networks and the shared reward function of multiple agents. In particular, Lepus pre-trains the policy networks via an adversarial process, improving its collaborative decision-making capability and further the stability of car driving. Moreover, for alleviating the problem of sparse rewards, Lepus learns an approximate reward function from expert trajectories by combining a random network and a distillation network. We conduct extensive experiments on the MADRaS simulation plat-form. The experimental results show that multiple agents trained by Lepus can avoid collisions as many as possible while driving simultaneously and outperform the other four methods, that is, DDPG-FDE, PSDDPG, MADDPG, and MAGAIL(DDPG) in terms of stability. Ruiyang Yang, Siheng Li, Beihong Jin |
IJCNN | 2 |
| 2022 | EmpHi: Generating Empathetic Responses with Human-like IntentsabstractIn empathetic conversations, humans express their empathy to others with empathetic intents.However, most existing empathetic conversational methods suffer from a lack of empathetic intents, which leads to monotonous empathy.To address the bias of the empathetic intents distribution between empathetic dialogue models and humans, we propose a novel model to generate empathetic responses with humanconsistent empathetic intents, EmpHi for short.Precisely, EmpHi learns the distribution of potential empathetic intents with a discrete latent variable, then combines both implicit and explicit intent representation to generate responses with various empathetic intents.Experiments show that EmpHi outperforms state-ofthe-art models in terms of empathy, relevance, and diversity on both automatic and human evaluation.Moreover, the case studies demonstrate the high interpretability and outstanding performance of our model.Our code are avaliable at https://github.com/mattc95/EmpHi. Mao Yan Chen, Siheng Li, Yujiu Yang 0001 |
NAACL-HLT | 2 |
| 2021 | Exploiting Passive Beamforming of Smart Speakers to Monitor Human Heartbeat in Real TimeabstractCurrently, cardiac diseases have become one of the biggest health concerns. Existing heartbeat monitoring methods either require dedicated intrusive devices (e.g., ECG devices) that suffer high costs or leverage video camera analyses that are light-sensitive. In this paper, leveraging the acoustic signals sent by a speaker and received by a microphone array, we develop a prototype system to achieve the contactless and low-cost heartbeat monitoring. In particular, while we exploit the passive beamforming to enhance the user's heartbeat signal, we design a filtering method in frequency domain to remove the line-of-sight (LoS) impact and retain the target-reflected signals, and propose a wideband time-delay method to estimate the direction of arrival of target-reflected signal. Thus, our prototype is able to robustly estimate the human heartbeat and push the limit of acoustic sensing range. The experimental results show that our prototype achieves a heart rate monitoring at 1.7 m with the estimation error of 0.5 bpm, which is comparable to ECG or other contact-based solutions. Zhi Wang 0016, Fusang Zhang, Siheng Li, Beihong Jin |
GLOBECOM | 3 |
| 2021 | Fine-Grained Respiration Monitoring During Overnight Sleep Using IR-UWB Radar
Siheng Li, Zhi Wang 0016, Fusang Zhang, Beihong Jin |
MobiQuitous | 1 |