EDBT 2026 Demo / reviewers in the wild / expert
Bing Li 0002
dblp:13/2692-2
· DBLP profile ↗
27ranked-venue papers
11as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 7 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataabstractSensory Temporal Action Detection (STAD) aims to localize and classify human actions within long, untrimmed sequences captured by non-visual sensors such as WiFi or inertial measurement units (IMUs). Unlike video-based TAD, STAD poses unique challenges due to the low-dimensional, noisy, and heterogeneous nature of sensory data, as well as the real-time and resource constraints on edge devices. While recent STAD models have improved detection performance, their high computational cost hampers practical deployment. In this paper, we propose SlimSTAD, a simple yet effective framework that achieves both high accuracy and low latency for STAD. SlimSTAD features a novel Decoupled Channel Modeling (DCM) encoder, which preserves modality-specific temporal features and enables efficient inter-channel aggregation via lightweight graph attention. An anchor-free cascade predictor then refines action boundaries and class predictions in a two-stage design without dense proposals. Experiments on two real-world datasets demonstrate that SlimSTAD outperforms strong video-derived and sensory baselines by an average of 2.1 mAP, while significantly reducing GFLOPs, parameters, and latency, validating its effectiveness for real-world, edge-aware STAD deployment. Wei Cui 0002, Lukai Fan, Zhenghua Chen, Min Wu 0008, Shili Xiang, Haixia Wang 0003, Bing Li 0002 |
AAAI | 7 |
| 2026 | TIDE: Making Task-Agnostic Backdoors Harder to Erase in Pre-trained Language Models
Zhigang Lu 0001, Bing Li 0002, Anan Du, Shuchao Pang |
ACISP (2) | 3 |
| 2026 | TERM Model: Tensor Ring Mixture Model for Density EstimationabstractProbabilistic modeling is a core challenge in statistical machine learning. Tensor-based probabilistic graph methods address interpretability and stability concerns encountered in neural network approaches and allow tractable inference (e.g., marginal inference and conditional inference). In this paper, we introduce tensor ring decomposition for density estimation, which reduces the number of permutation candidates compared to existing methods, while simultaneously enhancing expressive power and maintaining tractable inference. Different non-negative strategies for density function results in two variants: Born TRDE offers simpler inference and sampling but with slightly lower accuracy, while Energy TRDE, though more complex, achieves superior performance. Furthermore, a mixture model that incorporates multiple permutation candidates with adaptive weights is designed, resulting in increased expressive flexibility and comprehensiveness. Unlike existing methods that focus on finding a single optimal permutation, our approach, inspired by ensemble learning, demonstrates that combining multiple suboptimal permutations can yield superior results. Experiments demonstrate that the proposed approach excels in estimating probability density functions and sampling, capturing intricate details with competitive or superior performance compared to existing state-of-the-art (SOTA) tractable density methods. Ruituo Wu, Jiani Liu 0002, Bing Li 0002, Anh Huy Phan 0001, Ivan V. Oseledets, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | WiFi CSI Based Temporal Activity Detection via Dual Pyramid NetworkabstractWe address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency components, using a novel Signed Mask-Attention mechanism to emphasize important areas and downplay unimportant ones, with the features fused using ContraNorm. The Local Sensitive Response Encoder captures fluctuations without learning. These feature pyramids are then combined using a new cross-attention fusion mechanism. We also introduce a dataset with over 2,114 activity segments across 553 WiFi CSI samples, each lasting around 85 seconds. Extensive experiments show our method outperforms challenging baselines. Le Zhang 0001, Bing Li 0002, Yingjie Zhou 0001, Zhenghua Chen, Ce Zhu |
AAAI | 3 |
| 2025 | Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFiabstractWiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor scenario via WiFi. Routers on different floors are utilized without major modifications to the existing router layout. Many previous works require a high sample rate and ignore the phase information. In this paper, we propose Codar, a complex-valued LSTM-CNN neural network. The LSTM effectively captures temporal dependencies at a low sample rate in harsh propagation environments. Moreover, amplitude and phase features are explored jointly by complex-valued operations. Experimental results demonstrate Codar achieves 95%, 94.5%, and 99% accuracy for intrusion detection, user identification, and intruded floor identification, surpassing competitive methods. The code and dataset are available at https://github.com/ouweiting/Codar. Weiting Ou, Yipeng Liu 0001, Bing Li 0002, Le Zhang 0001, Ce Zhu |
ICASSP | 4 |
| 2025 | One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention HeadabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive capabilities in tasks requiring multimodal understanding. However, recent studies indicate that LVLMs are more vulnerable than LLMs to unsafe inputs and prone to generating harmful content. Existing defense strategies primarily include fine-tuning, input sanitization, and output intervention. Although these approaches provide a certain level of protection, they tend to be resource-intensive and struggle to effectively counter sophisticated attack techniques. To tackle such issues, we propose One-head Defense (Oh Defense), a novel yet simple approach utilizing LVLMs' internal safety capabilities. Through systematic analysis of the attention mechanisms, we discover that LVLMs' safety capabilities are concentrated within specific attention heads that respond differently to safe or unsafe inputs. Further exploration reveals that a single critical attention head can effectively serve as a safety guard, providing a strong discriminative signal that amplifies the model's inherent safety capabilities. Hence, the Oh Defense requires no additional training or external modules, making it computationally efficient while effectively reactivating suppressed safety mechanisms. Extensive experiments across diverse LVLM architectures and unsafe datasets validate our approach, i.e., the Oh Defense achieves near-perfect defense success rates (> 98\%) for unsafe inputs while maintaining low false positive rates (< 5\%) for safe content. The source code is available at https://github.com/AIASLab/Oh-Defense. Junhao Xia, Shuchao Pang, Zhigang Lu 0001, Bing Li 0002, Yongbin Zhou, Minhui Xue 0001 |
NeurIPS | 5 |
| 2025 | STADe: Sensory Temporal Action Detection via Temporal-Spectral Representation LearningabstractTemporal action detection (TAD) is a vital challenge in computer vision and the Internet of Things, aiming to detect and identify actions within temporal sequences. While TAD has primarily been associated with video data, its applications can also be extended to sensor data, opening up opportunities for various real-world applications. However, applying existing TAD models to sensory signals presents distinct challenges such as varying sampling rates, intricate pattern structures, and subtle, noise-prone patterns. In response to these challenges, we propose a Sensory Temporal Action Detection (STADe) model. STADe leverages Fourier kernels and adaptive frequency filtering to adaptively capture the nuanced interplay of temporal and frequency features underlying complex patterns. Moreover, STADe embraces adaptability by employing deep fusion at varying resolutions and scales, making it versatile enough to accommodate diverse data characteristics, such as the wide spectrum of sampling rates and action durations encountered in sensory signals. Unlike conventional models with unidirectional category-to-proposal dependencies, STADe adopts a cross-cascade predictor to introduce bidirectional and temporal dependencies within categories. To extensively evaluate STADe and promote future research in sensory TAD, we establish three diverse datasets using various sensors, featuring diverse sensor types, action categories, and sampling rates. Experiments across one public and our three new datasets demonstrate STADe's superior performance over state-of-the-art TAD models in sensory TAD tasks. Bing Li 0002, Haotian Duan, Yun Liu 0011, Le Zhang 0001, Wei Cui 0002, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | DA-Flow: Dual Attention Normalizing Flow for Skeleton-Based Video Anomaly DetectionabstractCooperation between temporal convolutional networks (TCN) and graph convolutional networks (GCN) as a processing module has shown promising results in skeleton-based video anomaly detection (SVAD). However, to maintain a lightweight model with low computational and storage complexity, shallow GCN and TCN blocks are constrained by small receptive fields and a lack of cross-dimension interaction capture. To tackle this limitation, we propose a lightweight module called the Dual Attention Module (DAM) for capturing cross-dimension interaction relationships in spatio-temporal skeletal data. It employs the frame attention mechanism to identify the most significant frames and the skeleton attention mechanism to capture broader relationships across fixed partitions with minimal parameters and total Floating Point Operations (FLOPs). Furthermore, the proposed Dual Attention Normalizing Flow (DA-Flow) integrates the DAM as a post-processing unit after GCN within the normalizing flow framework. Simulations show that the proposed model is robust against noise and negative samples. Experimental results show that DA-Flow reaches competitive or better performance than the existing state-of-the-art (SOTA) methods in terms of the micro AUC metric with the fewest parameters and FLOPs. Moreover, we found that even without training, simply using random projection without dimensionality reduction on skeleton data enables substantial anomaly detection capabilities. Ruituo Wu, Bing Li 0002, Jicong Fan 0001, Frédéric Dufaux, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Veil Privacy on Visual Data: Concealing Privacy for Humans, Unveiling for DNNs
Shuchao Pang, Ruhao Ma, Bing Li 0002, Yongbin Zhou, Yazhou Yao |
ECCV (83) | 3 |
| 2024 | Democratizing Federated WiFi-Based Human Activity Recognition Using Hypothesis TransferabstractHuman activity recognition (HAR) is a crucial task in IoT systems with applications ranging from surveillance and intruder detection to home automation and more. Recently, non-invasive HAR utilizing WiFi signals has gained considerable attention due to advancements in ubiquitous WiFi technologies. However, recent studies have revealed significant privacy risks associated with WiFi signals, raising concerns about bio-information leakage. To address these concerns, the decentralized paradigm, particularly federated learning (FL), has emerged as a promising approach for training HAR models while preserving data privacy. Nevertheless, FL models may struggle in end-user environments due to substantial domain discrepancies between the source training data and the target end-user environment. This discrepancy arises from the sensitivity of WiFi signals to environmental changes, resulting in notable domain shifts. As a consequence, FL-based HAR approaches often face challenges when deployed in real-world WiFi environments. Albeit there are pioneer attempts on federated domain adaptation, they typically require non-trivial communication and computation cost, which is prohibitively expensive especially considering edge-based hardware equipment of end-user environment. In this paper, we propose a model to democratize the WiFi-based HAR system by enhancing recognition accuracy in unannotated end-user environments while prioritizing data privacy. Our model leverages the hypothesis transfer and a lightweight hypothesis ensemble to mitigate negative transfer. We prove a tighter theoretical upper bound compared to existing multi-source federated domain adaptation models. Extensive experiments shows our model improves the average accuracy by approximately 10 absolute percentage points in both cross-person and cross-environment settings comparing several state-of-the-art baselines. Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Min Wu 0008, Joey Tianyi Zhou |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | GrapHAR: A Lightweight Human Activity Recognition Model by Exploring the Sub-Carrier CorrelationsabstractHuman activity recognition (HAR) is an important task due to its far-reaching applications, such as surveillance, healthcare systems, and human-computer interaction. Recently, Channel State Information (CSI)-based HAR has attracted increasing attention in the research community due to its ubiquitous availability, good user privacy, and fewer constraints on working conditions. Most of the existing methods for CSI-based HAR use various deep learning models, such as Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM), and Transformers, to distinguish activities based on their temporal patterns. Despite their remarkable effectiveness, these methods solely focus on temporal patterns while ignoring the correlations among sub-carriers. This limitation prevents them from achieving further performance improvement. Moreover, recent works often involve advanced yet massive and inefficient neural architectures, like Transformers, to obtain satisfactory recognition accuracy. The performance gain is traded off with a steep increase in model complexity, which leads to low efficacy and high training/inference costs outsides the small time window. To address these issues, we propose a lightweight CSI-based HAR model. Our model makes the first effort to explore the graphical correlations of CSI sub-carriers, working in conjunction with a temporal causal convolution module. The high efficacy design enables our model to be highly effective without requiring excessive model complexity. Extensive experiments conducted on four real-world datasets demonstrate that our model outperforms state-of-the-art methods, including a strong Transformer-based baseline. It achieves an average improvement of 8 percentage points in recognition accuracy, with only 10% of the parameters compared to the Transformer-based method (4.95M vs. 49.24M). Additionally, our model is significantly faster, with empirical training and execution times at least 2.07 times faster than the baseline. Wei Meng 0002, Zhicong Liu, Bing Li 0002, Wei Cui 0002, Joey Tianyi Zhou, Le Zhang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | DifFormer: Multi-Resolutional Differencing Transformer With Dynamic Ranging for Time Series AnalysisabstractTime series analysis is essential to many far-reaching applications of data science and statistics including economic and financial forecasting, surveillance, and automated business processing. Though being greatly successful of Transformer in computer vision and natural language processing, the potential of employing it as the general backbone in analyzing the ubiquitous times series data has not been fully released yet. Prior Transformer variants on time series highly rely on task-dependent designs and pre-assumed "pattern biases", revealing its insufficiency in representing nuanced seasonal, cyclic, and outlier patterns which are highly prevalent in time series. As a consequence, they can not generalize well to different time series analysis tasks. To tackle the challenges, we propose DifFormer, an effective and efficient Transformer architecture that can serve as a workhorse for a variety of time-series analysis tasks. DifFormer incorporates a novel multi-resolutional differencing mechanism, which is able to progressively and adaptively make nuanced yet meaningful changes prominent, meanwhile, the periodic or cyclic patterns can be dynamically captured with flexible lagging and dynamic ranging operations. Extensive experiments demonstrate DifFormer significantly outperforms state-of-the-art models on three essential time-series analysis tasks, including classification, regression, and forecasting. In addition to its superior performances, DifFormer also excels in efficiency - a linear time/memory complexity with empirically lower time consumption. Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Ce Zhu, Wei Wang 0011, Ivor W. Tsang, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Privacy-Preserving Cross-Environment Human Activity RecognitionabstractRecent studies have demonstrated the success of using the channel state information (CSI) from the WiFi signal to analyze human activities in a fixed and well-controlled environment. Those systems usually degrade when being deployed in new environments. A straightforward solution to solve this limitation is to collect and annotate data samples from different environments with advanced learning strategies. Although workable as reported, those methods are often privacy sensitive because the training algorithms need to access the data from different environments, which may be owned by different organizations. We present a practical method for the WiFi-based privacy-preserving cross-environment human activity recognition (HAR). It collects and shares information from different environments, while maintaining the privacy of individual person being involved. At the core of our approach is the utilization of the Johnson-Lindenstrauss transform, which is theoretically shown to be differentially private. Based on that, we further design an adversarial learning strategy to generate environment-invariant representations for HAR. We demonstrate the effectiveness of the proposed method with different data modalities from two real-life environments. More specifically, on the raw CSI dataset, it shows 2.18% and 1.24% improvements over challenging baselines for two environments, respectively. Moreover, with the discrete wavelet transform features, it further yields 5.71% and 1.55% improvements, respectively. Le Zhang 0001, Wei Cui 0002, Bing Li 0002, Zhenghua Chen, Min Wu 0008, Sin G. Teo |
IEEE Trans. Cybern. | 3 |
| 2023 | Hierarchical filtering: improving similar substring matching under edit distance
Tao Qiu, Chuanyu Zong, Xiaochun Yang 0001, Bin Wang 0015, Bing Li 0002 |
World Wide Web (WWW) | 5 |
| 2022 | Robust cross-network node classification via constrained graph mutual information
Shuiqiao Yang, Borui Cai, Taotao Cai, Jiaojiao Jiang 0001, Bing Li 0002, Jianxin Li 0001 |
Knowl. Based Syst. | 6 |
| 2021 | Two-Stream Convolution Augmented Transformer for Human Activity RecognitionabstractRecognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e.g., RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the time-over-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency. Bing Li 0002, Wei Cui 0002, Wei Wang 0011, Le Zhang 0001, Zhenghua Chen, Min Wu 0008 |
AAAI | 1 |
| 2021 | Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionabstractBERT has set a new state-of-the-art performance on entity resolution (ER) task, largely owed to fine-tuning pre-trained language models and the deep pair-wise interaction. Albeit being remarkably effective, it comes with a steep increase in computational cost, as the deep-interaction requires to exhaustively compute every tuple pair to search for co-references. For ER task, it is often prohibitively expensive due to the large cardinality to be matched. To tackle this, we introduce a siamese network structure that independently encodes tuples using BERT but delays the pair-wise interaction via an enhanced alignment network. This siamese structure enables a dedicated blocking module to quickly filter out obviously dissimilar tuple pairs, and thus drastically reduces the cardinality of fine-grained matching. Further, the blocking and entity matching are integrated into a multi-task learning framework for facilitating both tasks. Extensive experiments on multiple datasets demonstrate that our model significantly outperforms state-of-the-art models (including BERT) in both efficiency and effectiveness. Bing Li 0002, Yukai Miao, Yaoshu Wang, Yifang Sun, Wei Wang 0011 |
AAAI | 1 |
| 2020 | Fine-Grained Named Entity Typing over Distantly Supervised Data Based on Refined RepresentationsabstractFine-Grained Named Entity Typing (FG-NET) is a key component in Natural Language Processing (NLP). It aims at classifying an entity mention into a wide range of entity types. Due to a large number of entity types, distant supervision is used to collect training data for this task, which noisily assigns type labels to entity mentions irrespective of the context. In order to alleviate the noisy labels, existing approaches on FG-NET analyze the entity mentions entirely independent of each other and assign type labels solely based on mention's sentence-specific context. This is inadequate for highly overlapping and/or noisy type labels as it hinders information passing across sentence boundaries. For this, we propose an edge-weighted attentive graph convolution network that refines the noisy mention representations by attending over corpus-level contextual clues prior to the end classification. Experimental evaluation shows that the proposed model outperforms the existing research by a relative score of upto 10.2% and 8.3% for macro-f1 and micro-f1 respectively. Muhammad Asif Ali, Yifang Sun, Bing Li 0002, Wei Wang 0011 |
AAAI | 3 |
| 2020 | GraphER: Token-Centric Entity Resolution with Graph Convolutional Neural NetworksabstractEntity resolution (ER) aims to identify entity records that refer to the same real-world entity, which is a critical problem in data cleaning and integration. Most of the existing models are attribute-centric, that is, matching entity pairs by comparing similarities of pre-aligned attributes, which require the schemas of records to be identical and are too coarse-grained to capture subtle key information within a single attribute. In this paper, we propose a novel graph-based ER model GraphER. Our model is token-centric: the final matching results are generated by directly aggregating token-level comparison features, in which both the semantic and structural information has been softly embedded into token embeddings by training an Entity Record Graph Convolutional Network (ER-GCN). To the best of our knowledge, our work is the first effort to do token-centric entity resolution with the help of GCN in entity resolution task. Extensive experiments on two real-world datasets demonstrate that our model stably outperforms state-of-the-art models. Bing Li 0002, Wei Wang 0011, Yifang Sun, Linhan Zhang, Muhammad Asif Ali, Yi Wang 0017 |
AAAI | 1 |
| 2020 | Recursively Binary Modification Model for Nested Named Entity Recognition
Bing Li 0002, Shifeng Liu 0002, Yifang Sun, Wei Wang 0011, Xiang Zhao 0002 |
AAAI | 1 |
| 2020 | HAMNER: Headword Amplified Multi-Span Distantly Supervised Method for Domain Specific Named Entity RecognitionabstractTo tackle Named Entity Recognition (NER) tasks, supervised methods need to obtain sufficient cleanly annotated data, which is labor and time consuming. On the contrary, distantly supervised methods acquire automatically annotated data using dictionaries to alleviate this requirement. Unfortunately, dictionaries hinder the effectiveness of distantly supervised methods for NER due to its limited coverage, especially in specific domains. In this paper, we aim at the limitations of the dictionary usage and mention boundary detection. We generalize the distant supervision by extending the dictionary with headword based non-exact matching. We apply a function to better weight the matched entity mentions. We propose a span-level model, which classifies all the possible spans then infers the selected spans with a proposed dynamic programming algorithm. Experiments on all three benchmark datasets demonstrate that our method outperforms previous state-of-the-art distantly supervised methods. Shifeng Liu 0002, Yifang Sun, Bing Li 0002, Wei Wang 0011, Xiang Zhao 0002 |
AAAI | 3 |
| 2020 | WiFi-Based Indoor Robot Positioning Using Deep Fuzzy ForestsabstractAddressing the positioning problem of a mobile robot remains challenging to date despite many years of research. Indoor robot positioning strategies developed in the literature either rely on sophisticated computer vision techniques to handle visual inputs or require strong domain knowledge for nonvisual sensors. Although some systems have been deployed, the former may be lacking due to the intrinsic limitation of cameras (such as calibration, data association, system initialization, etc.) and the latter usually only works under certain environment layouts and additional equipment. To cope with those issues, we design a lightweight indoor robot positioning system which operates on cost-effective WiFi-based received signal strength (RSS) and could be readily pluggable into any existing WiFi network infrastructures. Moreover, a novel deep fuzzy forest is proposed to inherit the merits of decision trees and deep neural networks within an end-to-end trainable architecture. Real-world indoor localization experiments are conducted and results demonstrate the superiority of the proposed method over the existing approaches. Le Zhang 0001, Zhenghua Chen, Wei Cui 0002, Bing Li 0002, Cen Chen 0002, Zhiguang Cao, Kai-Zhou Gao |
IEEE Internet Things J. | 4 |
| 2019 | An Efficient Method for High Quality and Cohesive Topical Phrase MiningabstractA phrase is a natural, meaningful, and essential semantic unit. In topic modeling, visualizing phrases for individual topics is an effective way to explore and understand unstructured text corpora. However, from phrase quality and topical cohesion perspectives, the outcomes of existing approaches remain to be improved. Usually, the process of topical phrase mining is twofold: phrase mining and topic modeling. For phrase mining, existing approaches often suffer from order sensitive and inappropriate segmentation problems, which make them often extract inferior quality phrases. For topic modeling, traditional topic models do not fully consider the constraints induced by phrases, which may weaken the cohesion. Moreover, existing approaches often suffer from losing domain terminologies since they neglect the impact of domain-level topical distribution. In this paper, we propose an efficient method for high quality and cohesive topical phrase mining. A high quality phrase should satisfy frequency, phraseness, completeness, and appropriateness criteria. In our framework, we integrate quality guaranteed phrase mining method, a novel topic model incorporating the constraint of phrases, and a novel document clustering method into an iterative framework to improve both phrase quality and topical cohesion. We also describe efficient algorithmic designs to execute these methods efficiently. The empirical verification demonstrates that our method outperforms the state-of-the-art methods from the aspects of both interpretability and efficiency. Bing Li 0002, Xiaochun Yang 0001, Rui Zhou 0001, Bin Wang 0015, Chengfei Liu, Yanchun Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | An Adaptive Hierarchical Compositional Model for Phrase EmbeddingabstractPhrase embedding aims at representing phrases in a vector space and it is important for the performance of many NLP tasks. Existing models only regard a phrase as either full-compositional or non-compositional, while ignoring the hybrid-compositionality that widely exists, especially in long phrases. This drawback prevents them from having a deeper insight into the semantic structure for long phrases and as a consequence, weakens the accuracy of the embeddings. In this paper, we present a novel method for jointly learning compositionality and phrase embedding by adaptively weighting different compositions using an implicit hierarchical structure. Our model has the ability of adaptively adjusting among different compositions without entailing too much model complexity and time cost. To the best of our knowledge, our work is the first effort that considers hybrid-compositionality in phrase embedding. The experimental evaluation demonstrates that our model outperforms state-of-the-art methods in both similarity tasks and analogy tasks. Bing Li 0002, Xiaochun Yang 0001, Bin Wang 0015, Wei Wang 0011, Wei Cui 0002, Xianchao Zhang 0001 |
IJCAI | 1 |
| 2018 | Received Signal Strength Based Indoor Positioning Using a Random Vector Functional Link NetworkabstractFingerprinting based indoor positioning system is gaining more research interest under the umbrella of location-based services. However, existing works have certain limitations in addressing issues such as noisy measurements, high computational complexity, and poor generalization ability. In this work, a random vector functional link network based approach is introduced to address these issues. In the proposed system, a subset of informative features from many randomized noisy features is selected to both reduce the computational complexity and boost the generalization ability. Moreover, the feature selector and predictor are jointly learned iteratively in a single framework based on an augmented Lagrangian method. The proposed system is appealing as it can be naturally fit into parallel or distributed computing environment. Extensive real-world indoor localization experiments are conducted on users with smartphone devices and results demonstrate the superiority of the proposed method over the existing approaches. Wei Cui 0002, Le Zhang 0001, Bing Li 0002, Jing Guo 0007, Wei Meng 0002, Haixia Wang 0003, Lihua Xie 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Efficiently Mining High Quality Phrases from TextsabstractPhrase mining is a key research problem for semantic analysis and text-based information retrieval. The existing approaches based on NLP, frequency, and statistics cannot extract high quality phrases and the processing is also time consuming, which are not suitable for dynamic on-line applications. In this paper, we propose an efficient high-quality phrase mining approach (EQPM). To the best of our knowledge, our work is the first effort that considers both intra-cohesion and inter-isolation in mining phrases, which is able to guarantee appropriateness. We also propose a strategy to eliminate order sensitiveness, and ensure the completeness of phrases. We further design efficient algorithms to make the proposed model and strategy feasible. The empirical evaluations on four real data sets demonstrate that our approach achieved a considerable quality improvement and the processing time was 2.3X - 29X faster than the state-of-the-art works. Bing Li 0002, Xiaochun Yang 0001, Bin Wang 0015, Wei Cui 0002 |
AAAI | 1 |
| 2016 | CITPM: A Cluster-Based Iterative Topical Phrase Mining Framework
Bing Li 0002, Bin Wang 0015, Rui Zhou 0001, Xiaochun Yang 0001, Chengfei Liu |
DASFAA (1) | 1 |