Qin Lv

dblp:11/808 · also Christine Lv · DBLP profile ↗
← Back
87ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-9437-1376ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 26 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 20 · 4 since 2021Computer networks · 14 · 1 since 2021Systems, architecture and hardware · 11 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2026 Common column identification for table similarity detection in electrified transportation data lakes
abstract
Electrified transportation often requires researchers and operators to interact with datasets from a wide range of sources and disciplines, such as transportation, power systems, public health, policies, and regulations. These datasets vary in quality and format, making it difficult to understand, preprocess, and identify key columns representing real-world entities or values for indexing and joining, which can negatively impact downstream analysis and operation. Existing solutions are limited, requiring extensive manual customization or data expertise to utilize. In this article, we propose a multi-layered approach to automatically identify key columns to expedite preprocessing and aid in analysis of electrified transportation data. Our method leverages a dynamic ontology to identify common fields and an information theory-based strategy for edge cases that are difficult to generalize. Evaluations on a number of datasets from data.gov and kaggle.com show improved performance of our methods over several baseline techniques, and our ablation analyses illustrate the efficacy of individual components of our method. Our case studies also demonstrate that our methods have the potential to improve analysis of electrified transportation data and aid in automatic integration of such datasets.
Spencer Paulissen, Matthew Bruchon, Jason Lustbader, Qin Lv
Expert Syst. Appl.4
2025 CommTox: Contextually-Aware Community Perceived Toxicity Classification
Ayan Chowdhury, Rhett Hanscom, Tamara Silbergleit Lehman, Qin Lv, Shivakant Mishra
ASONAM (2)4
2025 Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference
abstract
Mixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures cannot effectively fulfill this advantage and will yield intolerable inference latencies of LLMs on memory-constrained devices. Our investigation pinpoints the essential cause as the remarkable temporal inconsistencies of inter-token expert activations, which generate overly frequent expert swapping demands dominating the latencies. To this end, we propose a novel MoE architecture, Oracle-MoE, to fulfill the real on-device potential of MoE-based LLMs. Oracle-MoE route tokens in a highly compact space suggested by attention scores, termed the oracle space, to effectively maintain the semantic locality across consecutive tokens to reduce expert activation variations, eliminating massive swapping demands. Theoretical analysis proves that Oracle-MoE is bound to provide routing decisions with better semantic locality and, therefore, better expert activation consistencies. Experiments on the pretrained GPT-2 architectures of different sizes (200M, 350M, 790M, and 2B) and downstream tasks demonstrate that without compromising task performance, our Oracle-MoE has achieved state-of-the-art inference speeds across varying memory budgets, revealing its substantial potential for LLM deployments in industry.
Jixian Zhou, Ruijun Huang, Hengjie Cao, Mengyi Chen, Anrui Chen, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, David A. Clifton, Qin Lv, Rui Zhu 0006, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002
ICML12
2025 WiLife: Long-Term Daily Status Monitoring and Habit Mining of the Elderly Leveraging Ubiquitous Wi-Fi Signals
abstract
The global aging demographic underscores the imperative for continuous in-home monitoring of the empty-nest elderly, ensuring their safety and well-being. The widespread deployment of Wi-Fi infrastructure has paved the way to monitor the elderly in a non-intrusive and privacy-preserving manner. Numerous studies have explored the potential of utilizing Wi-Fi signals to address urgent life safety concerns such as fall detection and vital sign monitoring. However, apart from these acute safety issues, the early detection of potential disease symptoms and managing the progression of chronic diseases are also crucial for elderly care, which calls for long-term and continuous monitoring of the elderly’s daily routines. Unfortunately, challenges like continuous activity segmentation and location/orientation dependencies have hindered the implementation of a long-term, around-the-clock activity monitoring system for the elderly. This work introduces “WiLife,” a cutting-edge Wi-Fi-based framework for continuous monitoring of the elderly’s spatio-temporal daily status information. Specifically, WiLife adopts a strategy of partitioning living spaces into functional areas and categorizing daily activities into atomic states . By encapsulating daily life status into a unique series of triple unit format: \(\left\langle\textit{Time, Area, State}\right\rangle\) , WiLife is able to offer valuable insights into when, where, and how activities occur. Field implementations spanning 1,080 hours (45 days \(\times\) 24 hours) in real-world home environments highlight WiLife’s exceptional capability in understanding individual living habits and timely detection of irregularities.
Shengjie Li 0001, Zhaopeng Liu, Qin Lv, Yanyan Zou 0003, Daqing Zhang 0001
ACM Trans. Comput. Heal.3
2025 Denoising Reuse: Exploiting Inter-Frame Motion Consistency for Efficient Video Generation
abstract
Denoising-based diffusion models have attained impressive image synthesis; however, their applications on videos can lead to unaffordable computational costs due to the per-frame denoising operations. In pursuit of efficient video generation, we present a Diffusion Reuse MOtion (Dr. Mo) network to accelerate the video-based denoising process. Our crucial observation is that the latent representations in early denoising steps between adjacent video frames exhibit high consistencies with motion clues. Inspired by the discovery, we propose to accelerate the video denoising process by incorporating lightweight, learnable motion features. Specifically, Dr. Mo will only compute all denoising steps for base frames. For a non-based frame, Dr. Mo will propagate the pre-computed based latents of a particular step with inter-frame motions to obtain a fast estimation of its coarse-grained latent representation, from which the denoising will continue to obtain more sensitive and fine-grained representations. On top of this, Dr. Mo employs a meta-network named Denoising Step Selector (DSS) to dynamically determine the step to perform motion-based propagations for each frame, ensuring the correct transformation of multi-granularity visual features. Extensive evaluations on video generation and editing tasks indicate that Dr. Mo delivers widely applicable acceleration for diffusion-based video generations while effectively retaining the visual quality and style. Video generation and visualization results can be found athttps://drmo-denoising-reuse.github.io.
Yixuan Chen 0003, Yujiang Wang 0001, Mingzhi Dong, Dongsheng Li 0002, Rui Zhu 0006, David A. Clifton, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002
IEEE Trans. Circuits Syst. Video Technol.12
2024 Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain Knowledge
abstract
Language models (LMs) only pretrained on a general and massive corpus usually cannot attain satisfying performance on domain-specific downstream tasks, and hence, applying domain-specific pretraining to LMs is a common and indispensable practice. However, domain-specific pretraining can be costly and time-consuming, hindering LMs' deployment in real-world applications. In this work, we consider the incapability to memorize domain-specific knowledge embedded in the general corpus with rare occurrences and long-tail distributions as the leading cause for pretrained LMs' inferior downstream performance. Analysis of Neural Tangent Kernels (NTKs) reveals that those long-tail data are commonly overlooked in the model's gradient updates and, consequently, are not effectively memorized, leading to poor domain-specific downstream performance. Based on the intuition that data with similar semantic meaning are closer in the embedding space, we devise a Cluster-guided Sparse Expert (CSE) layer to actively learn long-tail domain knowledge typically neglected in previous pretrained LMs. During pretraining, a CSE layer efficiently clusters domain knowledge together and assigns long-tail knowledge to designate extra experts. CSE is also a lightweight structure that only needs to be incorporated in several deep layers. With our training strategy, we found that during pretraining, data of long-tail knowledge gradually formulate isolated, outlier clusters in an LM's representation spaces, especially in deeper layers. Our experimental results show that only pretraining CSE-based LMs is enough to achieve superior performance than regularly pretrained-finetuned LMs on various downstream tasks, implying the prospects of domain-specific-pretraining-free language models.
Mengyi Chen, Jixian Zhou, Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, Rui Zhu 0006, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002
NeurIPS12
2023 Over-parameterized Model Optimization with Polyak-Łojasiewicz Condition
Yixuan Chen 0003, Yubin Shi, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Ning Gu 0001, Li Shang 0002
ICLR8
2023 Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
abstract
Despite their prevalence in deep-learning communities, over-parameterized models convey high demands of computational costs for proper training. This work studies the fine-grained, modular-level learning dynamics of over-parameterized models to attain a more efficient and fruitful training strategy. Empirical evidence reveals that when scaling down into network modules, such as heads in self-attention models, we can observe varying learning patterns implicitly associated with each module's trainability. To describe such modular-level learning capabilities, we introduce a novel concept dubbed modular neural tangent kernel (mNTK), and we demonstrate that the quality of a module's learning is tightly associated with its mNTK's principal eigenvalue $\lambda_{\max}$. A large $\lambda_{\max}$ indicates that the module learns features with better convergence, while those miniature ones may impact generalization negatively. Inspired by the discovery, we propose a novel training strategy termed Modular Adaptive Training (MAT) to update those modules with their $\lambda_{\max}$ exceeding a dynamic threshold selectively, concentrating the model on learning common features and ignoring those inconsistent ones. Unlike most existing training schemes with a complete BP cycle across all network modules, MAT can significantly save computations by its partially-updating strategy and can further improve performance. Experiments show that MAT nearly halves the computational cost of model training and outperforms the accuracy of baselines.
Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002
NeurIPS8
2022 Recursive Disentanglement Network
Yixuan Chen 0003, Yubin Shi, Dongsheng Li 0002, Yujiang Wang 0001, Mingzhi Dong, Robert P. Dick, Qin Lv, Fan Yang 0001, Li Shang 0002
ICLR8
2022 Cross-modal Ambiguity Learning for Multimodal Fake News Detection
abstract
Cross-modal learning is essential to enable accurate fake news detection due to the fast-growing multimodal contents in online social communities. A fundamental challenge of multimodal fake news detection lies in the inherent ambiguity across different content modalities, i.e., decisions made from unimodalities may disagree with each other, which may lead to inferior multimodal fake news detection. To address this issue, we formulate the cross-modal ambiguity learning problem from an information-theoretic perspective and propose CAFE — an ambiguity-aware multimodal fake news detection method. CAFE consists of 1) a cross-modal alignment module to transform the heterogeneous unimodality features into a shared semantic space, 2) a cross-modal ambiguity learning module to estimate the ambiguity between different modalities, and 3) a cross-modal fusion module to capture the cross-modal correlations. CAFE improves fake news detection accuracy by judiciously and adaptively aggregating unimodal features and cross-modal correlations, i.e., relying on unimodal features when cross-modal ambiguity is weak and referring to cross-modal correlations when cross-modal ambiguity is strong. Experimental studies on two widely used datasets (Twitter and Weibo) demonstrate that CAFE outperforms state-of-the-art fake news detection methods by 2.2-18.9% and 1.7-11.4% on accuracy, respectively.
Yixuan Chen 0003, Dongsheng Li 0002, Peng Zhang 0060, Jie Sui, Qin Lv, Tun Lu
WWW5
2022 Understanding WiFi Signal Frequency Features for Position-Independent Gesture Sensing
abstract
Recent years have witnessed rapid development in the research area of WiFi sensing, which senses human activities in a contactless and non-intrusive manner. One major issue that hinders real-world deployment of these systems is position dependence, i.e., once the human target changes location and orientation, the sensing performance degrades significantly. Existing machine learning based methods aim to solve this problem by either generating high-dimensional features or transfer learning the environment knowledge. However, these methods require significant training effort and yet acquire limited improvement. In this paper, we start by understanding and analyzing the Doppler frequency shift in WiFi sensing. We then develop a WiFi frequency model to quantify the relationship between signal frequency and target position, motion direction and speed for human activities. Based on this theoretical model, we prove that the commonly-used movement speed and motion direction features are position dependent, and further identify movement fragments and relative motion direction changes as two position-independent features. Building upon the frequency model and the position-independent features, we design a suite of position-independent gestures and develop the gesture recognition system accordingly. Evaluation results show that under various conditions (i.e., different locations, orientations, environments, and persons), our system achieves more than 96 percent recognition accuracy without any training, significantly outperforming state-of-the-art machine learning based solutions.
Kai Niu 0003, Fusang Zhang, Xuanzhi Wang, Qin Lv, Haitong Luo, Daqing Zhang 0001
IEEE Trans. Mob. Comput.4
2022 A Reinforcement-Learning-Based Energy-Efficient Framework for Multi-Task Video Analytics Pipeline
abstract
Deep-learning-based video processing has yielded transformative results in recent years. However, the video analytics pipeline is energy-intensive due to high data rates and reliance on complex inference algorithms, which limits its adoption in energy-constrained applications. Motivated by the observation of high and variable spatial redundancy and temporal dynamics in video data streams, we design and evaluate an adaptive-resolution optimization framework to minimize the energy use of multi-task video analytics pipelines. Instead of heuristically tuning the input data resolution of individual tasks, our framework utilizes deep reinforcement learning to dynamically govern the input resolution and computation of the entire video analytics pipeline. By monitoring the impact of varying resolution on the quality of high-dimensional video analytics features, hence the accuracy of video analytics results, the proposed end-to-end optimization framework learns the best non-myopic policy for dynamically controlling the resolution of input video streams to globally optimize energy efficiency. Governed by reinforcement learning, optical flow is incorporated into the framework to minimize unnecessary spatio-temporal redundancy that leads to re-computation, while preserving accuracy. The proposed framework is applied to video instance segmentation which is one of the most challenging computer vision tasks, and achieves better energy efficiency than all baseline methods of similar accuracy on the YouTube-VIS dataset.
Mingzhi Dong, Yujiang Wang 0001, Da Feng, Qin Lv, Robert P. Dick, Dongsheng Li 0002, Tun Lu, Ning Gu 0001
IEEE Trans. Multim.5
2021 Analyzing behavioral changes of Twitter users after exposure to misinformation
abstract
Social media platforms have been exploited to disseminate misinformation in recent years. The widespread online misinformation has been shown to affect users' beliefs and is connected to social impact such as polarization. In this work, we focus on misinformation's impact on specific user behavior and aim to understand whether general Twitter users changed their behavior after being exposed to misinformation. We compare the before and after behavior of exposed users to determine whether the frequency of the tweets they posted, or the sentiment of their tweets underwent any significant change. Our results indicate that users overall exhibited statistically significant changes in behavior across some of these metrics. Through language distance analysis, we show that exposed users were already different from baseline users before the exposure. We also study the characteristics of two specific user groups, multi-exposure and extreme change groups, which were potentially highly impacted. Finally, we study if the changes in the behavior of the users after exposure to misinformation tweets vary based on the number of their followers or the number of followers of the tweet authors, and find that their behavioral changes are all similar.
Yichen Wang 0008, Richard Han 0001, Tamara Silbergleit Lehman, Qin Lv, Shivakant Mishra
ASONAM4
2021 Air Pollution Hotspot Detection and Source Feature Analysis using Cross-Domain Urban Data
abstract
Air pollution is a major global environmental health threat, in particular for people who live or work near air pollution sources. Areas adjacent to pollution sources often have high ambient pollution concentrations, and those areas are commonly referred to as air pollution hotspots. In this work, we explore the use of mobile sensing data to detect pollution hotspots. We propose a two-step approach to detect hotspots from unevenly sampled mobile sensing data. To contextualize the detected hotspots and discover potential pollution source characteristics, we explore a variety of cross-domain urban data and extract features from them for hotspot inference. Evaluation results using real-world mobile sensing air quality data as well as cross-domain urban data demonstrate the effectiveness of our approach in detecting and inferring pollution hotspots.
Michael Hannigan, Qin Lv
SIGSPATIAL/GIS3
2021 Understanding the Diverging User Trajectories in Highly-related Online Communities during the COVID-19 Pandemic
Jason Shuo Zhang, Brian Keegan, Qin Lv, Chenhao Tan
ICWSM3
2021 Analyzing Twitter Users' Behavior Before and After Contact by the Russia's Internet Research Agency
abstract
Social media platforms have been exploited to conduct election interference in recent years. In particular, the Russian-backed Internet Research Agency (IRA) has been identified as a key source of misinformation spread on Twitter prior to the 2016 U.S. presidential election. The goal of this research is to understand whether general Twitter users changed their behavior in the year following first contact from an IRA account. We compare the before and after behavior of contacted users to determine whether there were differences in their mean tweet count, the sentiment of their tweets, and the frequency and sentiment of tweets mentioning @realDonaldTrump or @HillaryClinton. Our results indicate that users overall exhibited statistically significant changes in behavior across most of these metrics, and that those users that engaged with the IRA generally showed greater changes in behavior.
Upasana Dutta, Rhett Hanscom, Jason Shuo Zhang, Richard Han 0001, Tamara Silbergleit Lehman, Qin Lv, Shivakant Mishra
Proc. ACM Hum. Comput. Interact.6
2020 Understanding How Readers Determine the Legitimacy of Online News Articles in the Era of Fake News
abstract
Internet users are routinely exposed to fake news in their social media feeds. The main goal of this paper is to identify the factors readers consider important in discriminating against fake news from true news when reading an online news article. We design and conduct three surveys using Amazon Mechanical Turk to identify the top factors and rate them under diverse scenarios. Our results suggest that people perceive news Source and Content to be the most important factors, in general, to distinguish fake news from true news, however, their importance reduces in practice when people actually read a news article. Furthermore, the importance of different factors in the credibility determination of a news article varies with people's political leanings. Our work is the first of its kind and offers new insights into how people determine the legitimacy of online news articles.
Srihaasa Pidikiti, Jason Shuo Zhang, Richard Han 0001, Tamara Silbergleit Lehman, Qin Lv, Shivakant Mishra
ASONAM5
2020 HyTasker: Hybrid Task Allocation in Mobile Crowd Sensing
abstract
Task allocation is a major challenge in Mobile Crowd Sensing (MCS). While previous task allocation approaches follow either the opportunistic or participatory mode, this paper proposes to integrate these two complementary modes in a two-phased hybrid framework called HyTasker. In the offline phase, a group of workers (called opportunistic workers) are selected, and they complete MCS tasks during their daily routines (i.e., opportunistic mode). In the online phase, we assign another set of workers (called participatory workers) and require them to move specifically to perform tasks that are not completed by the opportunistic workers (i.e., participatory mode). Instead of considering these two phases separately, HyTasker jointly optimizes them with a total incentive budget constraint. In particular, when selecting opportunistic workers in the offline phase of HyTasker, we propose a novel algorithm that simultaneously considers the predicted task assignment for the participatory workers, in which the density and mobility of participatory workers are taken into account. Experiments on two real-world mobility datasets demonstrate that HyTasker outperforms other methods with more completed tasks under the same budget constraint.
Jiangtao Wang 0001, Feng Wang 0040, Yasha Wang, Leye Wang, Zhaopeng Qiu, Daqing Zhang 0001, Bin Guo 0001, Qin Lv
IEEE Trans. Mob. Comput.8
2019 TFPN: Twin Feature Pyramid Networks for Object Detection
abstract
FPN (Feature Pyramid Networks) is one of the most popular object detection networks, which can improve small object detection by enhancing shallow features. However, limited attention has been paid to the improvement of large object detection via deeper feature enhancement. One existing approach merges the feature maps of different layers into a new feature map for object detection, but can lead to increased noise and loss of information. The other approach adds a bottom-up structure after the feature pyramid of FPN, which superimposes the information from shallow layers into the deep feature map but weakens the strength of FPN in detecting small objects. To address these challenges, this paper proposes TFPN (Twin Feature Pyramid Networks), which consists of (1) FPN+, a bottom-up structure that improves large object detection; (2) TPS, a Twin Pyramid Structure that improves medium object detection; and (3) innovative integration of these two with FPN, which can significantly improve the detection accuracy of large and medium objects while maintaining the advantage of FPN in small object detection. Extensive experiments using the MSCOCO object detection datasets and the BDD100K automatic driving dataset demonstrate that TFPN significantly improves over existing models, achieving up to 2.2 improvement in detection accuracy (e.g., 36.3 for FPN vs. 38.5 for TFPN on COCO Val-17). Our method can obtain the same accuracy as FPN with ResNet-101 based on ResNet-50 and needs fewer parameters.
Fangzhao Li, Yuxing Peng 0001, Qin Lv, Yuan Yuan 0034, Zhen Huang 0006
ICTAI5
2019 Multi-Group Encoder-Decoder Networks to Fuse Heterogeneous Data for Next-Day Air Quality Prediction
abstract
Accurate next-day air quality prediction is essential to enable warning and prevention measures for cities and individuals to cope with potential air pollution, such as vehicle restriction, factory shutdown, and limiting outdoor activities. The problem is challenging because air quality is affected by a diverse set of complex factors. There has been prior work on short-term (e.g., next 6 hours) prediction, however, there is limited research on modeling local weather influences or fusing heterogeneous data for next-day air quality prediction. This paper tackles this problem through three key contributions: (1) we leverage multi-source data, especially high-frequency grid-based weather data, to model air pollutant dynamics at station-level; (2) we add convolution operators on grid weather data to capture the impacts of various weather parameters on air pollutant variations; and (3) we automatically group (cross-domain) features based on their correlations, and propose multi-group Encoder-Decoder networks (MGED-Net) to effectively fuse multiple feature groups for next-day air quality prediction. The experiments with real-world data demonstrate the improved prediction performance of MGED-Net over state-of-the-art solutions (4.2% to 9.6% improvement in MAE and 9.2% to 16.4% improvement in RMSE).
Qin Lv, Duanfeng Gao, Si Shen, Robert P. Dick, Michael Hannigan, Qi Liu 0052
IJCAI2
2019 WiMorse: A Contactless Morse Code Text Input System Using Ambient WiFi Signals
abstract
Recent years have witnessed advances of Internet of Things (IoT) technologies and their applications to enable contactless sensing and human-computer interaction in smart homes. For people with motor neurone disease (MND), their motion capabilities are severely impaired and they have difficulties interacting with IoT devices and even communicating with other people. As the disease progresses, most patients lose their speech function eventually which makes the widely adopted voice-based solutions fail. In contrast, most of the patients can still move their fingers slightly even after they have lost the control of their arms and hands. Thus, we propose to develop a Morse code-based text input system, called WiMorse, which allows patients with minimal single-finger control to input and communicate with other people without attaching any sensor to their fingers. WiMorse leverages ubiquitous commodity WiFi devices to track subtle finger movements contactlessly and encode them as Morse code input. In order to sense the very subtle finger movements, we propose to employ the ratio of the channel state information (CSI) between two antennas to enhance the signal to noise ratio. To address the severe location dependency issue in wireless sensing with accurate theoretical underpinning and experiments, we propose a signal transformation mechanism to automatically convert signals based on the input position, achieving stable sensing performance. Comprehensive experiments demonstrate that WiMorse can achieve higher than 95% recognition accuracy for finger generated Morse code, and is robust against input position, environment changes, and user diversity.
Kai Niu 0003, Fusang Zhang, Jie Xiong 0001, Qin Lv, Youwei Zeng, Daqing Zhang 0001
IEEE Internet Things J.5
2019 Intergroup Contact in the Wild: Characterizing Language Differences between Intergroup and Single-group Members in NBA-related Discussion Forums
abstract
Intergroup contact has long been considered as an effective strategy to reduce prejudice between groups. However, recent studies suggest that exposure to opposing groups in online platforms can exacerbate polarization. To further understand the behavior of individuals who actively engage in intergroup contact in practice, we provide a large-scale observational study of intragroup behavioral differences between members with and without intergroup contact. We leverage the existing structure of NBA-related discussion forums on Reddit to study the context of professional sports. We identify fans of each NBA team as members of a group and trace whether they have intergroup contact. Our results show that members with intergroup contact use more negative and abusive language in their affiliated group than those without such contact, after controlling for activity levels. We further quantify different levels of intergroup contact and show that there may exist nonlinear mechanisms regarding how intergroup contact relates to intragroup behavior. Our findings provide complementary evidence to experimental studies in a novel context and also shed light on possible reasons for the different outcomes in prior studies.
Jason Shuo Zhang, Chenhao Tan, Qin Lv
Proc. ACM Hum. Comput. Interact.3
2019 Understanding Event Organization at Scale in Event-Based Social Networks
abstract
Understanding real-world event participation behavior has been a subject of active research and can offer valuable insights for event-related recommendation and advertisement. The emergence of event-based social networks (EBSNs), which attracts online users to host/attend offline events, has enabled exciting new research in this domain. However, most existing works focus on understanding or predicting individual users’ event participation behavior or recommending events to individual users. Few studies have addressed the problem of event popularity from the event organizer’s point of view. In this work, we study the latent factors for determining event popularity using large-scale datasets collected from the popular Meetup.com EBSN in five major cities around the world. We analyze and model four contextual factors: spatial factor using location convenience, quality, popularity density, and competitiveness; group factor using group member entropy and loyalty; temporal factor using temporal preference and weekly event patterns; and semantic factor using readability, sentiment, part of speech, and text novelty. In addition, we have developed a group-based social influence propagation network to model group-specific influences on events. By combining the COntextual features and Social Influence NEtwork, our integrated prediction framework COSINE can capture the diverse influential factors of event participation and can be used by event organizers to predict/improve the popularity of their events. Detailed evaluations demonstrate that our COSINE framework achieves high accuracy for event popularity prediction in all five cities with diverse cultures and user event behaviors.
Jason Shuo Zhang, Qin Lv
ACM Trans. Intell. Syst. Technol.2
2018 Multi-LCNN: A Hybrid Neural Network Based on Integrated Time-Frequency Characteristics for Acoustic Scene Classification
abstract
Acoustic scene classification (ASC) is an important task in audio signal processing and can be useful in many real-world applications. Recently, several deep neural network models have been proposed for ASC, such as LSTMs based on temporal analysis and CNNs based on frequency spectrum, as well as hybrid models of LSTM and CNN to further improve classification performance. However, existing hybrid models fail to properly preserve the temporal information when transferring data between different models. In this work, we first analyze the cause of such temporal information loss. We then propose Multi-LCNN, a new hybrid model with two important mechanisms: (1) a LCNN architecture to effectively preserve temporal information; and (2) a multi-channel feature fusion mechanism (MCFF) that combines enhanced temporal information and frequency spectrogram information to learn highly integrated and discriminative features for ASC. Evaluations on the TUT ASC 2016 dataset show that our model can achieve an improvement of 10.23% over the baseline method, and is currently the best-performing end-to-end model on this dataset.
Jin Lei, Boqing Zhu, Qin Lv, Zhen Huang 0006, Yuxing Peng 0001
ICTAI4
2018 An Algorithmic Method for Tampering-Proof and Privacy-Preserving Smart Metering
abstract
The fast-growing smart meter installation enables utility providers to better track and manage the electricity use of end customers. However, the adoption of smart meter faces several challenges. In particular, collecting fine-grained electricity usage data raises privacy concerns by end customers; and the possibility of data manipulation by malicious customers threatens the data integrity for utility providers. In this article, a tampering-proof and privacy-preserving smart metering method is proposed. The proposed work consists of a new secure multi-party computation method to protect customer privacy and detect malicious data manipulation during real-time data aggregation of smart meter applications. Theoretical analysis proves that privacy preservation for customers and data tamper-proof for utility providers can be guaranteed. Experiments using real-world data demonstrate that the proposed work significantly improves the computation efficiency over a state-of-the-art homomorphic encryption based privacy-preserving smart metering approach, with minimal communication and storage overhead.
Dongsheng Li 0002, Qin Lv
INDIN4
2018 Deriving Customer Privacy from Randomly Perturbed Smart Metering Data
abstract
Privacy has been one of the major concerns for customers in smart grids. Randomized perturbation based privacy-preserving smart metering methods, which are efficient and easy to implement, have recently become one of the commonly adopted solutions. However, it is a challenging task to meet utility companies’ data collection requirements while protecting customer privacy. This paper analyzes the privacy protection capability of randomized perturbation based privacy-preserving smart metering methods. Both theoretical analysis and empirical studies show that statistical information of individual customers can still be accurately obtained from these randomly perturbed data. Also, an appliance usage inference method is proposed to accurately identify appliance operations of individual customers using randomly perturbed smart metering data. Evaluations using real-world smart metering data demonstrate that the proposed method can identify appliance operations with an accuracy between 92% and 99%.
Dongsheng Li 0002, Qi Liu 0052, Qin Lv
INDIN4
2018 CapBand: Battery-free Successive Capacitance Sensing Wristband for Hand Gesture Recognition
abstract
We present CapBand, a battery-free hand gesture recognition wearable in the form of a wristband. The key challenges in creating such a system are (1) to sense useful hand gestures at ultra-low power so that the device can be powered by the limited energy harvestable from the surrounding environment and (2) to make the system work reliably without requiring training every time a user puts on the wristband. We present successive capacitance sensing, an ultra-low power sensing technique, to capture small skin deformations due to muscle and tendon movements on the user's wrist, which corresponds to specific groups of wrist muscles representing the gestures being performed. We build a wrist muscles-to-gesture model, based on which we develop a hand gesture classification method using both motion and static features. To eliminate the need for per-usage training, we propose a kernel-based on-wrist localization technique to detect the CapBand's position on the user's wrist. We prototype CapBand with a custom-designed capacitance sensor array on two flexible circuits driven by a custom-built electronic board, a heterogeneous material-made, deformable silicone band, and a custom-built energy harvesting and management module. Evaluations on 20 subjects show 95.0% accuracy of gesture recognition when recognizing 15 different hand gestures and 95.3% accuracy of on-wrist localization.
Hoang Truong 0002, Jason Shuo Zhang, Ufuk Muncuk, Phuc Nguyen 0002, Nam Bui, Anh Nguyen 0001, Qin Lv, Kaushik R. Chowdhury, Thang N. Dinh, Tam Vu 0001
SenSys7
2018 AdaError: An Adaptive Learning Rate Method for Matrix Approximation-based Collaborative Filtering
abstract
Gradient-based learning methods such as stochastic gradient descent are widely used in matrix approximation-based collaborative filtering algorithms to train recommendation models based on observed user-item ratings. One major difficulty in existing gradient-based learning methods is determining proper learning rates, since model convergence would be inaccurate or very slow if the learning rate is too large or too small, respectively. This paper proposes AdaError, an adaptive learning rate method for matrix approximation-based collaborative filtering. AdaError eliminates the need of manually tuning the learning rates by adaptively adjusting the learning rates based on the noisiness level of user-item ratings, using smaller learning rates for noisy ratings so as to reduce their impact on the learned models. Our theoretical and empirical analysis shows that AdaError can improve the generalization performance of the learned models. Experimental studies on the MovieLens and Netflix datasets also demonstrate that AdaError outperforms state-of-the-art adaptive learning rate methods in matrix approximation-based collaborative filtering. Furthermore, by applying AdaError to the standard matrix approximation method, we can achieve statistically significant improvements over state-of-the-art collaborative filtering methods in both rating prediction accuracy and top-N recommendation accuracy.
Dongsheng Li 0002, Chao Chen 0016, Qin Lv, Hansu Gu, Tun Lu, Li Shang 0001, Ning Gu 0001, Stephen M. Chu
WWW3
2018 Hybrid EGU-based group event participation prediction in event-based social networks
Jason Shuo Zhang, Qin Lv
Knowl. Based Syst.2
2018 "This is why we play": Characterizing Online Fan Communities of the NBA Teams
abstract
Professional sports constitute an important part of people's modern life. People spend substantial amounts of time and money supporting their favorite players and teams, and sometimes even riot after games. However, how team performance affects fan behavior remains understudied at a large scale. As almost every notable professional team has its own online fan community, these communities provide great opportunities for investigating this research question. In this work, we provide the first large-scale characterization of online fan communities of professional sports teams. Since user behavior in these online fan communities is inherently connected to game events and team performance, we construct a unique dataset that combines 1.5M posts and 43M comments in NBA-related communities on Reddit with statistics that document team performance in the NBA. We analyze the impact of team performance on fan behavior both at the game level and the season level. First, we study how team performance in a game relates to user activity during that game. We find that surprise plays an important role: the fans of the top teams are more active when their teams lose and so are the fans of the bottom teams in an unexpected win. Second, we study fan behavior over consecutive seasons and show that strong team performance is associated with fans of low loyalty, likely due to "bandwagon fans." Fans of the bottom teams tend to discuss their team's future such as young talents in the roster, which may help them stay optimistic during adversity. Our results not only contribute to understanding the interplay between online sports communities and offline context but also provide significant insights into sports management.
Jason Shuo Zhang, Chenhao Tan, Qin Lv
Proc. ACM Hum. Comput. Interact.3
2018 Multi-Task Allocation in Mobile Crowd Sensing with Individual Task Quality Assurance
abstract
Task allocation is a fundamental research issue in mobile crowd sensing. While earlier research focused mainly on single tasks, recent studies have started to investigate multi-task allocation, which considers the interdependency among multiple tasks. A common drawback shared by existing multi-task allocation approaches is that, although the overall utility of multiple tasks is optimized, the sensing quality of individual tasks may become poor as the number of tasks increases. To overcome this drawback, we re-define the multi-task allocation problem by introducing task-specific minimal sensing quality thresholds, with the objective of assigning an appropriate set of tasks to each worker such that the overall system utility is maximized. Our new problem also takes into account the maximum number of tasks allowed for each worker and the sensor availability of each mobile device. To solve this newly-defined problem, this paper proposes a novel multi-task allocation framework named MTasker. Different from previous approaches which start with an empty set and iteratively select task-worker pairs, MTasker adopts a descent greedy approach, where a quasi-optimal allocation plan is evolved by removing a set of task-worker pairs from the full set. Extensive evaluations based on real-world mobility traces show that MTasker outperforms the baseline methods under various settings, and our theoretical analysis proves that MTasker has a good approximation bound.
Jiangtao Wang 0001, Yasha Wang, Daqing Zhang 0001, Feng Wang 0040, Haoyi Xiong, Chao Chen 0004, Qin Lv, Zhaopeng Qiu
IEEE Trans. Mob. Comput.7
2018 Identifying On-Site Users for Social Events: Mobility, Content, and Social Relationship
abstract
The wide spread use of social network services, especially location based services, has transformed social networks into an important information source of real-world events. Many event detection systems using geo-tagged posts from social networks have been developed in recent years. Besides detecting real-world events, it is also desirable for government officials, news media, and police, etc., to identify on-site users of an event, from whom we could gather valuable information regarding the process of events and investigate suspects when an event is associated with crime or terrorist. However, due to the high uncertainty of human mobility patterns and the low probability of users sharing their location information, it is difficult to identify on-site users while a social event unfolds, and research work in this area is still in its infancy. In this paper, we propose a Fused fEature Gaussian prOcess Regression (FEGOR) model, which exploits three influential factors in social networks for on-site user identification: mobility influence, content similarity, and social relationship. By integrating these factors, we are able to estimate the distance between a user and a social event even when the user's location profile is unknown, thus identify on-site users. Experiments on a real-world Twitter dataset demonstrate the effectiveness of our model, achieving a minimum mean absolute error of 1.7km and outperforming state-of-the-art methods.
Zhiwen Yu 0001, Fei Yi, Qin Lv, Bin Guo 0001
IEEE Trans. Mob. Comput.3
2017 GLOMA: Embedding Global Information in Local Matrix Approximation Models for Collaborative Filtering
abstract
Recommender systems have achieved great success in recent years, and matrix approximation (MA) is one of the most popular techniques for collaborative filtering (CF) based recommendation. However, a major issue is that MA methods perform poorly at detecting strong localized associations among closely related users and items. Recently, some MA-based CF methods adopt clustering methods to discover meaningful user-item subgroups and perform ensemble on different clusterings to improve the recommendation accuracy. However, ensemble learning suffers from lower efficiency due to the increased overall computation overhead. In this paper, we propose GLOMA, a new clustering-based matrix approximation method, which can embed global information in local matrix approximation models to improve recommendation accuracy. In GLOMA, a MA model is first trained on the entire data to capture global information. The global MA model is then utilized to guide the training of cluster-based local MA models, such that the local models can detect strong localized associations shared within clusters and at the same time preserve global associations shared among all users/items. Evaluation results using MovieLens and Netflix datasets demonstrate that, by integrating global information in local models, GLOMA can outperform five state-of-the-art MA-based CF methods in recommendation accuracy while achieving descent efficiency.
Chao Chen 0016, Dongsheng Li 0002, Qin Lv, Junchi Yan, Li Shang 0001, Stephen M. Chu
AAAI3
2017 ERMMA: Expected Risk Minimization for Matrix Approximation-based Recommender Systems
abstract
Matrix approximation (MA) is one of the most popular techniques in today's recommender systems. In most MA-based recommender systems, the problem of risk minimization should be defined, and how to achieve minimum expected risk in model learning is one of the most critical problems to recommendation accuracy. This paper addresses the expected risk minimization problem, in which expected risk can be bounded by the sum of optimization error and generalization error. Based on the uniform stability theory, we propose an expected risk minimized matrix approximation method (ERMMA), which is designed to achieve better tradeoff between optimization error and generalization error in order to reduce the expected risk of the learned MA models. Theoretical analysis shows that ERMMA can achieve lower expected risk bound than existing MA methods. Experimental results on the MovieLens and Netflix datasets demonstrate that ERMMA outperforms six state-of-the-art MA-based recommendation methods in both rating prediction problem and item ranking problem.
Dongsheng Li 0002, Chao Chen 0016, Qin Lv, Li Shang 0001, Stephen M. Chu, Hongyuan Zha
AAAI3
2017 SDF-NN: A Deep Neural Network with Semantic Dropping and Fusion for Natural Language Inference
abstract
Natural language inference (NLI) is an important task in natural language processing (NLP), and recently, several deep neural network based models have been proposed for NLI. In this work, we first make two important observations regarding NLI: (1) the existence of extra/interfering semantics and its negative impact on the correctness of final inference; and (2) the unbalanced importance of local inference results and the need to combine all local results for aggregation. Motivated by these two observations, we have designed SDF-NN, a new NLI model with two novel components: (1) a Semantic Dropping Network (SDN) to automatically discard some of the interfering semantics; and (2) a Semantic Fusion Alignment (SFA) method to effectively fuse all local inference results. Our model has achieved 88.2% accuracy on the SNLI corpus, which is currently the best performing single model.
Ludan Tan, Qin Lv, Yuxing Peng 0001, Xiang Zhao 0002, Zhen Huang 0006
ICTAI3
2017 Event Organization 101: Understanding Latent Factors of Event Popularity
Jason Shuo Zhang, Qin Lv
ICWSM2
2017 Data quality screening for high-resolution satellite imagery via spectral clustering
abstract
High-resolution satellite imagery data have been widely used in geoscience and remote sensing research. Dealing with data quality issue is the first and most important step before truly making use of these high-resolution images. Scientific results derived from poor-quality data can be problematic and unreliable. In this work, we propose a novel data quality screening method to discover and filter anomalous images contaminated by systematic errors in a dataset. In particular, cumulative distribution function based pairwise similarity matrix and spectral clustering are adapted to accurately identify clusters of normal and anomalous images. Using the proposed method, we have discovered abnormal images in a collection of high-resolution satellite imagery over the Arctic sea ice, which has been used as ground truth in previous melt pond studies.
Qi Liu 0052, Qin Lv
IGARSS3
2017 Hierarchical context-aware anomaly diagnosis in large-scale PV systems using SCADA data
abstract
Accurate anomaly diagnosis is essential for reducing operation and maintenance (O&M) cost, while improving safety and reliability of large-scale photovoltaic (PV) systems. Although many methods have been proposed, they either require extra sensing devices or suffer from high false alarm rates. In this work, we present a cost-effective hierarchical context-aware method for string-level anomaly diagnosis in large-scale PV systems. The proposed approach is based on unsupervised machine learning techniques and requires no additional hardware support beyond widely adopted supervisory control and data acquisition (SCADA) systems. The effectiveness and efficiency of our proposed approach are evaluated with a 40 MW PV system located in East China. The experimental results demonstrate that the proposed approach can support string-level anomaly diagnosis with high accuracy and provide sufficient lead time for daily maintenance.
Qi Liu 0052, Dahai Kang, Qin Lv
INDIN5
2017 Efficient privacy-preserving content recommendation for online social communities
Dongsheng Li 0002, Qin Lv, Ning Gu 0001
Neurocomputing2
2017 Intelligent Probing for Locality Sensitive Hashing: Multi-Probe LSH and Beyond
abstract
The past decade has been marked by the (continued) explosion of diverse data content and the fast development of intelligent data analytics techniques. One problem we identified in the mid-2000s was similarity search of feature-rich data. The challenge here was achieving both high accuracy and high efficiency in high-dimensional spaces. Locality sensitive hashing (LSH), which uses certain random space partitions and hash table lookups to find approximate nearest neighbors, was a promising approach with theoretical guarantees. But LSH alone was insufficient since a large number of hash tables were required to achieve good search quality. Building on an idea of Panigrahy, our multi-probe LSH method introduced the idea of intelligent probing. Given a query object, we strategically probe its neighboring hash buckets (in a query-dependent fashion) by calculating the statistical probabilities of similar objects falling into each bucket. Such intelligent probing can significantly reduce the number of hash tables while achieving high quality. In this paper, we revisit the problem motivation, the challenges, the key design considerations of multi-probe LSH, as well as discuss recent developments in this space and some questions for further research.
Qin Lv, William K. Josephson, Moses Charikar, Kai Li 0001
Proc. VLDB Endow.1
2017 Gazelle: Energy-Efficient Wearable Analysis for Running
abstract
Running is one of the most popular sports with hundreds of millions of participants worldwide. Good running form is the key to fast, efficient, and injury-free running. Existing kinematic analysis technologies, such as high-speed camera systems, are expensive, difficult to operate, and exclusive to sports physiology laboratories and elite athletes. Miniature MEMS-based motion sensors enable portable high-precision kinematic analysis, but suffer from high energy consumption hence short battery lifetime, especially for continued online analysis for running. This paper presents Gazelle, a wearable online analysis system for running that is compact, lightweight, accurate, and highly energy efficient; intended for runners of all levels. To enable long-term maintenance-free mobile analysis for running, Sparse Adaptive Sensing (SAS) is proposed, which selectively identifies the best sampling points to maintain high accuracy while greatly reducing sensing and analysis energy overheads. Experimental results demonstrate 97.7 percent accuracy with 76.9 to 99 percent reduced energy consumption (83.6 percent average reduction under real-world testing)-a one-order-of-magnitude improvement over existing solutions. SAS enables > 200 days of continuous high-precision operation using only a coin-cell battery. Since 2014, Gazelle has been used by over 100 elite and recreational runners during daily training and at top-level races like the Kona Ironman World Championships and New York Marathon.
Qi Liu 0052, James Williamson, Wyatt Mohrman, Qin Lv, Robert P. Dick
IEEE Trans. Mob. Comput.5
2016 Prediction of cyberbullying incidents in a media-based social network
abstract
Cyberbullying is a major problem affecting more than half of all American teens. Prior work has largely focused on detecting cyberbullying after the fact. In this paper, we investigate the prediction of cyberbullying incidents in Instagram, a popular media-based social network. The novelty of this work is building a predictor that can anticipate the occurrence of cyberbullying incidents before they happen. The Instagram media-based social network is well-suited to such prediction since there is an initial posting of an image typically with an associated text caption, followed later by the text comments that form the basis of a specific cyberbullying incident. We extract several important features from the initial posting data for automated cyberbullying prediction, including profanity and linguistic content of the text caption, image content, as well as social graph parameters and temporal content behavior. Evaluations using a real-world Instagram dataset demonstrate that our method achieves high performance in predicting the occurrence of cyberbullying incidents.
Homa Hosseinmardi, Rahat Ibn Rafiq, Richard Han 0001, Qin Lv, Shivakant Mishra
ASONAM4
2016 CyberSafety 2016: The First International Workshop on Computational Methods in CyberSafety
abstract
The theme of cybersafety is an important emerging research topic on the Internet that manifests itself daily as users navigate the Web and networked applications. Examples of cybersafety issues include cyberbullying, cyberthreats, recruiting minors via Internet services for nefarious purposes, using deceptive means to dupe vulnerable populations, exhibiting misbehaving behaviors such as using profanity or flashing in online video chats, and many others. These issues have a direct negative impact on the social, psychological and in some cases physical well-being of the end users. An important characteristic of these issues is that they fall in a grey legal area, where perpetrators may claim freedom of speech or rights to free expression despite causing harm. The main goal of this inaugural workshop on cybersafety is to bring together the researchers and practitioners from academia, industry, government and research labs working in the area of cybersafety to discuss the unique challenges in addressing various cybersafety issues and to share experiences, solutions, tools, and techniques. The focus is on the detection, prevention and mitigation of various cybersafety issues, as well as education and promoting safe practices.
Shivakant Mishra, Qin Lv, Richard Han 0001, Jeremy Blackburn
CIKM2
2016 Low-Rank Matrix Approximation with Stability
abstract
Low-rank matrix approximation has been widely adopted in machine learning applications with sparse data, such as recommender systems. However, the sparsity of the data, incomplete and noisy, introduces challenges to the algorithm stability – small changes in the training data may significantly change the models. As a result, existing low-rank matrix approximation solutions yield low generalization performance, exhibiting high error variance on the training dataset, and minimizing the training error may not guarantee error reduction on the testing dataset. In this paper, we investigate the algorithm stability problem of low-rank matrix approximations. We present a new algorithm design framework, which (1) introduces new optimization objectives to guide stable matrix approximation algorithm design, and (2) solves the optimization problem to obtain stable low-rank approximation solutions with good generalization performance. Experimental results on real-world datasets demonstrate that the proposed work can achieve better prediction accuracy compared with both state-of-the-art low-rank matrix approximation methods and ensemble methods in recommendation task.
Dongsheng Li 0002, Chao Chen 0016, Qin Lv, Junchi Yan, Li Shang 0001, Stephen M. Chu
ICML3
2016 MPMA: Mixture Probabilistic Matrix Approximation for Collaborative Filtering
Chao Chen 0016, Dongsheng Li 0002, Qin Lv, Junchi Yan, Stephen M. Chu, Li Shang 0001
IJCAI3
2016 An algorithm for efficient privacy-preserving item-based collaborative filtering
Dongsheng Li 0002, Chao Chen 0016, Qin Lv, Li Shang 0001, Tun Lu, Ning Gu 0001
Future Gener. Comput. Syst.3
2016 SalesExplorer: Exploring sales opportunities from white-space customers in the enterprise market
Dongsheng Li 0002, Yaoping Ruan, Qin Lv
Knowl. Based Syst.3
2015 Careful what you share in six seconds: Detecting cyberbullying instances in Vine
abstract
As online social networks have grown in popularity, teenage users have become increasingly exposed to the threats of cyberbullying. The primary goal of this research paper is to investigate cyberbullying behaviors in Vine, a mobile based video-sharing online social network, and design novel approaches to automatically detect instances of cyberbullying over Vine media sessions. We first collect a set of Vine video sessions and use CrowdFlower, a crowd-sourced website, to label the media sessions for cyberbullying and cyberaggression. We then perform a detailed analysis of cyberbullying behavior in Vine. Based on the labeled data, we design a classifier to detect instances of cyberbullying and evaluate the performance of that classifier.
Rahat Ibn Rafiq, Homa Hosseinmardi, Richard Han 0001, Qin Lv, Shivakant Mishra, Sabrina Arredondo Mattson
ASONAM4
2015 Poster: Detection of Cyberbullying in a Mobile Social Network: Systems Issues
abstract
Cyberbullying is a major problem affecting more than half of all American teens, and has been attributed to suicidal behavior among teens. Instagram, a media-based mobile social network, is one of the most popular social networks used for cyberbullying. In this paper, we describe the development of classifiers to detect cyberbullying in Instagram. We identify systems issues that need to be considered in the design of a cyberbullying detection system.
Homa Hosseinmardi, Sabrina Arredondo Mattson, Rahat Ibn Rafiq, Richard Han 0001, Qin Lv, Shivakant Mishra
MobiSys5
2015 WEMAREC: Accurate and Scalable Recommendation through Weighted and Ensemble Matrix Approximation
abstract
Matrix approximation is one of the most effective methods for collaborative filtering-based recommender systems. However, the high computation complexity of matrix factorization on large datasets limits its scalability. Prior solutions have adopted co-clustering methods to partition a large matrix into a set of smaller submatrices, which can then be processed in parallel to improve scalability. The drawback is that the recommendation accuracy is lower as the submatrices only contain subsets of the user-item rating information. This paper presents WEMAREC, a weighted and ensemble matrix approximation method for accurate and scalable recommendation. It builds upon the intuition that (sub)matrices containing more frequent samples of certain user/item/rating tend to make more reliable rating predictions for these specific user/item/rating. WEMAREC consists of two important components: (1) a weighting strategy that is computed based on the rating distribution in each submatrix and applied to approximate a single matrix containing those submatrices; and (2) an ensemble strategy that leverages user-specific and item-specific rating distributions to combine the approximation matrices of multiple sets of co-clustering results. Evaluations using real-world datasets demonstrate that WEMAREC outperforms state-of-the-art matrix approximation methods in recommendation accuracy (0.5?11.9% on the MovieLens dataset and 2.2--13.1% on the Netflix dataset) with 3--10X improvement on scalability.
Chao Chen 0016, Dongsheng Li 0002, Qin Lv, Li Shang 0001
SIGIR4
2015 ISC: An Iterative Social Based Classifier for Adult Account Detection on Twitter
abstract
The widespread of adult content on online social networks (e.g., Twitter) is becoming an emerging yet critical problem. An automatic method to identify accounts spreading sexually explicit content (i.e., adult account) is of significant values in protecting children and improving user experiences. Traditional adult content detection techniques are ill-suited for detecting adult accounts on Twitter due to the diversity and dynamics in Twitter content. In this paper, we formulate the adult account detection as a graph based classification problem and demonstrate our detection method on Twitter by using social links between Twitter accounts and entities in tweets. As adult Twitter accounts are mostly connected with normal accounts and post many normal entities, which makes the graph full of noisy links, existing graph based classification techniques cannot work well on such a graph. To address this problem, we propose an iterative social based classifier (ISC), a novel graph based classification technique resistant to the noisy links. Evaluations using large-scale real-world Twitter data show that, by labeling a small number of popular Twitter accounts, ISC can achieve satisfactory performance in adult account detection, significantly outperforming existing techniques.
Hanqiang Cheng, Xinyu Xing 0001, Xue (Steve) Liu, Qin Lv
IEEE Trans. Knowl. Data Eng.4
2015 Supporting Healthy Grocery Shopping via Mobile Augmented Reality
abstract
Augmented reality (AR) applications have recently become popular on modern smartphones. We explore the effectiveness of this mobile AR technology in the context of grocery shopping, in particular as a means to assist shoppers in making healthier decisions as they decide which grocery products to buy. We construct an AR-assisted mobile grocery-shopping application that makes real-time, customized recommendations of healthy products to users and also highlights products to avoid for various types of health concerns, such as allergies to milk or nut products, low-sodium or low-fat diets, and general caloric intake. We have implemented a prototype of this AR-assisted mobile grocery shopping application and evaluated its effectiveness in grocery store aisles. Our application's evaluation with typical grocery shoppers demonstrates that AR overlay tagging of products reduces the search time to find healthy food items, and that coloring the tags helps to improve the user's ability to quickly and easily identify recommended products, as well as products to avoid. We have evaluated our application's functionality by analyzing the data we collected from 15 in-person actual grocery-shopping subjects and 104 online application survey participants.
Junho Ahn, James Williamson, Mike Gartrell, Richard Han 0001, Qin Lv, Shivakant Mishra
ACM Trans. Multim. Comput. Commun. Appl.5
2014 Towards understanding cyberbullying behavior in a semi-anonymous social network
abstract
Cyberbullying has emerged as an important and growing social problem, wherein people use online social networks and mobile phones to bully victims with offensive text, images, audio and video on a 24/7 basis. This paper studies negative user behavior in the Ask.fm social network, a popular new site that has led to many cases of cyberbullying, some leading to suicidal behavior.We examine the occurrence of negative words in Ask.fm's question+answer profiles along with the social network of “likes” of questions+answers. We also examine properties of users with “cutting” behavior in this social network.
Homa Hosseinmardi, Richard Han 0001, Qin Lv, Shivakant Mishra, Amir Ghasemianlangroodi
ASONAM3
2014 Multi-modal fusion for flasher detection in a mobile video chat application
abstract
This paper investigates the development of accurate and efficient classifiers to identify misbehaving users (i.e., “flashers”) in a mobile video chat application. Our analysis is based on video session data collected from a mobile client that we built that connects to a popular random video chat ser
Lei Tian 0004, Rahat Ibn Rafiq, Shaosong Li, David Chu, Richard Han 0001, Qin Lv, Shivakant Mishra
MobiQuitous6
2014 Item-based top-N recommendation resilient to aggregated information revelation
Dongsheng Li 0002, Qin Lv, Ning Gu 0001
Knowl. Based Syst.2
2013 AnchorMF: towards effective event context identification
abstract
Online social networks (OSNs) such as Twitter provide a good platform for event discussions. Recent research [26][25] as shown that event discussions in OSNs are diverse and innovative and encourage public engagement in events. Although much research has been conducted in OSNs to track and detect events, there has been limited research on detecting or understanding the event context. Event context helps to better predict users' participation in events, identify relations among events, and recommend friends who share similar event context.
Hansu Gu, Mike Gartrell, Qin Lv, Dirk Grunwald
CIKM4
2013 A Hybrid Sensor System for Indoor Air Quality Monitoring
abstract
Indoor air quality is important. It influences human productivity and health. Personal pollution exposure can be measured using stationary or mobile sensor networks, but each of these approaches has drawbacks. Stationary sensor network accuracy suffers because it is difficult to place a sensor in every location people might visit. In mobile sensor networks, accuracy and drift resistance are generally sacrificed for the sake of mobility and economy. We propose a hybrid sensor network architecture, which contains both stationary sensors (for accurate readings and calibration) and mobile sensors (for coverage). Our technique uses indoor pollutant concentration prediction models to determine the structure of the hybrid sensor network. In this work, we have (1) developed a predictive model for pollutant concentration that minimizes prediction error; (2) developed algorithms for hybrid sensor network construction; and (3) deployed a sensor network to gather data on the airflow in a building, which are later used to evaluate the prediction model and hybrid sensor network synthesis algorithm. Our modeling technique reduces sensor network error by 40.4% on average relative to a technique that does not explicitly consider the inaccuracies of individual sensors. Our hybrid sensor network synthesis technique improves personal exposure measurement accuracy by 35.8% on average compared with a stationary sensor network architecture.
Xiang Yun, Ricardo Piedrahita, Robert P. Dick, Michael Hannigan, Qin Lv
DCOSS5
2013 Hallway based automatic indoor floorplan construction using room fingerprints
abstract
People spend approximately 70% of their time indoors. Understanding the indoor environments is therefore important for a wide range of emerging mobile personal and social applications. Knowledge of indoor floorplans is often required by these applications. However, indoor floorplans are either unavailable or obtaining them requires slow, tedious, and error-prone manual labor.
Yifei Jiang, Xiang Yun, Qin Lv, Robert P. Dick, Michael Hannigan
UbiComp5
2013 Understanding user behavior at scale in a mobile video chat application
abstract
Online video chat services such as Chatroulette and Omegle randomly match users in video chat sessions and have become increasingly popular, with tens of thousands of users online at anytime during a day. Our interest is in examining user behavior in the growing domain of mobile video, and in particular how users behave in such video chat services as they are extended onto mobile clients. To date, over four thousand people have downloaded and used our Android-based mobile client, which was developed to be compatible with an existing video chat service. The paper provides a first-ever detailed large scale study of mobile user behavior in a random video chat service over a three week period. This study identifies major characteristics such as mobile user session durations, time of use, demographic distribution and the large number of brief sessions that users click through to find good matches. Through content analysis of video and audio, as well as analysis of texting and clicking behavior, we discover key correlations among these characteristics, e.g., normal mobile users are highly correlated with using the front camera and with the presence of a face, whereas misbehaving mobile users have a high negative correlation with the presence of a face.
Lei Tian 0004, Shaosong Li, Junho Ahn, David Chu, Richard Han 0001, Qin Lv, Shivakant Mishra
UbiComp6
2013 Personalized multi-modality image management and search for mobile devices
Changyun Zhu, Qin Lv, Robert P. Dick
Pers. Ubiquitous Comput.3
2013 Large-Scale Energy Storage System Design and Optimization for Emerging Electric-Drive Vehicles
abstract
Energy consumption and the associated environmental impact are a pressing challenge faced by the transportation sector. Emerging electric-drive vehicles have shown promises for substantial reductions in petroleum use and vehicle emissions. Their success, however, has been hindered by the limitations of energy storage technologies. Existing in-vehicle lithium-ion battery systems are bulky, expensive, and unreliable. Energy storage system (ESS) design and optimization is essential for emerging transportation electrification. This paper presents an integrated ESS modeling, design, and optimization framework targeting emerging electric-drive vehicles. A large-scale ESS modeling solution is first presented, which considers major runtime and long-term battery effects, and uses fast frequency-domain analysis techniques for efficient and accurate characterization of large-scale ESS. The proposed design framework unifies design-time optimization and runtime control. This conducts statistical optimization for ESS cost and lifetime, which jointly considers the variances of ESS due to manufacture tolerance and heterogeneous driver-specific runtime usage. This optimizes ESS design by incorporating complementary energy storage technologies, e.g., lithium-ion batteries and ultracapacitors. Using physical measurements of battery manufacture variation and real-world user driving profiles, our experimental study has demonstrated that the proposed framework effectively explores the statistical design space and produces cost-efficient ESS solutions with statistical system lifetime guarantees.
Jia Wang 0003, Hai Zhou 0001, Qin Lv, Yihe Sun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2013 SafeVchat: A System for Obscene Content Detection in Online Video Chat Services
abstract
Online video chat services such as Chatroulette, Omegle, and vChatter that randomly match pairs of users in video chat sessions are quickly becoming very popular, with over a million users per month in the case of Chatroulette. A key problem encountered in such systems is the presence of flashers and obscene content. This problem is especially acute given the presence of underage minors in such systems. This article presents SafeVchat, a novel solution to the problem of flasher detection that employs an array of image detection algorithms. A key contribution of the article concerns how the results of the individual detectors are fused together into an overall decision classifying a user as misbehaving or not, based on Dempster-Shafer theory. The article introduces a novel, motion-based skin detection method that achieves significantly higher recall and better precision. The proposed methods have been evaluated over real-world data and image traces obtained from Chatroulette.com. SafeVchat has been deployed in Chatroulette. A combination of SafeVchat with human moderation has resulted in banning as many as 50,000 inappropriate users per day on Chatoulette. Furthermore, offensive content on Chatoulette has dropped significantly from 33.08% (before SafeVchat installation) to 3.49% (after SafeVchat installation).
Yu-Li Liang, Xinyu Xing 0001, Hanqiang Cheng, Jianxun Dang, Sui Huang, Richard Han 0001, Xue (Steve) Liu, Qin Lv, Shivakant Mishra
ACM Trans. Internet Techn.8
2012 ARIEL: automatic wi-fi based room fingerprinting for indoor localization
abstract
People spend the majority of their time indoors, and human indoor activities are strongly correlated with the rooms they are in. Room localization, which identifies the room a person or mobile phone is in, provides a powerful tool for characterizing human indoor activities and helping address challenges in public health, productivity, building management, etc. Existing room localization methods, however, require labor-intensive manual annotation of individual rooms.
Yifei Jiang, Qin Lv, Robert P. Dick, Michael Hannigan
UbiComp4
2012 Attribute Based Content Sharing in Mobile Adhoc Networks of Smartphones over WiFi
abstract
Mobile phones are increasingly equipped with features that allow them to generate, store, and make use of a large amount of digital content. Such content can be associated with attributes obtained via context and/or user annotation. With the increased trend of information sharing and the success of peer-to-peer (P2P) in the wired world, the ability for users to search and share content on their mobile devices is engaging. However, infrastructure-based networks may not always be available or affordable, and attribute- based content sharing in mobile adhoc networks is desirable. This paper investigates the feasibility of such solutions and proposes a system design that incorporates efficient content discovery, robust content distribution, and adaptive energy management techniques. Detailed performance evaluation of a prototype system demonstrates that attribute-based content sharing can be effectively supported in mobile adhoc networks of smartphones over WiFi.
Thomas Georges Cyrille Kooh, Qin Lv, Shivakant Mishra
ICCCN2
2012 Collaborative calibration and sensor placement for mobile sensor networks
abstract
Mobile sensing systems carried by individuals or machines make it possible to measure position- and time-dependent environmental conditions, such as air quality and radiation. The low-cost, miniature sensors commonly used in these systems are prone to measurement drift, requiring occasional re-calibration to provide accurate data. Requiring end users to periodically do manual calibration work would make many mobile sensing systems impractical. We therefore argue for the use of collaborative, automatic calibration among nearby mobile sensors, and provide solutions to the drift estimation and placement problems posed by such a system.
Xiang Yun, Lan S. Bai, Ricardo Piedrahita, Robert P. Dick, Qin Lv, Michael Hannigan
IPSN5
2012 Scalable misbehavior detection in online video chat services
abstract
The need for highly scalable and accurate detection and filtering of misbehaving users and obscene content in online video chat services has grown as the popularity of these services has exploded in popularity. This is a challenging problem because processing large amounts of video is compute intensive, decisions about whether a user is misbehaving or not must be made online and quickly, and moreover these video chats are characterized by low quality video, poorly lit scenes, diversity of users and their behaviors, diversity of the content, and typically short sessions. This paper presents EMeralD, a highly scalable system for accurately detecting and filtering misbehaving users in online video chat applications. EMeralD substantially improves upon the state-of-the-art filtering mechanisms by achieving much lower computational cost and higher accuracy. We demonstrate EMeralD's improvement via experimental evaluations on real-world data sets obtained from Chatroulette.com.
Xinyu Xing 0001, Yu-Li Liang, Sui Huang, Hanqiang Cheng, Richard Han 0001, Qin Lv, Xue (Steve) Liu, Shivakant Mishra, Yi Zhu 0010
KDD6
2012 Demo: MVChat: flasher detection for mobile video chat
abstract
Online video chat services such as Chatroulette [1] and Omegle [2] that randomly match pairs of users in video chat sessions have become increasingly popular, with over twenty thousand online users at anytime during a day. A key problem encountered in such systems is the presence of misbehaving users ("flashers") and obscene content. Our previous works [3] [4] prove that using some image recognition methods (skin-detection, dense SIFT) and machine learning algorithms could achieve significantly higher recall and better precision for flasher detection. Nowadays, with the rapid development of advanced mobile phones with both front and back cameras, we expect mobile video chat to become a popular extension of online video chat services. However, because of the computation-intensive features used by our previous solutions and mobile phones' hardware limitations such as memory size and CPU capacity, it is difficult to directly apply our previous works to mobile platforms. As smartphones are increasingly equipped with diverse sensing capabilities, we plan to utilize this multi-dimensional sensor information to extend flasher detection on mobile platform. This project explores how we can mine accelerometer and other mobile sensor data to infer some clues to optimize flasher detection accuracy while reducing the computation demands of flasher detection on the mobile device.
Lei Tian 0004, Junho Ahn, Hanqiang Cheng, Xinyu Xing 0001, Yu-Li Liang, Shivakant Mishra, David Chu, Xue (Steve) Liu, Richard Han 0001, Qin Lv
MobiSys10
2012 Fusing Text and Frienships for Location Inference in Online Social Networks
abstract
Location information is becoming prevalent in today's online social networks (OSNs), which raises special privacy concerns with regard to both location sharing and its applications. Even when no explicit location is disclosed by a user, it is possible to geolocate the user through his/her social context, e.g., status updates and social relationships in OSNs. To demonstrate this, we propose GeoFind, which accurately identifies users' geographic regions through effective fusion (re-ranking) of (1) text-based ranking using geo-sensitive textual features and (2) structure-based ranking using maximum likelihood estimation (MLE) of geotagged friends. Evaluation results using 0.8 million geotagged Twitter users over a 3-month period demonstrate that GeoFind outperforms state-of-the-art techniques, with significant reduction of estimation error (25% of average error, 66% of median error). The potential of improving location accuracy through the fusion of multiple data types calls for a re-examination of existing privacy protection policies and mechanisms.
Hansu Gu, Haojie Hang, Qin Lv, Dirk Grunwald
Web Intelligence3
2012 Efficient misbehaving user detection in online video chat services
abstract
Online video chat services, such as Chatroulette, Omegle, and vChatter are becoming increasingly popular and have attracted millions of users. One critical problem encountered in such applications is the presence of misbehaving users ("flashers") and obscene content. Automatically filtering out obscene content from these systems in an efficient manner poses a difficult challenge. This paper presents a novel Fine-Grained Cascaded (FGC) classification solution that significantly speeds up the compute-intensive process of classifying misbehaving users by dividing image feature extraction into multiple stages and filtering out easily classified images in earlier stages, thus saving unnecessary computation costs of feature extraction in later stages. Our work is further enhanced by integrating new webcam-related contextual information (illumination and color) into the classification process, and a 2-stage soft margin SVM algorithm for combining multiple features. Evaluation results using real-world data set obtained from Chatroulette show that the proposed FGC based classification solution significantly outperforms state-of-the-art techniques.
Hanqiang Cheng, Yu-Li Liang, Xinyu Xing 0001, Xue (Steve) Liu, Richard Han 0001, Qin Lv, Shivakant Mishra
WSDM6
2012 Interest-based real-time content recommendation in online social communities
Dongsheng Li 0002, Qin Lv, Xing Xie 0003, Li Shang 0001, Huanhuan Xia, Tun Lu, Ning Gu 0001
Knowl. Based Syst.2
2011 YANA: an efficient privacy-preserving recommender system for online social communities
abstract
In online social communities, many recommender systems use collaborative filtering, a method that makes recommendations based on what are liked by other users with similar interests. Serious privacy issues may arise in this process, as sensitive personal information (e.g., content interests) may be collected and disclosed to other parties, especially the recommender server. In this paper, we propose YANA (short for "you are not alone"), an efficient group-based privacy-preserving collaborative filtering system for content recommendation in online social communities. We have developed a prototype system on desktop and mobile devices, and evaluated it using real world data. The results demonstrate that YANA can effectively protect users' privacy, while achieving high recommendation quality and energy efficiency.
Dongsheng Li 0002, Qin Lv, Ning Gu 0001
CIKM2
2011 MAQS: a personalized mobile sensing system for indoor air quality monitoring
abstract
Most people spend more than 90% of their time indoors; indoor air quality (IAQ) influences human health, safety, productivity, and comfort. This paper describes MAQS, a personalized mobile sensing system for IAQ monitoring. In contrast with existing stationary or outdoor air quality sensing systems, MAQS users carry portable, indoor location tracking sensors that provide personalized IAQ information. To improve accuracy and energy efficiency, MAQS incorporates three novel techniques: (1) an accurate temporal n-gram augmented Bayesian room localization method that requires few Wi-Fi fingerprints; (2) an air exchange rate based IAQ sensing method, which measures general IAQ using only CO2 sensors; and (3) a zone-based proximity detection method for collaborative sensing, which saves energy and enables data sharing among users. MAQS has been deployed and evaluated via user study. Detailed evaluation results demonstrate that MAQS supports accurate personalized IAQ monitoring and quantitative analysis with high energy efficiency.
Yifei Jiang, Lei Tian 0004, Ricardo Piedrahita, Xiang Yun, Omkar Mansata, Qin Lv, Robert P. Dick, Michael Hannigan
UbiComp7
2011 MAQS: a mobile sensing system for indoor air quality
abstract
Most people spend more than 90% of their time indoors. Indoor air quality (IAQ) influences human health, safety, productivity, and comfort. This demo introduces MAQS, a personalized mobile sensing system for IAQ monitoring. In contrast with existing stationary or outdoor air quality sensing systems, MAQS users carry portable, indoor location tracking sensors that provide personalized IAQ information. To improve accuracy and energy efficiency, MAQS incorporates three novel techniques: (1) an accurate temporal n-gram augmented Bayesian room localization method; (2) an air exchange rate based IAQ sensing method; and (3) a zone-based proximity detection method for collaborative sensing.
Yifei Jiang, Lei Tian 0004, Ricardo Piedrahita, Xiang Yun, Omkar Mansata, Qin Lv, Robert P. Dick, Michael Hannigan
UbiComp7
2011 Deliberation for intuition: a framework for energy-efficient trip detection on cellular phones
abstract
Trip detection is a fundamental issue in many context-sensitive information services on mobile devices. It aims to automatically recognize significant places and trips between them. The key challenge is how to minimize energy consumption while maintaining high accuracy. Previous works that use GPS/WiFi sampling are accurate but energy efficiency is low and does not improve over time. Learning from the human decision making process, we propose an energy-efficient trip detection framework that consists of two modes: The deliberation mode learns cell-id patterns using GPS/WiFi based localization methods; the intuition mode only uses cell-ids and learned patterns for trip detection; transition between the two modes is controlled by parameters that are also learned. We evaluated our framework using real-life traces of six people over five months. Our experiments demonstrate that its energy consumption decreases rapidly as users' activities manifest regularity over time.
Yifei Jiang, Du Li, Guang Yang 0001, Qin Lv, Zhigang Liu 0010
UbiComp4
2011 ETree: Effective and Efficient Event Modeling for Real-Time Online Social Media Networks
abstract
Outline social media networks (OSMNs) such as Twitter provide great opportunities for public engagement and event information dissemination. Event-related discussions occur in real time and at the worldwide scale. However, these discussions are in the form of short, unstructured messages and dynamically woven into daily chats and status updates. Compared with traditional news articles, the rich and diverse user-generated content raises unique new challenges for tracking and analyzing events. Effective and efficient event modeling is thus essential for real-time information-intensive OSMNs. In this work, we propose ETree, an effective and efficient event modeling solution for social media network sites. Targeting the unique challenges of this problem, ETree consists of three key components: (1) an n-gram based content analysis technique for identifying core information blocks from a large number of short messages, (2) an incremental and hierarchical modeling technique for identifying and constructing event theme structures at different granularities, and (3) an enhanced temporal analysis technique for identifying inherent causalities between information blocks. Detailed evaluation using 3.5 million tweets over a 5-month period demonstrates that ETree can efficiently generate high-quality event structures and identify inherent causal relationships with high accuracy.
Hansu Gu, Xing Xie 0003, Qin Lv, Yaoping Ruan, Li Shang 0001
Web Intelligence3
2011 Pistis: A Privacy-Preserving Content Recommender System for Online Social Communities
abstract
With the explosive growth of online social communities and massive user-generated content, privacy-preserving recommender systems, which identify information of interest to individual users without disclosing personal interests to other parties, have become increasingly important. Collaborative filtering (CF), a widely used recommendation technique, recommends content that similar users have liked. As a result, CF-based recommender systems may expose sensitive personal interest information. This is demonstrated by a privacy attack model we present that targets online social communities. To solve this problem, we propose an interest group based privacy-preserving recommender system called Pistis. By identifying inherent item-user interest groups and separating users' private interests from their public interests, Pistis can make recommendations based on aggregated judgments of group members and local personalization, thus avoiding the disclosure of personal interest information. Pistis has been deployed and evaluated in an online social community with over 63,000 users, 20,000 daily posts, and 180,000 daily reads. Compared with two representative CF-based methods, our evaluation results demonstrate that Pistis achieves better performance in privacy preservation, recommendation quality, and efficiency.
Dongsheng Li 0002, Qin Lv, Huanhuan Xia, Tun Lu, Ning Gu 0001
Web Intelligence2
2011 SafeVchat: detecting obscene content and misbehaving users in online video chat services
abstract
Online video chat services such as Chatroulette, Omegle, and vChatter that randomly match pairs of users in video chat sessions are fast becoming very popular, with over a million users per month in the case of Chatroulette. A key problem encountered in such systems is the presence of flashers and obscene content. This problem is especially acute given the presence of underage minors in such systems. This paper presents SafeVchat, a novel solution to the problem of flasher detection that employs an array of image detection algorithms. A key contribution of the paper concerns how the results of the individual detectors are fused together into an overall decision classifying the user as misbehaving or not, based on Dempster-Shafer Theory. The paper introduces a novel, motion-based skin detection method that achieves significantly higher recall and better precision. The proposed methods have been evaluated over real-world data and image traces obtained from Chatroulette.com.
Xinyu Xing 0001, Yu-Li Liang, Hanqiang Cheng, Jianxun Dang, Sui Huang, Richard Han 0001, Xue (Steve) Liu, Qin Lv, Shivakant Mishra
WWW8
2010 Enhancing group recommendation by incorporating social relationship interactions
abstract
Group recommendation, which makes recommendations to a group of users instead of individuals, has become increasingly important in both the workspace and people’s social activities, such as brainstorming sessions for coworkers and social TV for family members or friends. Group recommendation is a challenging problem due to the dynamics of group memberships and diversity of group members. Previous work focused mainly on the content interests of group members and ignored the social characteristics within a group, resulting in suboptimal group recommendation performance. In this work, we propose a group recommendation method that utilizes both social and content interests of group members. We study the key characteristics of groups and propose (1) a group consensus function that captures the social, expertise, and interest dissimilarity among multiple group members; and (2) a generic framework that automatically analyzes group characteristics and constructs the corresponding group consensus function. Detailed user studies of diverse groups demonstrate the effectiveness of the proposed techniques, and the importance of incorporating both social and content interests in group recommender systems.
Mike Gartrell, Xinyu Xing 0001, Qin Lv, Aaron Beach, Richard Han 0001, Shivakant Mishra, Karim Seada
GROUP3
2010 Large-scale battery system modeling and analysis for emerging electric-drive vehicles
abstract
Emerging electric-drive vehicles demonstrate the potential for significant reduction of petroleum consumption and greenhouse gas emissions. Existing electric-drive vehicles typi- cally include a battery system consisting of thousands of Lithium-ion battery cells. Therefore, large-scale battery-system modeling and analysis is essential for battery system performance analysis, next-generation battery system design, and transportation electrification.
Yifei Jiang, Zyad Hassan, Qin Lv, Dragan Maksimovic
ISLPED5
2010 Hybrid energy storage system integration for vehicles
abstract
Energy consumption and the associated environmental impact are a pressing challenge faced by the transportation sector. Emerging electric-drive vehicles have shown promises for substantial reductions in petroleum use and vehicle emissions. Their success, however, has been hindered by the limitations of energy storage technologies. Existing in-vehicle Lithium-ion battery systems are bulky, expensive, and unreliable. Energy storage system (ESS) design and optimization is essential for emerging transportation electrification. This paper presents an integrated ESS modeling, design and optimization framework targeting emerging electric-drive vehicles. Based on an ESS modeling solution that considers major run-time and long-term battery effects, the proposed framework unifies design-time optimization and run-time control. It conducts statistical optimization for ESS cost and lifetime, which jointly considers the variances of ESS due to manufacture tolerance and heterogeneous driver-specific run-time use. It optimizes ESS design by incorporating complementary energy storage technologies, e.g., Lithium-ion batteries and ultracapacitors. Using physical measurements of battery manufacture variation and real-world user driving profiles, our experimental study has demonstrated that the proposed framework can effectively explore the statistical design space, and produce cost-efficient ESS solutions with statistical system lifetime guarantee.
Jia Wang 0003, Qin Lv, Hai Zhou 0001
ISLPED3
2009 iScope: personalized multi-modality image search for mobile devices
abstract
Mobile devices are becoming a primary medium for personal information gathering, management, and sharing. Managing personal image data on mobile platforms is a difficult problem due to large data set size, content diversity, heterogeneous individual usage patterns, and resource constraints. This article presents a user-centric system, called iScope, for personal image management and sharing on mobile devices. iScope uses multi-modality clustering of both content and context information for efficient image management and search, and online learning techniques for predicting images of interest. It also supports distributed content-based search among networked devices while maintaining the same intuitive interface, enabling efficient information sharing among people. We have implemented iScope and conducted in-field experiments using networked Nokia N810 portable Internet tablets. Energy efficiency was a primary design focus during the design and implementation of the iScope search algorithms. Experimental results indicate that iScope improves search time and search energy by 4.1X and 3.8X on average, relative to browsing.
Changyun Zhu, Qin Lv, Robert P. Dick
MobiSys3
2007 Sizing sketches: a rank-based analysis for similarity search
abstract
Sketches are compact data structures that can be used to estimate properties of the original data in building large-scale search engines and data analysis systems. Recent theoretical and experimental studies have shown that sketches constructed from feature vectors using randomized projections can effectively approximate L1 distance on the feature vectors with the Hamming distance on their sketches. Furthermore, such sketches can achieve good filtering accuracy while reducing the metadata space requirement and speeding up similarity searches by an order of magnitude. However, it is not clear how to choose the size of the sketches since it depends ondata type, dataset size, and desired filtering quality. In real systems designs, it is necessary to understand how to choose sketch size without the dataset, or at least without the whole datase.
Wei Dong 0003, William K. Josephson, Qin Lv, Moses Charikar, Kai Li 0001
SIGMETRICS4
2007 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search
Qin Lv, William K. Josephson, Moses Charikar, Kai Li 0001
VLDB1
2006 Ferret: a toolkit for content-based similarity search of feature-rich data
abstract
Building content-based search tools for feature-rich data has been a challenging problem because feature-rich data such as audio recordings, digital images, and sensor data are inherently noisy and high dimensional. Comparing noisy data requires comparisons based on similarity instead of exact matches, and thus searching for noisy data requires similarity search instead of exact search.The Ferret toolkit is designed to help system builders quickly construct content-based similarity search systems for feature-rich data types. The key component of the toolkit is a content-based similarity search engine for generic, multi-feature object representations. To solve the similarity search problem in high-dimensional spaces, we have developed approximation methods inspired by recent theoretical results on dimension reduction. The search engine constructs sketches from feature vectors as highly compact data structures for matching, filtering and ranking data objects. The toolkit also includes several other components to help system builders address search system infrastructure issues. We have implemented the toolkit and used it to successfully construct content-based similarity search systems for four data types: audio recordings, digital photos, 3D shape models and genomic microarray data.
Qin Lv, William K. Josephson, Moses Charikar, Kai Li 0001
EuroSys1
2004 Image similarity search with compact data structures
abstract
The recent theoretical advances on compact data structures (also called "sketches") have raised the question of whether they can effectively be applied to content-based image retrieval systems. The main challenge is to derive an algorithm that achieves high-quality similarity searches while using compact metadata. This paper proposes a new similarity search method consisting of three parts. The first is a new region feature representation with weighted $=1 distance function, and EMD* match, an improved EMD match, to compute image similarity. The second is a thresholding and transformation algorithm to convert feature vectors into very compact data structures. The third is an EMD embedding based filtering method to speed up the query process. We have implemented a prototype system with the proposed method and performed experiments with a 10,000 image database. Our results show that the proposed method can achieve more effective similarity searches than previous approaches with metadata 3 to 72 times more compact than previous systems. The experiments also show that our EMD embedding based filtering technique can speed up the query process by a factor of 5 or more with little loss in query effectiveness.
Qin Lv, Moses Charikar, Kai Li 0001
CIKM1
2002 Search and replication in unstructured peer-to-peer networks
Qin Lv, Edith Cohen, Kai Li 0001, Scott Shenker
ICS1
2002 Search and replication in unstructured peer-to-peer networks
abstract
Decentralized and unstructured peer-to-peer networks such as Gnutella are attractive for certain applications because they require no centralized directories and no precise control over network topology or data placement. However, the flooding-based query algorithm used in Gnutella does not scale; each individual query generates a large amount of traffic and large systems quickly become overwhelmed by the query-induced load. This paper explores various alternatives to Gnutella's query algorithm and data replication strategy. We propose a query algorithm based on multiple random walks that resolves queries almost as quickly as Gnutella's flooding method while reducing the network traffic by two orders of magnitude in many cases. We also present a distributed replication strategy that yields close-to-optimal performance.
Qin Lv, Edith Cohen, Kai Li 0001, Scott Shenker
SIGMETRICS1