VLDB 2026 Research / reviewers in the wild / expert
Wenjun Zhou 0001
dblp:06/5230-1
· DBLP profile ↗
44ranked-venue papers
9as first author
4since 2021 · last 2026
0000-0003-2346-8151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 22 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorComputer networks · 3 · 1 since 2021Theory of computation · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Early Detection of Adverse Drug Reactions in Postmarket MonitoringabstractMost drugs are associated with some form of adverse drug reactions (ADRs). Understanding the connection between drugs and ADRs is crucial for minimizing patient harm and reducing public healthcare costs. Consequently, there has been sustained interest in correlation analysis within pharmacovigilance and drug development. In the postmarketing phase, the estimated correlation between drugs and their ADRs is affected by both the correlation degree and variability. Therefore, accounting for variability is particularly important when measuring correlations, particularly in the early stage with fewer data points, where variability is typically higher. In this study, we introduce a framework called error-controlled correlation (ECC), which provides correlation estimates while dynamically adjusting for variability. ECC offers a versatile framework that is applicable to any correlation measure. Using the five most widely used correlation measures, we demonstrate ECC’s efficacy in identifying highly correlated drug-ADR pairs while maintaining a controlled type 1 error rate. Experimental results on both real-world and simulated data show that ECC consistently outperforms benchmark methods. Notably, it achieves comparable performance to existing methods with only 1/10th of the data, enabling significantly earlier ADR detection. History: Accepted by J. Paul Brooks, Area Editor for Applications in Biology, Medicine, & Healthcare. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0585 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2024.0585 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Wenjun Zhou 0001, Yong Hu 0002 |
INFORMS J. Comput. | 2 |
| 2024 | Quality of Surveillance Analysis of LEO Satellite Constellations for Orbital Edge Computing in Precision Agriculture and Climate MonitoringabstractIn this paper, we study the Quality of Surveillance of Low Earth Orbit (LEO) satellite constellations in orbit edge computing platforms, with a focus on monitoring and detecting unpredictable intruding targets. We are primarily interested in targets for precision agriculture and climate monitoring purposes, ranging from natural events such as tornados and flooding, to human-induced phenomenon, such as forest fires and crop monitoring. Targets may be static or mobile. We develop an analytical model that seeks to predict performance attributes of the surveillance. The model takes into account tunable system parameters, such as the number of satellites and their altitude, thereby allowing us to decide the best constellation configuration for different purposes. The results from the analytical model validate the feasibility of using LEO constellation for surveillance purposes and illustrate how model-guided parameter tuning can significantly enhance the monitoring performance in precision agriculture and climate monitoring applications. Qing Cao 0001, Joshua Fu, Wenjun Zhou 0001 |
IWQoS | 4 |
| 2022 | Work from home during the COVID-19 pandemic: An observational study based on a large geo-tagged COVID-19 Twitter dataset (UsaGeoCov19)
Yunhe Feng, Wenjun Zhou 0001 |
Inf. Process. Manag. | 2 |
| 2022 | COSLE: Cost sensitive loan evaluation for P2P lending
Sen Wu 0001, Xiaonan Gao, Wenjun Zhou 0001 |
Inf. Sci. | 3 |
| 2020 | Multi-User Mobile Sequential Recommendation for Route OptimizationabstractWe enhance the mobile sequential recommendation (MSR) model and address some critical issues in existing formulations by proposing three new forms of the MSR from a multi-user perspective. The multi-user MSR (MMSR) model searches optimal routes for multiple drivers at different locations while disallowing overlapping routes to be recommended. To enrich the properties of pick-up points in the problem formulation, we additionally consider the pick-up capacity as an important feature, leading to the following two modified forms of the MMSR: MMSR-m and MMSR-d. The MMSR-m sets a maximum pick-up capacity for all urban areas, while the MMSR-d allows the pick-up capacity to vary at different locations. We develop a parallel framework based on the simulated annealing to numerically solve the MMSR problem series. Also, a push-point method is introduced to improve our algorithms further for the MMSR-m and the MMSR-d, which can handle the route optimization in more practical ways. Our results on both real-world and synthetic data confirmed the superiority of our problem formulation and solutions under more demanding practical scenarios over several published benchmarks. Keli Xiao, Zeyang Ye, Wenjun Zhou 0001, Yong Ge 0001, Yuefan Deng |
ACM Trans. Knowl. Discov. Data | 4 |
| 2019 | Chasing Total Solar Eclipses on Twitter: Big Social Data Analytics for Once-in-a-Lifetime EventsabstractWith the popularity of social networking services, big social data analytics emerged in various applications, such as discovering trending topics, monitoring public sentiment, and identifying human mobility patterns. In this paper, we take the opportunity of The 2017 Great American Eclipse, a once-in-a-lifetime event, to look into its potential social, emotional, and human movement impacts at the national level. Specifically, we collected more than five million English eclipse- mentioning tweets in a real-time manner using Twitter Streaming APIs. Then we profiled spatio- temporal distributions of the data, extracted both hashtagged and latent topics, analyzed emotions using polarized words, emojis and emoticons, and revealed both interstate and intrastate eclipse- chasing travel patterns. Our study provides a comprehensive example of understanding big social data and its associated influence from diverse perspectives. Yunhe Feng, Zheng Lu 0005, Zhonghua Zheng, Peng Sun 0003, Wenjun Zhou 0001, Qing Cao 0001 |
GLOBECOM | 5 |
| 2019 | Large-Scale Personalized Delivery for Guaranteed Display Advertising with Real-Time PacingabstractGuaranteed display (GD) has been a successful model for display advertising. Existing solutions usually model GD services as a crowd-level supply allocation problem. This formulation, however, not only ignores user heterogeneity within crowds, but also makes it difficult to incorporate individual-level constraints. In this paper, we present an large-scale system for personalized delivery in GD advertising services. A unique contribution is to model the allocation problem at the individual level that accounts for user-ad interactions. Therefore, our system can conveniently incorporate complex constraints, such as the priority of GD contracts, the display frequency of ads, and the effectiveness of ad slots arrangement. Moreover, we develop a real-time pacing strategy to fulfill GD contracts with smooth ad delivery and optimized ad performance, such as cost-per-click (CPC) and cost-per-action (CPA). Our system can be parallelized to efficiently compute the delivery solution with billions decision variables. Using both offline evaluation and online A/B tests, we demonstrate that our solution is effective in terms of both accuracy and efficiency. Yang Li 0198, Chuanren Liu, Wenxiang Zhu, Wenjun Zhou 0001 |
ICDM | 6 |
| 2019 | The World Wants Mangoes and Kangaroos: A Study of New Emoji Requests Based on Thirty Million TweetsabstractAs emojis become prevalent in personal communications, people are always looking for new, interesting emojis to express emotions, show attitudes, or simply visualize texts. In this study, we collected more than thirty million tweets mentioning the word “emoji” in a one-year period to study emoji requests on Twitter. First, we filtered out bot-generated tweets and extracted emoji requests from the raw tweets using a comprehensive list of linguistic patterns. Then, we examined patterns of new emoji requests by exploring their time, locations, and context. Finally, we summarized users' advocacy behaviors and identified expressions of equity, diversity, and fairness issues due to unreleased but expected emojis, and concluded the significance of new emojis on society. To the best of our knowledge, this paper is the first to conduct a systematic, large-scale study on new emoji requests. Yunhe Feng, Wenjun Zhou 0001, Zheng Lu 0005, Zhibo Wang 0001, Qing Cao 0001 |
WWW | 2 |
| 2019 | Exploiting intra-day patterns for market shock prediction: A machine learning approach
Jinwen Sun, Keli Xiao, Chuanren Liu, Wenjun Zhou 0001, Hui Xiong 0001 |
Expert Syst. Appl. | 4 |
| 2019 | Transit Pattern Detection Using Tensor FactorizationabstractUnderstanding citywide transit patterns is important for transportation management, including city planning and route optimization. The wide deployment of automated fare collection (AFC) systems in public transit vehicles has enabled us to collect massive amounts of transit records, which capture passengers’ traveling activities. Based on such transit records, origin–destination associations have been studied extensively in the literature. However, the identification of transit patterns that establish the origin–transfer–destination (OTD) associations, in spite of its importance, is underdeveloped. In this paper, we propose a framework based on transit tensor factorization (TTF) to identify citywide travel patterns. In particular, we create a transit tensor, which summarizes the citywide OTD information of all passenger trips captured in the AFC records. The TTF framework imposes spatial regularization in the formulation to group nearby stations into meaningful regions and uses multitask learning to identify traffic flows among these regions at different times of the day and days of the week. Evaluated with large-scale, real-world data, our results show that the proposed TTF framework can effectively identify meaningful citywide transit patterns. The online supplement is available at https://doi.org/10.1287/ijoc.2018.0824 . Bowen Du 0001, Wenjun Zhou 0001, Chuanren Liu, Yifeng Cui, Hui Xiong 0001 |
INFORMS J. Comput. | 2 |
| 2019 | Detecting Pickpocket Suspects from Large-Scale Public Transit RecordsabstractMassive data collected by automated fare collection (AFC) systems provide opportunities for studying both personal traveling behaviors and collective mobility patterns in urban areas. Existing studies on AFC data have primarily focused on identifying passengers' movement patterns. However, we creatively leveraged such data for identifying pickpocket suspects. Stopping pickpockets in the public transit system has been crucial for improving passenger satisfaction and public safety. Nonetheless, in practice, it is challenging to discern thieves from regular passengers. In this paper, we developed a suspect detection and surveillance system, which can identify pickpocket suspects based on their daily transit records. Specifically, we first extracted a number of useful features from each passenger's daily activities in the transit system. Then, we took a two-step approach that exploits the strengths of unsupervised outlier detection and supervised classification models to identify thieves, who typically exhibit abnormal traveling behaviors. Experimental results demonstrated the effectiveness of our method. We also developed a prototype system for potential uses by security personnel. Bowen Du 0001, Chuanren Liu, Wenjun Zhou 0001, Zhenshan Hou, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | Demand-Aware Charger Planning for Electric Vehicle SharingabstractCars of the future have been predicted as shared and electric. There has been a rapid growth in electric vehicle (EV) sharing services worldwide in recent years. For EV-sharing platforms to excel, it is essential for them to offer private charging infrastructure for exclusive use that meets the charging demand of their clients. Particularly, they need to plan not only the places to build charging stations, but also the amounts of chargers per station, to maximally satisfy the requirements on global charging coverage and local charging demand. Existing research efforts are either inapplicable for their different problem formulations or are at a coarse granularity. In this paper, we formulate the \underlineE lectric \underlineV ehicle \underlineC harger \underlineP lanning (EVCP) problem especially for EV-sharing. We prove that the \shortpro problem is NP-hard, and design an approximation algorithm to solve the problem with a theoretical bound of $1-\frac1 e $. We also devise some optimization techniques to speed up the solution. Extensive experiments on real-world datasets validate the effectiveness and the efficiency of our proposed solutions. Bowen Du 0001, Yongxin Tong, Zimu Zhou, Wenjun Zhou 0001 |
KDD | 5 |
| 2018 | Multi-User Mobile Sequential Recommendation: An Efficient Parallel Computing ParadigmabstractThe classic mobile sequential recommendation (MSR) problem aims to provide the optimal route to taxi drivers for minimizing the potential travel distance before they meet next passengers. However, the problem is designed from the view of a single user and may lead to overlapped recommendations and cause traffic problems. Existing approaches usually contain an offline pruning process with extremely high computational cost, given a large number of pick-up points. To this end, we formalize a new multi-user MSR (MMSR) problem that locates optimal routes for a group of drivers with different starting positions. We develop two efficient methods, PSAD and PSAD-M, for solving the MMSR problem by ganging parallel computing and simulated annealing. Our methods outperform several existing approaches, especially for high-dimensional MMSR problems, with a record-breaking performance of 180x speedup using 384 cores. Zeyang Ye, Keli Xiao, Wenjun Zhou 0001, Yong Ge 0001, Yuefan Deng |
KDD | 4 |
| 2018 | A multi-granularity perspective for spatial profiling of mobile apps
Yunhe Feng, Zheng Lu 0005, Wenjun Zhou 0001, Qing Cao 0001 |
Inf. Sci. | 3 |
| 2018 | Paradoxical Correlation Pattern MiningabstractGiven a large transactional database, correlation computing/association analysis aims at efficiently finding strongly correlated items. For traditional association analysis, relationships among variables are usually measured at a global level. In this study, we investigate confounding factors that can help to capture abnormal correlation behaviors at a local level. Indeed, many real-world phenomena are localized to specific markets or subpopulations. Such local relationships may not be visible or may be miscalculated when collectively analyzing the entire data. In particular, confounding effects that change the direction of correlation are a most severe problem because the global correlations alone leads to errant conclusions. To this end, we propose CONFOUND, an efficient algorithm to identify paradoxical correlation patterns (i.e., where controlling for a third item changes the direction of association for strongly correlated pairs) using effective pruning strategies. Moreover, we also provide an enhanced version of this algorithm, called CONFOUND±, which substantially speeds up the confounder search step. Finally, experimental results showed that our proposed CONFOUND and CONFOUND± algorithms can effectively identify confounders and the computational performance is orders of magnitude faster than benchmark methods. Wenjun Zhou 0001, Hui Xiong 0001, Keli Xiao, Robert W. Mee |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Approximate Cardinality Estimation (ACE) in large-scale Internet of Things deployments
Qing Cao 0001, Yunhe Feng, Zheng Lu 0005, Hairong Qi 0001, Leon M. Tolbert, Lipeng Wan 0001, Zhibo Wang 0001, Wenjun Zhou 0001 |
Ad Hoc Networks | 8 |
| 2017 | Efficient locality weighted sparse representation for graph-based learning
Xiaodong Feng 0001, Sen Wu 0001, Wenjun Zhou 0001, Min Quan |
Knowl. Based Syst. | 3 |
| 2017 | Optimizing the performance of sensor network programs through estimation-based code profiling
Lipeng Wan 0001, Qing Cao 0001, Wenjun Zhou 0001 |
Pervasive Mob. Comput. | 3 |
| 2017 | Multi-Hypergraph Consistent Sparse CodingabstractSparse representation has been a powerful technique for modeling high-dimensional data. As an unsupervised technique to extract sparse representations, sparse coding encodes the original data into a new sparse code space and simultaneously learns a dictionary representing high-level semantics. Existing methods have considered local manifold within high-dimensional data using graph/hypergraph Laplacian regularization, and more from the manifold could be utilized to improve the performance. In this article, we propose to further regulate the sparse coding so that the learned sparse codes can well reconstruct the hypergraph structure. In particular, we add a novel hypergraph consistency regularization term (HC) by minimizing the reconstruction error of the hypergraph incidence or weight matrix. Moreover, we extend the HC term to multi-hypergraph consistent sparse coding (MultiCSC) and automatically select the optimal manifold structure under the multi-hypergraph learning framework. We show that the optimization of MultiCSC can be solved efficiently, and that several existing sparse coding methods can fit into the general framework of MultiCSC as special cases. As a case study, hypergraph incidence consistent sparse coding is applied to perform semi-auto image tagging, demonstrating the effectiveness of hypergraph consistency regulation. We perform further experiments using MultiCSC for image clustering, which outperforms a number of baselines. Xiaodong Feng 0001, Sen Wu 0001, Wenjun Zhou 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | Quantifying skill relevance to job titlesabstractEliminating or reducing skill gaps in the job market is critical to putting people back to work, reducing the unemployment rate, and increasing the labor market participation rate. A key element in closing the skills gap is accurately identifying the mismatch between the skills expected by employers and those possessed by job seekers. In this study, our goal was to profile job titles by effectively quantifying the relevance of skills. We started by using a naive, frequency-based skill ranking approach, which resulted in the most generic skills ranked on the top. We then adapted a number of alternative metrics and compared their performances on a number of job titles. The outcome of this study can support CareerCoach, an analytical solution CareerBuilder has piloted to provide insights and data dashboards to job seekers. Wenjun Zhou 0001, Faizan Javed, Mahmudur Rahman, Janani Balaji, Matt McNair |
IEEE BigData | 1 |
| 2016 | Catch Me If You Can: Detecting Pickpocket Suspects from Large-Scale Transit RecordsabstractMassive data collected by automated fare collection (AFC) systems provide opportunities for studying both personal traveling behaviors and collective mobility patterns in the urban area. Existing studies on the AFC data have primarily focused on identifying passengers' movement patterns. In this paper, however, we creatively leveraged such data for identifying thieves in the public transit systems. Indeed, stopping pickpockets in the public transit systems has been critical for improving passenger satisfaction and public safety. However, it is challenging to tell thieves from regular passengers in practice. To this end, we developed a suspect detection and surveillance system, which can identify pick-pocket suspects based on their daily transit records. Specifically, we first extracted a number of features from each passenger's daily activities in the transit systems. Then, we took a two-step approach that exploits the strengths of unsupervised outlier detection and supervised classification models to identify thieves, who exhibit abnormal traveling behaviors. Experimental results demonstrated the effective- ness of our method. We also developed a prototype system with a user-friendly interface for the security personnel. Bowen Du 0001, Chuanren Liu, Wenjun Zhou 0001, Zhenshan Hou, Hui Xiong 0001 |
KDD | 3 |
| 2016 | Multi-hypergraph Incidence Consistent Sparse Coding for Image Data Clustering
Xiaodong Feng 0001, Sen Wu 0001, Wenjun Zhou 0001, Zhiwei Tang |
PAKDD (2) | 3 |
| 2015 | Discovery of College Students in Financial HardshipabstractCollege students with financial difficulties refer to those whose families can hardly afford their high tuition in universities, and should be supported by modern funding system. Indeed, students' economic plight negatively impact their mental health, academic performance, as well as their personal and social life. While funding students in financial hardship is widely accepted, there is limited understanding and research on effectively identification of the qualifying students. Traditional approaches relying on advisers' personal assessments are inefficient, and such subjective judgements may not reflect the truth. To this end, in this paper, we explore the data mining techniques for identifying students who are qualified for financial support. Specifically, we investigate students' complex behaviors on campus from multiple perspectives, and develop a learning framework, named Dis-HARD, by jointly incorporating the heterogeneous features to predict the portfolio of stipends a given student should be awarded. Our framework formalizes the above problem as a multi-label learning problem. Along this line, we first extract discriminative features from three perspectives: (i) smartcard usage behavior, (ii) internet usage behavior and (iii) trajectory on campus. Then, we develop a linear loss function with regularization to solve this multi-label classification problem. In addition, to effectively exploit the students' similarity and label dependency, we incorporate the graph Laplacian and composite l2,1-norm into the regularization of our model, and develop are-weighted algorithm to achieve effective optimization. Finally, experiments on real-world data demonstrate that our method consistently provides better performance compared to the existing state-of-the-art methods. Chu Guan, Xinjiang Lu, Enhong Chen, Wenjun Zhou 0001, Hui Xiong 0001 |
ICDM | 5 |
| 2015 | Estimation-based profiling for code placement optimization in sensor network programsabstractIn this work, we focus on applying profiling guided code placement to programs running on resource-constrained sensor motes. Specifically, we model the execution of sensor network programs under nondeterministic inputs as discrete-time Markov processes, and propose a novel approach named Code Tomography to estimate parameters of the Markov models that reflect sensor network programs' dynamic execution behavior by only using end-to-end timing information measured at start and end points of each procedure. The parameters estimated by Code Tomography are fed back to compilers to optimize the code placement so that branch misprediction rate can be reduced. Lipeng Wan 0001, Qing Cao 0001, Wenjun Zhou 0001 |
ISPASS | 3 |
| 2015 | Fuzzy Temporal Segmentation and Probabilistic Recognition of Continuous Human Daily ActivitiesabstractUnderstanding human activities is an essential capability for intelligent robots to help people in a variety of applications. Humans perform activities in a continuous fashion, and transitions between temporally adjacent activities are gradual. Our Fuzzy Segmentation and Recognition (FuzzySR) algorithm explicitly reasons about gradual transitions between continuous human activities. Our objective is to simultaneously segment a given video into a sequence of events and recognize the activity contained in each event. The algorithm uniformly segments the video into a sequence of nonoverlapping blocks, each lasting a short period of time. Then, a multivariable time series is formed by concatenating block-level human activity summaries that are computed using topic models over local spatiotemporal features extracted from each block. Through encoding an event as a fuzzy set with fuzzy boundaries to represent gradual transitions, our approach is capable of segmenting the continuous visual data into a sequence of fuzzy events. By incorporating all block summaries contained in an event, our algorithm determines the activity label for each event. To evaluate performance, we conduct experiments using six datasets. Our algorithm shows promising continuous activity segmentation results on these datasets and obtains the event-level activity recognition precision of 42.6%, 60.4%, 65.2%, and 78.9% on the Hollywood-2, CAD-60, ACT $4^2$, and UTK-CAP datasets, respectively. Hao Zhang 0011, Wenjun Zhou 0001, Lynne E. Parker |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2015 | Correlation range query for effective recommendations
Wenjun Zhou 0001, Hao Zhang 0011 |
World Wide Web | 1 |
| 2014 | Simplex-Based 3D Spatio-temporal Feature Description for Action RecognitionabstractWe present a novel feature description algorithm to describe 3D local spatio-temporal features for human action recognition. Our descriptor avoids the singularity and limited discrimination power issues of traditional 3D descriptors by quantizing and describing visual features in the simplex topological vector space. Specifically, given a feature's support region containing a set of 3D visual cues, we decompose the cues' orientation into three angles, transform the decomposed angles into the simplex space, and describe them in such a space. Then, quadrant decomposition is performed to improve discrimination, and a final feature vector is composed from the resulting histograms. We develop intuitive visualization tools for analyzing feature characteristics in the simplex topological vector space. Experimental results demonstrate that our novel simplex-based orientation decomposition (SOD) descriptor substantially outperforms traditional 3D descriptors for the KTH, UCF Sport, and Hollywood-2 benchmark action datasets. In addition, the results show that our SOD descriptor is a superior individual descriptor for action recognition. Hao Zhang 0011, Wenjun Zhou 0001, Christopher M. Reardon, Lynne E. Parker |
CVPR | 2 |
| 2014 | Towards approximate spatial queries for large-scale vehicle networksabstractWith advances in vehicle-to-vehicle communication, future vehicles will have access to a communication channel through which messages can be sent and received when two get close to each other. This enabling technology makes it possible for authenticated users to send queries to those vehicles of interest, such as those that are located within a geographic region, over multiple hops for various application goals. However, a naive method that requires flooding the queries to each active vehicle in a region will incur a total communication overhead that is proportional to the size of the area and the density of vehicles. In this paper, we study the problem of spatial queries for vehicle networks by investigating probabilistic methods, where we only try to obtain approximate estimates within desired confidence intervals using only sublinear overheads. We consider this to be particularly useful when spatial query results can be made approximate or not precise, as is the case with many potential applications. The proposed method has been tested on snapshots from real world vehicle network traces. Lipeng Wan 0001, Zhibo Wang 0001, Zheng Lu 0005, Hairong Qi 0001, Wenjun Zhou 0001, Qing Cao 0001 |
SIGSPATIAL/GIS | 5 |
| 2014 | Fuzzy segmentation and recognition of continuous human activitiesabstractMost previous research has focused on classifying single human activities contained in segmented videos. However, in real-world scenarios, human activities are inherently continuous and gradual transitions always exist between temporally adjacent activities. In this paper, we propose a Fuzzy Segmentation and Recognition (FuzzySR) algorithm to explicitly model this gradual transition. Our goal is to simultaneously segment a given video into events and recognize the activity contained in each event. Specifically, our algorithm uniformly partitions the video into a sequence of non-overlapping blocks, each of which lasts a short period of time. Then, a multi-variable time series is creatively formed through concatenating the block-level human activity summaries that are computed using topic models over each block's local spatio-temporal features. By representing an event as a fuzzy set that has fuzzy boundaries to model gradual transitions, our algorithm is able to segment the video into a sequence of fuzzy events. By incorporating all block summaries contained in an event, the proposed algorithm determines the most appropriate activity category for each event. We evaluate our algorithm's performance using two real-world benchmark datasets that are widely used in the machine vision community. We also demonstrate our algorithm's effectiveness in important robotics applications, such as intelligent service robotics. For all used datasets, our algorithm achieves promising continuous human activity segmentation and recognition results. Hao Zhang 0011, Wenjun Zhou 0001, Lynne E. Parker |
ICRA | 2 |
| 2014 | Spectral clustering of high-dimensional data exploiting sparse representation vectors
Sen Wu 0001, Xiaodong Feng 0001, Wenjun Zhou 0001 |
Neurocomputing | 3 |
| 2013 | Utilizing URLs Position to Estimate Intrinsic Query-URL RelevanceabstractQuery-URL relevance (QUR) is an important criterion to measure the quality of commercial search engines. However, the traditional way to collect high-quality QURs is time-consuming and labor-intensive since it is primarily based on human judges. To address these issues, numerous models have been studied to automatically infer the QURs. Unlike the prior studies in this literature, we first empirically analyze the correlation between multiple annotators' judgments on QURs and URL position in ranking lists. By doing so, we reveal and justify the potential impacts of URL position on inferring intrinsic QURs. Inspired by this finding, a position-sensitive model (PSM) is proposed to infer QURs more accurately. In contrast with most existing approaches that attempt to construct the direct relationship between QURs and the features characterizing query-URL pairs, PSM assumes that the QUR is connected with the features through URL position. We conducted the experiments in real search engine Baidu.com, and compared the experimental results to those of the typical methods used in similar tasks, reporting significant gains over click-through rate and the normalized discounted cumulative gains (NDCGs). Xiaogang Han, Wenjun Zhou 0001, Xing Jiang 0001, Hengjie Song, Toyoaki Nishida |
ICDM | 2 |
| 2013 | Correlation Range Query
Wenjun Zhou 0001, Hao Zhang 0011 |
WAIM | 1 |
| 2012 | Community discovery and profiling with social messagesabstractDiscovering communities from social media and collaboration systems has been of great interest in recent years. Existing work show prospects of modeling contents and social links, aiming at discovering social communities, whose definition varies by application. We believe that a community depends not only on the group of people who actively participate, but also the topics they communicate about or collaborate on. This is especially true for workplace email communications. Within an organization, it is not uncommon that employees multifunction, and groups of employees collaborate on multiple projects at the same time. In this paper, we aim to automatically discovering and profiling users' communities by taking into account both the contacts and the topics. More specifically, we propose a community profiling model called COCOMP, where the communities labels are latent, and each social document corresponds to an information sharing activity among the most probable community members regarding the most relevant community issues. Experiment results on several social communication datasets, including emails and Twitter messages, demonstrate that the model can discover users' communities effectively, and provide concrete semantics. Wenjun Zhou 0001, Hongxia Jin, Yan Liu 0002 |
KDD | 1 |
| 2011 | Enhancing investment decisions in P2P lending: an investor composition perspectiveabstractP2P lending, as a novel economic lending model, has imposed new challenges about how to make effective investment decisions. Indeed, a key challenge along this line is how to align the right information with the right people. For a long time, people have made tremendous efforts in establishing credit records for the borrowers. However, information from investors is still under-explored for improving investment decisions in P2P lending. To that end, we propose a data driven investment decision-making framework, which exploits the investor composition of each investment for enhancing decisions making in P2P lending. Specifically, we first build investor profiles based on quantitative analysis of past performances, risk preferences, and investment experiences of investors. Then, based on investor profiles, we develop an investor composition analysis model, which can be used to select valuable investments and improve the investment decisions. To validate the proposed model, we perform extensive experiments on the real-world data from the world's largest P2P lending marketplace. Experimental results reveal that investor composition can help us evaluate the profit potential of an investment and the decision model based on investor composition can help investors make better investment decisions. Chunyu Luo, Hui Xiong 0001, Wenjun Zhou 0001, Yanhong Guo, Guishi Deng |
KDD | 3 |
| 2011 | Dynamic Rank Correlation Computing for Financial Risk Analysis
Wenjun Zhou 0001, Keli Xiao |
KSEM | 1 |
| 2011 | Checkpoint evolution for volatile correlation computing
Wenjun Zhou 0001, Hui Xiong 0001 |
Mach. Learn. | 1 |
| 2011 | Multifocal learning for customer problem analysisabstractIn this study, we formalize a multifocal learning problem, where training data are partitioned into several different focal groups and the prediction model will be learned within each focal group. The multifocal learning problem is motivated by numerous real-world learning applications. For instance, for the same type of problems encountered in a customer service center, the problem descriptions from different customers can be quite different. Experienced customers usually give more precise and focused descriptions about the problem. In contrast, inexperienced customers usually provide diverse descriptions. In this case, the examples from the same class in the training data can be naturally in different focal groups. Therefore, it is necessary to identify those natural focal groups and exploit them for learning at different focuses. Along this line, the key development challenge is how to identify those focal groups in the training data. As a case study, we exploit multifocal learning for profiling customer problems. Also, we provide an empirical study about how the performance of multifocal learning is affected by the quality of focal groups. The results on real-world customer problem logs show that multifocal learning can significantly boost the performance of many existing classification algorithms, such as Support Vector Machines (SVMs), for classifying customer problems and there is strong correlation between the quality of focal groups and the learning performance. Yong Ge 0001, Hui Xiong 0001, Wenjun Zhou 0001, Siming Li, Ramendra K. Sahoo |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2009 | Efficient Discovery of Confounders in Large Data SetsabstractGiven a large transaction database, association analysis is concerned with efficiently finding strongly related objects. Unlike traditional associate analysis, where relationships among variables are searched at a global level, we examine confounding factors at a local level. Indeed, many real-world phenomena are localized to specific regions and times. These relationships may not be visible when the entire data set is analyzed. Specially, confounding effects that change the direction of correlation is the most significant. Along this line, we propose to efficiently find confounding effects attributable to local associations. Specifically, we derive an upper bound by a necessary condition of confounders, which can help us prune the search space and efficiently identify confounders. Experimental results show that the proposed CONFOUND algorithm can effectively identify confounders and the computational performance is an order of magnitude faster than benchmark methods. Wenjun Zhou 0001, Hui Xiong 0001 |
ICDM | 1 |
| 2009 | Multi-focal learning and its application to customer service supportabstractIn this study, we formalize a multi-focal learning problem, where training data are partitioned into several different focal groups and the prediction model will be learned within each focal group. The multi-focal learning problem is motivated by numerous real-world learning applications. For instance, for the same type of problems encountered in a customer service center, the problem descriptions from different customers can be quite different. The experienced customers usually give more precise and focused descriptions about the problem. In contrast, the inexperienced customers usually provide more diverse descriptions. In this case, the examples from the same class in the training data can be naturally in different focal groups. As a result, it is necessary to identify those natural focal groups and exploit them for learning at different focuses. The key developmental challenge is how to identify those focal groups in the training data. As a case study, we exploit multi-focal learning for profiling problems in customer service centers. The results show that multifocal learning can significantly boost the learning accuracies of existing learning algorithms, such as Support Vector Machines (SVMs), for classifying customer problems. Yong Ge 0001, Hui Xiong 0001, Wenjun Zhou 0001, Ramendra K. Sahoo, Xiaofeng Gao 0001, Weili Wu 0001 |
KDD | 3 |
| 2008 | Volatile correlation computation: a checkpoint viewabstractRecent years have witnessed increased interest in computing strongly correlated pairs in very large databases. Most previous studies have been focused on static data sets. However, in real-world applications, input data are often dynamic and must continually be updated. With such large and growing data sets, new research efforts are expected to develop an incremental solution for correlation computing. Along this line, in this paper, we propose a CHECK-POINT algorithm that can efficiently incorporate new transactions for correlation computing as they become available. Specifically, we set a checkpoint to establish a computation buffer, which can help us determine an upper bound for the correlation. This checkpoint bound can be exploited to identify a list of candidate pairs, which will be maintained and computed for correlations as new transactions are added into the database. However, if the total number of new transactions is beyond the buffer size, a new upper bound is computed by the new checkpoint and a new list of candidate pairs is identified. Experimental results on real-world data sets show that CHECK-POINT can significantly reduce the correlation computing cost in dynamic data sets and has the advantage of compacting the use of memory space. Wenjun Zhou 0001, Hui Xiong 0001 |
KDD | 1 |
| 2008 | Hypergraph partitioning for document clustering: a unified clique perspectiveabstractHypergraph partitioning has been considered as a promising method to address the challenges of high dimensionality in document clustering. With documents modeled as vertices and the relationship among documents captured by the hyperedges, the goal of graph partitioning is to minimize the edge cut. Therefore, the definition of hyperedges is vital to the clustering performance. While several definitions of hyperedges have been proposed, a systematic understanding of desired characteristics of hyperedges is still missing. To that end, in this paper, we first provide a unified clique perspective of the definition of hyperedges, which serves as a guide to define hyperedges. With this perspective, based on the concepts of hypercliques and shared (reverse) nearest neighbors, we propose three new types of clique hyperedges and analyze their properties regarding purity and size issues. Finally, we present an extensive evaluation using real-world document datasets. The experimental results show that, with shared (reverse) nearest neighbor based hyperedges, the clustering performance can be improved significantly in terms of various external validation measures without the need for fine tuning of parameters. Tianming Hu, Hui Xiong 0001, Wenjun Zhou 0001, Sam Yuan Sung, Hangzai Luo |
SIGIR | 3 |
| 2008 | Top-k Correlation ComputationabstractRecently, there has been considerable interest in efficiently computing strongly correlated pairs in large databases. Most previous studies require the specification of a minimum correlation threshold to perform the computation. However, it may be difficult for users to provide an appropriate threshold in practice because different data sets typically have different characteristics. To this end, in this paper, we propose an alternative task: finding the top-k strongly correlated pairs. Consequently, we identify a two-dimensional monotone property of an upper bound of ϕ correlation coefficient and develop an efficient algorithm, called TOP-COP, to exploit this property to effectively prune many pairs even without computing their correlation coefficients. Our experimental results show that TOP-COP can be an order of magnitude faster than alternative approaches for mining the top-k strongly correlated pairs. Finally, we show that the performance of the TOP-COP algorithm is tightly related to the degree of data dispersion. Indeed, the higher the degree of data dispersion, the larger the computational savings achieved by the TOP-COP algorithm. Hui Xiong 0001, Wenjun Zhou 0001, Mark Brodie, Sheng Ma |
INFORMS J. Comput. | 2 |
| 2007 | A Generalization of Proximity Functions for K-MeansabstractK-means is a widely used partitional clustering method. A large amount of effort has been made on finding better proximity (distance) functions for k-means. However, the common characteristics of proximity functions remain unknown. To this end, in this paper, we show that all proximity functions that fit k-means clustering can be generalized as k-means distance, which can be derived by a differentiable convex function. A general proof of sufficient and necessary conditions for k-means distance functions is also provided. In addition, we reveal that k-means has a general uniformization effect; that is, k-means tends to produce clusters with relatively balanced cluster sizes. This uniformization effect of k-means exists regardless of proximity functions. Finally, we have conducted extensive experiments on various real-world data sets, and the results show the evidence of the uniformization effect. Also, we observed that external clustering validation measures, such as entropy and variance of information (VI), have difficulty in measuring clustering quality if data have skewed distributions on class sizes. Junjie Wu 0002, Hui Xiong 0001, Jian Chen 0016, Wenjun Zhou 0001 |
ICDM | 4 |
| 2006 | Preserving Patterns in Bipartite Graph PartitioningabstractThis paper describes a new bipartite formulation for word-document co-clustering such that hyperclique patterns, strongly affiliated documents in this case, are guaranteed not to be split into different clusters. Our approach for pattern preserving clustering consists of three steps: mine maximal hyperclique patterns, form the bipartite, and partition it. With hyperclique patterns of documents preserved, the topic of each cluster can be represented by both the top words from that cluster and the documents in the patterns, which are expected to be more compact and representative than those in the standard bipartite formulation. Experiments with real-world datasets show that, with hyperclique patterns as starting points, we can improve the clustering results in terms of various external clustering criteria. Also, the partitioned bipartite with preserved topical sets of documents naturally lends itself to different functions in search engines Tianming Hu, Chao Qu, Chew Lim Tan, Sam Yuan Sung, Wenjun Zhou 0001 |
ICTAI | 5 |