Yue Wang 0007

dblp:33/4822-7 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0002-9648-2838ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 8 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Computer networks · 10 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
abstract
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Chen Gao 0001, Yue Wang 0007, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)8
2025 Bi-Dynamic Graph ODE for Opinion Evolution
abstract
Modeling opinion dynamics in social networks has been the focus of multiple disciplines in recent decades. Previous studies have often modeled the opinion dynamics as a discrete and homogeneous process, neglecting its continuous and complex nature. To fill this gap, we propose a Bi-Dynamics Graph Ordinary Differential Equation (BDG-ODE) framework, which models complex opinion dynamics as the result of two dynamical processes: the evolution of positive and negative opinions. The proposed model incorporates a dual opinion encoder that processes positive and negative opinions independently. Furthermore, the temporal opinion evolution is modeled through bidirectional graph ordinary differential equations, which allows the model to capture the changes in opinion in continuous time. We introduce an opinion synthesis decoder that effectively maps the evolved representations from the latent space back to the opinion space. Extensive experiments conducted on six datasets with varying characteristics highlight the superiority of BDG-ODE in forecasting opinion evolution within social networks. It achieved an average accuracy improvement of 23.16%, an average enhancement of 29.46% in the F1 score, and an average mean square error of difference improvement of 90. 30%, and an average correlation coefficient improvement of 45.93%, significantly outperforming eight state-of-the-art models. The code for reproduction is available: https://github.com/tsinghua-fib-lab/Bi-Dynamic-Graph-ODE-for-Opinion-Evolution.
Bowen Duan 0003, Henggang Deng, Jinghua Piao, Huandong Wang, Yue Wang 0007
KDD (1)5
2025 MSA-Net: A Multi-Scale Information Diffusion Model Awaring User Activity Level
abstract
Modeling information diffusion on social networks can be used to guide the prediction and control of information propagation and improve the structure and functionality of social networks. Existing information diffusion prediction methods can predict information diffusion paths and its volume by modeling social network structure and user behavior. However, none of the existing methods take user activity level, which is proved to be critical in modeling the information diffusion process, into account, thus weaken the prediction accuracy. To solve this problem, this article proposes a Multi-Scale Activity Network (MSA-Net) to capture topological and historical affect features for different scales and to predict the users who will be affected at a specific future timestamp with the help of user activity level. Specifically, we first learn the network representation of three scales or levels: micro-scale, meso-scale, and macro-scale, which refers to the user level, intra-community level, and inter-community level, respectively. Then, we introduce the user activity level for each user by using user degree and average number of tweets per time unit to model the individual differences of users to achieve a more accurate prediction. Extensive experiments based on real-world datasets show that MSA-Net achieves a 6.14% improvement in terms of precision, a 6.74% improvement in terms of recall metrics, a 4.26% improvement in terms of F1-score, a 3.15% improvement in terms of MAP, and a 25.78% improvement in terms of NRMSE over the best existing baseline. The code and data are available at https://github.com/tsinghua-fib-lab/MSA-Net.
Yinzhou Tang, Jinghua Piao, Huandong Wang, Yue Wang 0007, Yong Li 0008
ACM Trans. Web4
2024 Social Physics Informed Diffusion Model for Crowd Simulation
abstract
Crowd simulation holds crucial applications in various domains, such as urban planning, architectural design, and traffic arrangement. In recent years, physics-informed machine learning methods have achieved state-of-the-art performance in crowd simulation but fail to model the heterogeneity and multi-modality of human movement comprehensively. In this paper, we propose a social physics-informed diffusion model named SPDiff to mitigate the above gap. SPDiff takes both the interactive and historical information of crowds in the current timeframe to reverse the diffusion process, thereby generating the distribution of pedestrian movement in the subsequent timeframe. Inspired by the well-known social physics model, i.e., Social Force, regarding crowd dynamics, we design a crowd interaction encoder to guide the denoising process and further enhance this module with the equivariant properties of crowd interactions. To mitigate error accumulation in long-term simulations, we propose a multi-frame rollout training algorithm for diffusion modeling. Experiments conducted on two real-world datasets demonstrate the superior performance of SPDiff in terms of both macroscopic and microscopic evaluation metrics. Code and appendix are available at https://github.com/tsinghua-fib-lab/SPDiff.
Jingtao Ding, Yong Li 0008, Yue Wang 0007, Xiao-Ping Zhang 0002
AAAI4
2024 Learning from Hierarchical Structure of Knowledge Graph for Recommendation
abstract
Knowledge graphs (KGs) can help enhance recommendations, especially for the data-sparsity scenarios with limited user-item interaction data. Due to the strong power of representation learning of graph neural networks (GNNs), recent works of KG-based recommendation deploy GNN models to learn from both knowledge graph and user-item bipartite interaction graph. However, these works have not well considered the hierarchical structure of knowledge graph, leading to sub-optimal results. Despite the benefit of hierarchical structure, leveraging it is challenging since the structure is always partly-observed. In this work, we first propose to reveal unknown hierarchical structures with a supervised signal detection method and then exploit the hierarchical structure with disentangling representation learning. We conduct experiments on two large-scale datasets, of which the results well verify the superiority and rationality of the proposed method. Further experiments of ablation study with respect to key model designs have demonstrated the effectiveness and rationality of our proposed model. The code is available at https://github.com/tsinghua-fib-lab/HIKE .
Yingrong Qin, Chen Gao 0001, Shuangqing Wei, Yue Wang 0007, Depeng Jin, Lin Zhang 0001, Dong Li 0016, Jianye Hao, Yong Li 0008
ACM Trans. Inf. Syst.4
2023 Modeling Multi-Grained User Preference in Location Visitation
abstract
Location prediction acts as a fundamental service in today's location-based information platform, which helps users access locations satisfying their demands, improving both user experience and platform profit. Since users with unambiguous demands prefer specific locations while users with compound demands consider first regions and then specific locations, it is necessary to model multi-grained user preferences at different geographical scales. However, most of the existing works concentrate on user preferences at the location-scale only, which can not understand users traveling behaviors thoroughly. In this paper, we propose to model both the fine-grained user preferences at the location scale and the coarsegrained user preferences at the region scale. Specifically, the proposed model harnesses the efficient information extraction power of graph neural networks. Moreover, the proposed geographical calibration method also helps to capture multi-grained user preferences accurately. Experiments on datasets of two very large cities demonstrate the significant performance improvement using our approach over state-of-the-art models. We also conduct experiments to further demonstrate the effectiveness of each component in the proposed model. Source codes of this paper are available at https://github.com/tsinghua-fib-lab/SIGSPATIAL-MMGUP/.
Yingrong Qin, Chen Gao 0001, Zhen Tu, Hongsheng Wu, Shuangqing Wei, Yue Wang 0007, Lin Zhang 0001, Yong Li 0008
SIGSPATIAL/GIS6
2023 Meta-Learning-Based Spatial-Temporal Adaption for Coldstart Air Pollution Prediction
abstract
Air pollution is a significant public concern worldwide, and accurate data‐driven air pollution prediction is crucial for developing alerting systems and making urban decisions. As more and more cities establish their monitoring networks, there is a pressing need for coldstart model training with limited data accumulation in new cities. However, traditional spatial‐temporal modeling and transfer learning schemes have been challenged under this scenario because of insufficient usage of available source data and suboptimal transferring strategy. To address these issues, we propose a meta‐learning‐based spatial‐temporal adaptation solution for coldstart air pollution prediction. Our approach is a model‐agnostic framework that enables a given backbone predictor with adaption ability across different space and time locations. Specifically, it learns a factorization of the available source data distribution and recognizes the target city as one of its components, greatly reducing the data accumulation requirement and providing coldstart capability. Furthermore, we design a novel bidirectional meta‐learner that can simultaneously leverage task embeddings learned from data and features constructed based on prior knowledge. We conduct comprehensive experiments on both synthetic and real‐world air pollution datasets of four distinct pollutants. The results demonstrate that our proposed method achieves a 5.2% lower 24‐hour prediction mean absolute error (MAE) than pretraining and fine‐tuning solutions when facing a new city with only 200 hours of data, which empirically verifies the effectiveness of our approach as a coldstart training solution.
Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
Int. J. Intell. Syst.5
2023 Disentangling Geographical Effect for Point-of-Interest Recommendation
abstract
Point-of-Interest (POI) recommendation has drawn a lot of attention in both academia and industry. It utilizes user check-in data, aiming at recommending unvisited POIs to users. To address the data-sparsity problem, geographical information of POIs is often incorporated into recommender systems. However, most of the existing approaches model geographical impact in an implicit way, in which geographical information is encoded as auxiliary vectors for learning unified representations of users and POIs. Following this paradigm, the embedding of POIs can not reflect geographical similarity directly; thus, an explicit modeling approach is needed as geography is of great importance in POI recommendation. To address challenges in disentangling geographical effect, we proposed a disentangled representation learning method named DIG (short for Disentangled embedding of user Interest and POIs' Geographical information). Aiming at decoupling the geographical factor and the user interest factor thoroughly, we first proposed a geo-constrained negative sampling strategy, which helps to find reliable negative samples for the two factors. Second, a geo-enhanced soft-weighted loss function was proposed to quantify the trade-off between the two factors in loss computation. Extensive experiments have been conducted on two real-world datasets, and results have demonstrated the significant improvement of DIG at 3.92% - 20.32% 3.92% - 20.32% on recall, and 2.53% - 11.48% 2.53% - 11.48% on hit ratio, compared with other state-of-the-art approaches.
Yingrong Qin, Chen Gao 0001, Yue Wang 0007, Shuangqing Wei, Depeng Jin, Lin Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2022 Multi-Task Learning Based Blind Calibration for Low-Cost Air Quality Sensor Deployments
abstract
Air pollution problem has caught much attention globally. In addition to the national air quality monitoring stations deployed by the government, the number of low-cost air quality sensors increases rapidly as a supplement to support fine-grained monitoring. In-field calibration methods are necessary for these low-cost sensor nodes to assure the data quality. However, it is costly to collect enough reference data after deployment to train the in-field calibration model and many sensors even have no synchronized reference in the real application scenarios. To address the above challenge, we propose a multi-task learning based blind calibraiton method for air quality sensors after deployments. Our method introduces not only the reference data of the target location to formulate calibration task, but also reference measurements collected from highly accurate stations already deployed by the government in other geographical locations to formulate prediction task. To utilize the reference measurements which are not in the same location with our target sensors, e.g., in other cities, we combine the proposed calibration task and prediction task under a multi-task learning scheme. The introduced references in other locations alleviate our few-reference challenge. Furthermore, we elaborate on the choices of different tasks to have better effect of the target calibraiton task. Evaluations on the real-world collected datasets show that our proposed algorithm has better calibraiton effect.
Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
SenSys4
2022 MAIC: Metalearning-Based Adaptive In-Field Calibration for IoT Air Quality Monitoring System
abstract
Air pollution has become a global threat to human health. Fine-grained air quality monitoring has attracted much attention in recent years. Low-cost calibrated sensors make it possible for the large-scale deployment of IoT air quality monitoring systems. In practice, the calibration performance degrades after deployment due to the dynamic and diversity of system conditions. However, it is infeasible to collect sufficient in-field reference data to train calibration models for these new conditions. To address themulticonditionandfew-datachallenge, we proposed metalearning-based adaptive in-field calibration (MAIC), a metalearning-based adaptive in-field calibration algorithm. Specifically, MAIC adopts metalearning to learn how to adapt to new conditions quickly. To effectively leverage historical data, we first develop task generation strategies for sensor calibration under this scheme. Then, task-oriented optimization is introduced to train a model with superior adaptability in the offline training phase. Furthermore, an adaptation method is presented to learn the task-specific data distribution without forgetting the metaknowledge, enabling continual learning to utilize the temporal dependencies between multiple conditions. Our evaluations on synthetic and real-world data sets show that MAIC has high robustness and adaptability under multiple complicated conditions. Our proposed method outperforms the state-of-the-art calibration algorithms by 4.23%–29.46% in the real-world deployment data set, but with fewer requirements for the available in-field reference data.
Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
IEEE Internet Things J.5
2022 From Anticipation to Action: Data Reveal Mobile Shopping Patterns During a Yearly Mega Sale Event in China
abstract
The online retail market shows a sharp increase in traffic during holiday sales. The ability to distinguish customers who will likely purchase is critical for provisioning traffic and for providing cost-effective promotions. This paper uniquely studies the browsing and purchasing behaviors of online shoppers during a yearly sale event in China, the world’s largest online marketplace. Based on 31 million action logs gathered from wide residential areas, we characterize the steps leading to purchases and determine their precursors. We investigate the effect of time (e.g., date, time of date), environment (e.g., platform, viewed category), and action (e.g., session time, clicks, sequence) on purchases. Action cues from shopping behaviors can be used for early detection. While most shoppers start with strong intentions to purchase, yet the moment of ordering comes rather impulsively within 30 seconds to several minutes of browsing. The predictive accuracy reaches as a high AUC of 0.924. The findings in this paper provide an understanding of traffic during mega sale events that can help online shops plan and provide a better user experience for upcoming shopping festivals.
Muzhi Guan, Meeyoung Cha, Yue Wang 0007, Yong Li 0008
IEEE Trans. Knowl. Data Eng.3
2021 A Variational Bayesian Approach for Fast Adaptive Air Pollution Prediction
abstract
Air pollution problem has been a worldwide environmental concern in recent years. Accurate air pollution prediction can effectively protect public health and help government decisions. However, strong instability and frequent pattern shift in air pollution data challenges the conventional time series prediction paradigm and attracts interests in adaptive prediction algorithm. Recent progress in deep learning community, such as attention mechanism and meta learning algorithm, both use handcrafted adaptive strategy and lack sufficient usage of supporting observation data. In this paper, we adopt a variational Bayesian approach to enable fast adaption ability for a given air pollution predictor, which can make better use of recent observation data and adaptively inference task-specific parameters to achieve better adaption performance. Specifically, without explicitly designing a heuristic adaptive procedure, we formulate the adaptive prediction as a maximizing conditional likelihood problem on a generative graphic model, where a variational approximation to the intractable likelihood is further derived for end-to-end training. Experiments on real-world air pollution datasets show significant improvements of the proposed method compared to previous works.
Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
IEEE BigData5
2020 Poster Abstract: Robust Calibration for Low-Cost Air Quality Sensors using Historical Data
abstract
As pollution problems become increasingly prominent nowadays, urban air quality monitoring has attracted more and more attention. In recent years, sensing systems based on low-cost sensors are proposed to achieve fine-grained monitoring with larger amount of deployment as supplyment to conventional monitoring stations. Calibration is critical to guarantee the accuracy and consistency of these sensing systems to fight against sensor drift. While conventional field calibration approaches often rely on real-time data from a nearby standard station, they are not applicable to low-cost sensors which cannot receive the latest reference data from nearby stations after deployment. In reality, it is very difficult for sensors to get access to nearby standard stations deployed sparsely. To reduce the dependency on real-time and nearby reference data, we present a Robust Calibration approach based on Historical data (RCH) for the low-cost air pollution sensor calibration. Our method corrects the sensor drift by adapting sensitivity and offset based on estimating the probability distribution of pollutant's concentration. Experiments with real-world NO2data in Foshan, China show that our proposed method acheives close performance to conventional field calibration methods but addresses above challenges. Moreover, our method can use historical data collected from the sensors in more distant geographic locations than the compared method.
Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
IPSN4
2020 Consistent State Updates for Virtualized Network Function Migration
abstract
Combining Network Functions Virtualization (NFV) with Software-Defined Networking (SDN) is an emerging and promising solution to provide scalable and elastic network control and service. In such a system, virtualized Network Functions (NFs) need to be consistently migrated from one instance to another for various purposes, such as resource optimization, fault tolerance, load balancing, etc. These migrations involve simultaneously coordinating updates to the NF state and SDN forwarding state. To solve this problem, we design two consistent NF state update schemes: a controller-forwarding based scheme and a tagging-based scheme. Through analysis of the update process, we demonstrate that they both guarantee loss-free and order-preserving migrations. We further implement a prototype and carry out experiments with diverse traffic settings. Results demonstrate that the controller-forwarding based solution achieves 77 percent migration time compared with the state-of-the-art solution OpenNF, while correcting an error of it. Moreover, the tagging-based solution not only achieves 4.4 percent migration time, but also reduces up to 75 percent controller overhead compared with OpenNF at the cost of adding a tag in the unused fields of packet header.
Yujie Liu 0010, Jiaqiang Liu, Yong Li 0008, Haoyu Song 0001, Yue Wang 0007
IEEE Trans. Serv. Comput.6
2019 MSSTN: Multi-Scale Spatial Temporal Network for Air Pollution Prediction
abstract
Air pollution has become an important factor constraining city development and threatening public health in recent years. Air pollution prediction has been considered as the key part for the early warning of pollution event. Considering the multi-scale nature of geo-sensory data such as air pollution signal, in this paper we adopt a multi-level graph data structure for better utilization of multi-scale spatio-temporal information. We further present a novel deep convolutional neural network model, named Multi-Scale Spatial Temporal Network (MSSTN), for the learning task on this data structure. The MSSTN is specially designed to better discover multi-scale spatial temporal patterns and their high-level interactions, by explicitly using multi-scale neural network structure in both spatial and temporal component. We conduct extensive experiments and ablation studies on Urban Air Pollution Datasets in North China, where the MSSTN can make hourly PM2.5 concentration predictions jointly for a number of cities. And our results shows an outstanding prediction accuracy as well as high computational efficiency compared to existing works.
Yue Wang 0007, Lin Zhang 0001
IEEE BigData2
2019 λOpt: Learn to Regularize Recommender Models in Finer Levels
abstract
Recommendation models mainly deal with categorical variables, such as user/item ID and attributes. Besides the high-cardinality issue, the interactions among such categorical variables are usually long-tailed, with the head made up of highly frequent values and a long tail of rare ones. This phenomenon results in the data sparsity issue, making it essential to regularize the models to ensure generalization. The common practice is to employ grid search to manually tune regularization hyperparameters based on the validation data. However, it requires non-trivial efforts and large computation resources to search the whole candidate space; even so, it may not lead to the optimal choice, for which different parameters should have different regularization strengths. In this paper, we propose a hyperparameter optimization method, lambdaOpt, which automatically and adaptively enforces regularization during training. Specifically, it updates the regularization coefficients based on the performance of validation data. With lambdaOpt, the notorious tuning of regularization hyperparameters can be avoided; more importantly, it allows fine-grained regularization (i.e. each parameter can have an individualized regularization coefficient), leading to better generalized models. We show how to employ lambdaOpt on matrix factorization, a classical model that is representative of a large family of recommender models. Extensive experiments on two public benchmarks demonstrate the superiority of our method in boosting the performance of top-K recommendation.
Bei Chen 0008, Xiangnan He 0001, Chen Gao 0001, Yong Li 0008, Jian-Guang Lou, Yue Wang 0007
KDD7
2019 Understanding air pollution patterns in city based on minute-level event detection: poster abstract
abstract
Air pollution is a serious urban problem that threatens human health. Therefore, fine-grained pollution events detection has become a concerned issue for environmental management. Algorithms in previous studies identify pollution events as uptrend intervals at hour level. However, a significant part of pollution events caused by traffic and industry can be brief but frequent, which may be neglected under traditional coarse-grained detection. In this paper, we propose a fine-grained analysis of air pollution pattern based on minute-level event detection. Over the real-world deployment in Foshan, these events are analyzed according to their geographical contexts and temporal features. Results show insightful findings and this case study provides a practical reference for government inspection and pollution control.
Rui Ma 0014, Xinyu Liu 0003, Yue Wang 0007, Lin Zhang 0001
SenSys5
2019 Enhanced air quality inference with mobile sensing attention mechanism: poster abstract
abstract
Mobile sensor networks are widely deployed for air quality monitoring. However, fine-grained pollution inference based on these systems is challenging. Specifically, diverse geospatial attributes in urban areas bring great spatial variations of the pollution field. Besides, the preprocessing on raw samples, such as discretization and averaging, leads to the lost of fine-grained information of mobile sensing. In this paper, we propose an inference algorithm with the attention mechanism to better capture high-frequency information in the pollution field. Furthermore, we introduce the sensing gradients in the attention network to utilize the high-granularity information from the mobile sensors. Evaluations on real-world dataset show that our model outperforms the state-of-the-art method by 13.15% ~ 27.04%.
Yue Wang 0007, Rui Ma 0014, Lin Zhang 0001
SenSys2
2018 Guiding the Data Learning Process with Physical Model in Air Pollution Inference
abstract
The surveillance of air pollution is becoming a highly concerned issue for city residents and urban administrators. Fixed air quality stations as well as mobile gas sensors have been deployed for air quality monitoring but with sparse observations over the entire temporal-spatial space. Therefore, an inference algorithm is essential for comprehensive fine-grained air pollution sensing. Conventional physically-based models can hardly be applied to all the scenarios, while pure data-driven methods suffer from sampling bias and overfitting problems. This paper presents a hybrid algorithm for air pollution inference by guiding the data learning process with physical model. The quantitative combination of knowledge from observed dataset and a discretized convective-diffusion model is performed within a multi-task learning scheme. Evaluations show that, benefited from physical guidance, our hybrid method obtains higher extrapolation ability and more robustness, achieving the same performance with 1/8 sample amount and obtaining 31.9% less error in noisy synthesized environment. In a real-world deployment in Tianjin, our algorithm outperforms the pure data-driven model with 9.69% less inference error over a 9-day PM2.5data collection.
Rui Ma 0014, Xiangxiang Xu 0001, Yue Wang 0007, Hae Young Noh, Pei Zhang 0001, Lin Zhang 0001
IEEE BigData3
2018 Learning-to-Ask: Knowledge Acquisition via 20 Questions
abstract
Almost all the knowledge empowered applications rely upon accurate knowledge, which has to be either collected manually with high cost, or extracted automatically with unignorable errors. In this paper, we study 20 Questions, an online interactive game where each question-response pair corresponds to a fact of the target entity, to acquire highly accurate knowledge effectively with nearly zero labor cost. Knowledge acquisition via 20 Questions predominantly presents two challenges to the intelligent agent playing games with human players. The first one is to seek enough information and identify the target entity with as few questions as possible, while the second one is to leverage the remaining questioning opportunities to acquire valuable knowledge effectively, both of which count on good questioning strategies. To address these challenges, we propose the Learning-to-Ask (LA) framework, within which the agent learns smart questioning strategies for information seeking and knowledge acquisition by means of deep reinforcement learning and generalized matrix factorization respectively. In addition, a Bayesian approach to represent knowledge is adopted to ensure robustness to noisy user responses. Simulating experiments on real data show that LA is able to equip the agent with effective questioning strategies, which result in high winning rates and rapid knowledge acquisition. Moreover, the questioning strategies for information seeking and knowledge acquisition boost the performance of each other, allowing the agent to start with a relatively small knowledge set and quickly improve its knowledge base in the absence of constant human supervision.
Bei Chen 0008, Xuguang Duan, Jian-Guang Lou, Yue Wang 0007, Wenwu Zhu 0001
KDD5
2018 A Hybrid Air Pollution Reconstruction by Adaptive Interpolation Method
abstract
Air pollution in a city is the major environmental risk to health. Mobile sensing has become a popular solution in recent years. However, it still suffers from problems such as lack of data and high system uncertainty. This is because that the data amount and distribution vary over time. To address the problems, this paper combines two classic data driven models -- Kriging and Inverse Distance Weighting (IDW). We adopt the Random Forest Algorithm (RF) to adaptively choose the more accurate models (Kriging or IDW) according to the features we extracted. The experiment based on real world testbed shows our adaptive method achieves up to 30.6% error reduction.
Rui Ma 0014, Yue Wang 0007, Lin Zhang 0001
SenSys5
2018 Chernoff information between Gaussian trees
Shuangqing Wei, Yue Wang 0007
Inf. Sci.3
2016 Co-location social networks: Linking the physical world and cyberspace
abstract
Various dedicated web services in the cyberspace, e.g., social networks, e-commerce, and instant communications, play a significant role in people's daily-life. Billions of people around the world access them through multiple online identifiers (IDs), and interact with each other in both the cyberspace and the physical world. These two kinds of interactions are highly relevant to each other. In order to link between the cyberspace and the physical world, we propose a new type of social network, i.e., co-location social network (CLSN). A CLSN contains online IDs describing people's online presence and offline interactions when people come across each other. By analyzing real data collected from a mainstream ISP in China, which contains 32.7 million IDs across most popular web services, we build a large-scale CLSN, and evaluate its unique properties. The results verify that the CLSN is quite different from existing online and offline social networks in terms of different classic graph metrics. This paper is the first research to study CLSN at scale and paves the way for future studies of this new type of social network.
Huandong Wang, Yong Li 0008, Yang Chen 0001, Yue Wang 0007, Depeng Jin
ASONAM4
2016 Chernoff information of bottleneck Gaussian trees
abstract
In this paper, our objective is to find out the determining factors of Chernoff information in distinguishing a set of Gaussian trees. In this set, each tree can be attained via a subtree removal and grafting operation from another tree. This is equivalent to asking for the Chernoff information between the most-likely confused, i.e. “bottleneck”, Gaussian trees, as shown to be the case in ML estimated Gaussian tree graphs lately. We prove that the Chernoff information between two Gaussian trees related through a subtree removal and grafting operation is the same as that between two three-node Gaussian trees, whose topologies and edge weights are subject to the underlying graph operation. In addition, such Chernoff information is shown to be determined only by the maximum generalized eigenvalue of the two Gaussian covariance matrices. The Chernoff information of scalar Gaussian variables as a result of linear transformation (LT) of the original Gaussian vectors is also uniquely determined by the same maximum generalized eigenvalue. What is even more interesting is that after incorporating the cost of measurements into a normalized Chernoff information, Gaussian variables from LT have larger normalized Chernoff information than the one based on the original Gaussian vectors, as shown in our proved bounds.
Shuangqing Wei, Yue Wang 0007
ISIT3
2015 United Channel Assignments in Residential Environments
abstract
Residential wireless networks have grown rapidly in the past decade. Meanwhile, dense deployments and autonomous managements of home Access Points (APs) greatly increase channel congestion levels and degrade the user experience. To address such problems, this paper proposes an architecture in residential environments called Wi-Fi Union (WU), where home APs can voluntarily join WU and become "member APs". WU helps member APs decrease their congestion levels by assigning channels in a coordinated manner with incentive considerations. First, we propose congestion level metric normalized airtime which can be passively and independently measured by member APs. Normalized airtime is further used to classify APs into heavily congested APs and lightly congested APs. A tabu search based channel assignment algorithm is presented which can decrease the congestion level of heavily congested member APs, and guarantee lightly congested member APs to be still lightly congested. Extensive NS-3 simulations driven by actual Wi-Fi data show that the united channel assignments has a 1.5 times throughput than that in the default setting on average.
Chunxiao Jiang, Yue Wang 0007, Jiannong Cao 0001
GLOBECOM3
2015 Detection of graph structures via communications over a multiaccess Boolean channel
abstract
In this paper, we propose a novel model to study the efficiency of detecting latent connection relationships, represented by a given set of graphs, among N users. A subset of active nodes transmit following a common codebook over a multiple access Boolean channel. To maximize the error exponent of the structure detection, we formulate an optimization problem whose objective is to max-minimize the pairwise Chernoff information, and the constraint is a probability simplex due to the users' multiple dependency relationships, which are further shown to have close relationship to the internal connectivity of graphs. Case studies are provided to show certain inherent properties of the optimal solution. In addition, we present a particular case with two equally weighted complementary Paley graphs of prime square order, whose optimal solution for the codebook is proved and the resulting exponent is shown to be O(1/N). The case study demonstrates how the fundamental graph discrepancy property affects the solution to the problem.
Shuhang Wu, Shuangqing Wei, Yue Wang 0007, Ramachandran Vaidyanathan, Xiqin Wang
ISIT3
2015 Optimal scheduling for multi-flow update in Software-Defined Networks
Yujie Liu 0010, Yong Li 0008, Yue Wang 0007
J. Netw. Comput. Appl.3
2015 Partition Information and its Transmission Over Boolean Multi-Access Channels
abstract
In this paper, we propose a novel reservation system to study partition information and its transmission over a noise-free Boolean multiaccess channel. The objective of transmission is not to restore the message, but to partition active users into distinct groups so that they can, subsequently, transmit their messages without collision. We first calculate (by mutual information) the amount of information needed for the partitioning without channel effects, and then propose two different coding schemes to obtain achievable transmission rates over the channel. The first one is the brute force method, where the codebook design is based on centralized source coding; the second method uses random coding, where the codebook is generated randomly and optimal Bayesian decoding is employed to reconstruct the partition. Both methods shed light on the internal structure of the partition problem. A novel formulation is proposed for the random coding scheme, in which a sequence of channel operations and interactions induces a hypergraph. The formulation intuitively describes the transmitted information in terms of a strong coloring of this hypergraph. An extended Fibonacci structure is constructed for the simple, but nontrivial, case with two active users. A comparison between these methods and group testing is conducted to demonstrate the potential of our approaches.
Shuhang Wu, Shuangqing Wei, Yue Wang 0007, Ramachandran Vaidyanathan
IEEE Trans. Inf. Theory3
2015 Asymptotic Error Free Partitioning Over Noisy Boolean Multiaccess Channels
abstract
In this paper, we consider the problem of partitioning active users in a manner that facilitates multi-access without collision. The setting is of a noisy, synchronous, Boolean, and multi-access channel, where K active users (out of a total of N users) seek channel access. A solution to the partition problem places each of the N users in one of K groups (or blocks), such that no two active nodes are in the same block. We consider a simple, but non-trivial and illustrative, case of K = 2 active users and study the number of steps T used to solve the partition problem. By random coding and a suboptimal decoding scheme, we show that for any T ≥ (C1+ ξ1) log N, where C1and ξ1are positive constants (independent of N), and where ξ1can be arbitrary small, the partition problem can be solved with error probability Pe(N)→ 0, for large N. Under the same scheme, we also bound T from the other direction, establishing that, for any T ≤ (C2- ξ2) log N, the error probability Pe(N)→ 1 for large N; again, C2and ξ2are constants, and ξ2can be arbitrarily small. These bounds on the number of steps are lower than the tight achievable lower bound in terms of T ≥ (Cg+ ξ) log N for group testing (in which all active users are identified, rather than just partitioned). Thus, partitioning may prove to be a more efficient approach for multi-access than group testing.
Shuhang Wu, Shuangqing Wei, Yue Wang 0007, Ramachandran Vaidyanathan
IEEE Trans. Inf. Theory3
2014 Achievable partition information rate over noisy multi-access Boolean channel
abstract
In this paper, we formulate a novel problem to quantify the amount of information transferred to partition active users who transmit following a common codebook over noisy Boolean multi-access channels. The objective of transmission is to ultimately let each active user aware of its own group only, not others. To solve the problem, we propose a novel framework by considering the decoding as a process of removing hyperedges of a complete hypergraph. For a particular, but non-trivial, case with two active users, an achievable bound for the defined partition information rate is found by using strong typical set decoding, as well as a large deviation technique for an induced Markov chain.
Shuhang Wu, Shuangqing Wei, Yue Wang 0007, Ramachandran Vaidyanathan
ISIT3
2014 Multi-AS cooperative incoming traffic engineering in a transit-edge separate internet
Yue Wang 0007, Dan Pei
Comput. Networks2
2014 Track-to-Track Association for Biased Data Based on the Reference Topology Feature
abstract
In this letter, we propose a novel track-to-track association (TTTA) algorithm based upon the reference topology (RET) feature, which is insensitive to sensor biases. A rigorous mathematical definition of RET is presented. The insensitivity of RET to sensor biases is analyzed theoretically. In order to construct the association cost matrix, we make use of the optimal subpattern assignment (OSPA) metric to measure the distance between two RETs. Simulation results demonstrate the advantages of the proposed algorithm.
Wei Tian 0006, Yue Wang 0007, Xiuming Shan, Jian Yang 0011
IEEE Signal Process. Lett.2
2013 A topology control algorithm based on D-region fault tolerance
Ruozi Sun, Yue Wang 0007, Xiuming Shan, Yong Ren 0001
Sci. China Inf. Sci.2
2013 A DHT-based fast handover management scheme for mobile identifier/locator separation networks
Yue Wang 0007, Yong Ren 0001
Sci. China Inf. Sci.3
2010 Mobility Prediction in Cellular Network Using Hidden Markov Model
abstract
In next generation networks, mobile communication calls for service with higher quality, which brings new challenge for mobility management. Thereinto, utilization and improvement of mobility prediction helps for preserving resource and providing better performance. So this paper aims to propose a theoretical and factual method to perform mobility prediction in cellular network. By analyzing the demand and character of this kind of personal mobility prediction in large spacial and temporal scale, it is concluded that Hidden Markov Model fits for system modeling. However, classical HMM algorithm will meet with numerical calculation problem when adopted to practical communication system. An improved algorithm is put forward to overcome possible calculating defects. Three different scenarios are set to testify HMM's efficiency and accuracy, using factual measurement data in cellular network.
Hongbo Si, Yue Wang 0007, Xiuming Shan
CCNC2
2010 A PCA-based approach for exploring space-time structure of urban mobility dynamics
abstract
Understanding of urban mobility dynamics benefits both aggregate human mobility in wireless communications, and the planning and provision of urban facilities and services. Due to the high penetration of cell phones, the cellular networks provide information for urban dynamics with large spatial extent and continuous temporal coverage. In this paper, a novel approach is proposed to explore the space-time structure of urban dynamics, based on the original data collected by cellular networks in a southern city of China, recording population distribution by dividing the city into thousands of pixels. By applying principal component analysis, the intrinsic dimensionality is revealed. The structure of all the pixel population variations could be well captured by a small set of eigen pixel population variations. According to the classification of eigen pixel population variations, each pixel population variation can be decomposed into three constitutions: deterministic trends, short-lived spikes, and noise. Moreover, the most significant eigen pixel population variations are utilized in the applications of forecasting and anomaly detection.
Yue Wang 0007, Hongbo Si, Xia Mao, Xiuming Shan
IWCMC2
2006 A Novel Fuzzy Pattern Recognition Data Association Method for Biased Sensor Data
abstract
Data association is one of the most important problems in multiple-sensor multiple-target tracking systems. A novel approach for biased-data association based on patterns extracted from topologies of measurements is presented in this paper. The introduction of patterns reveals a new paradigm for exploiting redundant spatial information in association. Fuzzy pattern recognition method has been applied to analyze the similarity of target patterns, and the criterion of association can be set up accordingly. By virtue of the inherent character of patterns, the proposed pattern-based association algorithm is robust to large registration errors between systems. Furthermore, the association results can provide feedback information for sensor registration. Simulation results show that the proposed algorithm is effective and feasible in solving association problems under the conditions of sensor bias
Yue Wang 0007, Xiuming Shan
FUSION2
2005 A proxy-based framework to enhance user level performance of GPRS/UMTS networks
abstract
This paper presents a proxy-based framework for Internet access through GPRS/UMTS networks. The proxy, deployed at the network boundary, is able to employ some scheduling strategies to assist the wireless network in charge of incoming traffic according to system capacity and user demand. In particular, we investigate the system performance on flow level and consider the impact of some user behaviors. The objective is to maximize utilization of the network by exploiting the property of delay tolerance, and to improve user-level performance such as response time. In addition, the proxy splits TCP connections into two halves, the wired and wireless sides. The TCP stacks on the wireless-facing side of the proxy can be modified to enhance the performance over the wireless link, while keeping compatibility with the conventional TCP implementations on mobile hosts. The effectiveness of this approach is illustrated by analysis and simulations.
Yue Wang 0007, Junxiu Lu, Yong Ren 0001, Xiuming Shan, Yong-Hua Song
WCNC1