Jia Xu 0005

dblp:95/3616-5 · DBLP profile ↗
← Back
41ranked-venue papers
20as first author
22since 2021 · last 2026
0000-0003-4061-8262ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021Computer networks · 9 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 TQR : Modelling user temporal preference for effective question routing
Jia Xu 0005, Zhengkai Li, Hengyu Liu 0001, Zulong Chen, Ning Wang 0026, Tiancheng Zhang 0001
Neural Networks1
2025 HeterRec: Heterogeneous Information Transformer for Scalable Sequential Recommendation
abstract
Transformer-based sequential recommendation (TSR) models have shown superior performance in recommendation systems, where the quality of item representations plays a crucial role. Classical representation methods integrate item features using concatenation or neural networks to generate homogeneous representation sequences. While straightforward, these methods overlook the heterogeneity of item features, limiting the transformer's ability to capture fine-grained patterns and restricting scalability. Recent studies have attempted to integrate user-side heterogeneous features into item representation sequences, but item-side heterogeneous features, which are vital for performance, remain excluded. To address these challenges, we propose a Heterogeneous Information Transformer model for Sequential Recommendation (HeterRec), which incorporates Heterogeneous Token Flatten Layer (HTFL) and Hierarchical Causal Transformer Layer (HCT). Our HTFL is a novel item tokenization method that converts items into a heterogeneous token set and organizes these tokens into heterogeneous sequences, effectively enhancing performance gains when scaling up the model. Moreover, HCT introduces token-level and item-level causal transformers to extract fine-grained patterns from the heterogeneous sequences. Experiments on offline and online datasets show that the HeterRec model achieves superior performance.
Hao Deng 0011, Haibo Xing, Kanefumi Matsuyama, Yulei Huang, Jinxin Hu, Hong Wen 0002, Jia Xu 0005, Zulong Chen, Yu Zhang 0206, Xiaoyi Zeng, Jing Zhang 0037
SIGIR7
2025 Coherence-Enhanced language representation learning for sequential Recommendations
Jia Xu 0005, Haiwei Wu
Knowl. Based Syst.1
2024 Tracing Human Stress From Physiological Signals Using UWB Radar
abstract
Stress tracing is an important research domain that supports many applications, such as health care and stress management; and its closest related works are derived from stress detection. However, these existing works cannot well address two important challenges facing stress detection. First, most of these studies involve asking the users to wear physiological sensors to detect their stress states, which has a negative impact on the user experience. Second, these studies have failed to effectively utilize the multimodal physiological signals, which results in less satisfactory detection results. This article formally defines the stress tracing problem, which emphasizes the continuous detection of human stress states. A novel deep stress tracing (DST) method, named DST, is presented. Note that, DST proposes tracing human stress based on the physiological signals collected by a noncontact ultrawideband radar, which is more friendly to users when collecting their physiological signals. In DST, a signal extraction module is carefully designed at first to robustly extract the multimodal physiological signals from the raw RF data of the radar, even in the presence of body movement. Afterward, a multimodal fusion module is proposed in DST to ensure that the extracted multimodal physiological signals can be effectively fused and utilized. Extensive experiments are conducted on the three real-world data sets, including one self-collected data set and two publicity data sets. Experimental results show that the proposed DST method significantly outperforms all the baselines in terms of tracing human stress states. On average, DST averagely provides a 6.31% increase in detection accuracy on all the data sets, compared with the best baselines.
Jia Xu 0005, Teng Xiao, Zhe Chen 0015, Chao Cai 0001, Yang Zhang 0025, Zehui Xiong
IEEE Internet Things J.1
2024 Ske-Fi: Estimating Hand Poses via RF Vision Under Low Contrast and Occlusion
abstract
Hand pose estimation (HPE), which aims to identify and recover the keypoints of a hand, is essential to many potential applications. Conventional computer vision (CV) methods extract visible features from images or videos captured by cameras. However, they are heavily affected by low image contrast, fail to work under occluded scenarios, and inevitably incur privacy concerns. Fortunately, CV leveraging widely available radio frequency (RF) signals (also known as RF vision) can fully address the problem with much lower computational complexity. In this article, we propose Ske-Fi as an avatar of hand pose estimation (HPE) enabled by RF vision, which uses the emerging impulse radio ultrawide band (IR-UWB) available on smart devices (e.g., Apple air tag) to sense the reflected RF signals of a hand to extract the hand skeleton features for pose estimation. Whereas Ske-Fi is apparently immune to low contrast and occlusion, its substantially reduced resolution provided by IR-UWB signal makes the resulting RF image incomprehensible by human eyes and thus negating offline labeling. To address the challenge, Ske-Fi involves a deep complex-valued neural network Ske-Net trained via a cross-modal supervision framework; it uses a synchronized camera assisted by a state-of-the-art vision network as a teacher to teach Ske-Net as a student in independently performing HPE afterward. Furthermore, for occlusion cases, Ske-Fi adopts an adversarial learning scheme to distill HPE features regardless of diversified occlusions. Our extensive evaluations evidently demonstrate that Ske-Fi outperforms conventional CV solutions which achieves a comparable HPE accuracy under normal circumstances and maintains this accuracy under adverse scenarios.
Jia Xu 0005, Zhe Chen 0015, Jun Luo 0001
IEEE Internet Things J.1
2024 Collaboration or Competition: An Infomax-Based Period-Aware Transformer for Ticket-Grabbing Prediction
abstract
Helping users to grab train tickets during a travel peak is a very important service provided by many mainstream online travel platforms (OTPs), e.g., booking.com, Ctrip.com, and Alibaba Fliggy, which greatly enriches the experience for platform users. To optimize such train ticket-grabbing service, a vital accompanying task is to predict the train ticket-grabbing success rates for users during their train ticket-grabbing process to help them make decisions. Although many endeavours have been made towards the traffic prediction problem, none of them was dedicated to solving the ticket-grabbing issue. That is, prior methods ignored the unique properties exhibited in the ticket-grabbing scenario, such as the specific spatial relationship between stations and trains, the collaboration and competition relationships between different routes, and the temporal periodic pattern in ticket-grabbing. In this paper, we propose a novel Infomax-based Period-aware Transformer (IPT) tailored for predicting the success rate of train ticket-grabbing that will be displayed on OTPs, which is to our best knowledge the first attempt along this line. IPT contains three modules: i) a multi-view node embedding module, which serves to model the special spatial relationships between stations and trains by employing the intra- and inter-graph aggregation layers; ii) an infomax-based graph representation learning module, which aims to learn a high-level node embedding by training a discriminator to distinguish different types of edges in the route graph; iii) a period-aware Transformer module, which intends to discover the ticket-grabbing temporal periodic dependencies by designing a periodic activation function. Extensive offline and online evaluations on a real-world dataset show that IPT substantially outperforms state-of-the-art baselines.
Wanjie Tao, Jia Xu 0005, Qun Dai, Hong Wen 0002, Zulong Chen
IEEE Trans. Intell. Transp. Syst.3
2023 PlanRanker: Towards Personalized Ranking of Train Transfer Plans
abstract
Train transfer plan ranking has become the core business of online travel platforms (OTPs), due to the flourish development of high- speed rail technology and convenience of booking trains online. Currently, mainstream OTPs adopt rule-based or simple preference- based strategies to rank train transfer plans. However, the insuf- ficient emphasis on the costs of plans and the negligence of con- sidering reference transfer plans make these existing strategies less effective in solving the personalized ranking problem of train transfer plans. To this end, a novel personalized deep network (Plan- Ranker) is presented in this paper to better address the problem. In PlanRanker, a personalized learning component is first proposed to capture both of the query semantics and the target transfer plan- relevant personalized interests of a user over the user's behavior log data. Then, we present a cost learning component, where both of the price cost and the time cost of a target transfer plan are emphasized and learned. Finally, a reference transfer plan learning component is designed to enable the whole framework of PlanRanker to learn from reference transfer plans which are pieced together by plat- form users and thus reflect the wisdom of crowd. PlanRanker is now successfully deployed at Alibaba Fliggy, one of the largest OTPs in China, serving millions of users every day for train ticket reservation. Offline experiments on two production datasets and a country-scale online A/B test at Fliggy both demonstrate the superiority of the proposed PlanRanker over baselines.
Jia Xu 0005, Wanjie Tao, Zulong Chen, Jin Huang 0001, Hong Wen 0002, Shenghua Ni, Qun Dai, Yu Gu 0002
KDD1
2023 Improving knowledge tracing via a heterogeneous information network enhanced by student interactions
Jia Xu 0005, Teng Xiao
Expert Syst. Appl.1
2023 OCro: Open-Set Cross-Domain Human Activity Recognition Based on Radio Frequency
abstract
With the help of machine learning, models are trained to recognize human activity based on radio frequency signals, which are widely used in human-computer interaction, healthcare, etc. Cross-domain human activity recognition (HAR) aims to adapt a model trained in a specific source domain (including environment and user) for another target domain. Most existing cross-domain recognition methods are proposed under an ideal closed-set assumption, which means the training set and the testing set contain the same categories of human activities. However, when a model is applied in practice for HAR, it often encounters new categories of activities which are not contained in the training set. Under such open-set condition, the traditional closed-set cross-domain recognition model usually incorrectly identifies the new activity as a known activity, which decline the recognition accuracy. In this article, a model is proposed for open-set cross-domain HAR. The model is established based on generative adversarial network, and a unique generation module is designed to generate confusing samples whose features are similar to known classes. Thanks to such design, the proposed model can autonomously select appropriate unlabeled samples under the open-set condition to improve the open-set recognition ability of the model in the target domain. Extensive experiments are conducted based on four real data sets, one is collected by ourselves and the other three are public. The results show that the proposed model outperforms other state-of-the-art methods in the open-set and the cross-domain contexts.
Shuyu Luo, Jia Xu 0005, Zhe Chen 0015
IEEE Internet Things J.3
2023 AdaML: An Adaptive Meta-Learning model based on user relevance for user cold-start recommendation
Jia Xu 0005, Hongming Zhang 0005, Xin Wang 0173
Knowl. Based Syst.1
2023 Leveraging user itinerary to improve personalized deep matching at Fliggy
Jia Xu 0005, Zulong Chen, Wanjie Tao, Ziyi Wang 0008, Detao Lv, Chuanfei Xu
VLDB J.1
2022 SQL-DP: A Novel Difficulty Prediction Framework for SQL Programming Problems
Jia Xu 0005
EDM1
2022 ODNET: A Novel Personalized Origin-Destination Ranking Network for Flight Recommendation
abstract
Origin-Destination recommendation that recom-mends personalized origin city (O) and destination city (D) of flight itinerary is of great value for both Online Travel Platforms (OTPs) and users. Existing studies on next location recommendation propose to model the sequential regularity of users' check-in location sequences, but cannot well solve two new challenges facing OTPs, namely the necessity of exploring O&D and learning O&D as a whole. To this end, we propose a novel personalized Origin-Destination ranking NETwork (ODNET) for flight recommendation. In particular, a heterogeneous spatial graph (HSG) which models historical interactions between users and cities is designed at first. HSG is then deployed in ODNET to identify user preference Os and Ds by exploring the neighbor-hood information in HSG. To cope with the second challenge, the idea of multi-task learning is employed by ODNET to learn$O$and$D$jointly so as to capture their correlations. Moreover, temporal information of Os and Ds are also considered to further improve the accuracy of origin-destination recommendation. An offline experiment on multiple real-world datasets and an online A/B test both show the superiority of ODNET towards the state-of-the-art methods. Further, the implementation and deployment details of the proposed ODNET at Fliggy, one of the most popular OTPs in China, are also described. ODNET has now been successfully applied to provide high-quality flight recommendation service at Fliggy, serving tens of millions of users.
Jia Xu 0005, Jin Huang 0001, Zulong Chen, Wanjie Tao, Chuanfei Xu
ICDE1
2022 G2NET: A General Geography-Aware Representation Network for Hotel Search Ranking
abstract
Hotel search ranking is the core function of Online Travel Platforms (OTPs), while geography information of location entities involved in it plays a critically important role in guaranteeing its ranking quality. The closest line of works to the hotel search ranking problem is thus the next POI (or location) recommendation problem, which has extensive works but fails to cope with two new challenges, i.e., consideration of two more location entities and effective utilization of geographical information, in a hotel search ranking scenario. To this end, we propose a General Geography-aware representation NETwork (G2NET for short) to better represent geography information of location entities so as to optimize the hotel search ranking. In G2NET, to address the first challenge, we first propose the concept of Geography Interaction Schema (GIS) which is a meta template for representing the arbitrary number of location entity types and their interactions. Then, a novel geography interaction encoder is devised providing general representation ability for an instance of GIS, followed by an attentive operation that aggregates representations of instances corresponding to all historically interacted hotels of a user in a weighted manner. The second challenge is handled by the combined application of three proposed geography embedding modules in G2NET, each of which focuses on computing embeddings of location entities based on a certain aspect of geographical information of location entities. Moreover, a self-attention layer is deployed in G2NET, to capture correlations among historically interacted hotels of a user which provides non-trivial functionality of understanding the user's behaviors. Both offline and online experiments show that G2NET outperforms the state-of-the-art methods. G2NET has now been successfully deployed to provide the high-quality hotel search ranking service at Fliggy, one of the most popular OTPs in China, serving tens of millions of users.
Jia Xu 0005, Zulong Chen, Mingyuan Tao, Liangyue Li
KDD1
2022 Modal decomposition-based hybrid model for stock index prediction
Yating Shu, Jia Xu 0005, Qinjuan Wu
Expert Syst. Appl.3
2022 Misbehavior Detection in Vehicular Ad Hoc Networks Based on Privacy-Preserving Federated Learning and Blockchain
abstract
As an irreversible trend, connected vehicles have become increasingly more popular. They depend on the generation and sharing of data between vehicles to improve safety and efficiency of the transportation system. Due to the open feature of the vehicular ad hoc network (VANET), it is possible for dishonest and misbehaving vehicles to disrupt traffic by transmitting false information. In recent years, misbehavior detection systems have been developed to detect the malicious behaviour, and machine learning methods have been employed to make the detection more accurately. However, existing misbehavior detection systems typically require a single entity (e.g., a central server) for centralized data collection and training. Model updates are restricted due to data privacy and high overhead of data communication, which reduces the defensive capability of misbehavior detection systems. In this paper, we propose a blockchain-based federated learning scheme to detect misbehavior, which is trained collaboratively by coordinating multiple distributed edge devices while ensuring data security and privacy. In addition, to further protect the privacy of the model on the blockchain, differential privacy with the Gaussian mechanism is leveraged to provide strict privacy protection. Common data falsification attacks are studied in this paper. The experimental results show that our proposed scheme is feasible and effective, and demonstrate that our scheme achieves satisfied accuracy and efficiency.
Linyan Xie, Jia Xu 0005, Taoshen Li
IEEE Trans. Netw. Serv. Manag.3
2021 Misbehavior Detection in VANET Based on Federated Learning and Blockchain
Linyan Xie, Jia Xu 0005, Taoshen Li
ICA3PP (3)3
2021 Improving Peer Assessment Accuracy by Incorporating Grading Behaviors
abstract
Peer assessment, which asks students to evaluate their peers’ submissions, has become the mainstream paradigm for solving the massive grading challenge of open-ended assignments faced by teachers at MOOC platforms. Since peer grades may be biased and unreliable, a group of probabilistic graph models are proposed to improve the estimation to the true scores of assignments derived based on peer grades, by explicitly modeling the bias and reliability of each grader. However, these models assume that graders’ reliability are only impacted by their knowledge/ability levels while ignoring their grading behaviors. In real life, graders’ grading behaviors (e.g., the time consumed for reviewing an assignment) reflect the seriousness of the graders in the assessment and greatly affect their reliability. Following this intuition, we propose two novel probabilistic graph models for cardinal peer assessment, which optimizes the modeling of the reliability of graders by incorporating various grading behaviors of them. In specific, a GBDT-based regressor is firstly built to quantify the grading seriousness of graders according to their behaviors. Second, the grading seriousness values together with knowledge/ability levels of graders are both employed to model their reliability. Finally, an algorithm based on Gibbs sampling is designed to infer true scores of assignments according to the models. Experimental results on a real peer assessment dataset show the superiority of the proposed models in improving the estimation accuracy to the true scores of assignments by leveraging grader grading behaviors.
Jia Xu 0005, Panyuan Yang
ICTAI1
2021 Objects Perceptibility Prediction Model Based on Machine Learning for V2I Communication Load Reduction
Yuebin He, Jinlei Han, Jia Xu 0005
WASA (3)4
2021 Blockchain Oracle-Based Privacy Preservation and Reliable Identification for Vehicles
Jia Xu 0005
WASA (3)5
2021 Itinerary-aware Personalized Deep Matching at Fliggy
abstract
Matching items for a user from a travel item pool of large cardinality have been the most important technology for increasing the business at Fliggy, one of the most popular online travel platforms (OTPs) in China. There are three major challenges facing OTPs: sparsity, diversity, and implicitness. In this paper, we present a novel Fliggy ITinerary-aware deep matching NETwork (FitNET) to address these three challenges. FitNET is designed based on the popular deep matching network, which has been successfully employed in many industrial recommendation systems, due to its effectiveness. The concept itinerary is firstly proposed under the context of recommendation systems for OTPs, which is defined as the list of unconsumed orders of a user. All orders in a user itinerary are learned as a whole, based on which the implicit travel intention of each user can be more accurately inferred. To alleviate the sparsity problem, users’ profiles are incorporated into FitNET. Meanwhile, a series of itinerary-aware attention mechanisms that capture the vital interactions between user’s itinerary and other input categories are carefully designed. These mechanisms are very helpful in inferring a user’s travel intention or preference, and handling the diversity in a user’s need. Further, two training objectives, i.e., prediction accuracy of user’s travel intention and prediction accuracy of user’s click behavior, are utilized by FitNET, so that these two objectives can be optimized simultaneously. An offline experiment on Fliggy production dataset with over 0.27 million users and 1.55 million travel items, and an online A/B test both show that FitNET effectively learns users’ travel intentions, preferences, and diverse needs, based on their itineraries and gains superior performance compared with state-of-the-art methods. FitNET now has been successfully deployed at Fliggy, serving major online traffic.
Jia Xu 0005, Ziyi Wang 0008, Zulong Chen, Detao Lv, Chuanfei Xu
WWW1
2021 Social-Enhanced Attentive Group Recommendation
abstract
With the proliferation of social networks, group activities have become an essential ingredient of our daily life. A growing number of users share their group activities online and invite their friends to join in. This imposes the need of an in-depth study on the group recommendation task, i.e., recommending items to a group of users. Despite its value and significance, group recommendation remains an unsolved problem due to 1) the weights of group members are crucial to the recommendation performance but are rarely learnt from data; 2) social followee information is beneficial to understand users' preferences but is rarely considered; and 3) user-item interactions are helpful to reinforce the performance of group recommendation but are seldom investigated. Toward this end, we devise neural network-based solutions by utilizing the recent developments of attention network and neural collaborative filtering (NCF). First of all, we adopt an attention network to form the representation of a group by aggregating the group members' embeddings, which allows the attention weights of group members to be dynamically learnt from data. Second, the social followee information is incorporated via another attention network to enhance the representation of individual user, which is helpful to capture users' personal preferences. Third, considering that many online group systems also have abundant interactions of individual users on items, we further integrate the modeling of user-item interactions into our method. Through this way, the recommendation for groups and users can be mutually reinforced. Extensive experiments on the scope of both macro-level performance comparison and micro-level analyses justify the effectiveness and rationality of our proposed approaches.
Da Cao, Xiangnan He 0001, Lianhai Miao, Guangyi Xiao 0001, Hao Chen 0051, Jia Xu 0005
IEEE Trans. Knowl. Data Eng.6
2019 Differentially private high-dimensional data publication via grouping and truncating techniques
Ning Wang 0026, Yu Gu 0002, Jia Xu 0005, Fangfang Li 0002, Ge Yu 0001
Frontiers Comput. Sci.3
2019 Graph Filter: Enabling Efficient Topology Calibration
abstract
The topology of a network may change inevitably, due to dynamic behaviors of nodes and links, and failures of hardware and software. Many protocols and applications must be aware of the up-to-date topology of the underlying network. This triggers the topology calibration problem, which means to deduce those different nodes and links between two topologies. The Bloom filter and its variants are efficient to represent and calibrate two general sets. They, however, fail to represent all links and nodes in a topology simultaneously, and thus remain inapplicable to the topology calibration problem. In this paper, we design the graph filter, a novel space-efficient data structure to record not only the node set but also the link set of any given topology. Accordingly, given two topologies we aim to represent them via two respective graph filters, and thereafter deduce those different links in an invertible manner. To this end, we design three essential operations for graph filter, i.e., encoding, subtracting and decoding. Although such operations are sufficient to solve the topology calibration problem, two challenging issues still remain open. First, the XOR traps which occur with low probability at the encoding stage may result in a few miscalculations at the decoding stage. Thus, we propose another augmented decoding algorithm to lessen the impact of XOR traps via terminating illegal decodings. Second, several different links may form cycles in the worst case; hence, we further design a cycle destruction algorithm to make such different links decodable. We implement the graph filter and the associated topology calibration method. Comprehensive evaluations indicate that our method finishes the topology calibration task efficiently with high probability, incurs the least space overhead, and supports invertible decoding reasonably.
Lailong Luo, Deke Guo, Jia Xu 0005, Xueshan Luo
IEEE Trans. Parallel Distributed Syst.3
2018 Semi-supervised multi-graph classification using optimal feature selection and extreme learning machine
Jun Pang 0002, Yu Gu 0002, Jia Xu 0005, Ge Yu 0001
Neurocomputing3
2017 Caching-Aware Techniques for Query Workload Partitioning in Parallel Search Engines
abstract
In this work, we propose efficient query workload partition techniques to reduce processing times of queries in parallel search engines. Existing methods cannot offer both high cache hit ratios and caching-aware load balance of the system. Aiming to solve this problem, we propose effective solutions to capture tradeoff between the cache hit ratio and load balance to reduce the total query processing time. The performance of the proposed algorithms are demonstrated by extensive experiments on real datasets, and the experimental results demonstrate that our algorithms have an efficiency improvement of up to at least 30% compared to extending current methods such as the roundrobin based algorithm and so on.
Chuanfei Xu, Yanqiu Wang, Jia Xu 0005
WISA4
2017 Topology calibration in data centers
abstract
The topology of data centers changes dynamically due to link malpositions, hardware failures or software crushes. However, many topology enabled protocols or applications must know the current topology of data center precisely, which triggers the topology calibration problem. Topology calibration needs to deduce the different nodes and links between two given topologies effectively. Based on the existing method, deriving the different nodes is relatively simple, since they can be uniquely identified by their IP or MAC addresses. On the contrary, picking the different links from the massive links can be costly. Therefore, we envision a method to locate the different links with respect to the following rationales: 1) efficient, the caused storage cost or communication overhead should be low; 2) without priori knowledge, there is no support information, thus the different links should be decoded inversely. However, the existing strategies based on Bloom filter, Hash table, or Search trees fail to achieve the two rationales simultaneously. Thus, we propose graph filter, a space-efficient data structure to represent and deduce the different links in an invertible manner. To this end, the associated encoding, subtracting and decoding algorithms are proposed. The simulations highlight the strength of graph filter reasonably.
Lailong Luo, Deke Guo, Jia Xu 0005, Xueshan Luo
IWQoS3
2017 Delay updating in Software-Defined Datacenter networks
abstract
Software-Define Datacenter networks (SDDCs) are constantly changing due to various network update events. A set of involved flows of each update event should be migrated to the feasible paths without congestion. To tackle a single update event, prior methods find and execute a migration sequence so as to transform the initial traffic distribution to the final traffic distribution. However, when handling a queue of multiple update events, the head-of-line blocking problem always appears and considerably lower the efficiency of network update. To tackle this problem, we employ the idea of delay updating which schedules other queued events instead of the blocked head-event, aiming to break the blocking state and provide updating opportunities for other events. We formulate this problem as an optimization problem and propose partial delay updating (PDU) strategy. PDU tackles the queued update events based on their arrival order to preserve the fairness. Moreover, it prefers to just delay those blocked flows, lacking enough bandwidth resources, instead of all flows in the head-event. The evaluation results show that our delay updating strategy achieves 60%–80% reduction in average completion time of update events, compared to FIFO scheduling strategy, in four types of flow size.
Ting Qu 0003, Deke Guo, Jia Xu 0005, Zhong Liu 0002
IWQoS3
2017 Parallel multi-graph classification using extreme learning machine and MapReduce
Jun Pang 0002, Yu Gu 0002, Jia Xu 0005, Xiaowang Kong, Ge Yu 0001
Neurocomputing3
2017 Differentially Private Event Histogram Publication on Sequences over Graphs
Ning Wang 0026, Yu Gu 0002, Jia Xu 0005, Fangfang Li 0002, Ge Yu 0001
J. Comput. Sci. Technol.3
2016 Efficient similarity join based on Earth Mover's Distance using MapReduce
abstract
Earth Mover's Distance (EMD) evaluates the similarity between probability distributions, known as a robust measure more consistent with human similarity perception than traditional similarity functions. EMD similarity join retrieves pairs of probability distributions with EMD below a specified threshold, supporting many important applications, such as duplicate image retrieval and sensor pattern recognition. This paper studies the possibility of using MapReduce to improve the scalability of EMD similarity join. Utilizing the dual-program mapping technique, we present a new general data partition framework to facilitate effective workload decomposition using MapReduce, ensuring similar distributions in terms of EMD are mapped to the same reduce task for further verification. New optimization strategies are also proposed to balance the workloads among reduce tasks and eliminate large unnecessary EMD evaluations. Our experiments verify the superiority of our proposal on system efficiency, with a huge advantage of at least one order of magnitude than the state-of-the-art solution, and on system effectiveness, with a real case study towards the abused image phenomenon on C2C website in China. Further details are reported in [4].
Jia Xu 0005, Yu Gu 0002, Marianne Winslett, Ge Yu 0001
ICDE1
2016 CATS: Cooperative Allocation of Tasks and Scheduling of Sampling Intervals for Maximizing Data Sharing in WSNs
abstract
Data sharing among multiple sampling tasks significantly reduces energy consumption and communication cost in low-power wireless sensor networks (WSNs). Conventional proposals have already scheduled the discrete point sampling tasks to decrease the amount of sampled data. However, less effort has been expended for applications that generate continuous interval sampling tasks. Moreover, most pioneering work limits its view to schedule sampling intervals of tasks on a single sensor node and neglects the process of task allocation in WSNs. Therefore, the gained efforts in prior work cannot benefit a large-scale WSN because the performance of a scheduling method is sensitive to the strategy of task allocation. Broadening the scope to an entire network, this article is the first work to maximize data sharing among continuous interval sampling tasks by jointly optimizing task allocation and scheduling of sampling intervals in WSNs. First, we formalize the joint optimization problem and prove it NP-hard. Second, we present the COMBINE operation, which is the crucial ingredient of our solution. COMBINE is a 2-factor approximate algorithm for maximizing data sharing among overlapping tasks. Furthermore, our heuristic named CATS is proposed. CATS is 2-factor approximate algorithm for jointly allocating tasks and scheduling sampling intervals so as to maximize data sharing in the entire network. Extensive empirical study is conducted on a testbed of 50 sensor nodes to evaluate the effectiveness of our methods. In addition, the scalability of our methods is verified by utilizing TOSSIM, a widely used simulation tool. The experimental results indicate that our methods successfully reduce the volume of sampled data and decrease energy consumption significantly.
Deke Guo, Jia Xu 0005, Tao Chen 0013, Jianping Yin
ACM Trans. Sens. Networks3
2015 Efficient Similarity Join Based on Earth Mover's Distance Using MapReduce
abstract
Earth Mover's Distance (EMD) evaluates the similarity between probability distributions, known as a robust measure more consistent with human similarity perception than traditional similarity functions. EMD-based similarity join retrieves pairs of probability distributions with EMD below a specified threshold, supporting many important applications, such as duplicate image retrieval and sensor pattern recognition. This paper studies the possibility of using MapReduce to improve the scalability of EMD similarity join. While existing MapReduce optimization techniques mainly aim to minimize the communication overhead, such methods are not applicable to our problem, due to the high computational cost of EMD. Utilizing the dual-program mapping technique, we present a new general data partition framework to facilitate effective workload decomposition using MapReduce, ensuring similar distributions in terms of EMD are mapped to the same reduce task for further verification. New optimization strategies are also proposed to balance the workloads among reduce tasks and eliminate large unnecessary EMD evaluations. Our experiments verify the superiority of our proposal on system efficiency, with a huge advantage of at least one order of magnitude than the state-of-the-art solution, and on system effectiveness, with a real case study towards the abused image phenomenon on the most popular C2C Web site in China.
Jia Xu 0005, Yu Gu 0002, Marianne Winslett, Ge Yu 0001
IEEE Trans. Knowl. Data Eng.1
2014 Efficient Graph Similarity Join with Scalable Prefix-Filtering Using MapReduce
Jun Pang 0002, Yu Gu 0002, Jia Xu 0005, Yubin Bao, Ge Yu 0001
WAIM3
2013 Differentially private histogram publication
Jia Xu 0005, Xiaokui Xiao, Yin Yang 0001, Ge Yu 0001, Marianne Winslett
VLDB J.1
2012 EUDEMON: A System for Online Video Frame Copy Detection by Earth Mover's Distance
abstract
The Earth Mover's Distance, or EMD for short, has been proven to be effective for content-based image retrieval. However, due to the cubic complexity of EMD computation, it remains difficult to use EMD in applications with stringent requirement for efficiency. In this paper, we present our new system, called EUDEMON, which utilizes new techniques to support fast Online Video Frame Copy Detection based on the EMD. Given a group of registered frames as queries and a set of targeted detection videos, EUDEMON is capable of identifying relevant frames from the video stream in real time. The significant improvement on efficiency mainly relies on the primal-dual theory in linear programming and well-designed B+tree filters for adaptive candidate pruning. Generally speaking, our system includes a variety of new features crucial to the deployment of EUDEMON in real applications. First, EUDEMON achieves high throughput even when a large number of queries are registered in the system. Second, EUDEMON contains self-optimization component to automatically enhance the effectiveness of the filters based on the recent content of the video stream. Finally, EUDEMON provides a user-friendly visualization interface, named EMD Flow Chart, to help the users to better understand the alarm with the perspective of the EMD.
Jia Xu 0005, Qiushi Bai, Yu Gu 0002, Anthony K. H. Tung, Guoren Wang, Ge Yu 0001
ICDE1
2012 Differentially Private Histogram Publication
abstract
Differential privacy (DP) is a promising scheme for releasing the results of statistical queries on sensitive data, with strong privacy guarantees against adversaries with arbitrary background knowledge. Existing studies on DP mostly focus on simple aggregations such as counts. This paper investigates the publication of DP-compliant histograms, which is an important analytical tool for showing the distribution of a random variable, e.g., hospital bill size for certain patients. Compared to simple aggregations whose results are purely numerical, a histogram query is inherently more complex, since it must also determine its structure, i.e., the ranges of the bins. As we demonstrate in the paper, a DP-compliant histogram with finer bins may actually lead to significantly lower accuracy than a coarser one, since the former requires stronger perturbations in order to satisfy DP. Moreover, the histogram structure itself may reveal sensitive information, which further complicates the problem. Motivated by this, we propose two novel algorithms, namely Noise First and Structure First, for computing DP-compliant histograms. Their main difference lies in the relative order of the noise injection and the histogram structure computation steps. Noise First has the additional benefit that it can improve the accuracy of an already published DP-complaint histogram computed using a naiive method. Going one step further, we extend both solutions to answer arbitrary range queries. Extensive experiments, using several real data sets, confirm that the proposed methods output highly accurate query answers, and consistently outperform existing competitors.
Jia Xu 0005, Xiaokui Xiao, Yin Yang 0001, Ge Yu 0001
ICDE1
2012 Adaptive Update Workload Reduction for Moving Objects in Road Networks
Yu Gu 0002, Jia Xu 0005, Ge Yu 0001
WAIM3
2012 Efficient and effective similarity search over probabilistic data based on Earth Mover's Distance
Jia Xu 0005, Anthony K. H. Tung, Ge Yu 0001
VLDB J.1
2010 Efficient and Effective Similarity Search over Probabilistic Data based on Earth Mover's Distance
abstract
Probabilistic data is coming as a new deluge along with the technical advances on geographical tracking, multimedia processing, sensor network and RFID. While similarity search is an important functionality supporting the manipulation of probabilistic data, it raises new challenges to traditional relational database. The problem stems from the limited effectiveness of the distance metric supported by the existing database system. On the other hand, some complicated distance operators have proven their values for better distinguishing ability in the probabilistic domain. In this paper, we discuss the similarity search problem with the Earth Mover's Distance , which is the most successful distance metric on probabilistic histograms and an expensive operator with cubic complexity. We present a new database approach to answer range queries and k-nearest neighbor queries on probabilistic data, on the basis of Earth Mover's Distance. Our solution utilizes the primal-dual theory in linear programming and deploys B + tree index structures for effective candidate pruning. Extensive experiments show that our proposal dramatically improves the scalability of probabilistic databases.
Jia Xu 0005, Anthony K. H. Tung, Ge Yu 0001
Proc. VLDB Endow.1
2008 Modeling and Service Capability Evaluation for RFID Complex Event Processing
abstract
Radio Frequency Identification (RFID) has gained a lot of attention as a promising technology that facilitates pervasive computing and the gradual proliferation of RFID-based applications will produce a large amount of urgent and complicated event streams. The RFID middleware system is expected to offer the further complex event processing service to satisfy the high-level query semantics and performance demands, and therefore, how to evaluate the service capability will be a key problem in mission-critical monitoring scenarios. In this paper, we analyze the RFID primitive event arrival and complex event processing model in a novel perspective. Furthermore, based on our proposed models, service capability evaluation method is introduced to estimate whether and how the RFID middleware system can afford the query requirements in a deterministic or statistical manner, which is believed to be very helpful in RFID-based applications. Our experiments demonstrate the utility and feasibility of our models and methods.
Yu Gu 0002, Yanfei Lv, Ge Yu 0001, Jia Xu 0005
WAIM4