Hong Yao

dblp:98/914 · DBLP profile ↗
← Back
68ranked-venue papers
16as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 10 since 2021Computer networks · 15 · 3 first-author · 2 since 2021Systems, architecture and hardware · 11 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Display Ads Contextual Relevance Modeling with LLM Labels
Chao Gan, Fangping Huang, Weijie Yuan 0007, Nahid Anwar, Musen Wen, Konstantin Shmakov, Hong Yao, Kuang-chih Lee
ECIR (4)8
2026 Unified Supervision for Walmart's Sponsored Search Retrieval via Joint Semantic Relevance and Behavioral Engagement Modeling
Shasvat Desai, Md Omar Faruk Rokon, Jhalak Nilesh Acharya, Isha Shah, Hong Yao, Utkarsh Porwal, Kuang-chih Lee
SIGIR5
2026 Fine-grained geographic named entity recognition with few-shot learning
abstract
Geographic Named Entity Recognition (GNER) focuses on extracting geographic entity names from text and classifying them into pre-defined categories. Previous methods have not paid much attention to identifying fine-grained categories of geographic entities in sparse data situations, thus remaining limited in serving various geographic applications. To address this limitation, this paper presents a fine-grained GNER task and proposes a fine-grained GNER model, LH-FGNER, which incorporates a prototype network and hierarchical contrastive learning to improve fine-grained GNER. Specifically, the model designs label-guided sentence-level prototypes to capture the contextual semantics of geographic entities. It introduces a hierarchy tree to guide the construction of prototypes in vector space, which utilizes the hierarchy as a priori knowledge to improve the discrimination of fine-grained categories. In addition, two datasets are constructed to support the study of the fine-grained GNER task. Experimental results show that the proposed model is superior to the baseline and is robust. This work provides a methodological reference for few-shot GNER, which can be used to facilitate various geographic applications with text.
Shengwen Li, Yuxing Wu, Chaofan Fan, Yaqin Ye, Hong Yao
Int. J. Geogr. Inf. Sci.5
2026 Spatial context-enhanced temporal knowledge graph reasoning
Tailong Li, Renyao Chen, Yilin Duan, Yongbin Xie, Shengwen Li, Hong Yao
Inf. Process. Manag.6
2025 C-KGE: Curriculum learning-based Knowledge Graph Embedding
Diange Zhou, Shengwen Li, Lijun Dong, Renyao Chen, Xiaoyue Peng, Hong Yao
Comput. Speech Lang.6
2025 Incorporating hydrological constraints with deep learning for streamflow prediction
Yilin Duan, Hong Yao, Xinchuan Li, Shengwen Li
Expert Syst. Appl.3
2025 Mineral resource prediction based on representation learning
Diange Zhou, Shengwen Li, Hong Yao
Expert Syst. Appl.3
2025 Synthesizing event semantics for geographical entity representation
abstract
Geographical entity representation learning (GERL) is the emerging approach that manages and represents geographical entities, which advances a wide range of geographical intelligent applications by mapping entities into a latent vector space. However, previous GERL methods ignore the semantics carried by events that are closely associated with geographical entities, resulting in partially missing semantics in the learned vectors of geographical entities. To fill this gap, this paper proposes an event-enhanced geographical entity representation learning (EGERL) method to incorporate geographical events toward improving GERL. Specifically, EGERL designs an event integration strategy to bridge geographical events and geographical entities. And, it develops an anchor identification algorithm to recognize geographical anchor entities with rich event information. In addition, it augments the connections between anchor entities and non-anchor entities with an enhanced graph to enrich the information dissemination of event semantics. Finally, EGERL derives a relation-aware encoding module to encode geographical entities and their complex relations into vectors. Experimental results show that EGERL outperforms the state-of-the-art methods on knowledge representation and graph representation learning, and is robust. The study provides a new exploration of learning entity representations by fusing external semantics, and provides methodological references for various intelligent applications.
Shengwen Li, Renyao Chen, Junye Lei, Tailong Li, Hong Yao
Int. J. Geogr. Inf. Sci.6
2025 The joint extraction of fact-condition statement and super relation in scientific text with table filling method
Hong Yao, Diange Zhou
Inf. Process. Manag.2
2025 Open-world semi-supervised relation extraction
Diange Zhou, Yilin Duan, Shengwen Li, Hong Yao
Neural Networks4
2025 MixSong: Diverse and Strictly Formatted Chinese Poetry Generation
abstract
Chinese poetry, renowned for its elegance and simplicity, is a hallmark of Chinese culture. While neural networks have made significant advancements in generating poetry, balancing diversity with adherence to rigid structural formats remains a challenge. Research indicates that factors such as themes, emotions (e.g., happiness, sadness), and sentiments (e.g., positive, negative) play a crucial role in poetic creation, influencing both the diversity and quality of the generated content. In this paper, we propose MixSong, an autoregressive language model based on the Transformer architecture, designed to incorporate a wide range of conditional factors. MixSong utilizes adversarial training to integrate these factors, enabling the model to implicitly learn distributional information in the latent space. Additionally, we introduce several uniquely customized symbol sets, including paragraph identifiers, position identifiers, rhyme identifiers, tune identifiers, and conditional distinctive identifiers. These symbols help MixSong effectively capture and enforce the constraints necessary for generating high-quality poetry. Extensive experimental results demonstrate that MixSong significantly outperforms existing models in both automatic metrics and human evaluations, achieving notable improvements in both diversity and quality of the generated poetry.
Xinglong Song, Changlin Song, Haolu Yu, Yonghua Zhu, Hong Yao
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2025 Anchor-Enhanced Geographical Entity Representation Learning
abstract
Geographical entity representation learning (GERL) aims to embed geographical entities into a low-dimensional vector space, which provides a generalized approach for utilizing geographical entities to serve various geographical intelligence applications. In practice, the spatial distribution of geographical entities is highly unbalanced; thus, it is challenging to embed them accurately. Previous GERL models treated all geographical entities uniformly, resulting in insufficient entity representations. To address this issue, this article proposes an anchor-enhanced GERL (AE-GERL) model, which utilizes the key informative entities as anchors to improve the representations of geographical entities. Specifically, AE-GERL develops an anchor selection algorithm to identify anchors from large-scale geographical entities based on their spatial distribution and entity types. To utilize anchors to guide geographical entities, AE-GERL constructs an anchor-enhanced graph to establish explicit connections between anchors and nonanchor entities. Finally, a graph neural network (GNN) based anchor to nonanchor node learning model is designed to impute missing information of nonanchor entities. Extensive experiments are conducted on four datasets, and the experimental results demonstrate that AE-GERL outperforms the baseline models in both sparse and dense scenarios. This study provides a methodological reference for embedding geographical entities in various geographical applications and also provides an effective approach to improve the performance of message-passing-based GNN models.
Renyao Chen, Junye Lei, Hong Yao, Tailong Li, Shengwen Li
IEEE Trans. Neural Networks Learn. Syst.3
2024 A Clustering-Oriented Method for Open-Domain Named Entity Recognition
abstract
Named Entity Recognition (NER), as a fundamental task in natural language understanding, has garnered widespread attention. However, most existing research assumes that labeled data is available, limiting the ability of NER models to perform in open-domain scenarios. Especially in real-world applications, the absence of labeled data for novel entity classes is an inevitable scenario. To address this issue, this paper proposes CONER, a Clustering-Oriented Named Entity Recognition method, which mainly consists of two modules: Label Center Clustering (LCC) and Pseudo Label Learning (PLL). LCC uses label classes as clustering centers to guide the learning process of entity representations, thereby achieving high cohesion of entities of the same type in the embedding space. PLL generates pseudo labels based on the cluster-friendly embeddings generated by LCC and uses pairwise similarity learning for the discriminative representation of the novel classes. To balance the learning pace for both seen classes and novel classes, LCC is employed as a pre-training procedure to initialize the model, and then LCC is jointly optimized with PLL. The experimental results on OntoNotes and AnatEM datasets demonstrate that the proposed model outperforms the current zero-shot NER models in open-domain NER, which validates the effectiveness of the approach presented in this paper in open-domain scenarios.
Diange Zhou, Yilin Duan, Xinchuan Li, Hong Yao
CCGrid5
2024 Improving few-shot named entity recognition via Semantics induced Optimal Transport
Diange Zhou, Shengwen Li, Hong Yao
Neurocomputing4
2023 Neural-Symbolic Reasoning with External Knowledge for Machine Reading Comprehension
Yilin Duan, Xiaoyue Peng, Xiaojun Kang, Hong Yao
ICONIP (11)5
2023 Vehicles driving behavior recognition based on transfer learning
Hong Yao, Fengxiang Qiao, Yongfeng Ma
Expert Syst. Appl.2
2023 Anchors-Based Incremental Embedding for Growing Knowledge Graphs
abstract
Knowledge graph embedding aims to transform the entities and relations of triplets into the low-dimensional vectors. Previous methods are oriented towards the static knowledge graphs, in which all entities and relations are assumed to be known and only some unknown triplets need to be predicted. However, the real-world knowledge graphs can grow dynamically, and some new knowledge are often added. To embed the new knowledge into the space of original knowledge graph, the classic models have to perform the entire re-embedding with including the new and original knowledge. This causes heavy computational burden for embedding. To address this problem, this study proposes a new model of anchors-based incremental embedding (ABIE) to implement the dynamical embedding for the growing knowledge graph. According to ABIE, every knowledge graph has some key entities, called anchors, which can fix the embedding space of knowledge graph. When some new knowledge is added into the graph, only a few updated entities and relations are embedded into the embedding space with the help of anchors, and the entire re-embedding on the whole graph is not necessary. By this way, the computational burden of embedding caused by the growth of knowledge graph is reduced significantly.
Lijun Dong, Dongyang Zhao, Xiaoai Zhang, Xinchuan Li, Xiaojun Kang, Hong Yao
IEEE Trans. Knowl. Data Eng.6
2022 Location-aware neural graph collaborative filtering
abstract
Collaborative filtering (CF) is initiated by representing users and items as vectors and seeks to describe the relationship between users and items at a profound level, thus predicting users’ preferred behavior. To address the issue that previous research ignored higher-order geographical interactions hidden in users’ historical behaviors, this paper proposes a location-aware neural graph collaborative filtering model (LA-NGCF), which incorporates location information of items for improving prediction performance. The model characterizes the interactions between items based on spatial decay law from a graph perspective and designs two strategies to capture the interaction effects of users and items considering node heterogeneity. An optimized loss function with spatial distances of items is also developed in the model. Extensive experiments are conducted on three publicly available real-world datasets to examine the effectiveness of our model. Results show that LA-NGCF achieves competitive performances compared with several state-of-the-art models, which suggests that location information of items is beneficial for improving the performance of personalized recommendations. This paper offers an approach to incorporate weighted interactions between items into CF algorithms and enriches the methods of utilizing geographical information for artificial intelligence applications.
Shengwen Li, Chenpeng Sun, Renyao Chen, Xinchuan Li, Qingzhong Liang, Junfang Gong, Hong Yao
Int. J. Geogr. Inf. Sci.7
2022 Dynamic hypergraph neural networks based on key hyperedges
Xiaojun Kang, Xinchuan Li, Hong Yao, Xiaoyue Peng, Tiejun Wu, Shihua Qi, Lijun Dong
Inf. Sci.3
2022 Graph Neural Networks with Information Anchors for Node Representation Learning
Chao Liu 0007, Xinchuan Li, Dongyang Zhao, Shaolong Guo, Xiaojun Kang, Lijun Dong, Hong Yao
Mob. Networks Appl.7
2021 Class-specific information measures and attribute reducts for hierarchy and systematicness
Xianyong Zhang, Hong Yao, Zhiying Lv, Duoqian Miao 0001
Inf. Sci.2
2021 Function-level obfuscation detection method based on Graph Convolutional Networks
Hong Yao, Cai Fu, Yekui Qian, Lansheng Han
J. Inf. Secur. Appl.2
2021 Improving graph neural network via complex-network-based anchor structure
Lijun Dong, Hong Yao, Shengwen Li, Qingzhong Liang
Knowl. Based Syst.2
2021 A Network Calculus Based Delay and Backlog Analysis for Cloud Radio Access Networks
Muzhou Xiong, Lin Gu 0002, Deze Zeng, Hong Yao, Zhuzhong Qian
Mob. Networks Appl.5
2020 An efficient iterative graph data processing framework based on bulk synchronous parallel model
abstract
Summary Graph data processing has been widely applied in a variety of domains such as industry, science, social network, and so on. It therefore has stimulated many efforts devoted to this area. To embrace the fast development trend of big graph data, graph data processing based on Pregel‐like systems has been regarded as one of the most promising ways and has widely attracted the attention of researchers. However, it still remains in its early stage and there still exist many challenges. In Pregel, the superstep synchronization is time consuming as the graph data iteration operation requires multiple synchronizations. Furthermore, the graph data partition strategy adopted by Pregel fails to support load balancing, therefore causing the increase of network I/O overhead as the scale of graph data grows. To address these issues, this paper presents an efficient computational framework for graph data processing based on the bulk synchronous parallel model. The global synchronization control mechanism is improved by determining the start time of the next round of superstep through counting the number of global message files. Furthermore, an improved graph data partition mechanism based on a balanced hash method is proposed to reduce the communication overhead between different partitions of sub‐graph computational tasks. We also re‐design the PageRank algorithm to verify the effectiveness of the proposed framework. Experimental results on different real‐world datasets verify the efficiency of our proposed framework as it outperforms Giraph (an open source Pregel‐like system) by 58%−69%, and achieves 10×−17× performance improvement over Hadoop.
Chao Liu 0007, Deze Zeng, Hong Yao, Xuesong Yan 0001, Linchen Yu, Zhangjie Fu 0001
Concurr. Comput. Pract. Exp.3
2020 Joint optimization of function mapping and preemptive scheduling for service chains in network function virtualization
Hong Yao, Muzhou Xiong, Lin Gu 0002, Deze Zeng
Future Gener. Comput. Syst.1
2020 Towards energy efficient service composition in green energy powered Cyber-Physical Fog Systems
Deze Zeng, Lin Gu 0002, Hong Yao
Future Gener. Comput. Syst.3
2019 Distant-Supervised Relation Extraction with Hierarchical Attention Based on Knowledge Graph
abstract
Relation Extraction is concentrated on finding the unknown relational facts automatically from the unstructured texts. Most current methods, especially the distant supervision relation extraction (DSRE), have been successfully applied to achieve this goal. DSRE combines knowledge graph and text corpus to corporately generate plenty of labeled data without human efforts. However, the existing methods of DSRE ignore the noisy words within sentences and suffer from the noisy labelling problem; the additional knowledge is represented in a common semantic space and ignores the semantic-space difference between relations and entities. To address these problems, this study proposes a novel hierarchical attention model, named the Bi-GRU-based Knowledge Graph Attention Model (BG2KGA) for DSRE using the Bidirectional Gated Recurrent Unit (Bi-GRU) network. BG2KGA contains the word-level and sentence-level attentions with the guidance of additional knowledge graph, to highlight the key words and sentences respectively which can contribute more to the final relation representations. Further-more, the additional knowledge graph are embedded in the multi-semantic vector space to capture the relations in 1-N, N-1 and N-N entity pairs. Experiments are conducted on a widely used dataset for distant supervision. The experimental results have shown that the proposed model outperforms the current methods and can improve the Precision/Recall (PR) curve area by 8% to 16% compared to the state-of-the-art models; the AUC of BG2KGA can reach 0.468 in the best case.
Hong Yao, Lijun Dong, Shiqi Zhen, Xiaojun Kang, Xinchuan Li, Qingzhong Liang
ICTAI1
2019 A-GNN: Anchors-Aware Graph Neural Networks for Node Embedding
Chao Liu 0007, Xinchuan Li, Dongyang Zhao, Shaolong Guo, Xiaojun Kang, Lijun Dong, Hong Yao
QSHINE7
2018 Mining multiple spatial-temporal paths from social media data
Hong Yao, Muzhou Xiong, Deze Zeng, Junfang Gong
Future Gener. Comput. Syst.1
2017 Minimize Coflow Completion Time via Joint Optimization of Flow Scheduling and Processor Placement
abstract
The recent progress in big data has inspired lots of data- parallel applications deployed in the datacenters. Although how to optimize the data flow scheduling in datacenters has been extensively studied, traditional per-flow based optimizations usually do not perform well in dealing with the transferring of a collection of parallel flows, i.e., coflow. Consequently, how to schedule the coflow towards various objectives, e.g., minimizing the coflow completion time, has attracted much attention recently. We notice that existing coflow scheduling studies usually suggest a fixed destination for each coflow. Taking the advantage of virtualization technology, we argue that the destination can be flexibly placed in the cloud. Therefore, it is essential to jointly optimize the coflow scheduling and data processor placement. In this paper, we are motivated to investigate the problem of coflow completion time minimization with joint consideration of coflow scheduling and data processor placement. We first formally describe the problem into a mixed integer non-linear programming (MINLP) problem. By linearizing the MINLP, we further propose a relaxation based heuristic algorithm. Via extensive simulation studies, the high efficiency of our heuristic algorithm is validated.
Deze Zeng, Jie Zhang 0076, Lin Gu 0002, Peng Li 0017, Hong Yao
GLOBECOM5
2017 Joint Optimization of Virtual Function Migration and Rule Update in Software Defined NFV Networks
abstract
Emerging technologies such as Software-Defined Networks (SDN) and Network Function Virtualization (NFV) promise to address cost reduction and flexibility in network operation while enabling innovative network service delivery. To catch up with the time- varying traffic demands, the network changes frequently. We should come up with a sequence of instructions to manipulate the starting network into the goal network, while preserving the network semantics correctness (e.g., freedom of loops, bandwidth guaranteeing). In this case, how to migrate the virtual network functions (VNF) and update the flow forwarding rules efficiently is an important and challenging problem. In this paper, we are motivated to address the migration of VNF and flow update rule problem with joint consideration of migration cost and update delay. The problem is first formulated into a mixed integer non-linear programming (MINLP). By linearizing and relaxing the MINLP, we then present a polynomial-time two-stage heuristic algorithm. The high efficiency of our algorithm is extensively validated by simulation based studies by the fact that it performs much closer to the optimal solution.
Jie Zhang 0076, Deze Zeng, Lin Gu 0002, Hong Yao, Muzhou Xiong
GLOBECOM4
2017 Heterogeneous cloudlet deployment and user-cloudlet association toward cost effective fog computing
abstract
Summary Both mobile computing and cloud computing have experienced rapid development in recent years. Although centralized cloud computing exhibits abundant resources for computation‐intensive tasks, the unpredictable and unstable communication latency between the mobile users and the cloud makes it challenging to handle latency‐sensitive mobile computing tasks. To address this issue, fog computing recently was proposed by pushing the cloud computing to the network edge closer to the users. To realize such vision, we can augment existing access points in wireless networks with cloudlet servers for hosting various mobile computing tasks. In this paper, we investigate how to deploy the servers in a cost‐effective manner without violating the predetermined quality of service. In particular, we practically consider that the available cloudlet servers are heterogeneous, ie, with different cost and resource capacities. The problem is formulated into an integer linear programming form, and a low‐complexity heuristic algorithm is invented to address it. Extensive simulation studies validate the efficiency of our algorithm by it performs much close to the optimal solution.
Hong Yao, Changmin Bai, Muzhou Xiong, Deze Zeng, Zhangjie Fu 0001
Concurr. Comput. Pract. Exp.1
2017 Stochastic Time Series Analysis for Energy System Based on Markov Chain Model
Zhengshun Ruan, Aihua Luo, Hong Yao
Mob. Networks Appl.3
2017 Encounter Probability Aware Task Assignment in Mobile Crowdsensing
Hong Yao, Muzhou Xiong, Chao Liu 0007, Qingzhong Liang
Mob. Networks Appl.1
2016 Joint optimization on switch activation and flow routing towards energy efficient software defined data center networks
abstract
The rapid development of cloud computing has raised big concerns over the high energy consumption of modern data centers. To satisfy the ever increasing data traffic needs, the energy consumption of data center network (DCN) also takes a significant proportion. The newly emerging technology, Software Defined Networking (SDN), which allows flexible control of network devices, brings a new opportunity towards DCN energy optimization. In this paper, we investigate how to design an energy-efficient network management strategy with guaranteed satisfaction of network traffic demands in Software Defined Data Center Networks (SD-DCNs). To this end, three issues will be tackled: 1) the subset of switches that shall be activated, i.e., switch activation, 2) multi-path routing scheduling for all flows and 3) forwarding rule placement in SDN switches. They are jointly considered and formulated as an integer linear programming (ILP) problem. A heuristic algorithm to deal with its high computational complexity is proposed. Extensive simulation-based evaluations are conducted to validate the high efficiency of our algorithm.
Deze Zeng, Lin Gu 0002, Song Guo 0001, Hong Yao
ICC5
2016 A Crowd Simulation Based UAV Control Architecture for Industrial Disaster Evacuation
abstract
In past decades, we have witnessed lots of gas leakage diasters all over the world, causing serious casualties, property damage and severe negative social impact. During evacuation after gas leakage incident, the poisonous gas shall be accurately detected. Accordingly, the evacuation routine shall be carefully planned and warned to the evacuating people. Unmanned aerial vehicle (UAV) has been widely regarded as a promising tool to support crowd evacuation. In this paper, aiming at providing an efficient UAV control system for crowd evacuation, we propose a crowd simulation based UAV control system to direct crowd evacuation from the polluted area. The architecture mainly consists of a UAV fleet management module, UAV trajectory planning module, and a crowd simulation module. The architecture forms a closed control loop, emphasizing the inter-operation between the real scenario and the simulation scenario. A case study on UAV gas leakage detection scheduling is given. The results show that the proposed architecture can provide efficient way to help pedestrian in the environment to evacuate and avoid to be infected by the toxic gas.
Muzhou Xiong, Deze Zeng, Hong Yao, Yong Li 0045
VTC Spring3
2016 MEMoMR: Accelerate MapReduce via reuse of intermediate results
abstract
Summary MapReduce has been widely regarded as a flexible, scalable, and easy‐to‐use distributed programming paradigm for big data processing such as social network data analysis on cloud computing platforms. To embrace the upcoming of big data era, many efforts have been devoted to accelerating the MapReduce performance from different aspects, especially intermediate result reusing like Dache. In this paper, we observe that existing intermediate result reusing mechanism is not efficient enough as many I/O operations are wasted. Efficient reusing of the intermediate results could potentially improve the MapReduce performance. Inspired by such fact, we propose a framework named MEMoMR (more efficient intermediate result reusing for MapReduce) by introducing a novel reusing mechanism that can substantially reduce the I/O overhead. To this end, we invent a new metadata description method and apply it in the reusing phase. We practically realize MEMoMR and evaluate its performance by implementing it in a real cluster. The experiment results show that MEMoMR can improve the system performance as high as 23.4%, comparing against Dache. Copyright © 2015 John Wiley & Sons, Ltd.
Hong Yao, Jinlai Xu, Zhongwen Luo, Deze Zeng
Concurr. Comput. Pract. Exp.1
2016 On Cost-Efficient Sensor Placement for Contaminant Detection in Water Distribution Systems
abstract
In recent years, water pollution or contamination incidents happened frequently, causing serious disasters and negative social impact. To reduce the water contamination risk, water quality monitoring sensors should be deployed in water distribution system (WDS) to enable real-time pollution detection. It is desirable to deploy sensors everywhere so that any contamination event can be detected and reported in a timely manner. Unfortunately, this is a luxury and unrealistic vision because of high deployment cost. It is significant to lower the deployment cost provided that the quality-of-sensing, e.g., coverage and contamination detection time, can be guaranteed for effective depollution action. In this paper, we consider a water quality monitoring sensor network consisting of two kinds of sensors with different prices. The expensive one is of cellular communication capability and therefore is able to send sensing information to control center directly, while the cheaper one is of only sensor-to-sensor communication capability. We investigate a cost-efficient sensor deployment problem on how to deploy these two kinds of sensors in a given WDS to minimize the deployment cost, without violating the quality-of-sensing requirement. We first formulate the problem into a mixed integer quadratically constrained programming problem, which is then linearized into an equivalent mixed integer linear programming. We further propose a polynomial two-stage heuristic algorithm and evaluate its efficiency via extensive simulation-based studies.
Deze Zeng, Lin Gu 0002, Lu Lian, Song Guo 0001, Hong Yao, Jiankun Hu
IEEE Trans. Ind. Informatics5
2015 On Participant Selection for Minimum Cost Participatory Urban Sensing with Guaranteed Quality of Information
Hong Yao, Changkai Zhang, Chao Liu 0007, Qingzhong Liang, Xuesong Yan 0001, Chengyu Hu 0002
CollaborateCom1
2015 On Rule Placement for Multi-path Routing in Software-Defined Networks
Jie Zhang 0076, Deze Zeng, Lin Gu 0002, Hong Yao
CollaborateCom4
2015 MR-COF: A Genetic MapReduce Configuration Optimization Framework
Chao Liu 0007, Deze Zeng, Hong Yao, Chengyu Hu 0002, Xuesong Yan 0001
ICA3PP (4)3
2015 Flow setup time aware minimum cost switch-controller association in Software-Defined Networks
Deze Zeng, Chao Teng, Lin Gu 0002, Hong Yao, Qingzhong Liang
QSHINE4
2015 An Abelian group model of commutative data dependence relations for the iteration space slicing
abstract
In loop parallelization, data dependence relations are used to decide which pair of statement instances should be allocated to a same processor or should have a synchronization communication. However, in existing researches, little attention has been paid to the widespread symmetrical patterns of data dependence implied in the loop iteration. These patterns are usually induced by the regular expressions as array indices. If these expressions are all of the same type, the transitive calculations of them are always commutative. In this paper, we introduce a permutation group model to represent data dependences and discuss the application of the model. We focus on three issues: 1) the basic permutation model and the symmetrical patterns, 2) the application of Abelian group theory for commutative relations such as some uniform (addition) relations, multiplication relations and hybrids relations, and 3) an approach to obtaining the iteration slices for parallelization based on previous analyses.
Hong Yao, Huifang Deng
SNPD1
2015 Migrate or not? Exploring virtual machine migration in roadside cloudlet-based vehicular cloud
abstract
Summary Vehicle Ad‐Hoc Networks (VANET) enable all components in intelligent transportation systems to be connected so as to improve transport safety, relieve traffic congestion, reduce air pollution, and enhance driving comfort. The vision of all vehicles connected poses a significant challenge to the collection, storage, and analysis of big traffic‐related data. Vehicular cloud computing, which incorporates cloud computing into vehicular networks, emerges as a promising solution. Different from conventional cloud computing platform, the vehicle mobility poses new challenges to the allocation and management of cloud resources in roadside cloudlet. In this paper, we study a virtual machine (VM) migration problem in roadside cloudlet‐based vehicular network and unfold that (1) whether a VM shall be migrated or not along with the vehicle moving and (2) where a VM shall be migrated, in order to minimize the overall network cost for both VM migration and normal data traffic. We first treat the problem as a static off‐line VM placement problem and formulate it into a mixed‐integer quadratic programming problem. A heuristic algorithm with polynomial time is then proposed to tackle the complexity of solving mixed‐integer quadratic programming. Extensive simulation results show that it produces near‐optimal performance and outperforms other related algorithms significantly. Copyright © 2015 John Wiley & Sons, Ltd.
Hong Yao, Changmin Bai, Deze Zeng, Qingzhong Liang
Concurr. Comput. Pract. Exp.1
2015 Opportunistic Offloading of Deadline-Constrained Bulk Cellular Traffic in Vehicular DTNs
abstract
The ever-growing cellular traffic demand has laid a heavy burden on cellular networks. The recent rapid development in vehicle-to-vehicle communication techniques makes vehicular delay-tolerant network (VDTN) an attractive candidate for traffic offloading from cellular networks. In this paper, we study a bulk traffic offloading problem with the goal of minimizing the cellular communication cost under the constraint that all the subscribers receive their desired whole content before it expires. It needs to determine the initial offloading points and the dissemination scheme for offloaded traffic in a VDTN. By novelly describing the content delivery process via a contact-based flow model, we formulate the problem in a linear programming (LP) form, based on which an online offloading scheme is proposed to deal with the network dynamics (e.g., vehicle arrival/departure). Furthermore, an offline LP-based analysis is derived to obtain the optimal solution. The high efficiency of our online algorithm is extensively validated by simulation results.
Hong Yao, Deze Zeng, Huawei Huang, Song Guo 0001, Ahmed Barnawi, Ivan Stojmenovic
IEEE Trans. Computers1
2014 SAPSN: A Sensor Network for Signal Acquisition and Processing
abstract
Software defined wireless sensor network can be adapted to different application needs through dynamic programming. In this paper, we propose a signal acquisition and processing wireless sensor network (SAPSN). SAPSN consists of sampling nodes, processing nodes and remote controllers. At first, the sampling node completes the local signal sampling by analog-digital conversion. Next, according to the different demand from the remote controller, processing node completes time domain or frequency domain analysis of signal processing, and transfers the results back to the remote controller. Finally, the application in remote controller will display the results according to different user's needs. SAPSN is capable of time domain or frequency domain signal analysis and processing, depending on different application requirements. In this paper, we present the concept underlying SAPSN, its architecture. We also present preliminary experimental results.
Qingzhong Liang, Xuesong Yan 0001, Chengyu Hu 0002, Hong Yao
DASC6
2014 Joint optimization of task mapping and routing for service provisioning in distributed datacenters
abstract
Service provisioning has been widely regarded as a critical issue to quality-of-service (QoS) of cloud services in datacenters. Conventional studies on service provisioning mainly focus on task mapping, i.e., how to distribute the service-oriented tasks onto the servers to achieve different goals, e.g., makespan minimization. In distributed datacenters, a task is usually routed from its generation point (i.e., control room) to the designated server within a datacenter network. Since the routing delay also has a deep influence on the task makespan, we are motivated to study how to minimize the maximum makespan of all tasks in a duty period by joint optimization of both task mapping and routing. It is formulated as an integer programming with quadratic constraints (IPQC) problem and proved as NP-hard. To tackle the computational complexity of solving IPQC, a heuristic algorithm with polynomial time is proposed. Extensive simulation results show that it performs close to the optimal one and outperforms existing algorithms significantly.
Huawei Huang, Deze Zeng, Song Guo 0001, Hong Yao
ICC4
2014 A trace-driven analysis on the user behaviors in social e-commerce network
abstract
E-commerce has become one of the common commercial activities in people's daily lives. The major advantage of e-commerce over conventional commercial activities is the information transparency while people can freely share their opinions and comments. Such information has profound influence on user behaviors in e-commerce activities. Meanwhile, social network service (SNS) has also become the most popular way to get and share information on the Internet. Therefore, it is quite natural to put e-commerce and SNS together. Recently, there emerge many online social e-commerce network (SECON) services, which not only allow users to conduct e-commerce transactions but also enable users to share information as in the other SNS like Twitter. Although conventional SNS has been widely investigated, little is known about SECON. To address this problem, we conduct a trace-driven analysis on a successful SECON called Jumei, with millions of users. Our analysis is based on shared information and all activities they created, all these data are crawled from the website of Jumei. We shed light on the user activity characteristics in SECON. By analyzing the crawled data, we discover that the social ties have an important influence on commercial activities. However, to our surprise, there are many differences between SECON and SNS: (a) the network topology structure is greatly different from SNS, (b) strong ties play a more crucial role than in SNS, same to viral marketing intuition, and (c) the behavior of adoption activity is influenced weakly by peers in the social network, for example, nearly 60% users influenced by only one information propagated from social links before they decided to buy it. Furthermore, we find that it takes a long time to adopt what their followees have bought.
Zhongwen Luo, Huanhuan Zhu, Deze Zeng, Hong Yao
ICC4
2014 An energy-aware deadline-constrained message delivery in delay-tolerant networks
Hong Yao, Huawei Huang, Deze Zeng, Bo Li 0001, Song Guo 0001
Wirel. Networks1
2013 Stochastic analysis on epidemic dissemination of lifetime-controlled messages in DTNs
abstract
To understand the delivery performance of message dissemination in Disruption Tolerant Networks (DTNs), various methods have been proposed in the literature. However, existing work shares a common simplification that the pairwise meeting rate between any two mobile nodes is exponentially distributed. In this paper, instead of relying on such assumption, we jointly consider the transmission range and Random Direction Mobility (RDM) model to stochastically analyze delivery performance of epidemic routing in terms of percolation ratio and delivery delay. Furthermore, we study a controlled epidemic routing, in which any message stays at a mobile node longer than a predefined lifetime should be removed from the node. It can be considered as an age-structure process described by the Susceptible-Infectious-Recovered (SIR) model. To the best of our knowledge, we are the first to characterize the message propagation process by applying the Delay Differential Equations (DDEs) in DTNs. The correctness of our analysis is validated by extensive simulations.
Huawei Huang, Deze Zeng, Song Guo 0001, Hong Yao, Toshiaki Miyazaki
IWCMC4
2013 The design model of evolutionary antenna with finite reflector
abstract
In practical applications, the interference between the antenna and the metal structure nearby often cause deviation of the electromagnetic performance of the antenna. The factor of finite reflector should be considered in the evolutionary antenna to avoid the influence so that the optimal antenna structure can be better applicable to real-world conditions. Therefore, a kind of design model of evolutionary antenna with a finite reflector is proposed in this paper, which includes the corresponding chromosome coding and the design workflow. In this design model, the antenna is optimized with its finite reflector during the evolutionary process. Meanwhile aiming at the characteristics of the model, the method of invoking the simulation software and setting the finite reflector is discussed to evaluate the antenna individuals effectively. This design model with a finite reflector was applied in a practical design problem, and the experimental result shows that it is superior to the one with a infinite reflector both on gain and VSWR. It illustrates that it is very important to take the reflector factor into consideration in the design model of evolutionary antenna.
Qingzhong Liang, Xuesong Yan 0001, Chengyu Hu 0002, Chao Liu 0007, Hong Yao
MSN6
2013 Multi-label Classification based on Particle Swarm Algorithm
abstract
Multi-label classification is a generalization of single-label classification, and its samples belong to multiple labels. The K-nearest neighbor algorithm can solve this problem as an optimization problem. It finds the optimum solution by caculating the distance between each sample in general. But in fact, the distance of K-nearest neighbor algorithm may be miscalculated due to the caused by the redundant or irrelevant characteristic value. In order to solve this problem, in this paper, we propose a novel method that uses the particle swarm algorithm to optimize the feature weights to improve the accuracy of distance calculation. As a result, it can improve classification accuracy further. The experimental results show that applying particle swarm algorithm's optimization technique to improving K-nearest neighbor algorithm for multi-label classification problem, can improve the accuracy of classification effectively.
Qingzhong Liang, Chao Liu 0007, Xuesong Yan 0001, Chengyu Hu 0002, Hong Yao
MSN7
2013 The Evaluation of Security Algorithms on Mobile Platform
abstract
The rapid development of mobile platform leads to growing demand for network communications, which suffer from increasingly severe security threats as well. Considering the urgent security demand and the limitation of battery capacity, the power consumption of the encryption algorithms plays a pivotal role for mobile device. Aiming at the above problem, this paper proposes a performance evaluation system on Android platform, which can evaluate the state parameters of device battery. Our work provides elaborate analysis and exploration of 10 security algorithms' energy consumption through extensive simulation results. With the popularity of mobile network and the rapid increase of related applications, it is significant to reduce the power consumption of security encryption algorithms on mobile devices in both academic field and industry.
Hong Yao, Lu Lian, Qingzhong Liang
MSN1
2012 Deadline-constrained content distribution in vehicular delay tolerant networks
abstract
Content distribution in vehicular networks is essential to many emerging applications. The issues such as content distribution from road side units (RSUs) to vehicles or the cooperation between vehicles have drawn a lot of interests in the literature. However, little work is on packets distribution from content providers to RSUs and many related issues are still under-investigated. In this paper, we consider the problem of minimizing the distribution cost, which is defined as the number of packets that shall be dispatched to RSUs, for deadline-constrained content distribution in vehicular networks. The problem is first formulated as an integer programming problem, based on a link-coloring concept. Then, a heuristic algorithm with low computational complexity is proposed. The high efficiency of the proposed algorithm is extensively validated by the fact that it performs close to the optimal solution obtained by the CPLEX solver.
Deze Zeng, Lei Cong, Huawei Huang, Song Guo 0001, Hong Yao
IWCMC5
2010 Join tree propagation with prioritized messages
abstract
Current join tree propagation algorithms treat all propagated messages as being of equal importance. On the contrary, it is often the case in real-world Bayesian networks that only some of the messages propagated from one join tree node to another are relevant to subsequent message construction at the receiving node. In this article, we propose the first join tree propagation algorithm that identifies and constructs the relevant messages first. Our approach assigns lower priority to the irrelevant messages as they only need to be constructed so that posterior probabilities can be computed when propagation terminates. Experimental results, involving the processing of evidence in four real-world Bayesian networks, empirically demonstrate an improvement over the state-of-the-art method for exact inference in discrete Bayesian networks. © 2009 Wiley Periodicals, Inc. NETWORKS, 2010
Cory J. Butz, Shan Hua, Ken Konkel, Hong Yao
Networks4
2009 Measuring web feature impacts in Peer-to-Peer file sharing systems
Sirui Yang, Hai Jin 0001, Bo Li 0001, Xiaofei Liao, Hong Yao, Xuping Tu
Comput. Commun.5
2009 A simple graphical approach for understanding probabilistic inference in Bayesian networks
Cory J. Butz, Shan Hua, Hong Yao
Inf. Sci.4
2009 A join tree probability propagation architecture for semantic modeling
abstract
We propose the first join tree (JT) propagation architecture that labels the probability information passed between JT nodes in terms of conditional probability tables (CPTs) rather than potentials. By modeling the task of inference involving evidence, we can generate three work schedules that are more time-efficient for LAZY propagation. Our experimental results, involving five real-world or benchmark Bayesian networks (BNs), demonstrate a reasonable improvement over LAZY propagation. Our architecture also models inference not involving evidence. After the CPTs identified by our architecture have been physically constructed, we show that each JT node has a sound, local BN that preserves all conditional independencies of the original BN. Exploiting inference not involving evidence is used to develop an automated procedure for building multiply sectioned BNs. It also allows direct computation techniques to answer localized queries in local BNs, for which the empirical results on a real-world medical BN are promising. Screen shots of our implemented system demonstrate the improvements in semantic knowledge.
Cory J. Butz, Hong Yao, Shan Hua
J. Intell. Inf. Syst.2
2008 TCPBridge: A software approach to establish direct communications for NAT hosts
abstract
Traversing Network Address Translation (NAT) for Peer-to-Peer (P2P) communication has become a hot topic recently. Compared to UDP, establishing TCP connections for hosts behind different NATs is more complex. Thus, many TCP-based applications do not address TCP traversal through NATs. Some solutions suggest using delegates to relay all communications, or tunneling TCP over UDP. However, they require a big reform to network architecture, or using a non-standard TCP/IP stack. In this paper, we present a novel idea called TCPBridge. TCPBridge converts TCP traversal to UDP traversal without modifying any binaries of the TCP-based applications. Our design can be integrated with those P2P applications which have not solved TCP traversal problem, and extends them to support direct communications between NAT hosts. It deals with the problem of TCP traversal, so as to improve the usability of applications. We have implemented TCPBridge in several existing P2P systems. Statistics prove that TCPBridge is scalable and robust, and we believe it will benefit many other existing P2P applications.
Sanmin Liu, Hai Jin 0001, Xiaofei Liao, Hong Yao, Deze Zeng
AICCSA4
2008 The Content Pollution in Peer-to-Peer Live Streaming Systems: Analysis and Implications
abstract
There has been significant progress in the development and deployment of peer-to-peer (P2P) live video streaming systems. However, there has been little study on the security aspect in such systems. Our prior experiences in Anysee exhibit that existing systems are largely vulnerable to intermediate attacks, in which the content pollution is a common attack that can significantly reduce the content availability, and consequently impair the playback quality. This paper carries out a formal analysis of content pollution and discusses its implications in P2P live video streaming systems. Specifically, we establish a probabilistic model to capture the progress of content pollution. We verify the model using a real implementation based on Anysee system; we evaluate the content pollution effect through extensive simulations. We demonstrate that (1) the number of polluted peers can grow exponentially, similar to random scanning worms. This is vital that with 1% polluters, the overall system can be compromised within minutes; (2) the effective bandwidth utilization can be sharply decreased due to the transmission of polluted packets; (3) Augmenting the number of polluters does not imply a faster progress of content pollution, in which the most influential factors are the peer degree and access bandwidth. We further examine several techniques and demonstrate that a hash-based signature scheme can be effective against the content pollution, in particular when being used during the initial phase.
Sirui Yang, Hai Jin 0001, Bo Li 0001, Xiaofei Liao, Hong Yao, Xuping Tu
ICPP5
2008 Measuring web feature impacts in BitTorrent-like systems
abstract
In Peer-to-Peer (P2P) file sharing systems, the attributes of resource description can influence the user behavior, especially on resource selection. However, this has been only qualitatively speculated but lacks of quantitative analysis. In this paper, we carry out a systematically quantitative stu
Sirui Yang, Hai Jin 0001, Bo Li 0001, Xiaofei Liao, Hong Yao
QSHINE5
2008 Mining functional dependencies from data
Hong Yao, Howard J. Hamilton
Data Min. Knowl. Discov.1
2006 Mining itemset utilities from transaction databases
Hong Yao, Howard J. Hamilton
Data Knowl. Eng.1
2006 An Approach for Solving Fuzzy Games
abstract
This paper is to compute a Nash equilibrium in a fuzzy environment, which is represented by a fuzzy approximate Nash equilibrium in a space of discrete mixed strategies. For discrete mixed strategies, the relationship between the discrete degree and the approximate degree is discussed. Based on the fuzzy regret degree, a genetic algorithm for computing a fuzzy Nash equilibrium is given.
Jin Li 0007, Kun Yue, Hong Yao
Int. J. Uncertain. Fuzziness Knowl. Based Syst.5
2004 A Foundational Approach to Mining Itemset Utilities from Databases
abstract
Most approaches to mining association rules implicitly consider the utilities of the itemsets to be equal. We assume that the utilities of itemsets may differ, and identify the high utility itemsets based on information in the transaction database and external information about utilities. Our theoretical analysis of the resulting problem lays the foundation for future utility mining algorithms.
Hong Yao, Howard J. Hamilton, Cory J. Butz
SDM1
2002 FD_Mine: Discovering Functional Dependencies in a Database Using Equivalences
abstract
The discovery of FDs from databases has recently become a significant research problem. In this paper, we propose a new algorithm, called FD-Mine. FD-Mine takes advantage of the rich theory of FDs to reduce both the size of the dataset and the number of FDs to be checked by using discovered equivalences. We show that the pruning does not lead to loss of information. Experiments on 15 UCI datasets show that FD-Mine can prune more candidates than previous methods.
Hong Yao, Howard J. Hamilton, Cory J. Butz
ICDM1
1997 A logical design method for relational databases based on generalization and aggregation semantics
Hong Yao
J. Comput. Sci. Technol.2