Zhi Cai

dblp:81/6349 · DBLP profile ↗
← Back
43ranked-venue papers
13as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories
abstract
Monte Carlo Tree Search (MCTS) has been widely used for automated reasoning data exploration, but current supervision extraction methods remain inefficient.Standard approaches retain only the single highest-reward trajectory, discarding the comparative signals present in the many explored paths.Here we introduce Contrastive Reasoning Path Synthesis (CRPS), a framework that transforms supervision extraction from a filtering process into a synthesis procedure.CRPS uses a structured reflective process to analyze the differences between high-and low-quality search trajectories, extracting explicit information about strategic pivots and local failure modes.These insights guide the synthesis of reasoning chains that incorporate success patterns while avoiding identified pitfalls.We show empirically that models fine-tuned on just 60K CRPS-synthesized examples match or exceed the performance of baselines trained on 590K examples derived from standard rejection sampling, a 20× reduction in dataset size.Furthermore, CRPS improves generalization on out-of-domain benchmarks, demonstrating that learning from the contrast between success and failure produces more transferable reasoning capabilities than learning from success alone.
Peiyang Liu, Zhirui Chen 0001, Di Liang, Youru Li, Zhi Cai, Wei Ye 0004
ACL (1)6
2026 Spatiotemporal data imputation based on spatiotemporal feature fusion network
Xing Su 0001, Zhi Cai, Yongping Du
Neurocomputing3
2026 Towards unbiased source-free object detection via vision foundation models
Zhi Cai, Yingjie Gao 0001, Yanan Zhang 0005, Xinzhu Ma, Di Huang 0001
Pattern Recognit.1
2026 AST-Adapter: Parameter-Efficient Video-to-Video Transfer Learning With Adaptive Spatiotemporal Information Bias
abstract
Leveraging video pre-trained models for video downstream tasks has recently emerged with promising performance. Except for the full fine-tuning paradigm, parameter-efficient transfer learning (PETL) exists as a promising way and has not yet been fully explored in video-to-video transfer learning. While current PETL approaches succeed to reduce parameter quantity and computation cost, they overlook the critical spatiotemporal property in video modality. In this paper, we first propose a novel metric to quantify the spatiotemporal information bias across video datasets and uncover its impact on transfer preferences through systematic analysis. Based on the analysis, we introduce an innovative parameter-efficient transfer learning method, named Adaptive SpatioTemporal Adapter (AST-Adapter). Our approach automatically adjusts layer-wise architectures with different spatiotemporal adapter modules to exploit the intrinsics of downstream tasks to achieve adaptive spatiotemporal learning, thus delivering robustness and generalization. Extensive experiments on five datasets across action recognition and action detection task show that AST-Adapter surpasses both video-to-video and image-to-video approaches, whilst keeping the advantage of parameter efficiency. Notably, AST-Adapter achieves 89.9% on Kinectics-400, 80.3% on HMDB51 and 97.7% on UCF101 while introduces only 1% to 10% tunable parameters. Our code is available at https://github.com/hhhhhpy/AST-Adapter.
Puyue Hou, Guohao Li 0010, Zhi Cai, Di Huang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 A Novel Vector Processing-Based Online Trajectory Data Indexing Approach
abstract
With the rapid development of geolocation technology, the volume of spatio-temporal trajectory data has surged. This data is widely used in fields such as geographic information systems and mobile computing, but its storage and query processing present significant challenges. Current methods of offline indexing are inefficient and cannot be updated in real-time. To address this issue, this paper proposes a concept of the online index that supports real-time storage and indexing of trajectory data and significantly reduces indexing time and storage space requirements. Based on this concept, two vector-based online trajectory indexing methods are proposed in this paper. The first is an online trajectory indexing method based on vector extraction (VBIndex), which offers the advantages of high efficiency and less storage space. The second is an online trajectory indexing method based on road-network matching (RAIndex), which further improves the vector-based indexing efficiency when road network involved. Through experiments with real datasets, the proposed algorithms were evaluated, confirming their superiority in terms of indexing construction time and storage space. Furthermore, we have theoretically proven that queries based on this index are accurate, and statistical analysis is feasible. Both algorithms have a time complexity of$O(N)$in indexing construction, demonstrating good performance.
Zhi Cai, Mengxiao Liu, Shuaibing Lu, Meihui Shi, Xing Su 0001, Limin Guo 0002
IEEE Trans. Intell. Transp. Syst.1
2026 A Generative Graph Augmentation Neural Network for Traffic Risk Prediction on Online Crowd Queries
abstract
In the existing traffic prediction scenarios, the lack of accompanying event data, noise interference and insufficient supervised signals seriously restrict the effect of actual traffic prediction. Meanwhile, currently prevalent graph neural networks often struggle to capture effective semantic structures when dealing with learning tasks involving diverse specific events, consequently exhibiting limited generalization and transfer capabilities. This study focuses on crowd gathering events in transportation scenarios and conducts quantitative analysis of their potential risks to traffic network. Relying on the massive online crowd query data produced in Location Based Services (LBS), we propose a generative strategy for node and edge augmentation based on event-traffic interactions, which seeks to generate richer supervised signals. Furthermore, in response to the generative graph structure derived from event chains that fail to match the contextual semantic information, we utilize comparative learning for self-supervised training as the auxiliary proxy task of time series prediction. Experiments on the benchmark datasets of real road networks show that the proposed method is effective in identifying the traffic risk of road segments, especially when the breakdown probability is greater than 50%.
Mengmeng Chang, Zhiming Ding, Zhi Cai, Zilin Zhao, Yafei Sun
IEEE Trans. Knowl. Data Eng.3
2026 Knowledge Graph-Based Debiasing for Trustworthy Recommendation Systems
abstract
These years have witnessed remarkable progress in modeling user behaviour from personalized online services, especially knowledge graph-based recommendation systems. Meanwhile, more studies are focusing on aspects beyond recommendation performance, since such an observational data-driven paradigm is posing threats to both users and society in terms of trustworthiness. In fact, existing problem-oriented solutions still face significant challenges, as almost all of them suffer from the generality limitations to improve their trustworthiness in a uniform fashion. To address these issues, we propose a plug-and-playDebiasing framework forKnowledgeGraph-basedRecommendationSystems, also known as DiKGRS. Specifically, the Knowledge-augmented Pseudo-Samples Generation (KPSG) method, a novel data augmentation perspective, is proposed to explore more auxiliary information beyond observational user behaviors. Furthermore, the Debiasing Value Networks (DVN), is also developed to evaluate the reliability of generated pseudo-samples by modeling both the item popularity and user demographic bias in the platform. Moreover, an adaptive weighting coordination module is performed to coordinate the proposed DiKGRS framework and its backbones. Experimental results on four real-world datasets from different online service personalization scenarios have illustrated that the proposed framework can significantly improve the trustworthiness of existing knowledge graph-based recommendation systems. The code has been released public available at:https://github.com/alipay/A-Knowledge-augmented-Method-DiKGRS.
Youru Li, Xuying Ning, Zhenfeng Zhu, Hanqiu Wang, Zhi Cai, Minnan Luo, Yao Zhao 0001
IEEE Trans. Knowl. Data Eng.5
2026 Practical Efficient Deployment and Updating for Microservice With Dependencies in Multi-Access Edge Computing
abstract
As mobile edge computing technology advances rapidly, latency-sensitive and resource-intensive applications are being offloaded to edge servers to enhance Quality of Service (QoS) for users. Traditional monolithic architectures, however, struggle to meet the escalating service and traffic requirements of distributed users due to their inherent inflexibility. In response to these challenges, microservices architecture, characterized by scalability and flexibility, has been adopted for dynamic deployment at the network edge. However, the deployment of these lightweight, dependency-rich components in a way that minimally impacts the makespan and maximizes quality of service is complex. Current studies often overlook the deployment of microservices with specific dependencies within constrained environments of edge server clusters and communication links. This paper introduces practical and effective strategies for the deployment and updating of microservices, tailored to various application contexts. Initially, two scenarios are analyzed: one constrained by bandwidth with unlimited storage, and the other by storage with unlimited bandwidth. For each scenario, optimal solutions are developed using a novel enhanced graph construction method. The study progresses to a more intricate scenario involving comprehensive constraints on storage, computation, and communication resources. An optimized deployment method is proposed, utilizing main path embedding followed by an innovative simulated annealing algorithm for iterative refinement. This method is validated by demonstrating that the main path coincides with the critical path. Furthermore, the dynamic reallocation of edge resources is explored through a critical path-based updating algorithm that optimizes microservice locations to reduce overall makespan. Extensive experiments demonstrate that our strategies outperform existing representative benchmark approaches in terms of overall performance and microservice deployment efficiency.
Shuaibing Lu, Jie Wu 0001, Zhi Cai, Jackson Yang, Shuyang Zhou, Juan Fang 0004
IEEE Trans. Serv. Comput.4
2025 CoSDH: Communication-Efficient Collaborative Perception via Supply-Demand Awareness and Intermediate-Late Hybridization
abstract
Multi-Agent collaborative perception enhances perceptual capabilities by utilizing information from multiple agents and is considered a fundamental solution to the problem of weak single-vehicle perception in autonomous driving. However, existing collaborative perception methods face a dilemma between communication efficiency and perception accuracy. To address this issue, we propose a novel communication-efficient collaborative perception framework based on supply-demand awareness and intermediate-late hybridization, dubbed as CoSDH. By modeling the supply-demand relationship between agents, the framework refines the selection of collaboration regions, reducing unnecessary communication cost while maintaining accuracy. In addition, we innovatively introduce the intermediate-late hybrid collaboration mode, where late-stage collaboration compensates for the performance degradation in collaborative perception under low communication bandwidth. Extensive experiments on multiple datasets, including both simulated and real-world scenarios, demonstrate that CoSDH achieves state-of- the-art detection accuracy and optimal bandwidth tradeoffs, delivering superior detection precision under real communication bandwidths, thus proving its effectiveness and practical applicability. The code will be released at https://github.com/Xu2729/CoSDH.
Yanan Zhang 0005, Zhi Cai, Di Huang 0001
CVPR3
2025 Test-Time Adaptive Object Detection with Foundation Model
abstract
In recent years, test-time adaptive object detection has attracted increasing attention due to its unique advantages in online domain adaptation, which aligns more closely with real-world application scenarios. However, existing approaches heavily rely on source-derived statistical characteristics while making the strong assumption that the source and target domains share an identical category space. In this paper, we propose the first foundation model-powered test-time adaptive object detection method that eliminates the need for source data entirely and overcomes traditional closed-set limitations. Specifically, we design a Multi-modal Prompt-based Mean-Teacher framework for vision-language detector-driven test-time adaptation, which incorporates text and visual prompt tuning to adapt both language and vision representation spaces on the test data in a parameter-efficient manner. Correspondingly, we propose a Test-time Warm-start strategy tailored for the visual prompts to effectively preserve the representation capability of the vision branch. Furthermore, to guarantee high-quality pseudo-labels in every test batch, we maintain an Instance Dynamic Memory (IDM) module that stores high-quality pseudo-labels from previous test samples, and propose two novel strategies-Memory Enhancement and Memory Hallucination-to leverage IDM's high-quality instances for enhancing original predictions and hallucinating images without available pseudo-labels, respectively. Extensive experiments on cross-corruption and cross-dataset benchmarks demonstrate that our method consistently outperforms previous state-of-the-art methods, and can adapt to arbitrary cross-domain and cross-category target data. Code is available at https://github.com/gaoyingjay/ttaod_foundation.
Yingjie Gao 0001, Yanan Zhang 0005, Zhi Cai, Di Huang 0001
NeurIPS3
2025 Enhanced Multi-Stage Optimization of Dynamic QoS-Aware Service Caching and Updating in Mobile Edge Computing
abstract
In the context of mobile edge computing, achieving dynamic service caching and updating to guarantee the QoS of users and reduce system costs is a challenging problem. However, existing research still has certain deficiencies in considering the dynamic behavior of users and the limited storage resources of edge servers. To address this problem, this paper investigates optimizing the service caching and updating problem within multi-stage and proposes a novel framework with three proposed strategies for the different stages to jointly optimize the delay and cost. At the initial service caching stage, we propose a basic caching strategy based on dynamic programming for the single-area scenario, taking into account the constraint of limited memory resources. To improve the caching strategy, we extend our consideration to the multiple-area scenario and design an improved algorithm based on tabu search. Given the dynamic behavior of users, we formulate the joint optimization problem as a Markov Decision Process (MDP) and design a service extension strategy based on reinforcement learning at the service updating decision-making stage and a replacement strategy taking both the distribution of service replications and service access frequency into account at the service updating replacement stage to guarantee the QoS of users. We effectively tackle the challenges arising from the dynamic behavior of users and limited storage resources. Through extensive comparative experiments, our approach outperforms traditional strategies by significantly reducing user latency and system cost.
Shuaibing Lu, Jie Wu 0001, Shuyang Zhou, Jackson Yang, Zhi Cai
IEEE Trans. Netw. Serv. Manag.8
2025 Online Elastic Resource Provisioning With QoS Guarantee in Container-Based Cloud Computing
abstract
In cloud data centers, the exponential growth of data places increasing demands on computing, storage, and network resources, especially in multi-tenant environments. While this growth is crucial for ensuring Quality of Service (QoS), it also introduces challenges such as fluctuating resource requirements and static container configurations, which can lead to resource underutilization and high energy consumption. This article addresses online resource provisioning and efficient scheduling for multi-tenant environments, aiming to minimize energy consumption while balancing elasticity and QoS requirements. To address this, we propose a novel optimization framework that reformulates the resource provisioning problem into a more manageable form. By reducing the original multi-constraint optimization to a container placement problem, we apply the interior-point barrier method to simplify the optimization, integrating constraints directly into the objective function for efficient computation. We also introduce elasticity as a key parameter to balance energy consumption with autonomous resource scaling, ensuring that resource consolidation does not compromise system flexibility. The proposed Energy-Efficient and Elastic Resource Provisioning (EEP) framework comprises three main modules: a distributed resource management module that employs vertical partitioning and dynamic leader election for adaptive resource allocation; a prediction module using$\omega$-step prediction for accurate resource demand forecasting; and an elastic scheduling module that dynamically adjusts to tenant scaling needs, optimizing resource allocation and minimizing energy consumption. Extensive experiments across diverse cloud scenarios demonstrate that the EEP framework significantly improves energy efficiency and resource utilization compared to established baselines, supporting sustainable cloud management practices.
Shuaibing Lu, Jie Wu 0001, Jackson Yang, Xinyu Deng, Zhi Cai, Juan Fang 0004
IEEE Trans. Parallel Distributed Syst.7
2024 Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
Zhi Cai, Guodong Wang 0006, Zheng Ge, Xiangyu Zhang 0005, Di Huang 0001
BMVC1
2024 Crowd-SAM: SAM as a Smart Annotator for Object Detection in Crowded Scenes
Zhi Cai, Yingjie Gao 0001, Yaoyan Zheng, Di Huang 0001
ECCV (69)1
2024 Double Layer A*: An Emergency Path Planning Model Based on Map Grid and Double Layer Search Structure
abstract
With the vigorous development of transportation infrastructure in various countries, the traffic network within the city is becoming more and more complex, and when an emergency occurs in one or more areas of the city, it will inevitably cause traffic congestion in the area and keep spreading. There are still many challenges to solve the urban emergency route planning problem. In this paper, we have employed a double layer search structure, where we have empowered the traditional A* model with a neural network, to construct a region-level dynamic path planning model known as “Double Layer A*”. The model divides the road network into two layers, and implements the outer layer and inner layer search. In the outer layer search, we use the historical cab travel data for training to achieve the general direction planning; in the inner layer search, we update the original planning according to the changes of the road condition characteristics of the regional nodes, and perform the re-planning in real time. We conducted experimental evaluations using the road network data of Beijing, and the results showed that compared to a single-layer search structure path planning model, our Double layer A* model planned paths with higher similarity in land characteristics, connectivity, and average connectivity between adjacent nodes, which demonstrates the effectiveness and reasonableness of the Double layer A* model in emergency path planning.
Zhi Cai, Zhihao Hou, Meihui Shi, Xing Su 0001, Limin Guo 0002, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2024 Heterogeneous Modular Traffic Prediction Based on Multilayer Graph Convolutional Network
abstract
Traffic patterns in the spatiotemporal network are affected by temporal dynamics and spatial correlations. The network flows have different strengths interacting at various implicit layers, and this dynamic process needs to be further explored. Predicting future traffic based on historical data from transportation IoT has been well studied, however, most of the works focus on traffic dynamics in the homogeneous spatial or temporal structure. When the spatiotemporal graph structure turns complex, it becomes a challenge to capture the deep traffic patterns on it. In this paper, a heterogeneous modular flows graph is constructed to characterize the implied spatiotemporal correlations within the traffic data. Then, we proposed a Multilayer Graph Skip Temporal Convolution Network (MGSTCN) which extracts skip aggregated representations of node status to the modular flows graph. And an extended random walk on diverse modular graphs is used to learn the spatial dependencies. The experiments based on real traffic networks confirmed that the MGSTCN has a better performance compared to the spatiotemporal homogeneous methods.
Mengmeng Chang, Zhiming Ding, Zilin Zhao, Zhi Cai
IEEE Trans. Intell. Transp. Syst.4
2024 RL-CoPref: a reinforcement learning-based coordinated prefetching controller for multiple prefetchers
abstract
Abstract Modern processors employ data prefetchers to alleviate the impact of long memory access latency. However, current prefetchers are designed for specific memory access patterns, which perform poorly on mixed applications with multiple memory access patterns. To address these issues, RL-CoPref, a reinforcement learning (RL)-based coordinated prefetching controller for multiple prefetchers, is proposed in this paper. RL-CoPref takes diverse program context information as the input, learns to maximize cumulative rewards, and evaluates prefetch quality based on prefetch hits/misses and memory bandwidth utilization. It can dynamically adjust the prefetch activation and prefetch degree, enabling multiple prefetchers to complement each other on mixed applications. Our extensive evaluation, utilizing the ChampSim simulator, demonstrates that RL-CoPref can effectively adapt to various workloads and system configurations, optimizing prefetch control. On average, RL-CoPref achieves 76.15% prefetch coverage, having 35.50% IPC improvement, outperforming state-of-the-art individual prefetchers by 5.91–16.54% and outperforming SBP, a state-of-the-art (non-RL) prefetch controller, by 4.64%.
Huijing Yang, Juan Fang 0004, Xing Su 0001, Zhi Cai, Yuening Wang
J. Supercomput.4
2023 A Prefetch-Adaptive Intelligent Cache Replacement Policy Based on Machine Learning
Huijing Yang, Min Cai, Zhi Cai
J. Comput. Sci. Technol.4
2023 What makes a readable code? A causal analysis method
abstract
Abstract Context Code readability is one of the most important quality attributes for software source code. To investigate which features affect code readability, most existing studies rely on correlation‐based methods. However, spurious correlations (a mathematical relationship wherein two variables appear to be causal but are not) involved in correlation‐based methods may affect research conclusions. Objective In order to remove spurious correlations and obtain conclusions from the perspective of causation as to what makes a readable code, we propose a causal theory‐based approach to analyze the relationship between code features and code readability scores. Method First, we adopt the PC algorithm and additive noise models to construct the causal graph on the basis of the selected code features. Then, we use the linear regression algorithm based on the back‐door criterion to obtain the causal effect of different features on code readability. Result We conduct a set of experiments on readability data labeled by human annotators. The experimental results show that the average number of comments positively impacts code readability, with each additional unit increasing the code readability score by 0.799 points. Whereas the average number of assignments, identifiers, and periods have a negative impact, with each additional unit decreasing the code readability score by 0.528, 0.281, and 0.170 points respectively. Conclusion We believe that our findings will provide developers with a better understanding of the patterns behind code readability, and guide developers to optimize their code as the ultimate goal.
Qing Mi, Zhi Cai, Xibin Jia
Softw. Pract. Exp.3
2023 VOLTCom: A Novel Online Trajectory Compression Method Based on Vector Processing
abstract
With the widespread use of the Global Positioning System (GPS) in the fields such as traffic monitoring, sports navigation, and track recording, the trajectory data recording users’ spatial and temporal information has grown dramatically. The huge volume of trajectory data causes high cost and poses a great challenge to data storage, network transmission, query and analysis. Therefore, the compression of trajectory data becomes a crucial issue. This paper proposes an online trajectory compression algorithm based on vector extraction (VOLTCom), which aims to achieve efficient data compression while retaining more effective information, and is mainly applied to trajectory recording and analysis in the traffic field. VOLTCom first generates vectors for trajectory data according to customized vector features, and then performs real-time vector extraction to achieve online trajectory compression. The vector extraction of the trajectory data ensures the stability of the compression time per unit and achieves efficient compression. Experiments on real datasets show that VOLTCom can retain the information of object velocity variation by vector density and outperforms traditional algorithms in terms of error, compression rate, and execution time. The algorithm is$O(1)$in compression time complexity and has better compression performance.
Zhi Cai, Meihui Shi, Xing Su 0001, Limin Guo 0002, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2023 A Q-Learning-Based Routing Approach for Energy Efficient Information Transmission in Wireless Sensor Network
abstract
Nowadays, wireless sensor networks have played an important role in many applications. In these applications, a large number of wireless sensors are deployed in an environment to collect information and form network to transmit collected information to the base station or sink. Because wireless sensors have to work for a long period of time without maintenance, the energy consumption of wireless sensors has a great impact on the lifetime of wireless sensor network. To reduce and balance the energy consumption of wireless sensors, many information transmission routing approaches have been proposed. However, most of them do not consider all energy consumption factors of wireless sensors. To this end, an innovative information transmission routing approach based on Q-learning is proposed in this paper, which enables wireless sensors to adaptively select suitable neighboring sensors to achieve energy efficient information transmission in a decentralized manner. Based on the proposed approach, a wireless sensor first collects state and action information of its neighboring sensors, which includes information transmission direction and distance, the remaining energy, the energy consumption for information transmission and information transmission action. Then, all this information is used to update the Q-values of neighboring sensors, so as to enable the wireless sensor to select suitable neighboring sensor to transmit information according to Q-values. From simulation experiments, it can be seen that the proposed approach enables wireless sensors to reduce and balance the energy consumption of wireless sensors and extend the lifetime of the entire wireless sensor network.
Xing Su 0001, Yiting Ren, Zhi Cai, Limin Guo 0002
IEEE Trans. Netw. Serv. Manag.3
2022 Speed and Direction Aware Skyline Query for Moving Objects
abstract
The skyline query is one of the most important supporting technologies for the location-based query services in the road network. Usually, when a user queries the skyline points in the road network, the query area is a user-centered circle or rectangle area, without considering the impact of the current movement speed and direction of the user on the formation of the query area. In this context, a speed and direction aware skyline query method is proposed, which can provide the skyline query area for the users by considering their moving speed and direction. Since the efficiency to directly obtain points of interest from speed and direction aware query area is not high, a Voronoi based speed and direction query area generation algorithm is proposed to approximate the query area, so as to improve the obtaining efficiency of points of interest in the area. The experiments on road networks and points of interest data of Beijing show the performance of the proposed method in terms of query efficiency and quality.
Zhi Cai, Xuerui Cui, Xing Su 0001, Limin Guo 0002, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2022 Prediction of Evolution Behaviors of Transportation Hubs Based on Spatiotemporal Neural Network
abstract
With the deterioration of the transportation ecosystem and traffic congestion, traffic graph representations become more complex and lack intelligibility. The deep learning method provides a new way for traffic prediction by mining the spatiotemporal relations of historical states in traffic network. Recurrent neural networks based on gate control can avoid the gradient vanishing and graph convolutional networks provide theoretical support for spatial feature extraction of traffic graph. However, in the prediction of traffic evolution behaviors, the statuses of units and groups distributed in the traffic network have strong and continuous spatiotemporal correlations. Modeling only based on the temporal or spatial perspective usually lacks comprehensiveness and semantic relevance which results in poor prediction performance. In this paper, we propose a spatiotemporal hierarchical propagation graph convolutional network(SHPGCN) to get the whole spatiotemporal evolution features of traffic graph, which can be used to predict the propagation effects and state changes of traffic flow. Considered the inadequacy of graph convolution in the learning of hierarchical features, we construct a low-dimensional propagation graph representation that projects complex node relationships into first-order neighborhoods to capture dynamic changes at different spatial and temporal scales. The SHPGCN uses graph convolution to encode the traffic propagation features and embeds them into a bidirectional recurrent sequence for traffic prediction. The experiment results show that SHPGCN can provide good prediction accuracy and robustness for the traffic evolution behaviors.
Mengmeng Chang, Zhiming Ding, Zhi Cai, Zilin Zhao
IEEE Trans. Intell. Transp. Syst.3
2021 Application of Improved YOLOv3 Algorithm in Mask Recognition
Fanxing Meng, Weimin Wei, Zhi Cai
ACIIDS3
2021 The effectiveness of data augmentation in code readability classification
Qing Mi, Yan Xiao 0002, Zhi Cai, Xibin Jia
Inf. Softw. Technol.3
2021 Continuous Road Network-Based Skyline Query for Moving Objects
abstract
With the development of location-based services and smart terminals, skyline query technique has been used widely in intelligent transportation systems. In skyline queries, the areas and keywords queried by users have a great impact on the quality of the query results and users may only be interested in the closer results of the skyline query. However, current approaches for the continuous skyline query limit the area of the skyline query to a specific area in the road network, which leads to that many useful query results cannot be retrieved. To this end, an innovative continuous skyline query approach in city range is proposed in this paper, where a multi-scale area divisions of the urban road network are provided to find the optimize query scale and area. In our approach, first, the dominant area of each intersection node in the road network is established based on the Voronoi. Then, all Points of Interest ($POI\text{s}$) are divided into the dominant area of each intersection node. After that, the intersection node aggregation algorithm ($INAA$), link remolding algorithm ($LMA$) and link fitting algorithm ($LFA$) are proposed to reduce the number of intersection nodes in the road network, so as to increase the dominant area of the remaining intersection nodes and the number of POIs in these nodes. Finally, a better query scale by considering the efficiency and quality of the query is given through the studies.
Zhi Cai, Xuerui Cui, Xing Su 0001, Limin Guo 0002, Zhining Liu 0003, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2021 Visual Analysis of Land Use Characteristics Around Urban Rail Transit Stations
abstract
Urban rail transit stations are the key nodes of urban rail transit network. Identifying and analyzing land use characteristics around urban rail transit stations can significantly contribute to urban rail transportation operation and management. Therefore, a visualization method of land use characteristics around urban rail transit stations based on POI is proposed in this paper. In the proposed method, first, the Voronoi diagram is used to determine coverage of urban rail transit stations and each POI is put in a coverage area based on their physical location. Then, topic-oriented hierarchical POIs of each urban rail transit station are extracted based on skyline idea. Finally, the land use characteristics around an urban rail transit station are visualized based on the extracted hierarchical POIs. We carried out two case studies and a quality evaluation. By using realistic data from Beijing rail transit in order to validate the method proposed in this paper. Results show that our method can clarify various situations of land use of urban rail transit stations and may provide support for the application of transportation model technology.
Zhi Cai, Gongyu Sun, Xing Su 0001, Tong Li 0001, Limin Guo 0002, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2021 Visual Navigation and Landing Control of an Unmanned Aerial Vehicle on a Moving Autonomous Surface Vehicle via Adaptive Learning
abstract
This article presents a visual navigation and landing control paradigm for an unmanned aerial vehicle (UAV) to land on a moving autonomous surface vehicle (ASV). Therein, an adaptive learning navigation rule with a multilayer nested guidance is designed to pinpoint the position of the ASV and to guide and control the UAV to fulfill horizontal tracking and vertical descending in a narrow landing region of the ASV by means of merely relative position feedback. To ensure the feasibility of the proposed control law, asymptotical stability conditions are derived based on Lyapunov stability theory. Landing experimental results are reported for a UAV-ASV system consisting of an M-100 UAV and a self-developed three-meters-long HUSTER-30 ASV on a lake to substantiate the efficacy of the proposed landing control method.
Hai-Tao Zhang, Binbin Hu, Zhecheng Xu, Zhi Cai, Tao Geng, Sheng Zhong 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 A Tensor-Based Approach for the QoS Evaluation in Service-Oriented Environments
abstract
Multi-agent technologies have been widely applied to many applications, such as in e-markets, cloud computing, service-oriented environments, etc. In real applications, service-oriented environments are open and dynamic, where loosely coupled agents interact to consume and provide services. How to accurately evaluate the potential performance (i.e., QoS) of service providers on the service requested by a service consumer in such open and dynamic environments is a challenging issue in both theory and practice. In this paper, an innovative approach is proposed to evaluate the QoS of service providers in service-oriented environments. The proposed approach first borrows the reference report mechanism from the certified reputation model, so as to efficiently collect reference reports (i.e., historical performance) of service providers in open and dynamic environments. Then, a tensor-based QoS model is proposed to construct multi-dimensional relationships between QoS evaluation factors and the QoS values of service providers based on the collected reference reports. The QoS evaluation factors include the type of services, the performance of service providers, the subjectivity of service consumers, the time slot of reference reports. Finally, a CANDECOMP/PARAFAC decomposition and gradient descent-based mechanism is used to evaluate the QoS values of service providers through completing the missing entry values in the constructed tensor. The uniform random simulation experiments indicate that the proposed approach can achieve efficient and accurate QoS evaluation in service-oriented environments with only limited collected reference reports, especially when some service providers do not have reference reports.
Xing Su 0001, Minjie Zhang 0001, Zhi Cai, Limin Guo 0002, Zhiming Ding
IEEE Trans. Netw. Serv. Manag.4
2020 CoCNN: RGB-D deep fusion for stereoscopic salient object detection
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Qixiang Ye
Pattern Recognit.5
2020 Research on Analysis Method of Characteristics Generation of Urban Rail Transit
abstract
With the development of society and economy, the urban rail transit has become one of the important components of urban transportation system, while the construction of the urban rail greatly improves the public transportation environments. Currently, there are many research focus on the passenger flow predictions according to their corresponding historical data, however, it is hard to assist transport models vary such volumes for a new station planning or being constructed. In view of this limitation, we provide a novel method for urban rail station characteristics analysis in intelligent transportation considering city land usages. Initially, point of interest (POIs) are divided by the proposed RC-tree (Colored R-tree)-based algorithm into the bounded areas for each station. Second, the Diversity and Proportion approaches are proposed to extract the top-k POIs from bounded areas based on their semantic and spatial characteristics. Then, classify the stations based on the similarity of the extracted top-k POIs. Moreover, we made a case study on real dataset, including a large volume of Automatic Fare Collection system (AFC) records for the experimental evaluations, and the results show that the proposed method can verify the rationality of land use and provide support for the application of transportation model technology.
Zhi Cai, Tong Li 0001, Xing Su 0001, Limin Guo 0002, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2020 Diversified spatial keyword search on RDF data
abstract
Abstract The abundance and ubiquity of RDF data (such as DBpedia and YAGO2) necessitate their effective and efficient retrieval. For this purpose, keyword search paradigms liberate users from understanding the RDF schema and the SPARQL query language. Popular RDF knowledge bases (e.g., YAGO2) also include spatial semantics that enable location-based search. In an earlier location-based keyword search paradigm, the user inputs a set of keywords, a query location, and a number of RDF spatial entities to be retrieved. The output entities should be geographically close to the query location and relevant to the query keywords. However, the results can be similar to each other, compromising query effectiveness. In view of this limitation, we integrate textual and spatial diversification into RDF spatial keyword search, facilitating the retrieval of entities with diverse characteristics and directions with respect to the query location. Since finding the optimal set of query results is NP-hard, we propose two approximate algorithms with guaranteed quality. Extensive empirical studies on two real datasets show that the algorithms only add insignificant overhead compared to non-diversified search, while returning results of high quality in practice (which is verified by a user evaluation study we conducted).
Zhi Cai, Georgios Kalamatianos, Georgios John Fakas, Nikos Mamoulis, Dimitris Papadias
VLDB J.1
2018 Thematic ranking of object summaries for keyword search
Georgios John Fakas, Yilun Cai, Zhi Cai, Nikos Mamoulis
Data Knowl. Eng.3
2018 Multi-vehicles dynamic navigating method for large-scale event crowd evacuations
Zhi Cai, Fujie Ren, Yuanying Chi, Xibin Jia, Lijuan Duan, Zhiming Ding
GeoInformatica1
2018 Stereoscopic saliency model using contrast and depth-guided-background prior
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Laiyun Qing
Neurocomputing5
2018 Vector-Based Trajectory Storage and Query for Intelligent Transport System
abstract
With the developing of smart sensors and mobile devices produces an increasing volume of data, and it captures the states of transportation infrastructures. Such data are collected and uploaded frequently, which forms the heavy data calculation and storage. Moreover, state of monitored object may be keeping the same or slight change according to a certain state during a period, such as moving vehicles on a certain path with a basic uniform speed. Therefore, if trajectory pattern of vehicles can be obtained through the state change mode, scale and update frequency of the data can be greatly reduced. Based on the above-mentioned ideas, we are aware that the trajectory data storage is divided into traceability storage and vector storage, where original sampled data from sensing device, and state vectors are extracted from the analysis of the original sample data. In this way, only a relatively small amount of vector data is stored. The system will not only effectively reduce the frequency of sampling data storage, but also reduce query and analysis operations involved with the amount of data. The vector function is used to represent road network with indexes, and the data query based on the road network is used to extract the semantic information. Our experimental results show that our proposed methodologies have significant improvements in intelligent transportation.
Zhi Cai, Fujie Ren, Juncheng Chen, Zhiming Ding
IEEE Trans. Intell. Transp. Syst.1
2016 Diverse and proportional size-l object summaries using pairwise relevance
Georgios John Fakas, Zhi Cai, Nikos Mamoulis
VLDB J.2
2015 Diverse and Proportional Size-l Object Summaries for Keyword Search
abstract
The abundance and ubiquity of graphs (e.g., Online Social Networks such as Google+ and Facebook; bibliographic graphs such as DBLP) necessitates the effective and efficient search over them. Given a set of keywords that can identify a Data Subject (DS), a recently proposed relational keyword search paradigm produces, as a query result, a set of Object Summaries (OSs). An OS is a tree structure rooted at the DS node (i.e., a tuple containing the keywords) with surrounding nodes that summarize all data held on the graph about the DS. OS snippets, denoted as size-l OSs, have also been investigated. Size-l OSs are partial OSs containing l nodes such that the summation of their importance scores results in the maximum possible total score. However, the set of nodes that maximize the total importance score may result in an uninformative size-l OSs, as very important nodes may be repeated in it, dominating other representative information. In view of this limitation, in this paper we investigate the effective and efficient generation of two novel types of OS snippets, i.e. diverse and proportional size-l OSs, denoted as DSize-l and PSize-l OSs. Namely, apart from the importance of each node, we also consider its frequency in the OS and its repetitions in the snippets. We conduct an extensive evaluation on two real graphs (DBLP and Google+). We verify effectiveness by collecting user feedback, e.g. by asking DBLP authors (i.e. the DSs themselves) to evaluate our results. In addition, we verify the efficiency of our algorithms and evaluate the quality of the snippets that they produce.
Georgios John Fakas, Zhi Cai, Nikos Mamoulis
SIGMOD Conference2
2014 Fast Orthogonal Haar Transform PatternMatching via Image Square Sum
abstract
Although using image strip sum, an orthogonal Haar transform (OHT) pattern matching algorithm may have good performance, it requires three subtractions to calculate each Haar projection value on the sliding windows. By establishing a solid mathematical foundation for OHT, this paper based on the concept of image square sum, proposes a novel fast orthogonal Haar transform (FOHT) pattern matching algorithm, from which a Haar projection value can be obtained by only one subtraction. Thus, higher speed-ups can be achieved, while producing the same results with the full search pattern matching. A large number of experiments show that the speed-ups of FOHT are very competitive with OHT in most cases of matching one single pattern, and generally higher than OHT in all cases of matching multiple patterns, exceeding other high-level full search equivalent algorithms.
Houjun Li, Zhi Cai
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Versatile Size-$l$ Object Summariesfor Relational Keyword Search
abstract
The Object Summary (OS)is a recently proposed tree structure, which summarizes all data held in a relational database about a data subject. An OS can potentially be very large in size and therefore unfriendly for users who wish to view synoptic information about the data subject. In this paper, we investigate the effective and efficient retrieval of concise and informative OS snippets (denoted as size-l OSs). We propose and investigate the effectiveness of two types of size- l OSs, namely size- l OS (t)s and size-l OS (a)s that consist of l tuple nodes and l attribute nodes respectively. For computing size-l OSs, we propose an optimal dynamic programming algorithm, two greedy algorithms and preprocessing heuristics. By collecting feedback from real users (e.g., from DBLP authors), we assess the relative usability of the two different types of snippets, the choice of the size- l parameter, as well as the effectiveness of the snippets with respect to the user expectations. In addition, via thorough evaluation on real databases, we test the speed and effectiveness of our techniques.
Georgios John Fakas, Zhi Cai, Nikos Mamoulis
IEEE Trans. Knowl. Data Eng.2
2013 Human eyebrow recognition in the matching-recognizing framework
Houjun Li, Zhi Cai
Comput. Vis. Image Underst.3
2011 Size-l Object Summaries for Relational Keyword Search
abstract
A previously proposed keyword search paradigm produces, as a query result, a ranked list of Object Summaries (OSs). An OS is a tree structure of related tuples that summarizes all data held in a relational database about a particular Data Subject (DS). However, some of these OSs are very large in size and therefore unfriendly to users that initially prefer synoptic information before proceeding to more comprehensive information about a particular DS. In this paper, we investigate the effective and efficient retrieval of concise and informative OSs. We argue that a good size-lOS should be a stand-alone and meaningful synopsis of the most important information about the particular DS. More precisely, we define a size-lOS as a partial OS composed oflimportant tuples. We propose three algorithms for the efficient generation of size-lOSs (in addition to the optimal approach which requires exponential time). Experimental evaluation on DBLP and TPC-H databases verifies the effectiveness and efficiency of our approach.
Georgios John Fakas, Zhi Cai, Nikos Mamoulis
Proc. VLDB Endow.2
2009 Ranking of Object Summaries
abstract
A previously proposed keyword search paradigm produces, as a query result, a ranked list of object summaries (OSs); each OS summarizes all data held in a relational database about a particular data subject (DS). This paper further investigates the ranking of OSs and their tuples as to facilitate (1) the top-k ranking of OSs and also (2) the generation of partial size-l OSs (i.e. comprised of the l most important tuples). Therefore, a global Importance score for each tuple of the database (denoted as Im(ti)) is investigated and quantified. For this purpose, ValueRank (an extension of ObjectRank) is introduced which facilitates the estimation of scores for arbitrary databases (in contrast to PageRank-style techniques that are only effective on bibliographic databases). In addition, a variation of Combined functions are investigated for assigning an Importance score to an OS (denoted as Im(OS)) and a local Importance score of their tuples (denoted as Im(OS, ti)). Preliminary experimental evaluation on DBLP and Northwind Databases is presented.
Georgios John Fakas, Zhi Cai
ICDE2