Qinghua Cao

dblp:62/9545 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-3314-9341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Network measurement and analytics · 100%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 50% Data stream processing · 50%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › cardinality estimation
distinct element counting
0.912025
TardySketch: A Framework for Cardinality Estimation Adaptable to Sliding Windows · ICDE 2025
Data stream processing › sketch
sketch-based estimation
0.912025
TardySketch: A Framework for Cardinality Estimation Adaptable to Sliding Windows · ICDE 2025
Network measurement and analytics
sketch data structures
0.912025
LocalSketch: An Accurate and Efficient Sketch for Range Spread Estimation · IEEE Trans. Dependable Secur. Comput. 2025
Network measurement and analytics › traffic measurement
spread estimation
0.912025
LocalSketch: An Accurate and Efficient Sketch for Range Spread Estimation · IEEE Trans. Dependable Secur. Comput. 2025
Network measurement and analytics
traffic measurement
0.912025
LocalSketch: An Accurate and Efficient Sketch for Range Spread Estimation · IEEE Trans. Dependable Secur. Comput. 2025
Network measurement and analytics › traffic measurement
traffic monitoring
0.312025
TardySketch: A Framework for Cardinality Estimation Adaptable to Sliding Windows · ICDE 2025

Methods — techniques the papers use, named apart from their topics

bitmap sketch · 1.7bidirectional pointer · 1.7sketch · 0.9local key aggregation · 0.9adaptive counter sizing · 0.9
YearPublicationVenuePosition
2025 TardySketch: A Framework for Cardinality Estimation Adaptable to Sliding Windows
abstract
Sliding cardinality estimation is crucial in many data analysis scenarios, e.g., detecting abnormal network behav-iors by monitoring unique connections in real time, detecting fraud in online transactions by monitoring unique user behavior patterns, and improving inventory management in supply chains by analyzing unique buyer behaviors. However, existing sliding cardinality estimation methods suffer from a cardinality barrel-down problem caused by unexpired item elimination in advance and item excessive removal, which remains unresolved so far. In this paper, we propose TardySketch, a sketch framework to make sliding cardinality estimation accurate and efficient by solving the above problem. The cornerstone of TardySketch is a Bidirectional Pointer-based Bitmap (BP-Bitmap), which stores the arrival sequence of items without timestamps. To prevent the premature elimination of unexpired items, we propose a Gap mechanism to enhance the accuracy of BP-Bitmap for identifying truly expired items through intermittent monitoring. To ensure an appropriate number of items are eliminated as the window moves, we design a Slow-Down mechanism to slacken the reset rate of bucket in BP- Bitmap to prevent over removal of items. Experimental results based on real-world datasets demonstrate that TardySketch significantly outperforms state-of-the-art methods, achieving a performance improvement of 5–40 times. The source code of TardySketch is available on GitHub.
Xuyang Jing, Qinghua Cao, Zheng Yan 0002, Wenxiu Ding, Witold Pedrycz, Pu Wang 0003
ICDE2
2025 LocalSketch: An Accurate and Efficient Sketch for Range Spread Estimation
abstract
Sketch demonstrates good properties in spread estimation over network measurements, providing fast processing and accurate estimation under limited memory usage. However, most current methods remain limited to single-flow spread estimation, resulting in suboptimal performance when applied to range spread estimation that requires measuring the spread of a range of flows. In this paper, we propose LocalSketch, a novel sketch that achieves both high estimation accuracy and memory efficiency for range spread estimation with provable theoretical guarantees. LocalSketch has two key innovations: (1) local key aggregation within predefined ranges that eliminates duplicate spread information through locality correlation, and (2) adaptive counter sizing that dynamically allocates memory resources for large-spread ranges while maintaining compact representations for low-spread ranges. LocalSketch also features an efficient abnormal bucket detection mechanism by comparing identification sign, avoiding exhaustive bucket traversal during super range detection. Moreover, the main idea of LocalSketch can be adapted to existing plug-in spread counters, which has been experimentally proved. We provide a theoretical analysis of estimation accuracy and conduct comprehensive evaluations using real-world network traffic datasets. Experimental results demonstrate that LocalSketch outperforms state-of-the-art methods by achieving 76× higher estimation accuracy for range spread estimation, while showing 15× better accuracy and 39× faster detection speed for super range identification across all datasets.
Xuyang Jing, Qinghua Cao, Zheng Yan 0002, Witold Pedrycz, Pu Wang 0003
IEEE Trans. Dependable Secur. Comput.2
2024 Dual-Track Aspect-Level Sentiment Analysis for Alleviating Cold Start in MOOC Course Reviews
abstract
In the realm of contemporary educational data mining, aspect-based sentiment analysis plays a crucial role in deciphering students’ nuanced perceptions of MOOC courses. However, sentiment analysis in educational context often encounters the prevalent challenge of cold start issues. This paper proposes a novel methodology for aspect-level sentiment analysis of course reviews, beginning with the identification of critical aspects in course reviews, followed by a comprehensive sentiment analysis at the aspect level. We introduce a Dual-Track Sentiment Analysis model (DTSA), which dynamically integrates two analytical tracks: one utilizing fine-tuned BERT model and the other employing sentiment dictionaries to effectively mitigate the cold start problem. Experimental results demonstrate the superiority of our approach over baseline models in various key metrics, particularly in addressing cold start challenges with limited review data. By incorporating a matching strategy, our model ensures reliable and timely sentiment analysis of course reviews, even with small amount of course reviews. This methodology effectively alleviates the cold start problem in aspect-level sentiment analysis in educational evaluation text, providing accurate insights when lacking sufficient initial learners’ review data and enhancing the robustness of MOOC course evaluation processes.
Bangqi Li, Qing Sun 0004, Haochun Xia, Qinghua Cao, Wenge Rong
HPCC4
2021 Cell Traffic Prediction Based on Convolutional Neural Network for Software-Defined Ultra-Dense Visible Light Communication Networks
abstract
With the explosive growth of ubiquitous mobile services and the advent of the 5G era, ultra-dense wireless network (UDN) architectures have entered daily production and life. However, the massive access capacity provided by 5G networks and the dense deployment of micro base stations also bring challenges such as high energy consumption, high maintenance costs, and inflexibility. Fiber-based visible light communication (FVLC) has the advantages of large bandwidth and high speed, which provides an efficient connection option for UDN. Thus, in order to make up for the poor flexibility of UDN, we propose a new FVLC-UDN architecture based on software-defined networks (SDNs). Specifically, SDN decouples the data plane and the control plane of the device and centralizes the control of the LED in the cell through a unified control plane, which can not only improve the resource allocation ability of the network but also transmit the data only as the data plane, reducing the manufacturing and implementation costs of the LED. In order to get a better resource allocation scheme, this paper proposes a model for predicting cell traffic based on convolutional neural networks. By predicting the traffic of each cell in the control domain, the traffic trend and cells’ status in the future period of time in the control domain can be obtained, so that a much more efficient resource allocation scheme can be formulated proactively to reduce energy consumption and balance communication loads. The experimental results show that on the real cell traffic dataset, this method is better than the existing prediction methods when the size of training dataset is limited.
Shanjun Zhan, Lisu Yu, Zhen Wang 0022, Yichen Du, Qinghua Cao, Shuping Dang, Zahid Khan
Secur. Commun. Networks6
2020 An open source engineering practice assistant training system based on virtual reality
abstract
The Engineering training course takes the responsibility of cultivating students' innovation, team cooperation and practical operation ability. Now there are two ways to improve the innovation training performance in the engineering practical course. The most popular way is to carry on a project-based curriculum which has the practical training as one section of the course. Another way is to create new practice teaching methods. Our research is a combination of the project curriculum and the innovative training methods. It is focused on the Virtual reality (VR) for the engineering practice teaching. First of all, the project-based course named "Comprehensive innovation training" was offered for the junior or senior students. VR methods applied for the practice course was one of the projects in the course. Instructors built the initial model from the real practice situation into the virtual environment by the software Unreal Engine (UE) 4, and guided students the basic operation of the software. Then the students in the class would design the motion of the equipment based on the reality rules, finish the operating system with the VR device, and debug the program for the final application. Thirdly, this project was applied as an assistant practice method for the sophomores in their practice training class, and the students would give feedbacks to the system. Finally, when those sophomore students become juniors and seniors, they could involve in the project-based course, and contribute to the VR training project. Now the project is carried on, and the model of drilling machine and lathe machine are established in the virtual environment. The engineering practice assistant training system based on VR has already applied for the junior students, and some of feedbacks have been received. There is 81 percent of the students shown interest in the new method. As the platform depends on the open source software, and the whole project is sustainable and the system will be abundant in the next few years.
Dan Zhang 0006, Xueling Luo, Qinghua Cao, Zhilong Li
FIE4
2017 An Improved Approach to Traceability Recovery Based on Word Embeddings
abstract
Software traceability recovery, which reconstructs links between software artifacts, has become more and more vital to maintaining a software life cycle with the increase of software scale and complexity of software architecture. However, existing approaches mainly rely on information retrieval (IR) techniques. These methods are not very efficient at complex software artifacts which are mixed with multilingual texts, code snippets and proper nouns. Moreover, it is hard to predict new traceability links with existing approaches when requirements are changed or software functions are added, since these methods have not made the most of the final ranked lists. In this paper, we propose a novel approach WELR, based on word embeddings and learning to rank to recover traceability links. We use word embeddings to calculate semantic similarities between software artifacts and bring in query expansion and a weighting strategy during calculation. Different from other work, we leverage learning to rank to build prediction models for traceability links. We conducted experiments on five public datasets and took account of traceability links among different kinds of software artifacts. The results show that our method outperforms the state-of-the-art method that works under the same conditions.
Qinghua Cao, Qing Sun 0004
APSEC2
2016 Course Relatedness Based on Concept Graph Modeling
Jingwen Pang, Qinghua Cao, Qing Sun 0004
CollaborateCom2
2012 Applying adaptive over-sampling technique based on data density and cost-sensitive SVM to imbalanced learning
abstract
Resampling method is a popular and effective technique to imbalanced learning. However, most resampling methods ignore data density information and may lead to overfitting. A novel adaptive over-sampling technique based on data density (ASMOBD) is proposed in this paper. Compared with existing resampling algorithms, ASMOBD can adaptively synthesize different number of new samples around each minority sample according to its level of learning difficulty. Therefore, this method makes the decision region more specific and can eliminate noise. What's more, to avoid over generalization, two smoothing methods are proposed. Cost- Sensitive learning is also an effective technique to imbalanced learning. In this paper, ASMOBD and Cost-Sensitive SVM are combined. Experiments show that our methods perform better than various state-of-art approaches on 9 UCI datasets by using metrics of G-mean and area under the receiver operation curve (AUC).
Senzhang Wang, Zhoujun Li 0001, Wen-Han Chao, Qinghua Cao
IJCNN4