Chao Li 0012

dblp:66/190-12 · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-6844-6127ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Edge-Optimized Voice Control with 0.26 M Parameters: Distilling 86M Adaptive Window Audio Transformer for Real-World Variable-Length Inputs
Pinze Ren, Zhen Chen 0001, Yinjun Wu, Weiran Lin, Qilong Shi, Chao Li 0012, Jianxin Yang
IEEE Big Data6
2025 Enhancing Chain-of-Thought Reasoning for Text-to-SQL with Effective Retrieval-Augmented Generation
Xuguang Zhu, Yong Zhang 0002, Chao Li 0012, Chunxiao Xing
DASFAA (1)3
2024 An Empirical Study on the Power Consumption of LLMs with Different GPU Platforms
abstract
This paper researches on the power consumption of AIGC applications based on LLM with different parameter scales across different hardware platforms. Artificial Intelligence Generated Content (AIGC) represents a leading-edge application of AI technology, primarily driven by large language models (LLMs) and their associated technologies. The deployment of LLM typically relies on critical facilities with three layers, i.e., the hardware, model, and application layers. This empirical study aims to identify key factors in power consumption when a large model is serving in the inference stage, which will hint the insights for improving the energy efficiency of computational infrastructures. In the context of the "dual carbon" goals, i.e., carbon peaking and carbon neutrality, this study aims to find an effective way to reduce the energy cost of AIGC applications, thereby supporting sustainable AI development in industry.
Zhen Chen 0001, Weiran Lin, Xinyu Xie, Yaodong Hu, Chao Li 0012, Qiaojuan Tong, Yinjun Wu, Shuangshou Li
IEEE Big Data5
2023 HuaBaseChain: An Extensible Blockchain With High Performance
abstract
Blockchain has been extensively used on the Internet of Things (IoT). However, several problems exist in current blockchain systems that limit their usability in IoT networks. First, Proof-of-Work (PoW) consensus most used in current blockchain systems generates excessive energy consumption. Second, the single digital token system limits the development of blockchain in multiple IoT applications. Finally, the large amount of data generated by IoT devices and stored in blockchains has huge storage requirements. To deal with these problems, we propose HuaBaseChain—an extensible blockchain with high performance. HuaBaseChain has three features. First, HuaBaseChain uses a novel Proof-of-Participation (PoP) consensus to reduce energy consumption. Second, HuaBaseChain extends the single digital token model to a multidimensional digital token model, enabling more varied applications that can be tailored to the needs of IoT device managers. Third, HuaBaseChain devises an underlying Merkle forest data structure to reduce storage requirements. HuaBaseChain is suited for IoT applications that contain a variety of IoT devices and generate massive amounts of data. We demonstrate the efficiency of HuaBaseChain on IoT data sets and the significantly larger Bitcoin data set. Our experiments show that HuaBaseChain runs efficiently, with significantly reduced energy consumption, greater efficiency, lower storage requirements, and higher query speed.
Xiangke Mao, Chao Li 0012, Yong Zhang 0002, Guigang Zhang, Jiafu Li, Mira Shah, Chunxiao Xing
IEEE Internet Things J.2
2022 A Research on the Theory and Technology of Trusted Transaction in Modern Service Industry
Guigang Zhang, Chao Li 0012, Yong Zhang 0002, Chunxiao Xing
WISA5
2022 CrowdMed-II: a blockchain-based framework for efficient consent management in health data sharing
abstract
The healthcare industry faces serious problems with health data. Firstly, health data is fragmented and its quality needs to be improved. Data fragmentation means that it is difficult to integrate the patient data stored by multiple health service providers. The quality of these heterogeneous data also needs to be improved for better utilization. Secondly, data sharing among patients, healthcare service providers and medical researchers is inadequate. Thirdly, while sharing health data, patients' right to privacy must be protected, and patients should have authority over who can access their data. In traditional health data sharing system, because of centralized management, data can easily be stolen, manipulated. These systems also ignore patient's authority and privacy. Researchers have proposed some blockchain-based health data sharing solutions where blockchain is used for consensus management. Blockchain enables multiple parties who do not fully trust each other to exchange their data. However, the practice of smart contracts supporting these solutions has not been studied in detail. We propose CrowdMed-II, a health data management framework based on blockchain, which could address the above-mentioned problems of health data. We study the design of major smart contracts in our framework and propose two smart contract structures. We also introduce a novel search contract for searching patients in the framework. We evaluate their efficiency based on the execution costs on Ethereum. Our design improves on those previously proposed, lowering the computational costs of the framework. This allows the framework to operate at scale and is more feasible for widespread adoption.
Chaochen Hu, Chao Li 0012, Guigang Zhang, Zhiwei Lei, Mira Shah, Yong Zhang 0002, Chunxiao Xing, Jinpeng Jiang, Renyi Bao
World Wide Web2
2021 DaaS: Internet-perception big data systems based on AI
abstract
The DaaS (Data as a Service) is an Internet-perception big data system based on AI, which is built by "Think Tank 2861 Project Team". This is an Internet-area, data-based, and neural feedback system for the Internet information in China. It takes Internet activities as the input, and processes through AI algorithms and machine learning framework to generate the output, based on which building the real-time macro economics and society big data for about 9.8 million grids in China and its intelligent applications. DaaS covers all 2,861 administrative districts and counties in the country and is refined to geography grid of one square kilometer granularity. The real-time objective information generated by distributed AI algorithms, that are constantly trained and calibrated, is of great value in scientific research and commercial applications.
Zexuan Lyu, Chao Li 0012, Guigang Zhang, Chunmei Huang, Mengyuan Du
IEEE BigData3
2020 An Experimental Study of Time Series Based Patient Similarity with Graphs
Kalkidan Fekadu Eteffa, Samuel Ansong, Chao Li 0012, Ming Sheng, Yong Zhang 0002, Chunxiao Xing
WISA3
2020 DSQA: A Domain Specific QA System for Smart Health Based on Knowledge Graph
Ming Sheng, Yuelin Bu, Yong Zhang 0002, Xin Li 0111, Chao Li 0012, Chunxiao Xing
WISA7
2019 How to Empower Disease Diagnosis in a Medical Education System Using Knowledge Graph
Samuel Ansong, Kalkidan Fekadu Eteffa, Chao Li 0012, Ming Sheng, Yong Zhang 0002, Chunxiao Xing
WISA3
2019 Application of Patient Similarity in Smart Health: A Case Study in Medical Education
Kalkidan Fekadu Eteffa, Samuel Ansong, Chao Li 0012, Ming Sheng, Yong Zhang 0002, Chunxiao Xing
WISA3
2019 Anti-money Laundering (AML) Research: A System for Identification and Multi-classification
Yixuan Feng, Chao Li 0012, Jian Wang 0029, Guigang Zhang, Chunxiao Xing, Zengshen Lian
WISA2
2019 CLMed: A Cross-lingual Knowledge Graph Framework for Cardiovascular Diseases
Ming Sheng, Han Zhang 0054, Yong Zhang 0002, Chao Li 0012, Chunxiao Xing, Yuyao Shao
WISA4
2019 Learning from User Social Relation for Document Sentiment Classification
Kangzhi Zhao, Yong Zhang 0002, Chunxiao Xing, Chao Li 0012
DASFAA (2)5
2015 A Package Generation and Recommendation Framework Based on Travelogues
abstract
Tourism has become the world's largest economy industry. More and more people share their travelogues on travel websites. Recommender system is an effective tool to provide travel services (e.g., Landscapes selection) for tourists. Many recommender systems are based on travel data that are supplied by travel agencies, and provide travel packages from a fixed package set, which bring two challenges for travel package recommender system. One is how to generate more travel packages. The other is how to measure more fine-grained user similarity. To address these challenges, we develop a package generation and recommendation framework to help travelers select landscapes. Firstly, we propose a Fuzzy Clustering based Package Generation algorithm (FCPG) to generate new travel packages to improve the overall recommendation effectiveness. Then, we develop a Dual Topic Model based Package Recommendation algorithm (DTMPR). It considers two user-related topics (travel seasons and areas), and provides more fine-grained user similarity measure. Experimental results show the superiority of our framework in comparison with the state-of-the-art methods.
Xinhuan Chen, Yong Zhang 0002, Chao Li 0012, Chunxiao Xing
COMPSAC4
2015 RMDN: New Approach to Maximize Influence Spread
abstract
Influence maximization modeling and analyzing is an important problem in Online Social Networks (OSNs). Influential nodes provide information on hot topics and forward interesting information, which can yield great influence on other nodes in OSNs. For word-of-mouth viral marketing, it is critical to find influential nodes within budgets. However, finding the optimal solution to maximize the influence spread in a given OSN has been proven to be NP-Hard. The existing algorithms suffer the following three defects: (1) need acquire the topological structure of the network, which is impractical for the continuously changing networks in real life, (2) are easy to get trapped in Rich Club phenomena, (3) can not balance very well between influence spread and running time. To solve this challenging problem, based on the randomly heuristic algorithm and the scale-free property of Complex Networks, we propose RMDN (Random Maximal Degree Neighbor) and its improved version RMDN++. We prove the feasibility and scalability of RMDNs in theory. Five real datasets are used to test the effectiveness and efficiency of RMDNs under two different influence diffusion models. The result shows that our methods have a comparable performance in terms of influence spread as state-of-the-art algorithms, but decrease the time by 1-2 orders of magnitude.
Qingcheng Hu, Yong Zhang 0002, Xinhui Xu, Chao Li 0012, Chunxiao Xing
COMPSAC4
2014 A LDA-Based Algorithm for Length-Aware Text Clustering
Xinhuan Chen, Yong Zhang 0002, Yanshen Yin, Chao Li 0012, Chunxiao Xing
APWeb4
2014 TL: A High Performance Buffer Replacement Strategy for Read-Write Splitting Web Applications
Zhiwen Jiang, Yong Zhang 0002, Jin Wang 0007, Chao Li 0012, Chunxiao Xing
APWeb4
2014 Continuous Temporal Top-k Query over Versioned Documents
Chao Lan, Yong Zhang 0002, Chunxiao Xing, Chao Li 0012
WAIM4
2012 Implementation of Space Optimized Bisecting K-Means (BKM) Based on Hadoop
abstract
This article is composed in the background of the study of scientific field of coauthors phenomenon factual basis. By the study of massive amounts of relational data, it provides us with major significances theoretically and practically on retrieving and obtaining professionally academic information and getting knowing of academic development trend of miscellaneous fields. In process of studying this type of project, the problem of cluttering for coauthors that are in the data is involved. However, it is hard to meet the need of implementing the analysis of massive amounts of data cluttering by the existing cluttering software and algorithms, for this reason, finding an approach to deal with this kind of question is toughly important. To solve this question, this article presents an optimized Bisecting K-Means (BKM) clustering algorithm based on Hadoop and states the fashion of how to optimize the algorithm and the key point of implementing in details after analyzing the status quo related to this study. Estimating the complexity of the algorithm by experiments indicates the current problems and the direction for the future study.
Yanshen Yin, Chengguang Wei, Guigang Zhang, Chao Li 0012
WISA4
2012 DataCloud: An Efficient Massive Data Mining and Analysis Framework on Large Clusters
abstract
With the development of cloud computing technologies, big data processing is becoming more and more important. How to mine and analyze massive data is facing a very big challenge. In this paper, we proposed an efficient massive data mining and analysis framework Data Cloud on large clusters. The most important part of Data Cloud is the Rabbit. It is a kind of massive data mining and analysis processing plan framework on the large clusters like the Pig and Hive. We make a detail analysis about the Rabbit plan.
Guigang Zhang, Chao Li 0012, Yong Zhang 0002, Chunxiao Xing
WISA2
2012 SemanMedical: A kind of semantic medical monitoring system model based on the IoT sensors
abstract
With the development of IoT technologies, more and more medial sensors have been used to monitor people's health. In this paper, we design a kind of semantic medical monitoring system model in the cloud based on the IoT sensors. Lots of IoT sensors will accept the massive sensor data every time. All these massive sensor data will be stored in the HDFS and some information will be stored into the HUADING-S, a column-based database. All these HUADING-S data will connect to the medical rule engine. When the users' or patients' health indicators' data is beyond the normal range, the medical rule engine will send the alert information to the users or patients. We design two algorithms: (1) massive semantic medical rules processing algorithm without external communication and (2) massive semantic medical rules processing algorithm with external communication. Our simulation experiment shows that the algorithm 2 will reduce the time cost a lot and improve the executive efficiency.
Guigang Zhang, Chao Li 0012, Yong Zhang 0002, Chunxiao Xing
Healthcom2
2012 SemanMR: big data processing framework based on semantics
abstract
In this paper, we design a kind of big data processing framework SemanMR (Semantic MapReduce). SemanMR is a programming framework based on the Hadoop MapReduce programming model. SemanMR provide a kind of bid data processing mechanism based on the metadata cluster of distributed file systems or cloud databases. In addition, we add some semantic index on the big data, and so it will improve our processing efficiency in SemanMR. SemanMR is a kind of big data processing internetware in the cloud environment.
Guigang Zhang, Chao Li 0012, Chunxiao Xing, Yong Zhang 0002
Internetware2
2012 A Packaging Approach for Massive Amounts of Small Geospatial Files with HDFS
Jifeng Cui, Yong Zhang 0002, Chao Li 0012, Chunxiao Xing
WAIM3
2010 A methodology for measuring the preservation durability of digital formats
abstract
It is now widely recognized that appropriate measures are required for digital preservation to ensure that digital data can be accessed and used currently and in the future. Among all the risks of digital preservation, format obsolescence is one of the most important. There have been several projects or initiatives dealing with the measurement method of format obsolescence risk, but there has been no mechanism to quantify the preservation risk or durability of digital formats based on a self-improving assessment model, executed with the aid of computers. This paper deals with a methodology for measuring the preservation durability of digital formats, especially for their risk assessment. This method is based on a quantitative assessment model for format risk, and can shift the non-quantifiable knowledge or experiences of field experts to a machine identifiable and processible form, or ‘risk scores’. Results can be recognized and communicated by computers automatically and formally, which can assist in the automatic/semi-automatic risk management for digital preservation, sharing this quantified knowledge among communities. Because technologies are changing quickly, the quantitative assessment model for risks will not be a status quo situation. Thus, also presented is a method to fine tune the quantitative assessment model for risk of formats through a self-learning and self-improving style.
Chao Li 0012, Xiaohui Zheng, Xing Meng, Li Wang 0042, Chunxiao Xing
J. Zhejiang Univ. Sci. C1