VLDB 2026 Research / reviewers in the wild / expert
Di Cai
dblp:63/3479
· DBLP profile ↗
9ranked-venue papers
6as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 87% GPUs and heterogeneous computing · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel scheduling
adaptive scheduling |
0.7 | 1 | 2023 | Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUs · SC 2023 |
Parallel and multicore computing › parallel algorithms › dynamic programming
parallel dynamic programming |
0.7 | 1 | 2023 | Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUs · SC 2023 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU scaling |
0.2 | 1 | 2023 | Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUs · SC 2023 |
Knowledge graphs
concept relatedness |
0.1 | 1 | 2010 | An Information-Theoretic Foundation for the Measurement of Discrimination Information · IEEE Trans. Knowl. Data Eng. 2010 |
Methods — techniques the papers use, named apart from their topics
factored dataflow · 0.7dynamic programming · 0.7information-theoretic measures · 0.2information-theoretic measure · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How enterprise social media facilitate proactive socialization behavior and compensate for mentoring: The role of communication visibility
Di Cai, Shengming Liu |
Inf. Manag. | 1 |
| 2023 | Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUsabstractGlobal ocean data assimilation is a crucial technique to estimate the actual oceanic state by combining numerical model outcomes and observation data, which is widely used in climate research. Due to the imbalanced distribution of observation data in global ocean, the parallel efficiency of recent methods suffers from workload imbalance. When massive GPUs are applied for global ocean data assimilation, the workload imbalance becomes more severe, resulting in poor scalability. In this work, we propose a novel adaptive workload-balance scheduling strategy, Bassimilation, which successfully estimates the total workload prior to execution and ensures a balanced workload assignment. Further, we design a parallel dynamic programming approach to accelerate the schedule decision, and develop a factored dataflow to exploit the parallel potential of GPUs. Evaluation demonstrates that our algorithm outperforms the state-of-the-art method by up to 9.1× speedup. This work is the first to scale global ocean data assimilation to 4, 000 GPUs. Junmin Xiao, Chaoyang Shui, Di Cai, Kangyu Wang, Yunfei Pang, Guangming Tan |
SC | 3 |
| 2020 | Recurrence Behavior Statistics of Blast Furnace Gas Sensor Data in Industrial Internet of ThingsabstractBlast furnace gas (BFG) produced from steel industries is generally one of the most important energy supplies in enterprises. Due to a great deal of output, fluctuation, and heterogeneity in data, it is very difficult to provide profound insights into its internal dynamic. In this article, a novel analysis framework is developed for the BFG data processing, considering the recurrence plot (RP) and the recurrence quantification analysis (RQA). The specific aim is to investigate the relationship between BFG output and its potential influencing factors. This framework can be deemed as a uniform and consistent system with functional components of qualitative visualization and quantitative analysis. Concretely, the BFG outputs related to five factors are separately projected to high-dimensional spaces, followed by that their internal dynamics can be embodied through a 2-D recurrence representation of states. Finally, five RQA parameters are used to quantify the influence of these factors on the BFG output. This is the first attempt revealing the relations among the BFG data from the qualitative and quantitative perspectives. The experimental results show that RP can discover the BFG output patterns of laminar state, chaos, and instability over given three states of influencing factors, and the ranked influential degree can be given by a two-stage standard deviation of all considered RQA measures. Besides, we also demonstrate that the temperature of the hot-blast stove is most relevant to the BFG output, while the influence of considered other factors directly depends on the selected series length. Yingqi Li, Di Cai, Jialin Wang 0001, Xiaochuan Sun, Haijun Zhang 0001, Ning Wang 0017 |
IEEE Internet Things J. | 2 |
| 2010 | Sentiment in short strength detection informal textabstractAbstract A huge number of informal messages are posted every day in social network sites, blogs, and discussion forums. Emotions seem to be frequently important in these texts for expressing friendship, showing social support or as part of online arguments. Algorithms to identify sentiment and sentiment strength are needed to help understand the role of emotion in this informal communication and also to identify inappropriate or anomalous affective utterances, potentially associated with threatening behavior to the self or others. Nevertheless, existing sentiment detection algorithms tend to be commercially oriented, designed to identify opinions about products rather than user behaviors. This article partly fills this gap with a new algorithm, SentiStrength, to extract sentiment strength from informal English text, using new methods to exploit the de facto grammars and spelling styles of cyberspace. Applied to MySpace comments and with a lookup table of term sentiment strengths optimized by machine learning, SentiStrength is able to predict positive emotion with 60.6% accuracy and negative emotion with 72.8% accuracy, both based upon strength scales of 1–5. The former, but not the latter, is better than baseline and a wide range of general machine learning approaches. Mike Thelwall, Kevan Buckley, Georgios Paltoglou, Di Cai, Arvid Kappas |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2010 | An Information-Theoretic Foundation for the Measurement of Discrimination InformationabstractHitherto, it has not been easy to interpret the meaning of the amount of discrimination information conveyed in a term rationally and explicitly within practical application contexts; it has not been simple to introduce the concept of the extent of semantic relatedness between two terms meaningfully and successfully into scientific discussions. This study is part of an attempt to do this. We attempt to answer two important questions: (1) What is the discrimination information conveyed by a term and how to measure it? (2) What is the relatedness between two terms and how to estimate it? We focus on the first question and present an in-depth investigation into the discrimination measures based on several information measures, which are widely used in a variety of applications. The relatedness measures are then naturally defined according to the individual discrimination measures. Some key points are made for clarifying potential problems arising from using the relatedness measures, and solutions are suggested. Two example applications in the contexts of text mining and information retrieval are provided. The aim of this study, of which this paper forms part, is to establish a unified theoretical framework, with measurement of discrimination information (MDI) at the core, for achieving effective measurement of semantic relatedness (MSR). Due to its generality, our method can be expected to be a useful tool with a wide range of application areas. Di Cai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Learning semantic relatedness from term discrimination information
Di Cai, C. J. van Rijsbergen |
Expert Syst. Appl. | 1 |
| 2009 | Determining semantic relatedness through the measurement of discrimination information using Jensen differenceabstractMeasurement of semantic relatedness has been addressed in a number of application tasks and by researchers in a variety of disciplines. Measurement of discrimination information of terms is a fundamental issue for many areas of science. In this study, we attempt to introduce relatedness measures based on discrimination measures, with the aim of making fundamental concepts accessible and usable to the broad community of data analysis practitioners. We present an in-depth investigation into the basic concept of discrimination information conveyed in a term based on Jensen difference. The discrimination measures can then naturally and conveniently be utilized to introduce two generic concepts of semantic relatedness. We also address the issue of estimating arguments embedded in the relatedness measures and then demonstrate how our method can be supported by empirical evidence drawn from performance experiments. © 2009 Wiley Periodicals, Inc. Di Cai |
Int. J. Intell. Syst. | 1 |
| 2008 | An algorithm for modelling key termsabstractThe ability to formally analyse and represent semantic relations of terms is a major challenge for many areas of computing science and an intriguing problem for other sciences. In applications of evidence theory to, for instance, information retrieval, the problem of analysis and representation becomes apparent because evidence theory is based on set theory and individual key terms have to be modelled as subsets of the frame of discernment. How to find the frame and model the key terms is a challenge. The problem leads to other practical problems, as pointed out repeatedly in the literature. In this study, we focus on such a problem, present a method for simplifying and normalizing a thesaurus, and propose an algorithm for establishing the frame of discernment and for modelling individual key terms as a subset of the frame. The key aim of this study is to treat semantic relations of terms by means of a normalized thesaurus. © 2008 Wiley Periodicals, Inc. Di Cai, C. J. van Rijsbergen |
Int. J. Intell. Syst. | 1 |
| 2001 | Automatic Query Expansion Based on DivergenceabstractIn this paper we are mainly concerned with discussion of a formal model, based on the basic concept of divergence from information theory, for automatic query expansion. The basic principles and ideas on which our study is based are described. A theoretical framework is established, which allows the comparison and evaluation of different term scoring functions for identifying good terms for query expansion. The approaches proposed in this paper have been implemented and evaluated on collections from TREC. Preliminary results show that our approaches are viable and worthy of continued investigation. Di Cai, C. J. van Rijsbergen, Joemon M. Jose |
CIKM | 1 |