Jie Song 0001

dblp:09/4756-1 · DBLP profile ↗
← Back
36ranked-venue papers
17as first author
18since 2021 · last 2026
0000-0003-0704-3217ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Syllogism-Inspired TableQA: Evidentialization Makes Decomposition Reasoning and Answer Verification More Reliable
abstract
Existing large language model (LLM)-based table question answering (TableQA) methods primarily involve decomposition reasoning and answer verification processes. However, decomposing questions solely at the semantic level, without considering the factual evidence in tables, fails to significantly reduce the difficulty for LLMs in understanding the key information in questions. Furthermore, reasoning and verification without supporting factual evidence are often arbitrary and unreliable. In light of these issues, this paper proposes a Syllogism-Inspired Reasoning and Verification method (SIRV), which performs reliable decomposition reasoning and answer verification based on the evidential concept of syllogism. Specifically, SIRV extracts question-relevant factual evidence from the table to construct the premises. Based on the constructed premises, SIRV plans reasoning paths and generates sub-questions that explicitly indicate relevant factual evidence, performing evidence-centered reasoning. Additionally, SIRV examines the consistency between the premises and the table to focus on factual evidence, thereby reliably identifying and correcting errors in the reasoning process. Compared to state-of-the-art methods, SIRV achieves performance improvements of up to 5.24% in single-mode and 2.89% in joint reasoning, while also demonstrating excellent generalization ability and efficiency.
Zhe Zhang 0023, Lili Bai, Chaopeng Guo, Jie Song 0001
AAAI4
2026 Dimension Expansion for Learning Money Laundering Activities Hidden in Transaction Network
abstract
Due to the anonymity and decentralization inherent in blockchains, they have gradually become a new breeding ground for money laundering. Anti-money laundering (AML) in blockchains is urgently needed. A conventional AML method involves analyzing transaction features derived from blockchain records that define money transfers within a transaction network. However, this method often fails to identify all illicit transactions, as some are obscured within the network due to their indistinguishable features compared to both known licit and illicit transactions. Fortunately, all transactions involve at least two addresses (accounts), and the account features of illicit transactions are distinguishable because they reflect the social information of the parties involved. This insight provides a pathway to uncover hidden illicit transactions. In this article, we propose the account enhanced transactions classification ($\mathsf{ATC}$) model to enhance the performance of illicit transaction detection. The$\mathsf{ATC}$expands the dimensionality of transaction feature vectors by concatenating the corresponding accounts, represented as a variable number of high-dimensional feature vectors. In addition,$\mathsf{ATC}$employs clustering, pooling, and persistent homology techniques to address the challenges associated with learning from these dimension-expanded features. Experimental results on the Elliptic++ dataset demonstrate the advantages of$\mathsf{ATC}$with an accuracy of 98.28%, precision of 99.64%, recall of 91.85%, and an F1 score of 95.59%, significantly outperforming other comparative methods.
Jie Song 0001, Yu Gu 0002, Ge Yu 0001
IEEE Trans. Comput. Soc. Syst.1
2026 Consensus and Computing Integration for Processing Transactional Graphs in Consortium Blockchain
Pengxuan Ma, Chaopeng Guo, Yu Gu 0002, Jie Song 0001
IEEE Trans. Knowl. Data Eng.5
2025 L2SM: a query-optimized linked LSM-tree for HTAP workloads
Xiaoyue Feng, Dashan Wei, Chaopeng Guo, Jie Song 0001
Frontiers Comput. Sci.4
2025 Adaptive container auto-scaling for fluctuating workloads in cloud
Xiaoyue Feng, Tianzhe Jiao, Chaopeng Guo, Jie Song 0001
Future Gener. Comput. Syst.5
2025 Enhancing Text-to-SQL generation with language sequential consistency
Zhe Zhang 0023, Chaopeng Guo, Jie Song 0001, Guangyu He
Neurocomputing4
2025 Illicit Social Accounts? Anti-Money Laundering for Transactional Blockchains
abstract
In recent years, blockchain anonymity has led to more illicit accounts participating in various money laundering transactions. Existing studies typically detect money laundering transactions, known as AML (Anti-money Laundering), through learning transaction features on transaction graphs of transactional blockchains. However, transaction graphs fail to represent the accounts’ social features within transactional organizations. Account graphs reveal such features well, and detecting illicit accounts on account graphs provides a new perspective on AML. For example, it helps uncover illegal transactions whose transaction features are not distinct in transaction graphs, with a loose assumption that illicit accounts are likely involved in illegal transactions. In this paper, we propose a Social Attention Graph Neural Network ($\textsf {SGNN}$) on account graphs converted from transaction graphs. To detect illicit accounts,$\textsf {SGNN}$learns the social features on two sub-graphs, a heterogeneous graph and a hypergraph, extracted from the account graph, and fuses these features into account attribute vectors through attention. The experimental results on the Elliptic++ dataset demonstrate$\textsf {SGNN}$’s advances. It outperforms the best baseline by 14.18% in precision, 7.37% in F1 score, 0.96% in accuracy, and 0.64% in recall when detecting illicit accounts on account graphs, as well as detects 20.3% more recall of illegal transactions through these illicit accounts than state-of-the-art methods based on transaction graphs when the mappings between illegal transactions and illicit accounts are provided. Moreover, thanks to social features,$\textsf {SGNN}$has a novel capability that works under many account scales and activity degrees. We release our code onhttps://github.com/CloudLab-NEU/SGNN.
Jie Song 0001, Yu Gu 0002, Ge Yu 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Learning code better through structural information of data flow
Zhe Zhang 0023, Tianzhe Jiao, Lili Bai, Chaopeng Guo, Jie Song 0001
J. Supercomput.6
2024 JAPO: learning join and pushdown order for cloud-native join optimization
Yuchen Yuan, Xiaoyue Feng, Jie Song 0001
Frontiers Comput. Sci.5
2024 Introducing on-chain graph data to consortium blockchain for commercial transactions
Yuchen Yuan, Jie Song 0001, Yu Gu 0002, Qiang Qu 0001, Yongjie Bai
Frontiers Comput. Sci.3
2024 CloudSimPer: Simulating Geo-Distributed Datacenters Powered by Renewable Energy Mix
abstract
Nowadays, studies on energy-efficient datacenters, especially the DataCenters powered by Renewable Energy mix (DCRE), have gained great attention. DCREs are large-scale, geo-distributed, and equipped with on-site renewable energy generators. For these features, it is expensive to perform empirical evaluations of proposed algorithms and solutions on the real-world DCREs, while the state-of-the-art datacenter simulators are not applicable for DCREs. In this paper, we present CloudSimPer (CLOUD SIMulator hybrid-Powered by rEnewable eneRgy), a general-purpose simulator that comprehensively supports the simulation of DCREs. Besides the functions such as renewable energy, geo-distribution, and long-term simulation, we also design evaluation metrics and an integrated simulation case for experimental studies in the future. The main challenge of CloudSimPer lies in designing a new model and software layer upon CloudSim, to solve the complexity of traceable and comparable simulations which connect renewable energies, datacenters, workloads, regions, and schedulers. We use the term schedulers broadly, encompassing any optimization approaches on DCREs for energy saving. We prove CloudSimPer and integrated case to be valid, so that simulation results are scientifically sound, by examining the expectation and the simulation results, and comparing the simulation results with selected competitors. CloudSimPer offers simulation services to datacenter designers, datacenter administrators, and academics.
Jie Song 0001, Peimeng Zhu, Yanfeng Zhang 0001, Ge Yu 0001
IEEE Trans. Parallel Distributed Syst.1
2024 A periodic requests dispatcher for energy optimization of hybrid powered data centers
Chaopeng Guo, Gujun Lu, Jie Song 0001
Wirel. Networks4
2023 Learning Optimal Tree-Based Index Placement for Autonomous Database
Xiaoyue Feng, Tianzhe Jiao, Chaopeng Guo, Jie Song 0001
DEXA (1)4
2023 Why blockchain needs graph: A survey on studies, scenarios, and solutions
Jie Song 0001, Qiang Qu 0001, Yongjie Bai, Yu Gu 0002, Ge Yu 0001
J. Parallel Distributed Comput.1
2023 Towards an Energy Complexity Model for Distributed Data Processing Algorithms
abstract
Modern data centers exist as infrastructure in the era of Big Data. Big data processing applications are the major computing workload of data centers. Electricity cost accounts for about 50% of data centers’ operational costs. Therefore, the energy consumed for running distributed data processing algorithms on a data center is starting to attract both academia and industry. Most works study the energy consumption from the hardware perspective and only a few of them from the algorithm perspective. A general and hardware-independent energy evaluation model for the algorithms is in demand. With the model, algorithm designers can evaluate the energy consumption, compare energy consumption features and facilitate energy consumption optimization of distributed data processing algorithms. Inspired by the time complexity model, we propose an energy complexity model for describing the trends that an algorithm's energy consumption grows with the algorithm's input size. We argue that a good algorithm, especially for processing Big Data, should have a ‘small’ energy complexity. We define$E(n)$to represent the functional relationship that associates an algorithm's input size$n$with its notional energy consumption$E$. Based on the well-known abstract Bulk Synchronous Parallel (BSP) computer and programming model, we present a complete$E(n)$solution, including abstraction, generalization, quantification, derivation, comparison, analysis, examples, verification, and applications. Comprehensive experimental analysis shows that the proposed energy complexity model is practical, interestingly, and not equivalent to time complexity.
Jie Song 0001, Xingchen Zhao, Chaopeng Guo, Yu Gu 0002, Ge Yu 0001
IEEE Trans. Big Data1
2022 A survey of visual analytics in urban area
abstract
Abstract Nowadays, the population has been overgrowing due to urbanization, yielding many severe problems in the urban area, including traffic congestion, unbalanced distribution of urban hotspots, air pollution and so on. Due to the uncertainty of the urban environment, it always needs to integrate experts' domain knowledge into solving these issues. In recent years, the visual analytics method has been widely used to assist domain experts in solving urban problems with its intuitiveness, interactivity and interpretability. In this survey, we first introduce the background of urban computing, present the motivation of visual analytics in the urban area and point out the characteristics of visual analytics methods. Second, we introduce the most frequently used urban data, analyse the main properties and provide an overview on how to use these data. Thereafter, we propose our taxonomy for visual analytics in the urban area and illustrate the taxonomy. The taxonomy provides four levels for visual analytics on urban data from a new perspective based on the four stages in data mining. Four levels from our taxonomy include: descriptive analytics, diagnostic analytics, predictive analytics and prescriptive analytics. Finally, we conclude this survey by discussing the limitations of the existing related works and the challenges to visual analytics in the urban area.
Zezheng Feng, Huamin Qu, Shuang-Hua Yang, Jie Song 0001
Expert Syst. J. Knowl. Eng.5
2022 Versatility or validity: A comprehensive review on simulation of Datacenters powered by Renewable Energy mix
Jie Song 0001, Peimeng Zhu, Yanfeng Zhang 0001, Ge Yu 0001
Future Gener. Comput. Syst.1
2022 Compress Blocks or Not: Tradeoffs for Energy Consumption of a Big Data Processing System
abstract
Currently, in addition to the performance, the energy consumption (hereinafter EC) of jobs running in a big data processing system is also of interest to academia and industry because it grows rapidly as an increasing amount of data is processed. Many studies focus on the EC optimization of jobs from the perspective of computation, which is specific to the algorithms in each job. However, the part of EC involved in I/O operations, which is general and universal, is mostly ignored in optimization. In this paper, we concentrate on the EC optimization of jobs from the perspective of I/O operations. To save energy, we argue that data compression could be exploited. On one hand, energy is saved by processing compressed data with less I/O cost. On the other hand, extra EC is incurred from the necessary data compression/decompression process, which may offset the saved energy. Therefore, there are tradeoffs to consider when determining whether to compress data for these jobs. In this paper, such tradeoffs and boundary conditions are studied. We first abstract a paradigm for the runtime environment of big data processing jobs. Then, we establish the power, jobs, compression, and I/O models in detail. Based on these models, we discuss the compression tradeoffs and derive the boundary conditions for EC optimization. Finally, we design and conduct experiments to validate our proposition. The experimental results confirm that the tradeoffs and boundary conditions exist for typical jobs in MapReduce and Spark. As explained, first, the EC of a job is reduced using data compression. Second, whether or not such optimization occurs is related to the specification of both the compression algorithm and the job and is determined by corresponding boundary conditions. Third, for a compression algorithm, the larger its compression/decompression speed and the better its compression ratio, the more likely it is to achieve EC optimization.
Jie Song 0001, Shengqiang Hu, Yubin Bao, Ge Yu 0001
IEEE Trans. Sustain. Comput.1
2020 Haery: A Hadoop Based Query System on Accumulative and High-Dimensional Data Model for Big Data
abstract
Column-oriented stores, known for their scalability and flexibility, are a common NoSQL database implementation and are increasingly used in big data management. In column-oriented stores, a “full-scan” query strategy is inefficient and the search space can be reduced if data is well partitioned or indexed; however, there is no pre-defined schema for building and maintaining partitions and indexes at lower cost. We leverage an accumulative and high-dimensional data model, a sophisticated linearization algorithm, and an efficient query algorithm, to solve the challenge of how a pre-defined and well-partitioned data model can be applied to flexible and time-varied key-value data. We adapt a high-dimensional array as the data model to partition the key-value data without additional storage and massive calculation; improve the Z-order linearization algorithm, which map multidimensional data to one dimension while preserving locality of the data points, for flexibility; efficiently build an expansion mechanism for the data model to support time-varied data. The result is Haery, a column-oriented store, based on a distributed file system and computing framework. In experiments, Haery is compared with Hive, HBase, Cassandra, MongoDB, PostgresXL, and HyperDex in terms of query performance. With results indicating Haery on average performs 4.57x, 4.23x, 3.55x, 1.79x, 1.82x, and 120.6x faster, respectively.
Jie Song 0001, HongYan He, Richard Thomas 0002, Yubin Bao, Ge Yu 0001
IEEE Trans. Knowl. Data Eng.1
2019 Hot-N-Cold model for energy aware cloud databases
Chaopeng Guo, Jean-Marc Pierson, Jie Song 0001, Christina Herzog
J. Parallel Distributed Comput.3
2019 Minimizing temperature and energy of real-time applications with precedence constraints on heterogeneous MPSoC systems
Tiantian Li 0003, Tianyu Zhang 0001, Ge Yu 0001, Jie Song 0001
J. Syst. Archit.4
2019 An optimized phase-shifting algorithm for depth image acquisition
Hui Liu 0012, Beilei Wang, Fusheng Yan, Jie Song 0001
Multim. Tools Appl.4
2019 Image saliency detection for multiple objects
Beilei Wang, Lu Meng, Jie Song 0001
Multim. Tools Appl.3
2018 Modulo Based Data Placement Algorithm for Energy Consumption Optimization of MapReduce System
Jie Song 0001, HongYan He, Ge Yu 0001, Jean-Marc Pierson
J. Grid Comput.1
2018 Minimizing energy by thermal-aware task assignment and speed scaling in heterogeneous MPSoC systems
Tiantian Li 0003, Ge Yu 0001, Jie Song 0001
J. Syst. Archit.3
2018 Rim: A reusable iterative model for big data
Jie Song 0001, Zhongyi Ma, Tiantian Li 0003, Ge Yu 0001
Knowl. Based Syst.1
2016 SaaS-based enterprise application integration approach and case study
BeiLie Wang, Hui Liu 0012, Jie Song 0001
J. Supercomput.3
2015 HaoLap: A Hadoop based OLAP system for big data
Jie Song 0001, Chaopeng Guo, Yichan Zhang, Ge Yu 0001, Jean-Marc Pierson
J. Syst. Softw.1
2014 Analyzing the Waiting Energy Consumption of NoSQL Databases
abstract
NoSQL database (NoSQL DB) covers the shortage of traditional database and has been widely used in recent years. Currently, researches on NoSQL DB mainly focus on performance issues, few of them are about energy consumption (EC) evaluation and optimization. Waiting Energy Consumption (WEC) is another critical reason causing energy waste besides computer idleness. Study the WEC regularities of NoSQL DB facilitates achieving real "green computing". This paper first analyzes the model and measurement approaches of EC, then designs test cases to study the WEC regularities, finally proposes approaches of "reducing WEC" for EC optimization. Plenty of experiments show that, despite that NoSQL DB is an application of "green cloud computing", the ECs of selected NoSQL DBs are widely divergent, and some of them remain to be further optimized.
Tiantian Li 0003, Ge Yu 0001, Xuebing Liu, Jie Song 0001
DASC4
2012 Introducing SaaS Capabilities to Existing Web-Based Applications Automatically
Jie Song 0001, Zhenxing Yan, Yubin Bao, Zhiliang Zhu 0001
APWeb1
2011 A SaaSify Tool for Converting Traditional Web-Based Applications to SaaS Application
abstract
Nowadays, SaaS is increasingly used by web-based applications for the benefits and profits it brings to both users and service providers. It is significative if service providers can automatically convert traditional applications into SaaS mode, a SaaSify tool is needed urgently. In this paper, we analyze and conclude the new challenges of automatically SaaSify web-based application, propose several key technologies for SaaSifying, and further propose SaaSify Flow Language (SFL) to model and implement SaaSify process, finally, we use a case study to show the effects of proposed tool, and the performance experiments prove that the proposed approach is efficient and effective.
Jie Song 0001, Zhenxing Yan, Guoqi Liu, Zhiliang Zhu 0001
IEEE CLOUD1
2010 Partitioned Dimension: Modeling the Numerical Dimension in Data Warehouse
abstract
In the traditional data warehouse modeling, a numerical attribute with continuous data are not suitable for modeling as a dimension, unless it is partitioned into a concept hierarchy according to some predefined mapping-rules. In some cases, such rules are flexible or unavailable, so the partitioned dimension is proposed in this paper as the solution of modeling the numerical dimension in these cases. Partitioned dimension is generated by clustering the frequent query conditions. The query and dynamic OLAP operations over partitioned dimensions are also introduced. At last, some experiments show that the proposed modeling approach is effective and efficient.
Jie Song 0001, Yubin Bao
APWeb1
2010 A Multilevel and Domain-Independent Duplicate Detection Model for Scientific Database
Jie Song 0001, Yubin Bao, Ge Yu 0001
WAIM1
2006 An Effective Web Page Layout Adaptation for Various Resolutions
Jie Song 0001, Tiezheng Nie, Daling Wang, Ge Yu 0001
APWeb1
2006 An Approach for Composing Web Services on Demand
abstract
The Web services composition plays a more and more important role in SOA environment nowadays. Composing Web services on user's demand proves to be essential for both B2B and B2C applications. However, existing composition approaches and services composition description languages, such as BPEL4WS, are insufficient for satisfying user requirements. In this paper, we present an approach for Web services composition on demand. Our approach defines a description language for describing user's demand and supporting dynamic services binding for flexible service composition, and applies SLA to service discovery that can automatically adapt the change of QoS constraints. The description language can be applied to any process-oriented composition language that supports executable business processes. Finally, our approach is applied to a service grid system where it provides a fully automatic, stable service composition on user's demand
Tiezheng Nie, Ge Yu 0001, Derong Shen, Yue Kou, Jie Song 0001
CSCWD5
2006 Offline Web Client: Approach, Design and Implementation Based on Web System
Jie Song 0001, Ge Yu 0001, Daling Wang, Tiezheng Nie
WISE1