Jialiang Tan

dblp:274/0565 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Translating to a Low-Resource Language with Compiler Feedback: A Case Study on Cangjie
abstract
In the rapidly advancing field of software development, the demand for practical code translation tools has surged, driven by the need for interoperability across different programming environments. Existing learning-based approaches often need help with low-resource programming languages that lack sufficient parallel code corpora for training. To address these limitations, we propose a novel training framework that begins with monolingual seed corpora, generating parallel datasets via back-translation and incorporating compiler feedback to optimize the translation model.As a case study, we apply our method to train a code translation model for a new-born low-resource programming language, Cangjie. We also construct a parallel test dataset forJava-to-Cangjietranslation and test cases to evaluate the effectiveness of our approach. Experimental results demonstrate that compiler feedback greatly enhances syntactical correctness, semantic accuracy, and test pass rates of the translatedCangjiecode. These findings highlight the potential of our method to support code translation in low-resource settings, expanding the capabilities of learning-based models for programming languages with limited data availability.
Jun Wang 0151, Chenghao Su, Yijie Ou, Yanhui Li 0001, Jialiang Tan, Lin Chen 0015, Yuming Zhou
IEEE Trans. Software Eng.5
2024 ThemeRec: Personalizing IDE Themes for Students
abstract
Prolonged screen time while working on programming assignments can often lead to student fatigue, boredom, and stress. In response to these issues and to enhance the overall programming environment for students, we introduce ThemeRec, a Visual Studio Code extension. ThemeRec provides students with personalized recommendations for code themes, with a particular focus on semantic coloring combinations. It boasts an extensive collection of over 150 code themes, each offering unique combinations of background colors, font styles, and semantic color arrangements. Using an integrated collaborative filtering recommendation system, ThemeRec automatically suggests and installs various themes based on the preferences of similar users whenever they want to change a theme. To assess the effectiveness of ThemeRec, we conducted an evaluation within an introductory-level programming course. The feedback from students has been positive and encouraging. This underscores the potential value that ThemeRec brings to programming courses.
Jialiang Tan, Yu Chen 0036, Shuyin Jiao
SIGCSE (2)1
2023 Presto: A Decade of SQL Analytics at Meta
abstract
Presto is an open-source distributed SQL query engine that supports analytics workloads involving multiple exabyte-scale data sources. Presto is used for low-latency interactive use cases as well as long-running ETL jobs at Meta. It was originally launched at Meta in 2013 and donated to the Linux Foundation in 2019. Over the last ten years, upholding query latency and scalability with the hyper growth of data volume at Meta as well as new SQL analytics requirements have raised impressive challenges for Presto. A top priority has been ensuring query reliability does not regress with the shift towards smaller, more elastic container allocation, which requires queries to run with substantially smaller memory headroom and can be preempted at any time. Additionally, new demands from machine learning, privacy, and graph analytics have driven Presto maintainers to think beyond traditional data analytics. In this paper, we discuss several successful evolutions in recent years that have improved Presto latency as well as scalability by several orders of magnitude in production at Meta. Some of the notable ones are hierarchical caching, native vectorized execution engines, materialized views, and Presto on Spark. With these new capabilities, we have deprecated or are in the process of deprecating various legacy query engines so that Presto becomes the single piece to serve interactive, ad-hoc, ETL, and graph processing workloads for the entire data warehouse.
Yutian Sun, Tim Meehan, Rebecca Schlussel, Wenlei Xie, Masha Basmanova, Orri Erling, Andrii Rosa, Shixuan Fan, Rongrong Zhong, Arun Thirupathi, Nikhil Collooru, Dionysios Logothetis, Kostas Xirogiannopoulos, Varun Gajjala, Rohit Jain, Ajay Palakuzhy, Prithvi Pandian, Sergey Pershin, Abhisek Saikia, Pranjal Shankhdhar, Neerad Somanchi, Swapnil Tailor, Jialiang Tan, Sreeni Viswanadha, Zac Wen, Biswapesh Chattopadhyay, Deepak Majeti, Aditi Pandit
Proc. ACM Manag. Data27
2021 Interactive Analytic DBMSs: Breaching the Scalability Wall
abstract
Analytic DBMSs optimized for query interactivity commonly push the computation down to storage nodes, thus avoiding large network transfers and keeping query execution wall-time to a minimum. In these systems, data is sharded and stored locally by cluster nodes, which must all participate in query execution. As the system scales-out, hardware failures and other non-deterministic sources of tail latency start to dominate, to a point where query latency and success ratio increasingly violate the system's SLA. We refer to this tipping point as the system's scalability wall, when sharding data between more nodes only worsens the problem.This paper describes how an analytic DBMS optimized for low-latency queries can breach the scalability wall by sharding different tables to different subsets of cluster nodes - a strategy we call partial sharding - and reduce the query fan-out. Because partial sharding requires the DBMS to implement many tedious and complex shard management tasks, such as shard mapping, load balancing and fault tolerance, this paper describes how a database system can leverage an external general-purpose shard management service for such tasks. We present a case study based on Cubrick, an in-memory analytic DBMS developed at Facebook, highlighting the integration points with a shard management framework called Shard Manager. Finally, we describe the many design decisions, pitfalls and lessons learned during this process, which eventually allowed Cubrick to scale to thousands of nodes.
Pedro Pedreira, Sergey Pershin, Sushant Shringarpure, Jialiang Tan, Brian Landers, Karen Pieper
ICDE6
2021 Toward efficient interactions between Python and native libraries
abstract
Python has become a popular programming language because of its excellent programmability. Many modern software packages utilize Python for high-level algorithm design and depend on native libraries written in C/C++/Fortran for efficient computation kernels. Interaction between Python code and native libraries introduces performance losses because of the abstraction lying on the boundary of Python and native libraries. On the one side, Python code, typically run with interpretation, is disjoint from its execution behavior. On the other side, native libraries do not include program semantics to understand algorithm defects.
Jialiang Tan, Yu Chen 0036, Zhenming Liu, Bin Ren 0002, Shuaiwen Song, Xipeng Shen, Xu Liu 0001
ESEC/SIGSOFT FSE1
2020 What every scientific programmer should know about compiler optimizations?
abstract
Compilers are an indispensable component in the software stack. Besides generating machine code, compilers perform multiple optimizations to improve code performance. Typically, scientific programmers treat compilers as a blackbox and expect them to optimize code thoroughly. However, optimizing compilers are not performance panacea. They can miss optimization opportunities or even introduce inefficiencies that are not in the source code. There is a lack of tool infrastructures and datasets that can provide such a study to help understand compiler optimizations.
Jialiang Tan, Shuyin Jiao, Milind Chabbi, Xu Liu 0001
ICS1