VLDB 2026 Research / reviewers in the wild / expert
Junwen Liu
dblp:49/9600
· DBLP profile ↗
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-1105-8104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSM-Net: Relation embedding for few shot learning optimized by cross memory attention
Junwen Liu, Xutao Sun, Yonggong Ren |
Neural Networks | 2 |
| 2025 | MFB-SAC: A Multi-Scale Frequency and Boundary-Enhanced SAM for Cell SegmentationabstractIn medical image analysis, precise cell segmentation is crucial for understanding cell morphology and diagnosing diseases. Nevertheless, traditional models typically trained for specific modalities or cell types often fail to generalize to unknown categories. The Segment Anything Model (SAM) was introduced to address this but relies heavily on precise prompts and lacks integration of frequency domain features alongside multi-scale and multi-level information. To overcome these limitations, we propose Multi-scale Frequency and Boundary-enhanced Segment Any Cells (MFB-SAC), which includes a Multi-scale Frequency Convolution (MFC) module for providing multi-scale information and texture detail and a Multi-level Boundary Awareness (MBA) module to enhance boundary retention. Additionally, a Dual-representation Collaborative Attention (DCA) module optimizes feature integration. We tested MFB-SAC on 8 public datasets with various microscopy techniques and cell types. The results reveal that MFB-SAC outperforms the current general and benchmark cell segmentation models. The code is available at https://github.com/Mrliujunwen/SAC. Xutao Sun, Xiaolu Xu, Junwen Liu, Yonggong Ren |
ICIP | 3 |
| 2024 | Segment Any Nuclei: A Prompt-free Segment Anything Model for Nuclei from Histology ImageabstractAccurate cell nuclei segmentation is crucial for characterizing cell morphology and elucidating disease types in medical image analysis. However, fully-supervised segmentation models often have limited generalization to unseen classes or domains due to being trained on specific modalities or cell types. The recent Segment Anything Model (SAM) has shown promise for interactive instance segmentation and zero-shot generalization. However, its effectiveness is hindered by the sparse and distributed nature of nuclei. It also requires a large number of user prompts that scale with the number of nuclei in the image. To overcome these challenges, we introduce Segment Any Nuclei (SAN), a novel foundation model tailored specifically for nuclei segmentation. SAN is trained on an extensive multi-modal dataset containing diverse nuclei instances from various imaging modalities, staining methods, and tissue types. Unlike previous SAM approaches that rely on manual prompts, SAN incorporates an innovative auto-prompting auxiliary segmentation network, which enables the model to make predictions without manual intervention while still allowing for manual interaction when needed. We evaluate SAN on a large dataset of 9,244 images and demonstrate state-of-the-art nuclei segmentation performance, surpassing both fully-supervised approaches and other SAM-like models. SAN represents a step towards more generalizable, efficient and interactive nuclei segmentation. Our code is available at https://github.com/Mrliujunwen/SAN. Xutao Sun, Junwen Liu, Yonggong Ren, Xiaolu Xu, Yiqing Shen 0003 |
BIBM | 2 |
| 2024 | Integration of Blockchain Technology in Collaborative Scientific WorkflowsabstractBlockchain technology has emerged as a transformative force across various sectors, especially in enhancing collaborative scientific workflows. This paper delves into the unique challenges and opportunities associated with integrating blockchain into these workflows. We conduct a comprehensive analysis of recent literature to identify key themes, categorize different approaches, and assess the potential of blockchain to improve data integrity, provenance, and collaborative research efforts among diverse stakeholders. Through this exploration, we aim to provide an insightful overview of the current landscape and propose directions for future research focused on the role of blockchain in facilitating effective collaboration in scientific endeavors. Shiyong Lu, Junwen Liu, Yong Zhao 0009, Changxin Bai |
IEEE Big Data | 2 |
| 2023 | A Generic Efficient Scientific Workflow Engine for the Optimizations of Run-Time ExecutionabstractWorkflow has proven to be a highly effective computing model for a variety of scientific applications, offering flexible data types and unstructured parallelism that surpasses simple parallel execution models such as MapReduce. However, current workflow management systems in cloud computing environments experience unnecessary delays in task execution due to the separation of task execution and data transfer processes, which causes a child task to wait until all its predecessor tasks complete, rather than waiting only for necessary input data becoming ready. The goal of this paper is to eliminate the unnecessary delay of child tasks in a workflow, which is achieved through a new workflow engine architecture that separates workflow planner from workflow executor in the general framework of the DATAVIEW scientific workflow management system. This new engine architecture can be generalized and applied to other workflow systems. Our design integrates a new task release mechanism based on a data dependency model with the workflow executor of DATAVIEW. This approach enables prompt task launching once input data becomes available, instead of waiting for all predecessor tasks to finish. The architecture employs distributed algorithms for implementing the workflow executor and the task executors, performing various optimization on data movement, task movement, and communication among different subsystems. The experiments show that our new architecture based on the new task release model can significantly reduce overall execution time of a workflow in DATAVIEW. Changxin Bai, Junwen Liu, Anik Tahabilder, M. M. Imran, Shiyong Lu, Dunren Che |
SSE | 2 |
| 2023 | Infrastructure-level Support for GPU-Enabled Deep Learning in DATAVIEW
Junwen Liu, Ziyun Xiao, Shiyong Lu, Dunren Che, Ming Dong 0001, Changxin Bai |
Future Gener. Comput. Syst. | 1 |
| 2022 | Apache ShardingSphere: A Holistic and Pluggable Platform for Data ShardingabstractTraditional relational databases are nowadays over-whelmed by the increasing data volume and concurrent access. NoSQL databases can manage large-scale data, but most of them do not support complete transactions and standard SQL languages. NewSQL is proposed for both high scalability and transactional properties with SQL languages support. One type of NewSQL builds distributed systems from scratch, which is too radical for some critical applications. The other type of NewSQL, i.e., data sharding among relational databases, is a better option for these scenarios. This paper presents Apache ShardingSphere, the first top-level open-source platform for data sharding in Apache, which enables developers to use sharded databases like one database. Specifically Apache, ShardingSphere integrates six databases and designs and implements a complete SQL engine to route requests correctly and intelligently. Additionally it encapsulates three types of distributed transactions and provides two adaptors for different scenarios. Moreover it proposes a novel AutoTable strategy and a query language i.e DistSQL allowing database maintainers to easily configure the sharded databases. Further-more it provides many other pluggable features to better shard data. Extensive experiments are conducted using two famous benchmarking tools proving that Apache ShardingSphere is more efficient than eight state-of-the-art systems in our settings. All experimental source codes are publicly released. More than 170 companies are currently using Apache ShardingSphere. Juan Pan, Junwen Liu, Nianjun Sun, Shanmin Wang, Chao Chen 0004, Fuqiang Gu, Songtao Guo |
ICDE | 4 |
| 2021 | Deep-Learning-as-a-Workflow (DLaaW): An Innovative Approach to Enabling Deep Learning in Scientific WorkflowsabstractScientific workflow has become a popular cyberinfrastructure paradigm to accelerate scientific discoveries by enabling scientists to formalize and structure complex scientific processes. With the recent success of deep learning models in many scientific applications, there is a rising need for infrastructure-level support for deep learning technologies in scientific workflow cyberinfrastructures. However, current scientific workflow cyberinfrastructures and GPU-enabled deep learning frameworks are developed separately, neither alone can be a satisfactory choice. In this paper, We propose the Deep-Learning-as-a-Workflow approach in DATAVIEW, which for the first time incorporates native infrastructure level support for GPU-enabled deep learning in a scientific workflow management system and enables the fast training and execution of neural networks as workflows (NNWorkflows) leveraging various types of GPU resource configurations. Our experiments demonstrate the salient usability feature of DATAVIEW in providing seamless infrastructure-level support to both scientific and deep learning workflows in one system, while delivering competitive (better in most cases) learning efficiency compared to the conventional implementations based on Keras. Junwen Liu, Ziyun Xiao, Shiyong Lu, Dunren Che |
IEEE BigData | 1 |
| 2021 | Distributed Spatio-Temporal k Nearest Neighbors JoinabstractThe rapid development of positioning technology produces an extremely large volume of spatio-temporal data with various geometry types such as point, line string, polygon, or a mixed combination of them. As one of the most basic but time-consuming operations, k nearest neighbors join (kNN join) has attracted much attention. However, most existing works for kNN join either ignore temporal information or consider point data only. Rubin Wang, Junwen Liu, Zisheng Yu, Huajun He, Tianfu He, Sijie Ruan, Jie Bao 0003, Chao Chen 0004, Fuqiang Gu, Liang Hong 0001, Yu Zheng 0004 |
SIGSPATIAL/GIS | 3 |
| 2020 | JUST: JD Urban Spatio-Temporal Data EngineabstractWith the prevalence of positioning techniques, a prodigious number of spatio-temporal data is generated constantly. To effectively support sophisticated urban applications, e.g., location-based services, based on spatio-temporal data, it is desirable for an efficient, scalable, update-enabled, and easy-to-use spatio-temporal data management system.This paper presents JUST, i.e., JD Urban Spatio-Temporal data engine, which can efficiently manage big spatio-temporal data in a convenient way. JUST incorporates the distributed NoSQL data store, i.e., Apache HBase, as the underlying storage, GeoMesa as the spatio-temporal data indexing tool, and Apache Spark as the execution engine. We creatively design two indexing techniques, i.e., Z2T and XZ2T, which accelerates spatio-temporal queries tremendously. Furthermore, we introduce a compression mechanism, which not only greatly reduces the storage cost, but also improves the query efficiency. To make JUST easy-to-use, we design and implement a complete SQL engine, with which all operations can be performed through a SQL-like query language, i.e., JustQL. JUST also supports inherently new data insertions and historical data updates without index reconstruction. JUST is deployed as a PaaS in JD with multi-users support. Many applications have been developed based on the SDKs provided by JUST. Extensive experiments are carried out with six state-of-the-art distributed spatio-temporal data management systems based on two real datasets and one synthetic dataset. The results show that JUST has a competitive query performance and is much more scalable than them. Huajun He, Rubin Wang, Yuchuan Huang, Junwen Liu, Sijie Ruan, Tianfu He, Jie Bao 0003, Yu Zheng 0004 |
ICDE | 5 |