Ran Jia

dblp:175/1500 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
abstract
Zhiyuan Peng, Xin Yin, Pu Zhao, Fangkai Yang, Lu Wang, Ran Jia, Xu Chen, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pu Zhao 0004, Fangkai Yang, Lu Wang 0029, Ran Jia, Xu Chen 0022, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001
ACL (1)6
2026 BD-TNet: Balanced Density-Aware Point Cloud Completion Method Assisted by Trie-Like Shape Prior Embedding Dictionary
abstract
Point cloud completion is a fundamental yet challenging task in 3D computer vision. While existing two-stage methods have made significant progress in shape completion, they often overlook the critical issue of point density imbalance, leading to geometrically unsound results with illusory completeness. In this paper, we propose BD-TNet, a novel balanced density-aware point cloud completion method that systematically addresses this overlooked problem. Specifically, we design a Partial-Connected U-Net structure that strategically reduces top-layer skip connections to rebalance the influence between explicit features of partial inputs and abstract features for predicting missing regions, thereby generating uniformly distributed seed points. Furthermore, we introduce EmbeddingTrie, a Trie-like hierarchical shape prior dictionary that enables adaptive multi-scale prior fusion through a divide-and-conquer strategy. To directly constrain local point density, we devise an Elastic Potential Energy Loss that models point clouds as spring systems, effectively guiding the optimization toward balanced distributions. We evaluate our method on PCN, ShapeNet-34/55, and KITTI datasets, and the results demonstrate that BD-TNet achieves state-of-the-art performance on Density-aware Chamfer Distance while producing visually superior, complete point clouds with balanced density distribution.
Ran Jia, Junpeng Xue, Zi Guo, Kelei Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 IE-PMMA: Point Cloud Completion Through Inverse Edge-aware Upsampling and Precise Multi-Modal Feature Alignment
abstract
Point cloud completion is a crucial task in 3D computer vision. Multi-modal completion approaches have gained attention among the popular two-stage point cloud completion methods. However, there is a notable lack of research focused on accurately aligning data from different modalities within these methods. Additionally, in other point cloud-based tasks, edge point information often provides unexpected positive contributions. In this paper, we propose a novel point cloud completion method that leverages edge point information for the first time in the completion task, which also addresses the precise alignment of multi-modal data. In particular, we implement a two-step local-to-global module to achieve better alignment of multi-modal data during the preliminary point cloud generation process. Besides, we introduce a new spatial representation structure capable of extracting a fixed number of edge points. Moreover, with the assistance of edge information, we further design an inverse edge-aware upsampler to refine the point cloud. We evaluate our method on three typical datasets, and the results demonstrate that our IE-PMMA outperforms the existing state-of-the-art methods quantitatively and visually.
Ran Jia, Junpeng Xue, Kelei Wang
IJCAI1
2025 Lightweight network research for robotic visual grasp for deep space exploration
Junpeng Xue, Ran Jia
Neural Comput. Appl.4
2024 Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries
abstract
Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities.
Mengyu Zhou, Xinrun Xu, Xiaojun Ma 0001, Rui Ding 0001, Lun Du, Yan Gao 0002, Ran Jia, Xu Chen 0022, Shi Han, Zejian Yuan, Dongmei Zhang 0001
AAAI8
2024 SCSMD: Single Cell Consistent Clustering based on Spectral Matrix Decomposition
abstract
Cluster analysis, a pivotal step in single-cell sequencing data analysis, presents substantial opportunities to effectively unveil the molecular mechanisms underlying cellular heterogeneity and intercellular phenotypic variations. However, the inherent imperfections arise as different clustering algorithms yield diverse estimates of cluster numbers and cluster assignments. This study introduces Single Cell Consistent Clustering based on Spectral Matrix Decomposition (SCSMD), a comprehensive clustering approach that integrates the strengths of multiple methods to determine the optimal clustering scheme. Testing the performance of SCSMD across different distances and employing the bespoke evaluation metric, the methodological selection undergoes validation to ensure the optimal efficacy of the SCSMD. A consistent clustering test is conducted on 15 authentic scRNA-seq datasets. The application of SCSMD to human embryonic stem cell scRNA-seq data successfully identifies known cell types and delineates their developmental trajectories. Similarly, when applied to glioblastoma cells, SCSMD accurately detects pre-existing cell types and provides finer sub-division within one of the original clusters. The results affirm the robust performance of our SCSMD method in terms of both the number of clusters and cluster assignments. Moreover, we have broadened the application scope of SCSMD to encompass larger datasets, thereby furnishing additional evidence of its superiority. The findings suggest that SCSMD is poised for application to additional scRNA-seq datasets and for further downstream analyses.
Ran Jia, Ying-Zan Ren, Ponian Li, Rui Gao 0006, Yusen Zhang 0002
Briefings Bioinform.1
2023 SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention Network
abstract
Spreadsheets are an important and unique type of business document for data storage, analysis and presentation. The distinction between spreadsheets and most other types of digital documents lies in that spreadsheets provide users with high flexibility of data organization on the grid. Existing related techniques mainly focus on the tabular data and are incompetent in understanding the entire sheet. On the one hand, spreadsheets have no explicit separation across tabular data and other information, leaving a gap for the deployment of such techniques. On the other hand, pervasive data dependence and semantic relations across the sheet require comprehensive modeling of all the information rather than only the tables. In this paper, we propose SheetPT, the first pre-training technique on spreadsheets to enable effective representation learning under this scenario. For computational effectiveness and efficiency, we propose the coherent chunk, an intermediate semantic unit of sheet structure; and we accordingly devise a hierarchical attention-based architecture to capture contextual information across different structural granularities. Three pre-training objectives are also designed to ensure sufficient training against millions of spreadsheets. Two representative downstream tasks, formula prediction and sheet structure recognition are utilized to evaluate its capability and the prominent results reveal its superiority over existing state-of-the-art methods.
Ran Jia, Qiyu Li 0001, Xiaoyuan Jin, Lun Du, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001
AAAI1
2023 HermEs: Interactive Spreadsheet Formula Prediction via Hierarchical Formulet Expansion
abstract
Wanrong He, Haoyu Dong, Yihuai Gao, Zhichao Fan, Xingzhuo Guo, Zhitao Hou, Xiao Lv, Ran Jia, Shi Han, Dongmei Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wanrong He, Haoyu Dong 0001, Yihuai Gao, Zhichao Fan, Xingzhuo Guo, Zhitao Hou, Ran Jia, Shi Han, Dongmei Zhang 0001
ACL (1)8
2023 GetPt: Graph-enhanced General Table Pre-training with Alternate Attention Network
abstract
Tables are widely used for data storage and presentation due to their high flexibility in layout. The importance of tables as information carriers and the complexity of tabular data understanding attract a great deal of research on large-scale pre-training for tabular data. However, most of the works design models for specific types of tables, such as relational tables and tables with well-structured headers, neglecting tables with complex layouts. In real-world scenarios, there are many such tables beyond their target scope that cannot be well supported. In this paper, we propose GetPt, a unified pre-training architecture for general table representation applicable even to tables with complex structures and layouts. First, we convert a table to a heterogeneous graph with multiple types of edges to represent the layout of the table. Based on the graph, a specially designed transformer is applied to jointly model the semantics and structure of the table. Second, we devise the Alternate Attention Network (AAN) to better model the contextual information across multiple granularities of a table including tokens, cells, and the table. To better support a wide range of downstream tasks, we further employ three pre-training objectives and pre-train the model on a large table dataset. We fine-tune and evaluate GetPt model on two representative tasks, table type classification, and table structure recognition. Experiments show that GetPt outperforms existing state-of-the-art methods on these tasks.
Ran Jia, Haoming Guo, Xiaoyuan Jin, Lun Du, Xiaojun Ma 0001, Tamara Stankovic, Marko Lozajic, Goran Zoranovic, Igor Ilic, Shi Han, Dongmei Zhang 0001
KDD1
2022 FORTAP: Using Formulas for Numerical-Reasoning-Aware Table Pretraining
abstract
Zhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu, Shi Han, Fan Cheng, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zhoujun Cheng, Haoyu Dong 0001, Ran Jia, Pengfei Wu 0006, Shi Han, Fan Cheng 0002, Dongmei Zhang 0001
ACL (1)3
2022 HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation
abstract
Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zhoujun Cheng, Haoyu Dong 0001, Zhiruo Wang 0001, Ran Jia, Yan Gao 0002, Shi Han, Jian-Guang Lou, Dongmei Zhang 0001
ACL (1)4
2021 TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular Data
abstract
Tabular data are ubiquitous for the widespread applications of tables and hence have attracted the attention of researchers to extract underlying information. One of the critical problems in mining tabular data is how to understand their inherent semantic structures automatically. Existing studies typically adopt Convolutional Neural Network (CNN) to model the spatial information of tabular structures yet ignore more diverse relational information between cells, such as the hierarchical and paratactic relationships. To simultaneously extract spatial and relational information from tables, we propose a novel neural network architecture, TabularNet. The spatial encoder of TabularNet utilizes the row/column-level Pooling and the Bidirectional Gated Recurrent Unit (Bi-GRU) to capture statistical information and local positional correlation, respectively. For relational information, we design a new graph construction method based on the WordNet tree and adopt a Graph Convolutional Network (GCN) based encoder that focuses on the hierarchical and paratactic relationships between cells. Our neural network architecture can be a unified neural backbone for different understanding tasks and utilized in a multitask scenario. We conduct extensive experiments on three classification tasks with two real-world spreadsheet data sets, and the results demonstrate the effectiveness of our proposed TabularNet over state-of-the-art baselines.
Lun Du, Xu Chen 0022, Ran Jia, Junshan Wang, Jiang Zhang 0006, Shi Han, Dongmei Zhang 0001
KDD4
2021 TUTA: Tree-based Transformers for Generally Structured Table Pre-training
abstract
We propose TUTA, a unified pre-training architecture for understanding generally structured tables. Noticing that understanding a table requires spatial, hierarchical, and semantic information, we enhance transformers with three novel structure-aware mechanisms. First, we devise a unified tree-based structure, called a bi-dimensional coordinate tree, to describe both the spatial and hierarchical information of generally structured tables. Upon this, we propose tree-based attention and position embedding to better capture the spatial and hierarchical information. Moreover, we devise three progressive pre-training objectives to enable representations at the token, cell, and table levels. We pre-train TUTA on a wide range of unlabeled web and spreadsheet tables and fine-tune it on two critical tasks in the field of table structure understanding: cell type classification and table type classification. Experiments show that TUTA is highly effective, achieving state-of-the-art on five widely-studied datasets.
Zhiruo Wang 0001, Haoyu Dong 0001, Ran Jia, Jia Li 0012, Zhiyi Fu, Shi Han, Dongmei Zhang 0001
KDD3
2019 Prediction for Student Academic Performance Using SMNaive Bayes Model
Baoting Jia, Ke Niu 0002, Xia Hou, Ning Li 0024, Xueping Peng, Peipei Gu, Ran Jia
ADMA7
2016 Distilling Word Embeddings: An Encoding Approach
abstract
Distilling knowledge from a well-trained cumbersome network to a small one has recently become a new research topic, as lightweight neural networks with high performance are particularly in need in various resource-restricted systems. This paper addresses the problem of distilling word embeddings for NLP tasks. We propose an encoding approach to distill task-specific knowledge from a set of high-dimensional embeddings, so that we can reduce model complexity by a large margin as well as retain high accuracy, achieving a good compromise between efficiency and performance. Experiments reveal the phenomenon that distilling knowledge from cumbersome embeddings is better than directly training neural networks with small embeddings.
Lili Mou, Ran Jia, Yan Xu 0013, Ge Li 0001, Lu Zhang 0023, Zhi Jin 0001
CIKM2
2016 Improved relation classification by deep recurrent neural networks with data augmentation
abstract
Nowadays, neural networks play an important role in the task of relation classification. By designing different neural architectures, researchers have improved the performance to a large extent in comparison with traditional methods. However, existing neural networks for relation classification are usually of shallow architectures (e.g., one-layer convolutional neural networks or recurrent networks). They may fail to explore the potential representation space in different abstraction levels. In this paper, we propose deep recurrent neural networks (DRNNs) for relation classification to tackle this challenge. Further, we propose a data augmentation method by leveraging the directionality of relations. We evaluated our DRNNs on the SemEval-2010 Task 8, and achieve an F1-score of 86.1%, outperforming previous state-of-the-art recorded results.
Yan Xu 0013, Ran Jia, Lili Mou, Ge Li 0001, Yunchuan Chen, Yangyang Lu, Zhi Jin 0001
COLING2