VLDB 2026 Research / reviewers in the wild / expert
Jia Chen 0002
dblp:99/6879-2
· DBLP profile ↗
12ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-2116-3610ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-authorSystems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor DecompositionabstractWhile Large Language Model-based Multi-Agent Systems (MAS) consistently outperform single-agent systems on complex tasks, their intricate interactions introduce critical reliability challenges arising from communication dynamics and role dependencies.Existing Uncertainty Quantification methods, typically designed for single-turn outputs, fail to address the unique complexities of the MAS.Specifically, these methods struggle with three distinct challenges: the cascading uncertainty in multistep reasoning, the variability of inter-agent communication paths, and the diversity of communication topologies.To bridge this gap, we introduce MATU, a novel framework that quantifies uncertainty through tensor decomposition.MATU moves beyond analyzing final text outputs by representing entire reasoning trajectories as embedding matrices and organizing multiple execution runs into a higher-order tensor.By applying tensor decomposition, we disentangle and quantify distinct sources of uncertainty, offering a comprehensive reliability measure that is generalizable across different agent structures.We provide comprehensive experiments to show that MATU effectively estimates holistic and robust uncertainty across diverse tasks and communication topologies. Tiejin Chen, Huaiyuan Yao, Jia Chen 0002, Evangelos E. Papalexakis, Hua Wei 0001 |
ACL (1) | 3 |
| 2026 | Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression
Paimon Goulart, Jordan Steinhauser, Kylene Shuler, Edward Korzus, Jia Chen 0002, Evangelos E. Papalexakis |
WSDM | 5 |
| 2026 | A Real-Time System to Populate FRA Form 57 from NewsabstractLocal railway committees need timely situational awareness after highway–rail grade crossing incidents, yet official Federal Railroad Administration (FRA) investigations can take days to weeks. We present a demo system that populates Highway–Rail Grade Crossing Incident Data (Form 57) from news in real time. Our approach addresses two core challenges: the form is visually irregular and semantically dense, and news is noisy. To solve these problems, we design a pipeline that first converts Form 57 into a JSON schema using a vision language model with sample aggregation, and then performs grouped question answering following the intent of the form layout to reduce ambiguity. In addition, we build an evaluation dataset by aligning scraped news articles with official FRA records and annotating retrievable information. We then assess our system against various alternatives in terms of information retrieval accuracy and coverage. Chansong Lim, Haz Sameen Shahgir, Yue Dong 0002, Jia Chen 0002, Evangelos E. Papalexakis |
WSDM | 4 |
| 2025 | TRAWL: Tensor Reduced and Approximated Weights for Large Language Models
Het Patel, Yu Fu 0009, Dawon Ahn, Jia Chen 0002, Yue Dong 0002, Evangelos E. Papalexakis |
PAKDD (7) | 5 |
| 2024 | Automating Data Science Pipelines with Tensor CompletionabstractHyperparameter optimization is an essential component in many data science pipelines and typically entails exhaustive time and resource-consuming computations in order to explore the combinatorial search space. Similar to this problem, other key operations in data science pipelines exhibit the exact same properties. Important examples are: neural architecture search, where the goal is to identify the best design choices for a neural network, and query cardinality estimation, where given different predicate values for a SQL query the goal is to estimate the size of the output. In this paper, we abstract away those essential components of data science pipelines and we model them as instances of tensor completion, where each variable of the search space corresponds to one mode of the tensor. Now the goal is to identify all missing entries of the tensor, corresponding to all combinations of variable values, starting from a very small sample of observed entries. In order to do so, we first conduct a thorough experimental evaluation of existing state-of-the-art tensor completion techniques. We also introduce domaininspired adaptations (such as smoothness across the discretized variable space) and an ensemble technique which is able to achieve state-of-the-art performance. We extensively evaluate existing and proposed methods in a number of generated datasets corresponding to (a) hyperparameter optimization for non-neural network models, (b) neural architecture search, and (c) variants of query cardinality estimation. By doing this, we demonstrate the effectiveness of tensor completion as a tool for automating data science pipelines. Furthermore, we release our generated datasets and code in order to provide benchmarks for future work on this topic. Shaan Pakala, Bryce Graw, Dawon Ahn, Tam Dinh, Mehnaz Tabassum Mahin, Vassilis J. Tsotras, Jia Chen 0002, Evangelos E. Papalexakis |
IEEE Big Data | 7 |
| 2022 | TENALIGN: Joint Tensor Alignment and Coupled FactorizationabstractMultimodal datasets represented as tensors oftentimes share some of their modes. However, even though there may exist a one-to-one (or perhaps partial) correspondence between the coupled modes, such correspondence/alignment may not be given, especially when integrating datasets from disparate sources. This is a very important problem, broadly termed as entity alignment or matching, and subsets of the problem such as graph matching have been extremely popular in the recent years. In order to solve this problem, current work computes the alignment based on existing embeddings of the data. This can be problematic if our end goal is the joint analysis of the two datasets into the same latent factor space: the embeddings computed separately per dataset may yield a suboptimal alignment, and if such an alignment is used to subsequently compute the joint latent factors, the computation will similarly be plagued by compounding errors incurred by the imperfect alignment. In this work, we are the first to define and solve the problem of joint tensor alignment and factorization into a shared latent space. By posing this as a unified problem and solving for both tasks simultaneously, we observe that the both alignment and factorization tasks benefit each other resulting in superior performance compared to two-stage approaches. We extensively evaluate our proposed method TENALIGN and conduct a thorough sensitivity and ablation analysis. We demonstrate that TENALIGN significantly outperforms baseline approaches where embedding and matching happen separately. Yunshu Wu, Uday Singh Saini, Jia Chen 0002, Evangelos E. Papalexakis |
ICDM | 3 |
| 2019 | Multiview Canonical Correlation Analysis over GraphsabstractMultiview canonical correlation analysis (MCCA) looks for shared low-dimensional representations hidden in multiple transformations of common source signals. Existing MCCA approaches do not exploit the geometry of common sources, which can be either given a priori, or constructed from do- main knowledge. In this paper, a novel graph-regularized (G) MCCA is developed to account for such geometry-bearing in- formation via graph regularization in the classical maximum- variance MCCA model. GMCCA minimizes the distance between the sought canonical variables and the common sources, while incorporating the graph-induced prior of these sources. To capture nonlinear dependencies, GMCCA is fur- ther broadened to the graph-regularized kernel (GK) MCCA. Numerical tests using real datasets document the merits of G(K)MCCA in comparison with competing alternatives. Jia Chen 0002, Gang Wang 0014, Georgios B. Giannakis |
ICASSP | 1 |
| 2018 | Dpca: Dimensionality Reduction for Discriminative Analytics of Multiple Large-Scale DatasetsabstractPrincipal component analysis (PCA) has well-documented merits for data extraction and dimensionality reduction. PCA deals with a single dataset at a time, and it is challenged when it comes to analyzing multiple datasets. Yet in certain setups, one wishes to extract the most significant information of one dataset relative to other datasets. Specifically, the interest may be on identifying or extracting features that are specific to a single target dataset but not the others. This paper presents a novel approach for such so-termed discriminative data analysis, and establishes its optimality in the least-squares sense under suitable assumptions. The criterion reveals linear combinations of variables by maximizing the ratio of the variance of the target data to that of the remainders. The novel approach solves a generalized eigenvalue problem by performing SVD just once. Numerical tests using synthetic and real datasets showcase the merits of the proposed approach relative to its competing alternatives. Gang Wang 0014, Jia Chen 0002, Georgios B. Giannakis |
ICASSP | 2 |
| 2018 | Exploiting Multi-Phase On-Chip Voltage Regulators as Strong PUF Primitives for Securing IoT
Weize Yu, Yiming Wen, Selçuk Köse, Jia Chen 0002 |
J. Electron. Test. | 4 |
| 2017 | Data-driven sensors clustering and filtering for communication efficient field reconstruction
Jia Chen 0002, Akshay Malhotra, Ioannis D. Schizas |
Signal Process. | 1 |
| 2016 | Distributed information-based clustering of heterogeneous sensor data
Jia Chen 0002, Ioannis D. Schizas |
Signal Process. | 1 |
| 2015 | Regularized canonical correlations for sensor data clusteringabstractThe task of determining informative sensors and clustering the sensor measurements according to their information content is considered. To this end, the standard canonical correlation analysis (CCA) framework is equipped with norm-one and norm-two regularization terms to estimate the unknown number of field sources and identify informative groups of sensors. Coordinate descent techniques are combined with the alternating direction method of multipliers to derive an algorithm that minimizes the regularized CCA framework. An efficient scheme to properly select the regularization coefficients associated with the norm-one and norm-two terms is also developed. Numerical tests corroborate that the novel scheme outperforms existing alternatives. Jia Chen 0002, Ioannis D. Schizas |
ICASSP | 1 |