VLDB 2026 Research / reviewers in the wild / expert
Juntong Ni
dblp:352/5492
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0006-7070-8137ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Efficient and distributed learning · 27% Graph learning · 22% Time series and sequential data · 14% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data stream processing · 62% Data mining · 38% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network › efficient graph neural network
graph condensation |
1.9 | 2 | 2026 | Scalable Graph Condensation with Evolving Capabilities · KDD (1) 2026 GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.8 | 2 | 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation · KDD (1) 2026 Muti-Modal Emotion Recognition via Hierarchical Knowledge Distillation · IEEE Trans. Multim. 2024 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
cross-architecture distillation |
1.0 | 1 | 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation · KDD (1) 2026 |
Machine learning › Time series and sequential data › time series analysis › time series forecasting
long-term time series forecasting |
1.0 | 1 | 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation · KDD (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation · KDD (1) 2026 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
1.0 | 1 | 2026 | STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning |
1.0 | 1 | 2026 | STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning · ACL (1) 2026 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
1.0 | 1 | 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation · KDD (1) 2026 |
Machine learning › Graph learning › graph neural network
graph neural network evaluation |
0.9 | 1 | 2025 | GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › emotion recognition
multimodal emotion recognition |
0.8 | 1 | 2024 | Muti-Modal Emotion Recognition via Hierarchical Knowledge Distillation · IEEE Trans. Multim. 2024 |
Graph algorithms and graph theory › graph simplification
graph coarsening |
0.8 | 1 | 2024 | A Comprehensive Survey on Graph Reduction: Sparsification, Coarsening, and Condensation · IJCAI 2024 |
Graph algorithms and graph theory › graph theory › graph transformation
graph reduction |
0.8 | 1 | 2024 | A Comprehensive Survey on Graph Reduction: Sparsification, Coarsening, and Condensation · IJCAI 2024 |
Graph algorithms and graph theory
graph sparsification |
0.8 | 1 | 2024 | A Comprehensive Survey on Graph Reduction: Sparsification, Coarsening, and Condensation · IJCAI 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.7 | 1 | 2023 | General Debiasing for Multimodal Sentiment Analysis · ACM Multimedia 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | General Debiasing for Multimodal Sentiment Analysis · ACM Multimedia 2023 |
Data mining
clustering |
0.3 | 1 | 2026 | Scalable Graph Condensation with Evolving Capabilities · KDD (1) 2026 |
Data mining
time series analysis |
0.3 | 1 | 2026 | STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning · ACL (1) 2026 |
Machine learning › Graph learning › graph neural network › efficient graph neural network
graph neural network acceleration |
0.3 | 1 | 2025 | GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
feature disentanglement |
0.2 | 1 | 2023 | General Debiasing for Multimodal Sentiment Analysis · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.0multi-agent data synthesis · 2.0feature aggregation · 2.0clustering · 2.0benchmark evaluation · 1.7mixup data augmentation · 1.0knowledge distillation · 1.0MLP · 1.0neural architecture search · 0.9graph neural network · 0.9survey · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement LearningabstractSpatio-temporal reasoning in time series involves the explicit synthesis of temporal dynamics, spatial dependencies, and textual context.This capability is vital for high-stakes decision-making in systems such as traffic networks, power grids, and disease propagation.However, the field remains underdeveloped because most existing works prioritize predictive accuracy over reasoning.To address the gap, we introduce ST-Bench, a benchmark consisting of four core tasks, including etiological reasoning, entity identification, correlation reasoning, and in-context forecasting, developed via a network SDE-based multi-agent data synthesis pipeline.We then propose STReasoner, which empowers LLM to integrate time series, graph structure, and text for explicit reasoning.To promote spatially grounded logic, we introduce S-GRPO, a reinforcement learning algorithm that rewards performance gains specifically attributable to spatial information.Experiments show that STReasoner achieves average accuracy gains between 17% and 135% at only 0.004× the cost of proprietary models and generalizes robustly to real-world data.Our code is available at https://github.com/LingFengGold/STReasoner. Juntong Ni, Shiyu Wang 0001, Qi He 0002, Ming Jin 0005, Wei Jin 0009 |
ACL (1) | 1 |
| 2026 | Scalable Graph Condensation with Evolving CapabilitiesabstractThe rapid growth of graph data creates significant scalability challenges as most graph algorithms scale quadratically with size. To mitigate these issues, Graph Condensation (GC) methods have been proposed to learn a small graph from a larger one, accelerating downstream tasks. However, existing approaches critically assume a static training set, which conflicts with the inherently dynamic and evolving nature of real-world graph data. This work introduces a novel framework for continual graph condensation, enabling efficient updates to the distilled graph that handle data streams without requiring costly retraining. This limitation leads to inefficiencies when condensing growing training sets. In this paper, we introduce GECC (\underline{G}raph \underline{E}volving \underline{C}lustering \underline{C}ondensation), a scalable graph condensation method designed to handle large-scale and evolving graph data. GECC employs a traceable and efficient approach by performing class-wise clustering on aggregated features. Furthermore, it can inherit previous condensation results as clustering centroids when the condensed graph expands, thereby attaining an evolving capability. This methodology is supported by robust theoretical foundations and demonstrates superior empirical performance. Comprehensive experiments including real world scenario show that GECC achieves better performance than most state-of-the-art graph condensation methods while delivering an around 1000$\times$ speedup on large datasets. Shengbo Gong, Juntong Ni, Carl Yang 0001, Wei Jin 0009 |
KDD (1) | 3 |
| 2026 | TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture DistillationabstractTransformer-based and CNN-based methods demonstrate strong performance in long-term time series forecasting. However, their high computational and storage requirements can hinder large-scale deployment. To address this limitation, we propose integrating lightweight MLP with advanced architectures using knowledge distillation (KD). Our preliminary study reveals different models can capture complementary patterns, particularly multi-scale and multi-period patterns in the temporal and frequency domains. Based on this observation, we introduce TimeDistill, a cross-architecture KD framework that transfers these patterns from teacher models (e.g., Transformers, CNNs) to MLP. Additionally, we provide a theoretical analysis, demonstrating that our KD approach can be interpreted as a specialized form of mixup data augmentation. TimeDistill improves MLP performance by up to 18.6%, surpassing teacher models on eight datasets. It also achieves up to 7X faster inference and requires 130X fewer parameters. Furthermore, we conduct extensive evaluations to highlight the versatility and effectiveness of TimeDistill. The code is available at Github Code Repo. Juntong Ni, Zewen Liu 0005, Shiyu Wang 0001, Ming Jin 0005, Wei Jin 0009 |
KDD (1) | 1 |
| 2026 | A comprehensive survey of AI agents in healthcareabstractOBJECTIVE: This survey aims to systematically map the rapidly evolving landscape of AI agents in healthcare. It addresses the critical need to adapt general-purpose agentic frameworks characterized by autonomy, planning, and tool use to the high-stakes, safety-critical constraints of medical decision-making and patient care. METHODS: We conducted a comprehensive review of over 200 recent studies, synthesizing literature from major academic databases. We developed a holistic taxonomy that traces the full lifecycle of healthcare agents, analyzing perception modalities, core technical architectures, and evaluation protocols specific to autonomous systems. RESULTS: The review presents a quantitative landscape analysis showing exponential growth in the field. We structure the domain into three pillars: (1) Perception of multi-modal clinical data (e.g., EHR, imaging, genomics); (2) Agent Capabilities, including tool use, reasoning, memory, and multi-agent collaboration; and (3) an Application Ecosystem organized by stakeholder roles (clinicians, patients, researchers, and administrators). Additionally, we categorize evaluation frameworks, and discuss the deployment readiness of current systems across technical, evidentiary, and governance dimensions. Finally, we identify challenges for advancing healthcare agents from controlled evaluation toward real-world clinical integration. A continuously updated repository of related papers is available at https://github.com/AgenticHealthAI/Awesome-AI-Agents-for-Healthcare. CONCLUSION: AI agents offer significant potential to enhance healthcare through autonomous reasoning and workflow integration. However, current research remains largely concentrated in benchmark and controlled evaluation settings, and the translation into clinical practice will require advances in reliability, privacy protection, governance, and operational integration. Gelei Xu, Yixiong Chen, Yuying Duan, Shuqing Wu, Haoxinran Yu, Ching-Hao Chiu, Juntong Ni, Ningzhi Tang, Toby Jia-Jun Li, Alan L. Yuille, Wei Jin 0009, Yiyu Shi 0001 |
J. Biomed. Informatics | 8 |
| 2025 | LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchabstractShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang, Daoyang Li, Liqiang Jing, Kaiyu He, Peilin Wu, Juntong Ni, George Michalopoulos, Yue Zhang, Ziyang Zhang, Mian Zhang, Zhiyu Chen, Xinya Du. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziming Luo, Daoyang Li, Liqiang Jing, Kaiyu He, Juntong Ni, George Michalopoulos, Zhiyu Chen 0002, Xinya Du |
EMNLP | 9 |
| 2025 | GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New InsightsabstractGraph condensation (GC) is an emerging technique designed to learn a significantly smaller graph that retains the essential information of the original graph. This condensed graph has shown promise in accelerating graph neural networks while preserving performance comparable to those achieved with the original, larger graphs. Additionally, this technique facilitates downstream applications like neural architecture search and deepens our understanding of redundancies in large graphs. Despite the rapid development of GC methods, particularly for node classification, a unified evaluation framework is still lacking to systematically compare different GC methods or clarify key design choices for improving their effectiveness. To bridge these gaps, we introduce GC4NC, a comprehensive framework for evaluating diverse GC methods on node classification across multiple dimensions including performance, efficiency, privacy preservation, denoising ability, NAS effectiveness, and transferability. Our systematic evaluation offers novel insights into how condensed graphs behave and the critical design choices that drive their success. These findings pave the way for future advancements in GC methods, enhancing both performance and expanding their real-world applications. The code is available at https://github.com/Emory-Melody/GraphSlim/tree/main/benchmark. Shengbo Gong, Juntong Ni, Noveen Sachdeva, Carl Yang 0001, Wei Jin 0009 |
NeurIPS | 2 |
| 2024 | A Comprehensive Survey on Graph Reduction: Sparsification, Coarsening, and Condensation
Shengbo Gong, Juntong Ni, Wenqi Fan, B. Aditya Prakash, Wei Jin 0009 |
IJCAI | 3 |
| 2024 | Muti-Modal Emotion Recognition via Hierarchical Knowledge DistillationabstractDue to its wide applications, multimodal emotion recognition has gained increasing research attention. Although existing methods have achieved compelling success with various multimodal fusion methods, they overlook that the dominated modality (e.g., text) may cause a shortcut and hence negatively affect the representation learning of other modalities (e.g., image and audio). To alleviate such a problem, we resort to the knowledge distillation to narrow the gap between different modalities. In particular, we develop a new hierarchical knowledge distillation model for multi-modal emotion recognition (HKD-MER), consisting of three components, feature extraction, hierarchical knowledge distillation, and attentive multi-modal fusion. As the major contribution in our proposed model, the hierarchical knowledge distillation is designed to transfer the knowledge from the dominant modality to the others at both the feature and label levels. It boosts the performance of non-dominated modalities by modeling the inter-modal relation between different modalities. We have justified the effectiveness of our proposed model over two benchmark datasets. Yinwei Wei, Juntong Ni, Xuemeng Song, Yaowei Wang 0001, Liqiang Nie |
IEEE Trans. Multim. | 3 |
| 2023 | General Debiasing for Multimodal Sentiment AnalysisabstractExisting work on Multimodal Sentiment Analysis (MSA) utilizes multimodal information for prediction yet unavoidably suffers from fitting the spurious correlations between multimodal features and sentiment labels. For example, if most videos with a blue background have positive labels in a dataset, the model will rely on such correlations for prediction, while "blue background'' is not a sentiment-related feature. To address this problem, we define a general debiasing MSA task, which aims to enhance the Out-Of-Distribution (OOD) generalization ability of MSA models by reducing their reliance on spurious correlations. To this end, we propose a general debiasing framework based on Inverse Probability Weighting (IPW), which adaptively assigns small weights to the samples with larger bias (i.e., the severer spurious correlations). The key to this debiasing framework is to estimate the bias of each sample, which is achieved by two steps: 1) disentangling the robust features and biased features in each modality, and 2) utilizing the biased features to estimate the bias. Finally, we employ IPW to reduce the effects of large-biased samples, facilitating robust feature learning for sentiment prediction. To examine the model's generalization ability, we keep the original testing sets on two benchmarks and additionally construct multiple unimodal and multimodal OOD testing sets. The empirical results demonstrate the superior generalization ability of our proposed framework. We have released the code to facilitate the reproduction https://github.com/Teng-Sun/GEAR. Juntong Ni, Wenjie Wang 0007, Liqiang Jing, Yinwei Wei, Liqiang Nie |
ACM Multimedia | 2 |