VLDB 2026 Research / reviewers in the wild / expert
Jingyan Jiang
dblp:45/7839
· DBLP profile ↗
23ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0003-4897-2645ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNormabstractTest-time adaptation (TTA) has proven effective in mitigating performance drops under single-domain distribution shifts by updating model parameters during inference. However, real-world deployments often involve mixed distribution shifts---where test samples are affected by diverse and potentially conflicting domain factors---posing significant challenges even for state-of-the-art TTA methods. A key limitation in existing approaches is their reliance on a unified adaptation path, which fails to account for the fact that optimal gradient directions can vary significantly across different domains. Moreover, current benchmarks focus only on synthetic or homogeneous shifts, failing to capture the complexity of real-world heterogeneous mixed distribution shifts. To address this, we propose MoETTA, a novel entropy-based TTA framework that integrates the Mixture-of-Experts (MoE) architecture. Rather than enforcing a single parameter update rule for all test samples, MoETTA introduces a set of structurally decoupled experts, enabling specialization along diverse gradient directions. This design allows the model to better accommodate heterogeneous shifts through flexible and disentangled parameter updates. To simulate realistic deployment conditions, we introduce two new benchmarks: potpourri and potpourri+. While classical settings focus solely on synthetic corruptions (i.e., ImageNet-C), potpourri encompasses a broader range of domain shifts—including natural, artistic, and adversarial distortions—capturing more realistic deployment challenges. On top of that, potpourri+ further includes source-domain samples to evaluate robustness against catastrophic forgetting. Extensive experiments across three mixed distribution shifts settings show that MoETTA consistently outperforms strong baselines, establishing new state-of-the-art performance and highlighting the benefit of modeling multiple adaptation directions via expert-level diversity. Jingyan Jiang, Zhaoru Chen, Fanding Huang, Qinting Jiang |
AAAI | 2 |
| 2026 | Balancing Uncertainty and Diversity in Active Learning through Multi-Stage Cluster-Representative Selection
Shinan Song, Yanting Sun, Lanlan Gao, Huayun Tang, Jingyan Jiang |
ICIC (16) | 5 |
| 2026 | Exploring Test-time Scaling via Prediction Merging on Large-Scale RecommendationabstractInspired by the success of language models (LM), scaling up deep learning recommendation systems (DLRS) has become a recent trend in the community. All previous methods tend to scale up the model parameters during training time. However, how to efficiently utilize and scale up computational resources during test time remains underexplored, which can prove to be a scaling-efficient approach and bring orthogonal improvements in LM domains. The key point in applying test-time scaling to DLRS lies in effectively generating diverse yet meaningful outputs for the same instance. We propose two ways: One is to explore the heterogeneity of different model architectures. The other is to utilize the randomness of model initialization under a homogeneous architecture. The evaluation is conducted across eight models, including both classic and SOTA models, on three benchmarks. Sufficient evidence proves the effectiveness of both solutions. We further prove that under the same inference budget, test-time scaling can outperform parameter scaling. Our test-time scaling can also be seamlessly accelerated with the increase in parallel servers when deployed online, without affecting the inference time on the user side. Code is available here. https://github.com/aTitye/TTS4CTR. Fuyuan Lyu, Zhentai Chen, Jingyan Jiang, Xing Tang 0007, Xiuqiang He 0001, Xue (Steve) Liu |
SIGIR | 3 |
| 2026 | MT2-CSD and LLM-CRAN: A new dataset and an LLM-based multi-semantic knowledge fusion model for conversational stance detection
Fuqiang Niu, Genan Dai, Yisha Lu, Jiayu Liao, Xiang Li 0130, Jingyan Jiang, Hu Huang 0009, Bowen Zhang 0005 |
Neural Networks | 6 |
| 2026 | Domain-invariant representation learning via SAM for blood cell classification
Lingcong Cai, Jingyan Jiang, Genan Dai, Bowen Zhang 0005, Jingzhou Cao, Xiangzhong Zhang, Xiaomao Fan |
Pattern Recognit. | 6 |
| 2025 | Adaptive Confidence Estimation for Data Distribution Shift Robustness in Cloud-Edge Collaborative Inference
Shinan Song, Wenjun Ma, Xiaomao Fan, Jinzhou Cao, Jingyan Jiang |
ADMA (2) | 6 |
| 2025 | Q-DiT: Accurate Post-Training Quantization for Diffusion TransformersabstractRecent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressive capabilities, the substantial computational costs of these large-scale models pose significant challenges for real-world deployment. Post-Training Quantization (PTQ) emerges as a promising solution, enabling model compression and accelerated inference for pretrained models, without the costly retraining. However, research on DiT quantization remains sparse, and existing PTQ frameworks, primarily designed for traditional diffusion models, tend to suffer from biased quantization, leading to notable performance degradation. In this work, we identify that DiTs typically exhibit significant spatial variance in both weights and activations, along with temporal variance in activations. To address these issues, we propose Q-DiT, a novel approach that seamlessly integrates two key techniques: automatic quantization granularity allocation to handle the significant variance of weights and activations across input channels, and sample-wise dynamic activation quantization to adaptively capture activation changes across both timesteps and samples. Extensive experiments conducted on ImageNet and VBench demonstrate the effectiveness of the proposed Q-DiT. Specifically, when quantizing DiT-XL/2 to W6A8 on ImageNet (256 × 256), Q-DiT achieves a remarkable reduction in FID by 1.09 compared to the baseline. Under the more challenging W4A8 setting, it maintains high fidelity in image and video generation, establishing a new benchmark for efficient, high-quality quantization in DiTs. Xinzhu Ma, Jingyan Jiang, Xin Wang 0019, Zhi Wang 0001, Wenwu Zhu 0001 |
CVPR | 5 |
| 2025 | COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time AdaptationabstractRecent vision-language models (VLMs) face significant challenges in test-time adaptation to novel domains. While cache-based methods show promise by leveraging historical information, they struggle with both caching unreliable feature-label pairs and indiscriminately using single-class information during querying, significantly compromising adaptation accuracy. To address these limitations, we propose COSMIC (Clique-Oriented Semantic Multi-space Integration for CLIP), a robust test-time adaptation framework that enhances adaptability through multi-granular, cross-modal semantic caching and graph-based querying mechanisms. Our framework introduces two key innovations: Dual Semantics Graph (DSG) and Clique Guided Hyper-class (CGH). The Dual Semantics Graph constructs complementary semantic spaces by incorporating textual features, coarse-grained CLIP features, and fine-grained DINOv2 features to capture rich semantic relationships. Building upon these dual graphs, the Clique Guided Hyper-class component leverages structured class relationships to enhance prediction robustness through correlated class selection. Extensive experiments demonstrate COSMIC’s superior performance across multiple benchmarks, achieving significant improvements over state-of-the-art methods: 15.81% gain on out-of-distribution tasks and 5.33% on cross-domain generation with CLIP RN-50. Code is available at github.com/hf618/COSMIC. Fanding Huang, Jingyan Jiang, Qinting Jiang, Hebei Li, Faisal Nadeem Khan |
CVPR | 2 |
| 2025 | Dynamic Model Fusion for Multi-Source Test-Time AdaptationabstractDeep Neural Networks suffer significant performance degradation when faced with distribution shifts between training and test data. Test-time adaptation (TTA) has emerged as a practical solution that enables models to adapt to the shifted test distribution. Currently, most existing TTA methods are designed around a single model, which incorporate limited information from a singular data distribution. In practice, pre-trained models derived from diverse source domains are readily accessible, each capturing a distinct data distribution and containing complementary information. To exploit this diversity, we propose Model Fusion-based multi-source Test-Time Adaptation (MFTTA), which constructs a target model by fusing the parameters of multiple source models. Drawing inspiration from deep model fusion, we introduce a fine-grained fusion mechanism governed by an off-policy reinforcement learning agent, which dynamically assigns fusion weights based on the current data distribution. Furthermore, we design a correlation-aware model update strategy that prioritizes the source model most relevant to the incoming test data. Extensive experiments on standard out-of-distribution benchmarks demonstrate that our method effectively integrates knowledge from multiple source models, adapts robustly to dynamic distribution shifts, and alleviates the problem of forgetting in long-term adaptation. Yuan Xue 0013, Qinting Jiang, Xingxuan Zhang, Jingyan Jiang, Zhi Wang 0001 |
ECAI | 6 |
| 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingabstractThe emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training. Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001 |
EMNLP | 2 |
| 2025 | Test-Time Adaptation via Dynamic Historical Knowledge Vector Fusion
Zhaoru Chen, Yongjin Wu, Qingsong Ye, Liusha Yang, Jingyan Jiang |
ICIC (18) | 5 |
| 2025 | DATTA: Domain Diversity Aware Test-Time Adaptation for Dynamic Domain Shift Data StreamsabstractTest-Time Adaptation (TTA) addresses domain shifts between training and testing. However, existing methods assume a homogeneous target domain (e.g., single domain) at any given time. They fail to handle the dynamic nature of real-world data, where single-domain and multiple-domain distributions change over time. We identify that performance drops in multiple-domain scenarios are caused by batch normalization errors and gradient conflicts, which hinder adaptation. To solve these challenges, we propose Domain Diversity Adaptive Test-Time Adaptation (DATTA), the first approach to handle TTA under dynamic domain shift data streams. It is guided by a novel domain-diversity score. DATTA has three key components: a domain-diversity discriminator to recognize single- and multiple-domain patterns, domain-diversity adaptive batch normalization to combine source and test-time statistics, and domain-diversity adaptive fine-tuning to resolve gradient conflicts. Extensive experiments show that DATTA significantly outperforms state-of-the-art methods by up to 13%. Code is available at https://github.com/DYW77/DATTA. Chuyang Ye, Dongyan Wei 0001, Yuanyi Pang, Yixi Lin, Qinting Jiang, Jingyan Jiang, Dongbiao He |
ICME | 7 |
| 2025 | Feature-Based Instance Neighbor Discovery: Advanced Stable Test-Time Adaptation in Dynamic WorldabstractDespite progress, deep neural networks still suffer performance declines under distribution shifts between training and test domains, leading to a substantial decrease in Quality of Experience (QoE) for applications. Existing test-time adaptation (TTA) methods are challenged by dynamic, multiple test distributions within batches. We observe that feature distributions across different domains inherently cluster into distinct groups with varying means and variances. This divergence reveals a critical limitation of previous global normalization strategies in TTA, which inevitably distort the original data characteristics. Based on this insight, we propose Feature-based Instance Neighbor Discovery (FIND), which comprises three key components: Layer-Wise Feature Disentanglement (LFD), Feature-Aware Batch Normalization (FABN) and Selective FABN (S-FABN). LFD stably captures features with similar distributions at each layer by constructing graph structures; while FABN optimally combines source statistics with test-time distribution-specific statistics for robust feature representation. Finally, S-FABN determines which layers require feature partitioning and which can remain unified, thus enhancing the efficiency of inference. Extensive experiments demonstrate that FIND significantly outperforms existing methods, achieving up to approximately 30\% accuracy improvement in dynamic scenarios while maintaining computational efficiency. The source code is available at https://github.com/Peanut-255/FIND. Qinting Jiang, Chuyang Ye, Dongyan Wei 0001, Bingli Wang, Jingyan Jiang |
NeurIPS | 6 |
| 2025 | Accelerating Parallel Diffusion Model Serving with Residual CompressionabstractDiffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However, parallel inference introduces significant communication overhead from exchanging large activations between devices, limiting efficiency and scalability. We present CompactFusion, a compression framework that significantly reduces communication while preserving generation quality. Our key observation is that diffusion activations exhibit strong temporal redundancy—adjacent steps produce highly similar activations, saturating bandwidth with near-duplicate data carrying little new information. To address this inefficiency, we seek a more compact representation that encodes only the essential information. CompactFusion achieves this via Residual Compression that transmits only compressed residuals (step-wise activation differences). Based on empirical analysis and theoretical justification, we show that it effectively removes redundant data, enabling substantial data reduction while maintaining high fidelity. We also integrate lightweight error feedback to prevent error accumulation. CompactFusion establishes a new paradigm for parallel diffusion inference, delivering lower latency and significantly higher generation quality than prior methods. On 4$\times$L20, it achieves $3.0\times$ speedup while greatly improving fidelity. It also uniquely supports communication-heavy strategies like sequence parallelism on slow networks, achieving $6.7\times$ speedup over prior overlap-based method. CompactFusion applies broadly across diffusion models and parallel settings, and integrates easily without requiring pipeline rework. Portable implementation demonstrated on xDiT is publicly available at https://github.com/Cobalt-27/CompactFusion Jiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You, Rongwei Lu, Jingyan Jiang, Zhi Wang 0001 |
NeurIPS | 7 |
| 2025 | LLM4Band: Enhancing Reinforcement Learning with Large Language Models for Accurate Bandwidth EstimationabstractReal-time communication (RTC) applications rely on accurate bandwidth estimation to ensure high-quality communication and user experience. Traditional heuristic and reinforcement learning (RL)-based methods often face challenges with the dynamic nature of real-time networks, leading to issues with generalization. Inspired by the success of Large Language Models (LLMs)---which, with billions of parameters pre-trained on massive datasets, have demonstrated exceptional capabilities in semantic representation, adaptability, and transfer learning---we propose LLM4Band, a novel framework that integrates LLMs with offline reinforcement learning to tackle bandwidth estimation in RTC scenarios. By leveraging the powerful feature extraction capabilities of LLMs and combining them with an offline RL algorithm, LLM4Band incorporates a Balanced Replay Buffer and an LLM-based policy network to significantly enhance robustness and adaptability. Extensive experiments demonstrate that LLM4Band surpasses state-of-the-art methods, achieving a 12.35% improvement in estimation accuracy and a 21% enhancement in communication quality. Rongwei Lu, Cédric Westphal, Dongbiao He, Jingyan Jiang |
NOSSDAV | 6 |
| 2024 | Potential Game Based Distributed IoV Service Offloading With Graph Attention Networks in Mobile Edge ComputingabstractVehicular services aim to provide smart and timely services (e.g., collision warning) by taking the advantage of recent advances in artificial intelligence and employing task offloading techniques in mobile edge computing. In practice, the volume of vehicles in the Internet of Vehicles (IoV) often surges at a single location and renders the edge servers (ESs) severely overloaded, resulting in a very high delay in delivering the services. Therefore, it is of practical importance and urgency to coordinate the resources of ESs with bandwidth allocation for mitigating the occurrence of a spike traffic flow. For this challenge, existing work sought the periodicities of traffic flow by analyzing historical traffic data. However, the changes in traffic flow caused by sudden traffic conditions cannot be obtained from these periodicities. In this paper, we propose a distributed traffic flow forecasting and task offloading approach named TFFTO to optimize the execution time and power consumption in service processing. Specifically, graph attention networks (GATs) are leveraged to forecast future traffic flow in short-term and the traffic volume is utilized to estimate the number of services offloaded to the ESs in the subsequent period. With the estimate, the current load of the ESs is adjusted to ensure that the services can be handled in a timely manner. Potential game theory is adopted to determine the optimal service offloading strategy. Extensive experiments are conducted to evaluate our approach and the results validate our robust performance. Qinting Jiang, Xiaolong Xu 0001, Muhammad Bilal 0003, Jon Crowcroft, Qi Liu 0001, Wan-Chun Dou, Jingyan Jiang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Fast-DRD: Fast decentralized reinforcement distillation for deadline-aware edge computing
Shinan Song, Zhiyi Fang, Jingyan Jiang |
Inf. Process. Manag. | 3 |
| 2021 | Joint Model and Data Adaptation for Cloud Inference ServingabstractReal-time deep learning inference serving systems often require prohibitive resources and diverse user requirements. The existing design of inference serving systems mainly focusing on computation resource efficiency, largely ignoring the trade-off between computation and bandwidth resources in need. Sub-optimal resource utilization usually leads to huge serving cost waste. In this paper, we tackle the dual challenge of computation-bandwidth trade-off and cost-effectiveness by proposing A2, an efficient joint Adaptive model, and Adaptive data deep learning serving solution across the geo-datacenters. Inspired by the insight that a trade-off between computational cost and bandwidth cost in achieving the same accuracy, we design a real-time inference serving framework, which selectively places different "versions" of the deep learning models at different geo-locations, and schedules different data sample versions to be sent to those model versions for inference. The goal is to minimize the total serving cost while meeting latency and accuracy demand for the serving requests. We formulate a joint placement and serving problem and propose an efficient approximation algorithm to solve it with a theoretical performance guarantee. We deploy A2on Amazon EC2 for experiments, which shows that A2achieves 30%-50% serving cost reduction under the same required latency and accuracy as compared to baselines. Jingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He, Zhi Wang 0001, Shutao Xia, Chuan Wu 0001 |
RTSS | 1 |
| 2021 | Reinforcement learning approach for resource allocation in humanitarian logistics
Canrong Zhang, Jingyan Jiang, Huasheng Yang, Huayan Shang |
Expert Syst. Appl. | 3 |
| 2019 | Dynamic pricing with traffic engineering for adaptive video streaming over software-defined content delivery networking
Pingting Hao, Liang Hu 0001, Kuo Zhao, Jingyan Jiang, Tong Li 0011, Xilong Che |
Multim. Tools Appl. | 4 |
| 2019 | Mobile Edge Provision with Flexible DeploymentabstractThe Mobile Edge Network (MEN) has emerged as the basic infrastructure to support fifth-generation networks, mobile edge computing and fog computing. The characteristics of mobility must be addressed to guarantee the quality of service in MENs. As one of the critical problems in MENs, flexible deployment plays a part in exploiting edge networks. Despite the abundance of recently proposed strategies, most concentrate on the change in user demands and inevitably ignore the influence for the user mobility, which is common in future networks. We propose Provision for Mobile Edge Computing (PMEC), a prototype that takes advantage of storage devices with flexible placement. In PMEC, we accommodate various considerations and select different storage devices to cache, deploying the cache with the relationship of a two-tiered structure in the MEN. Thus, we construct a flexible overlay network with the objective of minimizing the cost in a two-tiered edge network. Based on the analysis of the problem, we solve the two-tiered placement from bottom to top using the dynamic minimal spanning tree (MST) algorithm and design two algorithms for each tier including the basic algorithm and the improved algorithms. The simulation is conducted on realistic data to demonstrate the performance of our algorithms. Pingting Hao, Liang Hu 0001, Jingyan Jiang, Jiejun Hu, Xilong Che |
IEEE Trans. Serv. Comput. | 3 |
| 2018 | JALAD: Joint Accuracy-And Latency-Aware Deep Structure Decoupling for Edge-Cloud ExecutionabstractRecent years have witnessed a rapid growth of deep-network based services and applications. A practical and critical problem thus has emerged: how to effectively deploy the deep neural network models such that they can be executed efficiently. Conventional cloud-based approaches usually run the deep models in data center servers, causing large latency because a significant amount of data has to be transferred from the edge of network to the data center. In this paper, we propose JALAD, a joint accuracy- and latency-aware execution framework, which decouples a deep neural network so that a part of it will run at edge devices and the other part inside the conventional cloud, while only a minimum amount of data has to be transferred between them. Though the idea seems straightforward, we are facing challenges including i)how to find the best partition of a deep structure; ii)how to deploy the component at an edge device that only has limited computation power; and iii)how to minimize the overall execution latency. Our answers to these questions are a set of strategies in JALAD, including 1)A normalization based in-layer data compression strategy by jointly considering compression rate and model accuracy; 2)A latency-aware deep decoupling strategy to minimize the overall execution latency; and 3)An edge-cloud structure adaptation strategy that dynamically changes the decoupling for different network conditions. Experiments demonstrate that our solution can significantly reduce the execution latency: it speeds up the overall inference execution with a guaranteed model accuracy loss. Hongshan Li, Chenghao Hu, Jingyan Jiang, Zhi Wang 0001, Yonggang Wen 0001, Wenwu Zhu 0001 |
ICPADS | 3 |
| 2018 | Q-FDBA: improving QoE fairness for video streaming
Jingyan Jiang, Liang Hu 0001, Pingting Hao, Jiejun Hu, Hongtu Li |
Multim. Tools Appl. | 1 |