VLDB 2026 Research / reviewers in the wild / expert
Yi Jing
dblp:175/2516
· DBLP profile ↗
19ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MRACL: Multi-Reward Space Guided Adaptive Curriculum Reinforcement Learning for LLMsabstractReinforcement learning (RL) has recently become a powerful yet resource-intensive approach for post-training large language models (LLMs). Incorporating curriculum learning (CL) into RL has been shown to significantly improve training efficiency, particularly in reasoning tasks. However, existing CL methods face substantial challenges in multi-objective RL (MORL) settings, including: (1) difficulty in evaluating model capabilities online, (2) challenges in assessing sample importance under diverse objectives, and (3) inherent trade-offs between online training and offline inference in dynamically designing the curriculum. To address these issues, we propose a Multi-Reward space guided Adaptive Curriculum Learning framework (MRACL), which is the first to incorporate curriculum learning into multi-objective RL. MRACL first constructs a multi-dimensional reward space via offline inference to establish initial reward profiles for each training sample. During training, based on reward space, it estimates the evolving model capabilities by computing the centroid of the space and calculates the sample priority score through its capability distance, optimization direction, and historical evolution, which enables adaptive selection of the most informative training samples at each step, independent of the specific RL algorithm. After each RL training iteration, the reward space is dynamically updated to reflect the model's evolving capabilities and the shifting distribution of sample priorities. Experiments on multi-objective alignment tasks demonstrate that MRACL achieves 1.62× faster convergence compared to state-of-the-art curriculum methods and 2.55× faster than non-curriculum methods. Furthermore, it consistently outperforms all baselines in both win rate and rule-based evaluation. We further provide an in-depth analysis of the key factors contributing to \modelname's effectiveness, along with its advantages, scenarios, and generalization across diverse settings. Liangyu Huo, Yi Jing |
AAAI | 3 |
| 2026 | HistLens: Mapping Idea Change across Concepts and CorporaabstractLanguage change both reflects and shapes social processes, and the semantic evolution of foundational concepts provides a measurable trace of historical and social transformation.Despite recent advances in diachronic semantics and discourse analysis, existing computational approaches often (i) concentrate on a single concept or a single corpus, making findings difficult to compare across heterogeneous sources, and (ii) remain confined to surface lexical evidence, offering insufficient computational and interpretive granularity when concepts are expressed implicitly.We propose HistLens, a unified, SAE-based framework for multi-concept, multi-corpus conceptualhistory analysis.The framework decomposes concept representations into interpretable features and tracks their activation dynamics over time and across sources, yielding comparable conceptual trajectories within a shared coordinate system.Experiments on long-span press corpora show that HistLens supports crossconcept, cross-corpus computation of patterns of idea evolution and enables implicit concept computation.By bridging conceptual modeling with interpretive needs, HistLens broadens the analytical perspectives and methodological repertoire available to social science and the humanities for diachronic text analysis. Yi Jing, Weiyun Qiu, Yihang Peng, Zhifang Sui |
ACL (1) | 1 |
| 2026 | Two-Stage DQN Enabled Joint Communication-Computation Optimization in SAGIN for Multimodal Tasks
Zhenghao Lu, Yi Jing |
IWCMC | 3 |
| 2026 | Rhythmic Resource State Sensing in LEO-MEO Satellite Networks Using Deep Reinforcement LearningabstractSatellite-enabled Internet of Things (IoT) services such as maritime sensing, aviation tracking, emergency telemetry, and wide-area monitoring require timely network-state awareness to support access control, load balancing, and resource scheduling. In Low Earth orbit (LEO) constellations, conventional resource-state reporting via ground stations is limited by short and intermittent contact windows, while large-scale LEO-to-LEO relaying is constrained by inter-satellite capacity and multi-hop latency. We investigate a LEO-medium Earth orbit (MEO) architecture where MEO satellites act as persistent aggregation nodes that collect resource-state updates from many IoT-serving LEO satellites over cross-orbit links. Due to spatially non-uniform IoT traffic and heterogeneous service rhythms, LEO satellites exhibit different resource-evolution time scales, making fixed-period sensing either waste signaling for slowly varying satellites or produce stale information for rapidly varying satellites. This creates a coupled trade-off between reporting delay and Age of Information (AoI), under limited LEO–MEO sensing capacity. To address this issue, we propose a rhythmic resource-state sensing framework that adapts each LEO satellite’s reporting cadence using a dynamic frame structure and a deep reinforcement learning policy trained by proximal policy optimization to minimize average reporting delay subject to AoI and sensing constraints. Simulations across constellation scales and traffic patterns show that the proposed approach reduces the average reporting delay by 22.5% on average and up to 25.5% compared with a heuristic baseline, while reducing the composite delay–freshness cost by 16.7% on average and up to 25.7%. Yi Jing, Chunxiao Jiang, Jiawei Wang 0012, Yafeng Zhan |
IEEE Internet Things J. | 1 |
| 2026 | GCN combined with snake convolution for enhanced topological perception in thrombotic hepatic portal vein segmentation
Lijuan Ma, Xingshun Qi, Yi Jing |
Medical Image Anal. | 4 |
| 2026 | PLM-SynNet: A Pathology Large Model Synergy Network Based on Multi-Instance Learning for Whole Slide Imaging ClassificationabstractWhole slide imaging (WSI) provides rich tissue information at a gigapixel resolution, posing significant challenges for the development of pathology analysis algorithms. Mainstream approaches effectively analyze WSI but treat the pretrained feature extractor and the task-specific network as independent modules, thereby restricting downstream task accuracy due to the limitations of the pretrained model. Inspired by the knowledge complementarity mechanism in multi-agent collaboration, we propose PLM-SynNet, a pathology large model synergy network that integrates the strengths of multiple pathology large models (PLMs) by establishing a flexible collaborative structure and achieving information gain. Specifically, the PLM Synergy Block (PLM-SB) is designed based on Mixture of Experts (MoE), which flexibly generates and utilizes supplementary features by employing a feature generator as an expert in MoE and merging these outputs via pixel-wise summation for effective collaboration. Subsequently, a Synergy Reinforcement Loss (SRLoss) is defined to enhance the information gain of multiple PLMs by enforcing stricter constraints on both queried and generated features. Experiments on a private PCA-EPE (Extraprostatic Extension of Prostate Cancer) dataset and two public datasets demonstrate the effectiveness of the proposed method, yielding gains of 13.27% in F1-score, 6.30% in accuracy, and 15.16% in MCC on PCA-EPE. It further improves TCGA-CRC accuracy by 1.71% and enhances BRIGHT accuracy and AUC by 2.67% and 4.00%, respectively. The code repository is available at https://github.com/mathfyy/PLM-SynNet. Yingying Feng, Yi Jing, Moyu Xia, Xuanyi Zhang |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Automated Progressive Red TeamingabstractEnsuring the safety of large language models (LLMs) is paramount, yet identifying potential vulnerabilities is challenging. While manual red teaming is effective, it is time-consuming, costly and lacks scalability. Automated red teaming (ART) offers a more cost-effective alternative, automatically generating adversarial prompts to expose LLM vulnerabilities. However, in current ART efforts, a robust framework is absent, which explicitly frames red teaming as an effectively learnable task. To address this gap, we propose Automated Progressive Red Teaming (APRT) as an effectively learnable framework. APRT leverages three core modules: an Intention Expanding LLM that generates diverse initial attack samples, an Intention Hiding LLM that crafts deceptive prompts, and an Evil Maker to manage prompt diversity and filter ineffective samples. The three modules collectively and progressively explore and exploit LLM vulnerabilities through multi-round interactions. In addition to the framework, we further propose a novel indicator, Attack Effectiveness Rate (AER) to mitigate the limitations of existing evaluation metrics. By measuring the likelihood of eliciting unsafe but seemingly helpful responses, AER aligns closely with human evaluations. Extensive experiments with both automatic and human evaluations, demonstrate the effectiveness of APRT across both open- and closed-source LLMs. Specifically, APRT effectively elicits 54% unsafe yet useful responses from Meta’s Llama-3-8B-Instruct, 50% from GPT-4o (API access), and 39% from Claude-3.5 (API access), showcasing its robust attack capability and transferability across LLMs (especially from open-source LLMs to closed-source LLMs). The code and seed data are available at https://github.com/tjunlp-lab/APRT. Bojian Jiang, Yi Jing, Tianhao Shen, Deyi Xiong |
COLING | 2 |
| 2025 | LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-EncoderabstractLarge language models (LLMs) demonstrate exceptional performance on tasks requiring complex linguistic abilities, such as reference disambiguation and metaphor recognition/generation.Although LLMs possess impressive capabilities, their internal mechanisms for processing and representing linguistic knowledge remain largely opaque.Prior research on linguistic mechanisms is limited by coarse granularity, limited analysis scale, and narrow focus.In this study, we propose LINGUALENS, a systematic and comprehensive framework for analyzing the linguistic mechanisms of large language models, based on Sparse Auto-Encoders (SAEs).We extract a broad set of Chinese and English linguistic features across four dimensions-morphology, syntax, semantics, and pragmatics.By employing counterfactual methods, we construct a large-scale counterfactual dataset of linguistic features for mechanism analysis.Our findings reveal intrinsic representations of linguistic knowledge in LLMs, uncover patterns of cross-layer and cross-lingual distribution, and demonstrate the potential to control model outputs.This work provides a systematic suite of resources and methods for studying linguistic mechanisms, offers strong evidence that LLMs possess genuine linguistic knowledge, and lays the foundation for more interpretable and controllable language modeling in future research. Yi Jing, Zijun Yao 0002, Hongzhu Guo, Lingxu Ran, Xiaozhi Wang, Lei Hou 0001, Juan-Zi Li |
EMNLP | 1 |
| 2024 | Leave No Patient Behind: Enhancing Medication Recommendation for Rare Disease PatientsabstractMedication recommendation systems have gained significant attention in healthcare as a means of providing tailored and effective drug combinations based on patients' clinical information. However, existing approaches often suffer from fairness issues, as recommendations tend to be more accurate for patients with common diseases compared to those with rare conditions. In this paper, we propose a novel model called Robust and Accurate REcommendations for Medication (RAREMed), which leverages the pretrain-finetune learning paradigm to enhance accuracy for rare diseases. RAREMed employs a transformer encoder with a unified input sequence approach to capture complex relationships among disease and procedure codes. Additionally, it introduces two self-supervised pre-training tasks, namely Sequence Matching Prediction (SMP) and Self Reconstruction (SR), to learn specialized medication needs and interrelations among clinical codes. Experimental results on two real-world datasets demonstrate that RAREMed provides accurate drug sets for both rare and common disease patients, thereby mitigating unfairness in medication recommendation systems. The implementation is available via https://github.com/zzhUSTC2016/RAREMed. Zihao Zhao 0004, Yi Jing, Fuli Feng, Jiancan Wu, Chongming Gao, Xiangnan He 0001 |
SIGIR | 2 |
| 2024 | VCF2PCACluster: a simple, fast and memory-efficient tool for principal component analysis of tens of millions of SNPsabstractPrincipal component analysis (PCA) is an important and widely used unsupervised learning method that determines population structure based on genetic variation. Genome sequencing of thousands of individuals usually generate tens of millions of SNPs, making it challenging for PCA analysis and interpretation. Here we present VCF2PCACluster, a simple, fast and memory-efficient tool for Kinship estimation, PCA and clustering analysis, and visualization based on VCF formatted SNPs. We implemented five Kinship estimation methods and three clustering methods for its users to choose from. Moreover, unlike other PCA tools, VCF2PCACluster possesses a clustering function based on PCA result, which enabling users to automatically and clearly know about population structure. We demonstrated the same accuracy but a higher performance of this tool in performing PCA analysis on tens of millions of SNPs compared to another popular PLINK2 software, especially in peak memory usage that is independent of the number of SNPs in VCF2PCACluster. Weiming He, Lian Xu, JingXian Wang, Zhen Yue, Yi Jing, Shuaishuai Tai, Xiaodong Fang |
BMC Bioinform. | 5 |
| 2023 | Soft Language Clustering for Multilingual Model Pre-trainingabstractJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016 |
ACL (1) | 4 |
| 2023 | NGenomeSyn: an easy-to-use and flexible tool for publication-ready visualization of syntenic relationships across multiple genomesabstractSUMMARY: Large-scale comparative genomic studies have provided important insights into species evolution and diversity, but also lead to a great challenge to visualize. Quick catching or presenting key information hidden in the vast amount of genomic data and relationships among multiple genomes requires an efficient visualization tool. However, current tools for such visualization remain inflexible in layout and/or require advanced computation skills, especially for visualization of genome-based synteny. Here, we developed an easy-to-use and flexible layout tool, NGenomeSyn [multiple (N) Genome Synteny], for publication-ready visualization of syntenic relationships of the whole genome or local region and genomic features (e.g. repeats, structural variations, genes) across multiple genomes with a high customization. NGenomeSyn provides an easy way for its users to visualize a large amount of data with a rich layout by simply adjusting options for moving, scaling, and rotation of target genomes. Moreover, NGenomeSyn could be applied on the visualization of relationships on non-genomic data with similar input formats. AVAILABILITY AND IMPLEMENTATION: NGenomeSyn is freely available at GitHub (https://github.com/hewm2008/NGenomeSyn) and Zenodo (https://doi.org/10.5281/zenodo.7645148). Weiming He, Yi Jing, Lian Xu, Xiaodong Fang |
Bioinform. | 3 |
| 2023 | Inferring circadian gene regulatory relationships from gene expression data with a hybrid frameworkabstractBACKGROUND: The central biological clock governs numerous facets of mammalian physiology, including sleep, metabolism, and immune system regulation. Understanding gene regulatory relationships is crucial for unravelling the mechanisms that underlie various cellular biological processes. While it is possible to infer circadian gene regulatory relationships from time-series gene expression data, relying solely on correlation-based inference may not provide sufficient information about causation. Moreover, gene expression data often have high dimensions but a limited number of observations, posing challenges in their analysis. METHODS: In this paper, we introduce a new hybrid framework, referred to as Circadian Gene Regulatory Framework (CGRF), to infer circadian gene regulatory relationships from gene expression data of rats. The framework addresses the challenges of high-dimensional data by combining the fuzzy C-means clustering algorithm with dynamic time warping distance. Through this approach, we efficiently identify the clusters of genes related to the target gene. To determine the significance of genes within a specific cluster, we employ the Wilcoxon signed-rank test. Subsequently, we use a dynamic vector autoregressive method to analyze the selected significant gene expression profiles and reveal directed causal regulatory relationships based on partial correlation. CONCLUSION: The proposed CGRF framework offers a comprehensive and efficient solution for understanding circadian gene regulation. Circadian gene regulatory relationships are inferred from the gene expression data of rats based on the Aanat target gene. The results show that genes Pde10a, Atp7b, Prok2, Per1, Rhobtb3 and Dclk1 stand out, which have been known to be essential for the regulation of circadian activity. The potential relationships between genes Tspan15, Eprs, Eml5 and Fsbp with a circadian rhythm need further experimental research. Shuwen Hu, Yi Jing, You-Gan Wang, Jing Gao 0006, Yu-Chu Tian |
BMC Bioinform. | 2 |
| 2022 | ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence GenerationabstractBei Li, Quan Du, Tao Zhou, Yi Jing, Shuhan Zhou, Xin Zeng, Tong Xiao, JingBo Zhu, Xuebo Liu, Min Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Quan Du, Yi Jing, Shuhan Zhou, Tong Xiao 0001, Xuebo Liu 0002, Min Zhang 0005 |
ACL (1) | 4 |
| 2022 | Learning Multiscale Transformer Models for Sequence GenerationabstractMultiscale feature hierarchies have been witnessed the success in the computer vision area. This further motivates researchers to design multiscale Transformer for natural language processing, mostly based on the self-attention mechanism. For example, restricting the receptive field across heads or extracting local fine-grained features via convolutions. However, most of existing works directly modeled local features but ignored the word-boundary information. This results in redundant and ambiguous attention distributions, which lacks of interpretability. In this work, we define those scales in different linguistic units, including sub-words, words and phrases. We built a multiscale Transformer model by establishing relationships among scales based on word-boundary information and phrase-level prior knowledge. The proposed \textbf{U}niversal \textbf{M}ulti\textbf{S}cale \textbf{T}ransformer, namely \textsc{Umst}, was evaluated on two sequence generation tasks. Notably, it yielded consistent performance gains over the strong baseline on several test sets without sacrificing the efficiency. Yi Jing, Chengbo Jiao, Tong Xiao 0001 |
ICML | 3 |
| 2019 | GCNDA: Graph Convolutional Networks with Dual Attention Mechanisms for Aspect Based Sentiment Analysis
Hongxu Hou, Jing Gao 0006, Yatu Ji, Tiangang Bai, Yi Jing |
ICONIP (4) | 6 |
| 2016 | Dialog state tracking with attention-based sequence-to-sequence learningabstractWe present an advanced dialog state tracking system designed for the 5th Dialog State Tracking Challenge (DSTC5). The main task of DSTC5 is to track the dialog state in a human-human dialog. For each utterance, the tracker emits a frame of slot-value pairs considering the full history of the dialog up to the current turn. Our system includes an encoder-decoder architecture with an attention mechanism to map an input word sequence to a set of semantic labels, i.e., slot-value pairs. This handles the problem of the unknown alignment between the utterances and the labels. By combining the attention-based tracker with rule-based trackers elaborated for English and Chinese, the F-score for the development set improved from 0.475 to 0.507 compared to the rule-only trackers. Moreover, we achieved 0.517 F-score by refining the combination strategy based on the topic and slot level performance of each tracker. In this paper, we also validate the efficacy of each technique and report the test set results submitted to the challenge. Takaaki Hori, Chiori Hori, Shinji Watanabe 0001, Bret Harsham, Jonathan Le Roux, John R. Hershey, Yusuke Koji, Yi Jing, Zhaocheng Zhu, Takeyuki Aikawa |
SLT | 9 |
| 2015 | Cognitive MU-MIMO Scheduling in Circular Array Based Heterogeneous NetworksabstractFuture heterogeneous networks (HetNets) will have to face a great challenge of overwhelming demand of spectrum resource, due to the exponential increase in mobile internet traffic driven by a new generation of wireless devices. In this paper, we propose a spectrum sensing and scheduling scheme for circular array, in order to make better use of the spectrum resource and improve the performance of multi-user MIMO (MU-MIMO) in HetNets. The proposed scheme can effectively detect the users and frequency use based on angles, and schedule the users with optimized codebook. Simulation results show that our proposed scheme can achieve considerable gain in terms of throughput and users' data rate, with significantly reduced system complexity and increased efficiency. Na Chen 0004, Songlin Sun, Bo Rong, Yi Jing, Rose Qingyang Hu, Yi Qian 0001 |
GLOBECOM | 4 |
| 2015 | A Novel MBSFN Scheme for Vehicle-to-Vehicle Safety Communication Based on LTE NetworkabstractVehicle-to-Vehicle (V2V) Safety Communication is committed to reduce traffic accidents by information exchange between vehicles. This paper studies the V2V safety communication based on cellular systems, especially LTE, where vehicles send Cooperative Awareness Messages (CAM) periodically to Base Stations (BSs) and the BSs transmit the CAMs to target vehicles through appropriate strategies. For the case of vehicle to multi-vehicle communication(usually cross cell), Multimedia Broadcast multicast service Single Frequency Network(MBSFN) can be adopted to achieve better performance compared with unicast because every message is only transmitted once reaching all target vehicles which reduce network load. A novel MBSFN scheme is designed in this paper. Simulation results show that our proposed MBSFN scheme can support more cars than the existing scheme under the same latency and QoS constraints. Mengfei Xie, Yong Shang, Yi Jing, Haijun Zhou |
VTC Fall | 4 |